Maximum entropy models and subjective interestingness: an application to tiles in binary databases
From MaRDI portal
Abstract: Recent research has highlighted the practical benefits of subjective interestingness measures, which quantify the novelty or unexpectedness of a pattern when contrasted with any prior information of the data miner (Silberschatz and Tuzhilin, 1995; Geng and Hamilton, 2006). A key challenge here is the formalization of this prior information in a way that lends itself to the definition of an interestingness subjective measure that is both meaningful and practical. In this paper, we outline a general strategy of how this could be achieved, before working out the details for a use case that is important in its own right. Our general strategy is based on considering prior information as constraints on a probabilistic model representing the uncertainty about the data. More specifically, we represent the prior information by the maximum entropy (MaxEnt) distribution subject to these constraints. We briefly outline various measures that could subsequently be used to contrast patterns with this MaxEnt model, thus quantifying their subjective interestingness.
Recommendations
Cites work
- Discovery Science
- Emergence of Scaling in Random Networks
- Graphical models, exponential families, and variational inference
- scientific article; zbMATH DE number 3175697 (Why is no real title available?)
- scientific article; zbMATH DE number 3868459 (Why is no real title available?)
- scientific article; zbMATH DE number 3580314 (Why is no real title available?)
- scientific article; zbMATH DE number 3621626 (Why is no real title available?)
- Information Theory and Statistical Mechanics
- Itemset frequency satisfiability: complexity and axiomatization
- Krimp: mining itemsets that compress
- Testing Statistical Hypotheses
- The Average Distance in a Random Graph with Given Expected Degrees
- The budgeted maximum coverage problem
- The Structure and Function of Complex Networks
Cited in
(21)- The PRIMPING routine -- tiling through proximal alternating linearized minimization
- Summarizing categorical data by clustering attributes
- Discovering subjectively interesting multigraph patterns
- Subjectively interesting connecting trees and forests
- SIAS-miner: mining subjectively interesting attributed subgraphs
- Subjectively interesting alternative clusterings
- The blind men and the elephant: on meeting the problem of multiple truths in data from clustering and pattern mining perspectives
- A statistical significance testing approach to mining the most informative set of patterns
- Online summarization of dynamic graphs using subjective interestingness for sequential data
- Mining explainable local and global subgraph patterns with surprising densities
- Subjective interestingness of subgraph patterns
- Mining Compressing Sequential Patterns
- Interactive knowledge discovery from hidden data through sampling of frequent patterns
- Guided visual exploration of relations in data sets
- ROhAN: row-order agnostic null models for statistically-sound knowledge discovery
- What to expect from a set of itemsets?
- Constrained clustering: current and new trends
- Leveraging internal representations of GNNs with Shapley values
- SimHawNet: a modified Hawkes process for temporal network simulation
- Mining diverse sets of patterns with constraint programming using the pairwise Jaccard similarity relaxation
- Interesting pattern mining in multi-relational data
This page was built for publication: Maximum entropy models and subjective interestingness: an application to tiles in binary databases
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q408667)