Sampling frequent and minimal boolean patterns: theory and application in classification

Geng Li, Mohammed J. Zaki

Research output: Contribution to journalArticle

5 Citations (Scopus)

Abstract

We tackle the challenging problem of mining the simplest Boolean patterns from categorical datasets. Instead of complete enumeration, which is typically infeasible for this class of patterns, we develop effective sampling methods to extract a representative subset of the minimal Boolean patterns in disjunctive normal form (DNF). We propose a novel theoretical characterization of the minimal DNF expressions, which allows us to prune the pattern search space effectively. Our approach can provide a near-uniform sample of the minimal DNF patterns. We perform an extensive set of experiments to demonstrate the effectiveness of our sampling method. We also show that minimal DNF patterns make effective features for classification.

Original languageEnglish
Pages (from-to)181-225
Number of pages45
JournalData Mining and Knowledge Discovery
Volume30
Issue number1
DOIs
Publication statusPublished - 1 Jan 2016

    Fingerprint

Keywords

  • Classification
  • Disjunctive patterns
  • Frequent pattern mining
  • Markov chain monte carlo
  • Minimal boolean expressions
  • Minimal generators
  • Pattern sampling

ASJC Scopus subject areas

  • Information Systems
  • Computer Science Applications
  • Computer Networks and Communications

Cite this