Mining for outliers in sequential databases

Pei Sun, Sanjay Chawla, Bavani Arunasalam

Research output: Chapter in Book/Report/Conference proceedingConference contribution

59 Citations (Scopus)

Abstract

The mining of outliers (or anomaly detection) in large databases continues to remain an active area of research with many potential applications. Over the last several years many novel methods have been proposed to efficiently and accurately mine for outliers. In this paper we propose a unique approach to mine for sequential outliers using Probabilistic Suffix Trees (PST). The key insight that underpins our work is that we can distinguish outliers from non-outliers by only examining the nodes close to the root of the PST. Thus, if the goal is to just mine outliers, then we can drastically reduce the size of the PST and reduce its construction and query time. In our experiments, we show that on a real data set consisting of protein sequences, by retaining less than 5% of the original PST we can retrieve all the outliers that were reported by the full-sized PST. We also carry out a detailed comparison between two measures of sequence similarity: the normalized probability and the odds and show that while the current research literature in PST favours the odds, for outlier detection it is normalized probability which gives far superior results. We provide an information theoretic argument based on entropy to explain the success of the normalized probability measure. Finally, we describe a more efficient implementation of the PST algorithm, which dramatically reduces its construction time compared to the implementation of Bejerano [3].

Original languageEnglish
Title of host publicationProceedings of the Sixth SIAM International Conference on Data Mining
Pages94-105
Number of pages12
Publication statusPublished - 3 Jul 2006
EventSixth SIAM International Conference on Data Mining - Bethesda, MD, United States
Duration: 20 Apr 200622 Apr 2006

Publication series

NameProceedings of the Sixth SIAM International Conference on Data Mining
Volume2006

Conference

ConferenceSixth SIAM International Conference on Data Mining
CountryUnited States
CityBethesda, MD
Period20/4/0622/4/06

    Fingerprint

ASJC Scopus subject areas

  • Engineering(all)

Cite this

Sun, P., Chawla, S., & Arunasalam, B. (2006). Mining for outliers in sequential databases. In Proceedings of the Sixth SIAM International Conference on Data Mining (pp. 94-105). (Proceedings of the Sixth SIAM International Conference on Data Mining; Vol. 2006).