Finding local anomalies in very high dimensional space

Timothy De Vries, Sanjay Chawla, Michael E. Houle

Research output: Chapter in Book/Report/Conference proceedingConference contribution

44 Citations (Scopus)

Abstract

Time, cost and energy efficiency are critical factors for many data analysis techniques when the size and dimensionality of data is very large. We investigate the use of Local Outlier Factor (LOF) for data of this type, providing a motivating example from real world data. We propose Projection-Indexed Nearest-Neighbours (PINN), a novel technique that exploits extended nearest neighbour sets in the a reduced dimensional space to create an accurate approximation for k-nearest-neighbour distances, which is used as the core density measurement within LOF. The reduced dimensionality allows for efficient sub-quadratic indexing in the number of items in the data set, where previously only quadratic performance was possible. A detailed theoretical analysis of Random Projection (RP) and PINN shows that we are able to preserve the density of the intrinsic manifold of the data set after projection. Experimental results show that PINN outperforms the standard projection methods RP and PCA when measuring LOF for many high-dimensional real-world data sets of up to 300000 elements and 102600 dimensions.

Original languageEnglish
Title of host publicationProceedings - 10th IEEE International Conference on Data Mining, ICDM 2010
Pages128-137
Number of pages10
DOIs
Publication statusPublished - 1 Dec 2010
Event10th IEEE International Conference on Data Mining, ICDM 2010 - Sydney, NSW, Australia
Duration: 14 Dec 201017 Dec 2010

Publication series

NameProceedings - IEEE International Conference on Data Mining, ICDM
ISSN (Print)1550-4786

Conference

Conference10th IEEE International Conference on Data Mining, ICDM 2010
CountryAustralia
CitySydney, NSW
Period14/12/1017/12/10

    Fingerprint

Keywords

  • Anomaly detection
  • Dimensionality reduction

ASJC Scopus subject areas

  • Engineering(all)

Cite this

De Vries, T., Chawla, S., & Houle, M. E. (2010). Finding local anomalies in very high dimensional space. In Proceedings - 10th IEEE International Conference on Data Mining, ICDM 2010 (pp. 128-137). [5693966] (Proceedings - IEEE International Conference on Data Mining, ICDM). https://doi.org/10.1109/ICDM.2010.151