Don't be SCAREd: Use SCalable Automatic REpairing with maximal likelihood and bounded changes

Research output: Chapter in Book/Report/Conference proceedingConference contribution

64 Citations (Scopus)

Abstract

Various computational procedures or constraint-based methods for data repairing have been proposed over the last decades to identify errors and, when possible, correct them. However, these approaches have several limitations including the scalability and quality of the values to be used in replacement of the errors. In this paper, we propose a new data repairing approach that is based on maximizing the likelihood of replacement data given the data distribution, which can be modeled using statistical machine learning techniques. This is a novel approach combining machine learning and likelihood methods for cleaning dirty databases by value modification. We develop a quality measure of the repairing updates based on the likelihood benefit and the amount of changes applied to the database. We propose SCARE (SCalable Automatic REpairing), a systematic scalable framework that follows our approach. SCARE relies on a robust mechanism for horizontal data partitioning and a combination of machine learning techniques to predict the set of possible updates. Due to data partitioning, several updates can be predicted for a single record based on local views on each data partition. Therefore, we propose a mechanism to combine the local predictions and obtain accurate final predictions. Finally, we experimentally demonstrate the effectiveness, efficiency, and scalability of our approach on real-world datasets in comparison to recent data cleaning approaches.

Original languageEnglish
Title of host publicationSIGMOD 2013 - International Conference on Management of Data
Pages553-564
Number of pages12
DOIs
Publication statusPublished - 29 Jul 2013
Event2013 ACM SIGMOD Conference on Management of Data, SIGMOD 2013 - New York, NY, United States
Duration: 22 Jun 201327 Jun 2013

Publication series

NameProceedings of the ACM SIGMOD International Conference on Management of Data
ISSN (Print)0730-8078

Other

Other2013 ACM SIGMOD Conference on Management of Data, SIGMOD 2013
CountryUnited States
CityNew York, NY
Period22/6/1327/6/13

    Fingerprint

Keywords

  • Data cleaning
  • Inconsistent data

ASJC Scopus subject areas

  • Software
  • Information Systems

Cite this

Yakout, M., Berti-Équille, L., & Elmagarmid, A. K. (2013). Don't be SCAREd: Use SCalable Automatic REpairing with maximal likelihood and bounded changes. In SIGMOD 2013 - International Conference on Management of Data (pp. 553-564). (Proceedings of the ACM SIGMOD International Conference on Management of Data). https://doi.org/10.1145/2463676.2463706