• Media type: E-Article
  • Title: Dedoop : efficient deduplication with Hadoop : efficient deduplication with Hadoop
  • Contributor: Kolb, Lars; Thor, Andreas; Rahm, Erhard
  • imprint: Association for Computing Machinery (ACM), 2012
  • Published in: Proceedings of the VLDB Endowment
  • Language: English
  • DOI: 10.14778/2367502.2367527
  • ISSN: 2150-8097
  • Keywords: General Medicine
  • Origination:
  • Footnote:
  • Description: <jats:p>We demonstrate a powerful and easy-to-use tool called Dedoop (&lt;u&gt;De&lt;/u&gt;duplication with Ha&lt;u&gt;doop&lt;/u&gt;) for MapReduce-based entity resolution (ER) of large datasets. Dedoop supports a browser-based specification of complex ER workflows including blocking and matching steps as well as the optional use of machine learning for the automatic generation of match classifiers. Specified workflows are automatically translated into MapReduce jobs for parallel execution on different Hadoop clusters. To achieve high performance Dedoop supports several advanced load balancing strategies.</jats:p>