Purdue University Purdue e-Pubs Department of Computer Science Technical Reports Department of Computer Science 2003 Record Linkage: A Machine Learning Approach, A Toolbox, and a Digital Government Web Service Mohamed G. Elfeky Vassilios S. Verykios Ahmed K. Elmagarmid Purdue University,
[email protected] Thanaa M. Ghanem Ahmed R. Huwait Report Number: 03-024 Elfeky, Mohamed G.; Verykios, Vassilios S.; Elmagarmid, Ahmed K.; Ghanem, Thanaa M.; and Huwait, Ahmed R., "Record Linkage: A Machine Learning Approach, A Toolbox, and a Digital Government Web Service" (2003). Department of Computer Science Technical Reports. Paper 1573. https://docs.lib.purdue.edu/cstech/1573 This document has been made available through Purdue e-Pubs, a service of the Purdue University Libraries. Please contact
[email protected] for additional information. RECORD LINKAGE, A MACHINE LEARNING APPROACH, A TOOLBOX, AND A DIGITAL GOVERNMENT WEB SERVICE Mohamed G. Elfeky Vassilios S. Verykios Ahmed K. Elmagarmid Thanaa M. Ghanem Ahmed R. Huwait CSD TR #03-024 July 2003 Record Linkage: A Machine Learning Approach, A Toolbox, and A Digital Government Web Service' l Mohamed G. Elfeky 1 VassiliosS. Verykios Ahmed K. Elmagarrnid 1 Purdue University Drexel University Hewlett-Packard
[email protected] [email protected] [email protected] Thanaa M. Ghanem Ahmed R. Huwait 2 Purdue University Ain Shams University
[email protected],edu
[email protected] Abstract Data cleaning is a vital process mat ensures the quality of data smred in real·world databases. Data cleaning problems are frequently encountered in many research areas, such as kllowledge discovery in databases, data warellOusing. system imegratioll and e services.