Data Cleaning for XML Electronic Dictionaries via Statistical Anomaly Detection

Bloodgood, Michael; Strauss, Benjamin

Data Cleaning for XML Electronic Dictionaries via Statistical Anomaly Detection

dc.contributor.author	Bloodgood, Michael
dc.contributor.author	Strauss, Benjamin
dc.date.accessioned	2016-04-13T19:48:31Z
dc.date.available	2016-04-13T19:48:31Z
dc.date.issued	2016
dc.description.abstract	Many important forms of data are stored digitally in XML format. Errors can occur in the textual content of the data in the fields of the XML. Fixing these errors manually is time-consuming and expensive, especially for large amounts of data. There is increasing interest in the research, development, and use of automated techniques for assisting with data cleaning. Electronic dictionaries are an important form of data frequently stored in XML format that frequently have errors introduced through a mixture of manual typographical entry errors and optical character recognition errors. In this paper we describe methods for flagging statistical anomalies as likely errors in electronic dictionaries stored in XML format. We describe six systems based on different sources of information. The systems detect errors using various signals in the data including uncommon characters, text length, character-based language models, word-based language models, tied-field length ratios, and tied-field transliteration models. Four of the systems detect errors based on expectations automatically inferred from content within elements of a single field type. We call these single-field systems. Two of the systems detect errors based on correspondence expectations automatically inferred from content within elements of multiple related field types. We call these tied-field systems. For each system, we provide an intuitive analysis of the type of error that it is successful at detecting. Finally, we describe two larger-scale evaluations using crowdsourcing with Amazon’s Mechanical Turk platform and using the annotations of a domain expert. The evaluations consistently show that the systems are useful for improving the efficiency with which errors in XML electronic dictionaries can be detected.	en_US
dc.identifier	https://doi.org/10.13016/M2RT7D
dc.identifier.citation	Michael Bloodgood and Benjamin Strauss. Data cleaning for XML electronic dictionaries via statistical anomaly detection. In Proceedings of the 2016 IEEE Tenth International Conference on Semantic Computing (ICSC), pages 79-86, Laguna Hills, CA, USA, February 2016. IEEE.	en_US
dc.identifier.other	DOI 10.1109/ICSC.2016.38
dc.identifier.uri	http://ieeexplore.ieee.org/xpl/articleDetails.jsp?arnumber=7439308
dc.identifier.uri	http://hdl.handle.net/1903/17459
dc.language.iso	en_US	en_US
dc.publisher	IEEE	en_US
dc.relation.isAvailableAt	Center for Advanced Study of Language
dc.relation.isAvailableAt	Digitial Repository at the University of Maryland
dc.relation.isAvailableAt	University of Maryland (College Park, Md)
dc.subject	computer science	en_US
dc.subject	statistical methods	en_US
dc.subject	databases	en_US
dc.subject	Optical Character Recognition	en_US
dc.subject	artificial intelligence	en_US
dc.subject	machine learning	en_US
dc.subject	computational linguistics	en_US
dc.subject	natural language processing	en_US
dc.subject	human language technology	en_US
dc.subject	text processing	en_US
dc.subject	data cleaning	en_US
dc.subject	data cleansing	en_US
dc.subject	crowdsourcing	en_US
dc.subject	Amazon Mechanical Turk	en_US
dc.subject	XML	en_US
dc.subject	electronic lexicography	en_US
dc.subject	digital dictionaries	en_US
dc.subject	semantic computing	en_US
dc.subject	anomaly detection	en_US
dc.subject	error detection	en_US
dc.title	Data Cleaning for XML Electronic Dictionaries via Statistical Anomaly Detection	en_US
dc.type	Article	en_US

Files

Original bundle

Now showing 1 - 1 of 1

Name:: dataCleaningICSC2016.pdf
Size:: 324.59 KB
Format:: Adobe Portable Document Format
Description:

Download

License bundle

Now showing 1 - 1 of 1

Name:: license.txt
Size:: 1.57 KB
Format:: Item-specific license agreed upon to submission
Description:

Download

Collections

Center for Advanced Study of Language Research Works