A Modality Lexicon and its use in Automatic Tagging

Baker, Kathryn; Bloodgood, Michael; Dorr, Bonnie; Filardo, Nathaniel; Levin, Lori; Piatko, Christine

A Modality Lexicon and its use in Automatic Tagging

dc.contributor.author	Baker, Kathryn
dc.contributor.author	Bloodgood, Michael
dc.contributor.author	Dorr, Bonnie
dc.contributor.author	Filardo, Nathaniel
dc.contributor.author	Levin, Lori
dc.contributor.author	Piatko, Christine
dc.date.accessioned	2014-08-11T17:44:55Z
dc.date.available	2014-08-11T17:44:55Z
dc.date.issued	2010-05
dc.description.abstract	This paper describes our resource-building results for an eight-week JHU Human Language Technology Center of Excellence Summer Camp for Applied Language Exploration (SCALE-2009) on Semantically-Informed Machine Translation. Specifically, we describe the construction of a modality annotation scheme, a modality lexicon, and two automated modality taggers that were built using the lexicon and annotation scheme. Our annotation scheme is based on identifying three components of modality: a trigger, a target and a holder. We describe how our modality lexicon was produced semi-automatically, expanding from an initial hand-selected list of modality trigger words and phrases. The resulting expanded modality lexicon is being made publicly available. We demonstrate that one tagger—a structure-based tagger—results in precision around 86% (depending on genre) for tagging of a standard LDC data set. In a machine translation application, using the structure-based tagger to annotate English modalities on an English-Urdu training corpus improved the translation quality score for Urdu by 0.3 Bleu points in the face of sparse training data.	en_US
dc.description.sponsorship	This work is supported, in part, by the Johns Hopkins Human Language Technology Center of Excellence. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the sponsor.	en_US
dc.identifier.citation	Kathryn Baker, Michael Bloodgood, Bonnie J. Dorr, Nathaniel W. Filardo, Lori Levin, and Christine Piatko. 2010. A modality lexicon and its use in automatic tagging. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10), pages 1402-1407, Valletta, Malta, May. European Language Resources Association.	en_US
dc.identifier.uri	http://hdl.handle.net/1903/15554
dc.language.iso	en_US	en_US
dc.publisher	European Language Resources Association	en_US
dc.relation.isAvailableAt	Center for Advanced Study of Language
dc.relation.isAvailableAt	Digitial Repository at the University of Maryland
dc.relation.isAvailableAt	University of Maryland (College Park, Md)
dc.rights.license	Published with the permission of ELRA. This paper was published within the proceedings of the LREC 2010 Conference. © 1998-2012 ELRA - European Language Resources Association. All rights reserved.
dc.subject	computer science	en_US
dc.subject	computational linguistics	en_US
dc.subject	modality	en_US
dc.subject	modality lexicon	en_US
dc.subject	modality annotation	en_US
dc.subject	automatic modality tagging	en_US
dc.title	A Modality Lexicon and its use in Automatic Tagging	en_US
dc.type	Article	en_US

Files

Original bundle

Now showing 1 - 1 of 1

Name:: modalityTaggingLREC2010.pdf
Size:: 352.23 KB
Format:: Adobe Portable Document Format

Download

Collections

Center for Advanced Study of Language Research Works