Parsing Models for Identifying Multiword Expressions

Green, Spence;de Marneffe, Marie-Catherine;Manning, Christopher D.
(2013) Computational Linguistics (Print) — Vol. 39, n° 1, p. 195-227 (2013)

Files

J13-1009.pdf
  • Open Access
  • Adobe PDF
  • 224.83 KB

Details

Authors
Abstract
Multiword expressions lie at the syntax/semantics interface and have motivated alternative theories of syntax like Construction Grammar. Until now, however, syntactic analysis and multiword expression identification have been modeled separately in natural language processing. We develop two structured prediction models for joint parsing and multiword expression identification. The first is based on context-free grammars and the second uses tree substitution grammars, a formalism that can store larger syntactic fragments. Our experiments show that both models can identify multiword expressions with much higher accuracy than a state-of-theart system based on word co-occurrence statistics. We experiment with Arabic and French, which both have pervasive multiword expressions. Relative to English, they also have richer morphology, which induces lexical sparsity in finite corpora. To combat this sparsity, we develop a simple factored lexical representation for the context-free parsing model. Morphological analyses are automatically transformed into rich feature tags that are scored jointly with lexical items. This technique, which we call a factored lexicon, improves both standard parsing and multiword expression identification accuracy.
Affiliations

Citations

Green, S., de Marneffe, M.-C., & Manning, C. D. (2013). Parsing Models for Identifying Multiword Expressions. Computational Linguistics (Print), 39(1), 195-227. https://doi.org/10.1162/coli_a_00139 (Original work published 2013)