Partial parsing of spontaneous spoken French

Blanc, Olivier;Constant, Matthieu;Dister, Anne;Watrin, Patrick
(2010) International Conference on Language Resources and Evaluation — Location: Malte (17.May.2010)

Files

No attached file found for this publication.

Details

Authors
Abstract
This paper describes the process and the resources used to automatically annotate a French corpus of spontaneous speech transcriptions in super-chunks. Super-chunks are enhanced chunks that can contain lexical multiword units. This partial parsing is based on a preprocessing stage of the spoken data that consists in reformatting and tagging utterances that break the syntactic structure of the text, such as disfluencies. Spoken specificities were formalized thanks to a systematic linguistic study of a 40-hour-long speech transcription corpus. The chunker uses large-coverage and fine-grained language resources for general written language that have been augmented with resources specific to spoken French. It consists in iteratively applying finite-state lexical and syntactic resources and outputing a finite automaton representing all possible chunk analyses. The best path is then selected thanks to a hybrid disambiguation stage. We show that our system reaches scores that are comparable with state-of-the-art results in the field.
Affiliations

Citations

Blanc, O., Constant, M., Dister, A., & Watrin, P. (2010). Partial parsing of spontaneous spoken French. Proceedings of the 7th International Conference on Language Resources and Evaluation (LREC 2010), p. 2111-2116. https://hdl.handle.net/2078.5/223104