(en) Disfluencies, phenomenon specific to spoken texts, have the characteristic of breaking the syntactic linearity of speech. Multiword lexical units tend to form semantic units. In this paper, we show, through an automatic procedure, that multiword lexical units are less likely to contain a disfluency than a free sequence of words. The procedure consists first in locating disfluencies with a resource-free iterative algorithm, and then in recognizing multiword units thanks to a probabilistic model.
Constant, M., Dister, A., & et al. (2012). Mots composés et disfluences. 11es JADT, Université de Liège, Belgique. https://hdl.handle.net/2078.5/169813