The IdiomSearch Project: from theory to a web application

(2017) Computational and corpus-based Phraseology — Location: London (13.November.2017)

Files

JPColsonLondon17.pdf
  • Open Access
  • Adobe PDF
  • 1.81 MB

Details

Authors
Abstract
This paper reports the preliminary results of an experiment carried out on a large scale for the extraction of PUs (phraseological units, also called idioms) from large web corpora in four languages (English, Spanish, French, Chinese). The use of a new algorithm based on metric clustering techniques, of optimized database storage and of interaction with users and researchers by means of a web application, made it possible to reach high precision scores for most common PUs in the four languages, while further experimentation is still necessary for establishing recall levels with long n-grams. In the meantime, the freely accessible web application makes it possible to visualize the high proportion of phraseology in the broad sense (or of formulaic language): about 30 to 60 percent of the newspaper articles tested in the experiments consisted of PUs. The most surprising results, however, came from Chinese: as the algorithm had to be changed for taking into account the associations between morphemes, the methodology used made it possible to partly confirm, from a statistical point of view, one of the major claims of construction grammar: the existence of a probabilistic network of constructions, from morphemes to idiomatic phrases.
Affiliations

Citations

Colson, J.-P. (2017). The IdiomSearch Project: from theory to a web application. Computational and corpus-based Phraseology, London. https://hdl.handle.net/2078.5/174823