Set phrases around globalization : an experiment in corpus-based computational phraseology

(2016) Input a Word, Analyze the World. Selected Approaches to Corpus Linguistics — ISBN: [978-1-4438-8513-3], p. 141-152, published

Files

PaperJPColsonLP14Final.pdf
  • Closed Access
  • Adobe PDF
  • 228.75 KB

Details

Authors
Abstract
The aim of this study is to investigate whether the automatic extraction of set phrases around the term globalization yields interesting semantic information and may therefore contribute to a better understanding of how the notion is perceived by society in different languages. For the purpose of this experiment, a reasonable compromise had to be reached between the size of the corpora and the time needed for computational processing. We therefore used corpora of 200 million tokens: a portion of the WaCky corpora (Baroni et al. 2009) for English, Italian, French and German. We then created similar web corpora of the same size for Spanish and Dutch, by using the WebBootCat tool provided by the Sketch Engine. Apart from those 200 million word corpora in 6 languages, reflecting general language use, we also compiled 6 specialized corpora (one for each language), of 5 million words each, by using globalization and related terms as seed words with WebBootCat. The algorithm used here for the automatic extraction of phraseology from the corpora is non-parametric and does not directly derive from the theory of statistics. It is an adaptation of a well-known technique used in information retrieval (Baeza-Yates & Ribeiro-Neto 1999): metric clusters. The results of this experiment show that, in the 6 languages, the notion of globalization has already pervaded several layers of phraseology in the broad sense, especially clichés, routines and formulae, as illustrated by : in the era of globalization (English), à l'heure de la mondialisation (French), nell'epoca della globalizzazione (Italian), la era de la globalización (Spanish). In all 6 languages, set phrases around globalization contain references to a few common notions such as time, context and (neo-)liberal. The most striking finding, however, is that corpora of the same size, analyzed by the same algorithm within the framework of computational phraseology, reveal crucial differences between the 6 language communities: set phrases highlighting some negative aspects of globalization, for instance, are much clearer in French, Italian, and Spanish than in English, whereas German and Dutch use very few negative set phrases with globalization. The experiment also confirms that computational phraseology provides useful additional tools for deriving semantic and sociolinguistic information from large linguistic corpora.
Affiliations

Citations

Colson, J.-P. (2016). Set phrases around globalization : an experiment in corpus-based computational phraseology. In F. Alonso Almeida, I. Ortega Barrera, E. Quintana Toledo & M. E. Sanchez Cuervo (eds) (ed.), Input a Word, Analyze the World. Selected Approaches to Corpus Linguistics (p. p. 141-152). Cambridge Scholars Publishing. https://hdl.handle.net/2078.5/185387