Comparable corpora usually comprise a limited number of situation types, mostly defined by simplistic pre-conceived labels of genre, such as fiction or conversation. In the framework of a research project investigating disfluency in spoken language, I have assembled the English-French comparable corpus collection “DisFrEn” to address this issue of text categorization. “DisFrEn” represents a panel of eight interactional situations, defined as combinations of six features assumed to influence the investigated phenomena, namely discourse markers. Most of them are scalar categories, inspired by Koch & Oesterreicher’s (2001) seminal work on speech styles. For instance, the feature “degree of interactivity” counts three levels : interactive (equal access to turn-taking), semi-interactive (asymmetrical access) and non-interactive (one-way interaction). This categorizing system makes it possible to compare situations like face-to-face and radio interviews, but also seemingly unrelated situations like sports commentaries and political speech (which share all features but the degree of preparedness). I will present the method and data used to compile “DisFrEn” along with preliminary results to illustrate its applications that a subcorpus of interviews.
Crible, L. (2014). Reaching cross-linguistic comparability across eight speech situations : a challenge for corpus design. 7th Biennial IVACS Conference, Newcastle, UK. https://hdl.handle.net/2078.5/92915