Traditionally, corpora have been treated by corpus linguists as 'one big text', in which "the data obtained from (...) different speakers or writers are pooled" (Rietveld et al. 2004: 350). This approach relies on the assumption that well-sampled corpora allow for reliable and valid generalisations about a population as a whole (see Kennedy 1998: 74). However, the fact that "corpora are inherently variable internally" (Gries 2006: 110) suggests that, while generalisations about populations may still be valid and useful, interesting findings are also likely to emerge if we investigate corpus data as a series of individual texts rather than as an aggregate. This is particularly true of learner corpora, because of the "highly heterogeneous nature of learner language" (Granger et al. 2009: 3; see also Durrant & Schmitt 2009: 168). Following recent studies like Paquot (2010) which take account of the possible variance within a corpus, we will adopt both a global and individual approach to the study of discourse markers in native and non-native speech. Our data will come from the newly published Louvain International Database of Spoken English Interlanguage (LINDSEI, see Gilquin et al. 2010) and its native counterpart, the Louvain Corpus of Native English Conversation (LOCNEC, see De Cock 2004). Starting from the whole of LINDSEI and LOCNEC, we will show how the use and frequency of discourse markers such as 'you know' or 'I mean' differ in native and non-native English. This level of analysis reveals, for instance, an underuse of 'sort of' in non-native English as compared to native English. At the next level of analysis, we will make a distinction between the different learner populations represented in LINDSEI, that is, the groups of learners who share the same mother tongue. This will enable us to highlight features that seem to be transfer-related, such as for example the heavy overuse of 'in fact' in the French component of LINDSEI. Finally, we will consider individual speakers in LINDSEI and LOCNEC in an attempt to identify idiosyncratic features that are limited to just a few speakers. By adopting this threefold level of analysis, we hope to shed new light on the use of discourse markers by native speakers and learners of English and to distinguish between the characteristics that are typical of native or non-native speech in general, those that are limited to certain populations and those that are only found among certain speakers. More generally, we wish to advocate for a combined approach in (learner) corpus research which takes into consideration both the pooled data and the individual texts making up a corpus.
Gilquin, G., & Granger, S. (2011). The use of discourse markers in corpora of native and learner speech: From aggregate to individual data. Corpus Linguistics conference, Birmingham. https://hdl.handle.net/2078.5/229726