The first learner corpora that were compiled consisted of written data produced by foreign/second language learners, e.g. the Longman Learners’ Corpus or the International Corpus of Learner English. More recently, spoken learner corpora have became available too, thus opening up new possibilities for the analysis of interlanguage. In this talk, I will focus on spoken learner corpora, and how they can be developed and exploited for research and teaching purposes. I will describe some of the challenges that spoken learner corpora can present, starting with the compilation of such corpora, which is a time-consuming and painstaking process, and ending with their use in the language classroom, which should only be done cautiously and with proper guidance. I will also highlight the many opportunities that are offered by spoken learner corpora for the description of interlanguage as well as the learning and teaching of languages. This will involve a review of some of the ways in which spoken learner corpora can be further developed (e.g. through annotation) and the possible applications they can lead to. These considerations will be illustrated through concrete examples from the Louvain International Database of Spoken English Interlanguage (LINDSEI) and other spoken learner corpora.
Gilquin, G. (2015). Developing and exploiting spoken learner corpora: challenges and opportunities. Workshop on New Trends in Spoken Corpora, Universidade de Santiago de Compostela. https://hdl.handle.net/2078.5/188454