LINDSEI: A new resource on the market of learner corpora

(2010) The ICLE/LINDSEI Symposium — Location: Tokyo (30.October.2010)

Files

No attached file found for this publication.

Details

Authors
Abstract
Written learner corpora have become a relatively common resource on the market of corpora. Spoken learner corpora, on the other hand, are still very rare. The LINDSEI (Louvain International Database of Spoken English Interlanguage) project was meant to fill this gap. It was launched in 1995 at the University of Louvain as the spoken counterpart of the International Corpus of Learner English (ICLE). LINDSEI is a collection of spoken data produced by higher intermediate to advanced learners of English as a foreign language. The collaboration with several universities internationally made it possible to include data from learners with a wide variety of mother tongue backgrounds. To date, eleven mother tongues are represented: Bulgarian, Chinese, Dutch, French, German, Greek, Italian, Japanese, Polish, Spanish and Swedish. The data consist of informal interviews that were transcribed orthographically (with some prosodic and phonetic information like pauses or syllable lengthening), following guidelines which were specifically designed for the project and which were unified across the subcorpora to ensure perfect comparability of the data. The transcripts also include some special markup, for example for overlapping speech or foreign words. Each interview is accompanied by a learner profile recording a number of variables such as the learner’s age, his/her knowledge of other foreign languages, the interviewer’s mother tongue or the duration of the interview. In this workshop I will describe the genesis of LINDSEI and the steps involved in its compilation. I will also present its structure, the transcription and markup conventions it relies on, and the metalinguistic data it includes. I will then demo the LINDSEI software, which allows users to compile their own tailor-made corpus on the basis of a set of predefined variables and to extract useful statistics. It will be shown how the corpus thus compiled can be imported into a concordancer like WordSmith Tools for further analysis. Several methods of analysis of the LINDSEI data will be mentioned, in particular when used in combination with LOCNEC (Louvain Corpus of Native English Conversation), a corpus of spoken data produced by native British English students which is the exact replica of LINDSEI. Finally, I will give an overview of the contents of the handbook that accompanies the LINDSEI CD-ROM.
Affiliations

Citations

Gilquin, G. (2010). LINDSEI: A new resource on the market of learner corpora. The ICLE/LINDSEI Symposium, Tokyo. https://hdl.handle.net/2078.5/158603