Register variation is a crucial aspect of language production. Depending on the context in which it is used and the communicative purposes that it serves, language tends to display distinctive characteristics, which may have to do with lexis, but also phraseology or syntax, among others (e.g., Biber 2012). While register has important implications for any type of language production, it is particularly relevant to explore in learner language, because second language (L2) learners may not show the same register awareness as native or expert writers/speakers (Gilquin & Paquot 2008). Corpus-based explorations of language have raised awareness of register variation, and the study results have provided valuable insights into linguistic patterns associated with different registers (Biber et al. 1999). Specific methods have also been developed to study register variation, most notably multi-dimensional analysis (Biber 1988). In learner corpus research, it is common to focus on a specific register (e.g. telecollaborative discourse in Vyatkina 2012). However, studies comparing learner language registers are still relatively rare, with a few exceptions such as Fuchs et al. (2016) and Larsson et al. (2021). One reason for the lack of register studies in learner corpus research is that, until recently, the most widely used learner corpora have covered a limited range of registers, most notably argumentative essays for writing (as in ICLE, Granger et al. 2020) and interviews for speech (as in LINDSEI, Gilquin et al. 2010). In addition, when different registers have been compared, the analysis has mainly been based on texts produced by different groups of learners (e.g. argumentative essays produced by one group of students and interviews produced by another group). However, collecting texts written by the same learners across registers makes it possible to investigate how they adapt their language use to different communicative situations (e.g. Kerz et al. 2022). This paper sets out to describe the compilation of the STudent speech and writing Across Registers (STAR) corpus, a new corpus of student language productions that brings together texts from multiple registers produced by the same L2 English or L1 English students. The L2 English component of the STAR corpus contains data collected at UCLouvain (mainly) from French-speaking learners of English who are students in their second year of English major studies. The L1 English component comprises data collected from first-year students enrolled in a variety of degree programmes at Northern Arizona University (NAU). The corpus contains data from the following written registers: career readiness essays, cover letters, persuasive essays, critical literacy narratives and diary entries. To ensure comparability across the dataset, we controlled for task by assigning the same or very similar writing prompts to all students. As such, the STAR corpus offers a unique opportunity to examine how register influences written language use in learner and native productions while minimizing task-related variability. Maximising comparability across spoken registers proved challenging due to differences in course organisation at the two institutions. At UCLouvain, the spoken registers include a monologue about students’ future career, a debate and an informal conversation between two students. By contrast, the spoken data collected so far at NAU are connected with the persuasive essay. They include dialogues which are brainstorms for the essay and monologues in which students reflect on the essay. We made special efforts to collect rich metadata about the L1 and L2 students and the pedagogical tasks used to elicit language production across registers, relying on Paquot et al.’s (2024) Core Metadata Schema for Learner Corpora. We included information such as learners’ linguistic background (i.e. L1s and any possible additional language(s)), their exposure to English in different situations, but also their degree of literacy. Detailed metadata are also provided about the tasks (e.g. instructions, time constraints and use of language reference tools for writing) and the characteristics of the resulting registers (e.g. communicative purposes, settings, number of addressees). Comparable metadata about the L1 English students were also recorded. Once completed, the STAR corpus will enable researchers to make comparisons across registers while controlling for individual variables and styles, and to explore the effect of register on the linguistic features of L1 and L2 novice writers and speakers of English. STAR will be released in open access format.
De Cock, S., Egbert, J., Gilquin, G., Granger, S., Grixoni, F., Holmberg, A. J., Jadoulle, P., Larsson, T., & Paquot, M. (2025). The STAR corpus: Student speech and writing across registers. Register and task variation in Learner Corpus Research (VAR4LCR) conference, Université catholique de Louvain. https://hdl.handle.net/2078.5/274957