Enriching learner corpora with crowdsourced L2 proficiency assessment and the technique of adaptive comparative judgment

(2022) Linguists’ Day 2022 — Location: Liège, Belgium (21.October.2022)

Files

No attached file found for this publication.

Details

Authors
Abstract
Though proficiency is one of the most important constructs in Second Language Research, its measurement has not always received the attention it deserves, and practices of proficiency level assignment have been the subject of continued criticism (e.g. Hulstijn et al. 2010). In Learner Corpus Research, more particularly, Carlsen’s (2012) review of some of the most commonly used methods of proficiency-level assignments of texts showed that many learner corpora still rely on variables such as institutional status or year of study as a proxy for proficiency, despite the fact that these external criteria are largely regarded as unreliable (Thomas 1994). While a handful of learner corpora prove the exception by including reliable text-based proficiency scores (e.g. AndreSpråksKorpus, a learner corpus of Norwegian as a second language; Carlsen 2012), the time and cost difficulties typically associated with analytical scoring means that it is often absent from, or operationalized in unreliable ways, in learner corpora. This presentation has two main objectives. First, we will introduce the technique of adaptive comparative judgment (ACJ), coupled with a crowdsourcing approach, as a practical solution to the reliability issues as well as the time and cost difficulties associated with a text-based approach to proficiency assessment in learner corpus research. The method of CJ is based on Thurstone's (1927) ‘Law of Comparative Judgment’, which builds on the assumption that people are able to compare two performances more easily and reliably than to assign a score to an individual performance (Lesterhuis et al. 2017). The CJ approach involves the consensus of a panel of judges who are asked to compare two performances of any kind (be they dance performances, design portfolios, or, as in the present study, written learner productions) and to simply decide which of them is better. Under the ACJ framework, performances are paired adaptively, reducing the overall amount of comparisons required to achieve a reliable scale of performance abilities. A second critical assumption underpinning CJ is its reliance on holistic judgment: Judges do not receive criteria to guide their judgment process, but at best a general description regarding the competence to be assessed. We showcase this method by reporting on the methodological framework implemented in the Crowdsourcing Language Assessment Project (CLAP; Paquot et al., 2022) and presenting the results of a first pilot study that demonstrate that a crowd of 43 judges is able to assess (i.e. rank) 50 learner texts with high reliability (SSR = .95). No effect of language skills or language assessment experience was found on the assessment task, but there was a difference in the decisions made by judges who received formal language assessment training and those who did not. Nevertheless, the scores generated by the crowdsourced task exhibited a strong correlation with the rubric-based scores released with the learner corpus used (ETS Corpus of Non-Native Written English; Blanchard et al., 2014). The second objective of this presentation is to report on a newly launched collaborative initiative that aims to replicate and extend the pilot study described above by addressing some of the most pressing theoretical issues and avenues for future L2 research identified therein (Paquot et al., 2022). References: Paquot, M., Rubin, R. & Vandeweerd, N. (2022). Crowdsourced Adaptive Comparative Judgment: A community-based solution for proficiency rating. Language Learning 72(3), 853-885. https://onlinelibrary.wiley.com/doi/10.1111/lang.12498 Blanchard, D., Tetreault, J., Higgins, D., Cahill, A., & M. Chodorow (2014). ETS corpus of non-native written English. LDC2014T06. Philadelphia: Linguistic Data Consortium, 2014. https://catalog.ldc.upenn.edu/LDC2014T06 Carlsen, C. (2012). Proficiency level – a fuzzy variable in computer learner corpora. Applied Linguistics, 33, 161–183. Hulstijn, J., Alderson, C. & R. Schroonen (2010). Developmental stages in second-language acquisition and levels of second-language proficiency: Are there links between them? In Bartning, I., Martin, M. & I. Vedder (Eds.). Communicative Proficiency and Linguistic Development: Intersections between SLA and Language Testing Research. EUROSLA Monographs Series 1, 11-20. Lesterhuis, M., Verhavert, S., Coertjens, L., Donche, V., & S. De Maeyer (2017). Comparative judgement as a promising alternative to score competences. In Cano, E. & G. Ion (Eds.). Innovative Practices for Higher Education Assessment and Measurement. IGI Global, 119-138. Thomas, M. (1994). Assessment of L2 proficiency in second language acquisition research. Language Learning, 44, 307–336. Thurstone, L. L. (1927). A law of comparative judgment. Psychological Review, 34, 273-86.
Affiliations

Citations

Paquot, M., Thwaites, P., Rubin, R., & Vandeweerd, N. (2022). Enriching learner corpora with crowdsourced L2 proficiency assessment and the technique of adaptive comparative judgment. Linguists’ Day 2022, Liège, Belgium. https://hdl.handle.net/2078.5/106313