Introducing the CLAP project: Adaptive Comparative Judgment as a community-based solution for enriching learner corpora with crowdsourced L2 proficiency assessment

Paquot, Magali;Rubin, Rachel;Vandeweerd, Nathan
(2022) Learner Corpus Research 6 — Location: Padua, Italy (22.September.2022)

Files

No attached file found for this publication.

Details

Authors
Abstract
Though proficiency is one of the most important constructs in Second Language Research, its measurement has not always received the attention it deserves, and practices of proficiency level assignment have been the subject of continued criticism (e.g. Hulstijn et al. 2010). In Learner Corpus Research, more particularly, Carlsen’s (2012) review of some of the most commonly used methods of proficiency-level assignments of texts showed that many learner corpora still rely on variables such as institutional status or year of study as a proxy for proficiency, despite the fact that these external criteria are largely regarded as unreliable (Thomas 1994). While a handful of learner corpora prove the exception by including reliable text-based proficiency scores (e.g. AndreSpråksKorpus, a learner corpus of Norwegian as a second language; Carlsen 2012), the time and cost difficulties typically associated with analytical scoring means that it is often absent from, or operationalized in unreliable ways, in learner corpora. This presentation has two main objectives. First, we will introduce the technique of adaptive comparative judgment (ACJ), coupled with a crowdsourcing approach, as a practical solution to the reliability issues as well as the time and cost difficulties associated with a text-based approach to proficiency assessment in learner corpus research. The method of CJ is based on Thurstone's (1927) ‘Law of Comparative Judgment’, which builds on the assumption that people are able to compare two performances more easily and reliably than to assign a score to individual performance (Lesterhuis et al. 2017). The CJ approach involves the consensus of a panel of judges who are asked to compare two performances of any kind (be they dance performances, design portfolios, or, as in the present study, written learner productions) and to simply decide which of them is better. Under the ACJ framework, performances are paired adaptively, reducing the overall amount of comparisons required to achieve a reliable scale of performance abilities. A second critical assumption underpinning CJ is its reliance on holistic judgment: Judges do not receive criteria to guide their judgment process, but at best a general description regarding the competence to be assessed. We showcase this method by reporting on the methodological framework implemented in the CLAP project and presenting the results of a first pilot study that demonstrate that a crowd of 43 judges is able to assess (i.e. rank) 50 learner texts with high reliability (SSR = .95). No effect of language skills or language assessment experience was found on the assessment task, but there was a difference in the decisions made by judges who received formal language assessment training and those who did not. Nevertheless, the scores generated by the crowdsourced task exhibited a strong correlation with the rubric-based scores released with the learner corpus used (ETS Corpus of Non-Native Written English; Blanchard et al., 2014). The second objective of this presentation is to launch a collaborative initiative that aims to replicate and extend the pilot study described above by addressing some of the most pressing theoretical issues and avenues for future L2 research identified therein (Paquot et al. forthcoming). To that end, the project will be guided by the following three main research questions: - RQ1. To what extent do specific characteristics of learner texts (topic, length, homogeneity in terms of proficiency) have an effect on the reliability of an ACJ task? - RQ2. To what extent do specific characteristics of judges (language assessment training and expertise) have an effect on the reliability of an ACJ task? - RQ3. To what extent do specific characteristics of learner texts and characteristics of judges have an effect on the validity of an ACJ task? By the end of the project, we will be in a position to provide guidelines about the conditions in which the ACJ method can be used to enrich L2 data, with the hope that colleagues will replicate our work on other learner corpora, including learner corpora of other L2s than English. We will also distribute the comparative rank order for a set of ICLE texts as an open access resource. However, for this project to be successful, we will need to recruit a large crowd of judges. We hope that LCR participants will be as enthusiastic about this project as we are and contribute!
Affiliations

Citations

Paquot, M., Rubin, R., & Vandeweerd, N. (2022). Introducing the CLAP project: Adaptive Comparative Judgment as a community-based solution for enriching learner corpora with crowdsourced L2 proficiency assessment. Learner Corpus Research 6, Padua, Italy. https://hdl.handle.net/2078.5/106321