Comparative judgement (CJ) is an assessment method in which a large number of pairwise comparisons between learner productions are used to generate scales ranking each item from strongest to weakest. Recent research has suggested that combining CJ with various approaches to judge recruitment, including community-driven and crowdsourcing methods, holds promise as a method of assessing L2 writing, especially for research purposes. However, the majority of studies to date have tested CJ only using relatively simple, easily evaluated sets of texts. There remains insufficient evidence of the method’s potential for assessing more complex texts, particularly when the comparisons are being conducted by crowdsourced assessors. This study seeks to address this problem by testing the reliability and validity of a crowdsourced form of CJ for the assessment of texts which are longer, more topically diverse, and more homogeneous in proficiency than those used in earlier studies. The results suggest that CJ can be conducted with crowdsourced judges to generate reliable assessments of L2 writing, and provide initial evidence of concurrent validity. However, there are drawbacks in terms of efficiency.
Thwaites, P., & Paquot, M. (2025). Testing crowdsourcing as a means of recruitment for the comparative judgment of L2 argumentative essays. Journal of Second Language Writing, 68(forth), 101207. https://doi.org/10.1016/j.jslw.2025.101207 (Original work published 2025)