(2016) 37th International Computer Archive of Modern and Medieval English Conference (ICAME 37) — Location: The Chinese University of Hong Kong (25.May.2016)
As far back as 1944, Thompson noted that “[m]easuring accurately the quality of the public speaking performance is a perplexing problem, of importance to varied types of groups” (Thompson 1944: 65). Foreign language teachers, in particular, are regularly faced with this difficult but crucial task. In their attempt to assess students’ spoken performance, they have certain resources available to them, most notably the descriptors of the Common European Framework of Reference for Languages (CEFR, Council of Europe 2001). However, given the criticism that has been levelled at this assessment tool (see, e.g., Alderson 2007), and especially the observation that “the CEFR is not based on empirical evidence taken from L2 learner data” (Hulstijn 2007: 666), one may wonder to what extent it is a useful and reliable tool for teachers. With a view to testing this, we asked 27 French-speaking foreign language teachers from secondary education to evaluate audio samples from the Louvain International Database of Spoken English Interlanguage (LINDSEI, Gilquin et al. 2010). LINDSEI is a one-million word spoken corpus made up of informal interviews with EFL learners from different mother tongue backgrounds. Ten five-minute samples, taken from the French subcorpus, were extracted from the most spontaneous part of the interviews, corresponding to free discussion between the learner and the interviewer. The teachers received these samples, together with a copy of the CEFR grid for spoken language use and some instructions. Per interview, they were asked to assign a grade for each of five competences (range, accuracy, fluency, phonological control and coherence), as well as a grade for global assessment. As a benchmark, we also asked two native English-speaking expert raters to do the same exercise. The results of the rating experiment revealed a large discrepancy in the grades assigned, sometimes covering the whole spectrum from A2 to C2 for one and the same sample. It also turned out that the non-native teachers almost systematically assigned lower grades than the native experts, despite relying on the same descriptors, which confirms a tendency highlighted by e.g. Koster & Koet (1993) or Hyland & Anan (2006), namely that native speakers are generally more tolerant than non-native speakers. In a second step, we compared the teachers’ and experts’ ratings with the actual learner corpus data, using objective measures such as hesitation rate or lexical density. This comparison helped us establish which aspects of linguistic usage seemed to be more decisive when selecting a certain grade, as demonstrated by Isaacs et al. (2015) for the IELTS (International English Language Testing System) pronunciation scale. Thus, the type-token ratio was found to correlate quite well with the score given by the raters for lexical range. Surprisingly, the rating of fluency did not seem to be primarily affected by the frequency of (unfilled) pauses, despite the explicit reference to pauses in several CEFR descriptors (cf. “pausing for grammatical and lexical planning and repair is very evident”, “there are few noticeably long pauses”). In our study, we therefore investigated the impact of other factors, such as the frequency and type of discourse markers. It is interesting to note, in this respect, that the most highly rated sample in terms of fluency was the only one that included the discourse marker sort of. As shown by Thewissen (2015) for writing, learner corpus data can play a key role in fleshing out the descriptors of the CEFR. In our conclusion, we will make some corpus-informed suggestions for the assessment of speech, with a particular focus on fluency features.
Gilquin, G., Bestgen, Y., & Granger, S. (2016). Assessing the CEFR assessment grid for spoken language use: A learner corpus-based approach. 37th International Computer Archive of Modern and Medieval English Conference (ICAME 37), The Chinese University of Hong Kong. https://hdl.handle.net/2078.5/232273