A corpus-driven approach to native and learner spoken fluency: The contribution of pauses

(2016) Nouvelles approches du corpus en linguistique anglaise — Location: Avignon, France (9.June.2016)

Files

No attached file found for this publication.

Details

Authors
Abstract
Research into spoken fluency – the smooth, continuous delivery in oral production (Chambers 1997) – has traditionally focused on fine-grained qualitative analyses of a very limited set of speakers, data and/or phenomena. Recently, however, it has undergone profound changes. Thanks to the increasing number of large and representative spoken and speaking corpora (without/with access to the acoustic signal) (see Ballier and Martin 2015, 110), along with substantial improvements in speech technologies, it is now possible to combine quantitative and qualitative corpus analyses of (several aspects of) fluency. This association has yielded significant new insights into spoken behaviour (e.g. Götz 2013; Gut 2009; Osborne 2007). This paper adopts such a combined approach and draws on two comparable spoken corpora of native and advanced learner English, namely the Louvain Corpus of Native English Conversation (LOCNEC, De Cock 2004) and the French component of the Louvain International Database of Spoken English Interlanguage (LINDSEI, Gilquin, De Cock, and Granger 2010). Both contain 50 interviews with university students: English native speakers (LOCNEC) or foreign language learners of English (LINDSEI). The corpora (totalling c. 30 hours of recorded speech) have been time-aligned and annotated with EXMARaLDA (Schmidt and Wörner 2014) for a dozen fluency features, including filled and unfilled pauses (measured in milliseconds), restarts, repetitions and truncations (Dumont 2015). The study first presents the overall quantitative distribution of these fluency features. It then zooms in on filled (FPs) and unfilled pauses (UPs), two features that are claimed to be particularly discriminant between fluency levels (cf. CEFR descriptors; Council of Europe 2001), and examines their qualitative use in isolation as well as in larger cluster-like sequences, e.g. a FP used conjointly with a discourse marker (at first (0.310) well I was very hesitant), or an UP within a repetition (I think we’ll (0.540) we’ll leave together). Based on the hypothesis that both native speakers and advanced learners may achieve varying (and potentially overlapping) degrees of fluency, native and learner data are analysed together, making no a priori presupposition as to the intrinsic benchmarking property of native speaker data. Preliminary quantitative results indicate that among the dozen fluency features under investigation, FPs and UPs are by far the most frequent, making up together about 50% of the annotated features. However, a high rate of pausing is found in both native speakers and advanced learners, thereby shedding doubt on their use as level discriminator in the CEFR. The qualitative analysis of FPs and UPs reveals that, for all speakers (native and non-native), pauses are more frequently used in combination with (an)other fluency feature(s) than in isolation, and that, among the wide variety of possible clusters, the “bigrams” FP+UP and UP+FP are especially common.
Affiliations

Citations

Dumont, A. (2016). A corpus-driven approach to native and learner spoken fluency: The contribution of pauses. Nouvelles approches du corpus en linguistique anglaise, Avignon, France. https://hdl.handle.net/2078.5/175732