Comparing Online and Offline Methodologies for Spoken Corpus Annotation. The case of Filled Pauses and Hesitation-Related Lengthening in French

Duvivier, Emilie;Grosman, Iulia;Christodoulides, George;et.al.
(2017) ExLing 2017, 8th Tutorial and Research Workshop on Experimental Linguistics — Location: Heraklion, Crete, Greece (19.June.2017)

Files

Duvivier_al_2017-Euh-all_Exling2017_abstract.pdf
  • Open Access
  • Adobe PDF
  • 101.15 KB

Details

Authors
  • Duvivier, EmilieUCLouvain
    Author
  • Grosman, IuliaUCLouvain
    Author
  • Christodoulides, Georgeorcid-logoUCLouvain
    Author
  • et. al.
Abstract
In this contribution we compare the results of two methodologies for annotating phenomena related to hesitation in speech (e.g. Martin & Strange, 1968), namely filled pauses and lengthening. The first method is the off-line annotation of a spoken corpus by annotators who have access to additional information (e.g. a visual representation of the signal) and the possibility to listen multiple times to segments; most manual annotations in speech corpora are produced using this methodology. The second method is using software that allows for the online annotation of a phenomenon in real-time, without the possibility of stopping the playback or repeatedly listening to segments. In both cases, annotators were “experts” (they had training and experience in phonetics and prosody); we can envisage using the second methodology with “naïve” listeners (participants in an annotation campaign not having prior knowledge of the underlying research question). We also compare these two methodologies with the results of automatic annotation, which is based on statistical models trained on a different speech corpus. These comparisons are a form of triangulation, i.e. “cross-checking data from multiple sources to search for regularities in the research data” (O’Donoghue and Punch, 2003). Methodology All annotations were performed on the LOCAS-F corpus (Degand et al., 2014) which is a balanced corpus of spoken French with samples of 14 different speaking styles (discourse genres). The duration of the corpus is 3.5 hours and it contains 43.000 tokens. The corpus is aligned at the phone level (manually verified) and also contains an expert prosodic annotation (two levels of boundaries and four prosodic contours) as well as a part-of-speech and syntactical annotation. The transcription conventions used in the corpus readily identify autonomous filled pauses, and allow us to identify syllables ending in schwa (possibly perceived as epenthetic). With respect to off-line annotation methods, we used two different sets of data: the entire corpus has been annotated by the second and third authors, independently from one another, and using different annotation schemes. Both annotation protocols, however, include markers for filled pauses and for hesitation-related lengthening at the syllable level. With respect to the on-line annotation method, we used the Tapping Annotator software (previously used to study the perception of prosodic boundaries, cf. Simon & Christodoulides, 2016). A simple interface (Figure 1) allows the user to listen to corpus audio samples and press a key as soon as they perceive the phenomenon under study. All keystrokes (press and release times) are recorded for analysis. The first author used this interface to annotate the entire LOCAS-F corpus, in non-contiguous sessions of six corpus samples; she repeated the process twice. Analysis The annotation campaigns described above result in four datasets (two offline and two online). In the case of on-line annotation, we used a simple method to correlate a keypress to the syllables that triggered the perception of a hesitation (based in the average reaction time of the annotator to pure tones). An additional dataset was the result of an automatic detection of hesitation-related lengthening using the statistical models of the DisMo annotator (Christodoulides & Avanzi, 2015). We will present a comparison of the five datasets, along with an analysis of the correlates of syllables that triggered the perception of hesitation. These correlates include syllable duration, intonation contour, presence and duration of a subsequent silent pause, local speech rate, position within token and within syntactic constituent. References Christodoulides, G. & Avanzi, M. (2015). Automatic detection and annotation of disfluencies in spoken French corpora”, Proc. of the 16th Annual Conference of the International Speech Communication Association (Interspeech), September 6–10, Dresden, Germany, pp. 1849–1853. Degand, L., Martin, L.J., Simon, A.C. (2014). Unités discursives de base et leur périphérie gauche dans LOCAS-F, un corpus oral multigenres annoté. Actes du 4ème Congrès Mondial de Linguistique Française 2014, Berlin, Germany: EDP Sciences. Martin, J. G. & Strange, W. (1968). The perception of hesitation in spontaneous speech, Perception & Psychophysics, 3(6): 427-438, DOI: 10.3758/BF03205750 O’Donoghue, T. & Punch K. (2003). Qualitative Educational Research in Action: Doing and Reflecting, Routledge, p.78. Simon, A. C. & Christodoulides, G. (2016). Perception of Prosodic Boundaries by Naïve Listeners in French, Proc. of Speech Prosody 2016, 31 May – 3 June 2016, Boston, USA.
Affiliations

Citations

Duvivier, E., Grosman, I., Christodoulides, G., & et al. (2017). Comparing Online and Offline Methodologies for Spoken Corpus Annotation. The case of Filled Pauses and Hesitation-Related Lengthening in French. ExLing 2017, 8th Tutorial and Research Workshop on Experimental Linguistics, Heraklion, Crete, Greece. https://hdl.handle.net/2078.5/178769