Until fairly recently, annotation of continuous French speech has either been relatively rudimentary and approximate; either reserved to a bunch of specialists working within the framework of phonologic theories. On the one side indeed, specialists of spoken French [Blanche-Benveniste et al. 1990; Morel & Danon-Boileau 1998; Berrendonner et al. 2012] transcribe prosodic events using a reduced set of symbols, which does not reflect the actual complexity of prosodic phenomena. On the other side, phonologists such as [Verluyten 1982; Delais-Roussarie 1996; Jun & Fougeron 2000; Post 2000; Michelas & D’Imperio 2012; Delais-Roussarie et al. to app.] develop annotation systems which are hardly applicable to spontaneous speech, since the data they deal with mostly consist in laboratory speech, that is "light years ahead the complexity of spontaneous speech" [Lacheret 2003: 4]. Since a few years however, mostly thanks to automatic processing advances, the situation is changing [Avanzi 2012]. Protocols and tools designed to annotate (semi-)automatically French prosodic structure are emerging (see Delais-Roussarie et al. [2006], for an overview). In thi context, the purpose of this paper is not to make an inventory of the existing systems, but (i) to present a perceptually-based and computer-assisted procedure for the annotation of accentuation (calculation of the pitch accents position and strength within a given group of words) and phrasing (identification of the different prosodic groups of the prosodic hierarchy) in French corpora; (ii) to show how such annotations can be used to test some hypotheses regarding the syntax/prosody interface. The methodology was applied to a set of recordings extracted from the PFC database [Durand et al. 2009]. The set (almost 7-hours long) includes recordings of native French spoken in France (Paris and Lyon), in Switzerland (Geneva, Neuchâtel, Martigny and Nyon) and in Belgium (Tournai and Liège). For each of the 8 inquiry points, 8 speakers (4 males, 4 females, from 20 to 80 years old) were recorded in a reading task and in a spontaneous task. Each file (the entire text, which is approximately 400 words-long; 200sec of monologal stretch of speech extracted from the spontaneous task) was next transcribed and semi-automatically aligned in phones, syllables and words with the EasyAlign [Goldman 2011] script, which works under the Praat software [Boersma & Weeninck 2012]. Alignments were then manually corrected by one of the authors. Syllabic prominences and disfluencies were manually annotated by two of the authors, following the methodology presented in Avanzi et al. [2010]. A third expert intervened in cases of disagreement between the two annotators and decided the final value of the syllable (+/- prominent or +/ disfluent). Next, a tier indicating the Accentual Phrases boundaries was obtained in the following the way: each final syllable of a lexical word or of a polysyllabic functional word that was coded as prominent generates the boundary of a minor prosodic constituent , including every element without prominence on its left side. The Analor tool [Avanzi et al. 2011] was then used to identify automatically in another tier the syllables associated with a nuclear contour (syllables followed by a pause, associated with a major pitch rise and/or significant lengthening), i.e. the Intonational Phrase boundaries. Finally, a PoS tagging, providing classic categories of words (noun, adjectives, verbs, adverbs, prepositions, etc.) was automatically performed in a dedicated tier, and manually corrected. As this prosodically annotated database was designed to study stylistic and regional prosodic variation in French within the framework of formal models of French prosodic structure, it allows to test empirically the validity of the phonological rules involved in the recognition of the different units of the prosodic hierarchy; and to validate or reformulate these rules, if necessary. During the workshop, we will discuss the results of an ongoing study dealing with the effects of constituent length, syntactic branching, semantic type and local articulation rate on the phrasing of pre- and post-nominal adjectives in French. The aim is to test the empirical validity of the hypothesis according to which pre-nominal adjectives are predicted to be phrased in the same AP than the noun they complete, while post-nominal adjectives are claimed to generate an AP on their own [Selkirk 1978; Verluyten 1982; Nespor & Vogel 1986; Post 2000; Elordieta et al. 2003]
Avanzi, M., Goldman, J.-P., Bardiaux, A., & Schwab, S. (2013). What annotated corpora teach us about French prosodic phonology? Phonetics and Phonolgy in Iberia, Lisbon. https://hdl.handle.net/2078.5/48282