Cross-linguistic studies of discourse markers often face various methodological problems regarding the applicability of a single annotation protocol to very diverse data-driven sets of items from several languages. Ideally, such a protocol should overcome language-specific preferences and encompass all possible actualizations of discourse markers, in different contexts and speech situations. Given the great diversity of relations or functions that markers may perform, the list of possible values must be the largest possible, while remaining operational for the perspective of inter-coder agreement. Moreover, the criteria of what counts as a discourse marker must also be very explicit so that the onomasiological selection of items in context can be as stable and exhaustive as possible. We will present a proposal of a corpus-based annotation scheme (Crible, subm.) designed and tested on French and English spoken data across different speech situations (spontaneous and prepared, monologues and dialogues, etc.). This protocol was elaborated for both relational markers (or discourse connectives) and non-relational markers (other discourse markers such as you know or well). We will cover all the features that are accounted for in this annotation and which are mainly of two kinds: syntactic parameters on the one hand, such as a two-fold categorization of the position of the marker in the sentence; and functional features on the other. Position and relation are two parameters that were identified by Zufferey & Cartoni (2012) as “relevant criteria governing the choice of a causal connective” (2012: 233), and it may be assumed that it is the case for other types of markers. The present system also includes a definition of the segments related by the marker, with a fairly large understanding of the unit it may apply to. The functional categories of this protocol are revised versions of previous taxonomies such as Sweetser's (1990) or Gonzalez’ (2005). Functions are grouped into four domains: ideational, rhetorical, sequential and interpersonal. This first selection restricts the possible functional value(s) to an operational and closed list of domain-specific functions, elaborated from previous literature (a.o. Prasad et al. 2008, Cuenca 2013, Zufferey & Degand 2014) and empirical testing. In example (1), the English conjunction but works simultaneously in the rhetorical and the sequential domains, performing both a relation of opposition between the two segments, and a text-structuring function of topic-shifting. (1) “So that was the way the decisions were made. But what we also wanted, and we’re very keen on this here, …” (bb_en011) The efficiency and operationality of this annotation scheme will be assessed by an annotation experiment contrasting naive and expert coders (all native speakers of French) applying the present protocol to spoken conversations in English and French. Reports of inter-annotator agreement within and between groups will be illustrated and analyzed in order to investigate the source of disagreements, which we expect, at least for expert coders, to be linked to speech-specific and/or highly polysemous, under-specified uses of discourse markers, on the basis of previous studies (e.g. Crible & Zufferey 2015). This annotation experiment will also be an opportunity to explore patterns of combination between the different variables involved, and to raise more general issues regarding the annotation of the complex phenomenon of discourse markers in natural data.
Crible, L., & Degand, L. (2015). Functions and syntax of discourse connectives across languages and genres: towards a multilingual annotation scheme. International Pragmatics Association (IprA) conference, Antwerpen. https://hdl.handle.net/2078.5/241757