Starting from the observation that major discourse annotation frameworks (e.g. Penn Discourse TreeBank 2.0, henceforth PDTB, Prasad et al. 2008; Rhetorical Structure Theory, Mann & Thompson 1988) disagree on the top levels (number and type) and specific values for discourse relations which they include, we present and test a multilingual model where the two annotation levels, namely domains and functions, are independent. It distinguishes between four domains (ideational, rhetorical, sequential, interpersonal) and fifteen functions (e.g. cause, contrast, alternative) potentially expressed by discourse markers. This proposal is a revision of [Author] where the choice of a domain strongly restricts the range of functions and where each function is included in one domain only (e.g. “cause” is always ideational whereas “motivation” is always rhetorical). In the revised model, each function can, in principle, express each of the four domains: for instance, a contrastive relation can be ideational (connect objective facts), rhetorical (connect subjective arguments), sequential (target the structure of topics and turns) or interpersonal (manage the speaker-hearer relationship), as respectively illustrated by the following corpus-based examples. (1) I wasn’t looking forward to doing it but I am now (2) <speaker 1> a rebate is when they send the money back <speaker 2> yes but how do you define it in economic terms (3) (after a digression on the industries in Bristol area) but Bristol itself is a large metropolis (4) <speaker 1> I think the Marks is better <speaker 2> actually I’m not sure it is Independent annotation levels are assumed to enhance the reliability of the model, by allowing analysts to (dis)agree on one level but not necessarily on both. This adjustment stems from previous research [Author] showing that higher, more generic annotation levels trigger less disagreements. This new approach also provides empirical evidence for the need of integrating topic relations into taxonomies of discourse relations, as opposed to other frameworks (PDTB; Cognitive approach to Coherence Relations, Sanders et al. 1992), which exclude them. The inclusion of the sequential domain (topic relations and turn management) into the functional spectrum of discourse markers is in line with theories of the general functions of language (e.g. Halliday 1970) and makes a strong case for a more bottom-up approach to both identification and annotation of discourse markers, embracing the full heterogeneity of this complex category in authentic (spoken) data. The structure and operationalization of this taxonomy will be presented and discussed with respect to [Author]’s original model and other well-established proposals. So far, it has been applied to spoken English, French and Polish. We will report on quantitative findings of the distribution of domains and functions in each language, inter-rater agreement scores on the English and French data (annotated by two expert coders), and a qualitative-quantitative comparison with the original inter-dependent model on the French data where both versions of the taxonomy have been applied by the same annotator. In sum, our paper tackles both methodological and theoretical issues, striving towards reliability, exhaustivity and universality in the corpus-based study of discourse markers across languages and modalities.
Crible, L., Degand, L., & et al. (2017). Testing interdependent annotation levels for sense disambiguation in spoken English, French and Polish. 50th Annual Meeting of the Societas Linguistica Europaea, Zurich, Switzerland. https://hdl.handle.net/2078.5/241740