Discourse markers are the focus of an abundant research field investigating the many aspects of their behavior, either from a syntactic, semantic, prosodic or other approach. One crucial aspect is their polyfunctionality, which has been explained and modeled under several different theoretical frameworks (see Fischer 2006 for an overview). These models include, among others, the notion of multidimensionality in the Dynamic Interpretation Theory (Petukhova & Bunt 2009), the five “planes of talk” in Schiffrin (1987), the concept of “meaning potentials” (Norén & Linell 2006; Aijmer 2013), the constructionist approach by Fischer (2010, 2015) or the three components of discourse structure in Redeker (1990) (see also González 2005). Each approach provides a different (yet partially overlapping) account of the many dimensions of meaning that discourse markers can express, following different theories and agendas (e.g. discourse analysis, cognitive linguistics, computational applications). Combining theoretical and methodological considerations, we propose a corpus-based annotation scheme for (spoken) discourse markers, where their functional spectrum is seen as the interface between two independent dimensions, namely a domain and a function (Crible & Degand in press). Our four domains (viz. ideational, rhetorical, sequential and interpersonal) are rooted in the tradition of cognitive models of discourse structure (e.g. Redeker 1990; Sweetser 1990; Sanders 1997) and correspond to different layers of discourse which speakers (or writers) can address: content relations (ideational), subjective and metalinguistic meanings (rhetorical), discourse structure (sequential) and speaker-hearer relationship (interpersonal). Functions, on the other hand, are more specific interpretations of the type of operation which a discourse marker is performing in a given context (fifteen types, e.g. causal relation, topic-shift, turn-taking, face-saving, etc.). Following Bunt (2011), we distinguish between “general-purpose” functions, which can activate any of the four domains (e.g. a relation of contrast can be either ideational, rhetorical, sequential or interpersonal), and “dimension-specific”, here, domain-specific functions, which pertain to one domain only (e.g. topic-shift is always sequential). We consider domains and functions as two orthogonal dimensions of meaning which each correspond to a type of semantic variation, viz. polyfunctionality and polysemy, respectively. By polyfunctionality, we mean the possibility for a single invariant meaning to be expressed across several domains (e.g. ideational vs. rhetorical contrast). Polysemy, in turn, refers to the multiple functions a discourse marker can fulfil, regardless of the domain (e.g. so to express a consequence, an exemplification or a topic shift). In this sense, polysemy is different from “simultaneous multifunctionality” (Bunt 2011), which rather targets the joint expression of more than one meaning at a time in a given context, be it different functions (polysemy) or the same function in different domains (polyfunctionality). Our integrated approach is compatible with an inclusive definition of the discourse marker category as adopted and annotated by Crible (2017), especially since issues of categorization and functional classification are strongly interrelated (cf. Degand et al. 2013). The aim of this presentation is to test the assumed cross-linguistic validity of the approach and the ensuing annotation scheme by analyzing the variation in use and functions of a broad bottom-up selection of DMs across three languages from different typological families, namely French (Romance), English (Germanic) and Polish (Slavic). The taxonomy was applied to a sample of ca. 30 minutes (between 5000 and 6000 words) of spoken unplanned dialogues in each of the languages making use of available corpora (LOCAS-F corpus (Degand et al. 2014) for French, ICE-GB corpus (Nelson et al. 2002) for English, and Polish data (Pęzik 2015). Discourse markers were identified in a bottom-up approach, without any closed list of pre-selected items, following a broad definition: any expression, which is syntactically optional, has a (partly) procedural meaning and performs a discourse-level pragmatic function was selected. The items were then manually annotated according to the functional classification introduced above. The French data was double-coded for the purpose of inter-annotator agreement. Annotations were then extracted for contrastive analyses of distribution and variation of DMs and their functions. The data and annotation scheme will be briefly presented, focusing on challenges of multilingual annotation. In this discussion, we will include the results of an ongoing annotation campaign on additional languages (Slovenian, Spanish, L2 English, Brazilian Portuguese). We will then report on quantitative findings of the distribution of domains and functions across French, English and Polish conversations, looking for cross-linguistic differences and similarities in the functions DMs can express in each language, both from a categorical and a DM-specific point of view. In particular, our results show that the interpersonal domain is much more frequent in Polish than in the other two languages: this can be explained by the different nature and status of question tags (English isn’t it, Polish nie, French hein) in the three systems. This major result led us to further investigate interpersonal DMs. Each language has a different number of typically interpersonal DMs which display a different frequency in the data: for instance, the typical tu vois in French is highly infrequent in the sample. There are many more different types of interpersonal DMs in Polish, which suggests that these items are at the core of the DM category. In addition, interpersonal meanings can be expressed by other DMs as well: we identified a cline from “purely interpersonal” to “sometimes interpersonal” DMs: the latter is not attested in the English data. Interpersonal uses of otherwise adversative markers in French (mais) and Polish (przecież) will be discussed. This analysis shows that semantic equivalence of DMs attested in different languages does not necessarily lead to functional and distributional similarities between them. Overall, we observed a higher similarity between English and French than with Polish. However, we cannot conclude at this stage whether this observation is due to family resemblance or to annotators’ bias. Regarding the two-dimensional taxonomy, Crible & Degand’s (in press) revised version seems to reach higher inter-annotator agreement (compared to Crible’s (2017) original). It is applicable to a large range of languages from different typological families, and enables interesting analyses, both at a comprehensive level over the whole DM category and at a more DM-specific level.
Degand, E., Crible, L., & Grzech, K. (2018). Functions and domains of discourse markers across languages: Testing a two-dimensional annotation scheme. Cross-linguistic Discourse Annotation: Applications and Perspectives (TextLink2018), Toulouse, France. https://hdl.handle.net/2078.5/241737