We report on the results of an annotation experiment comparing naïve and expert coders in a sense disambiguation task consisting in the assignment of function labels to discourse markers (e.g. well, but, I mean) in spoken French and English using a taxonomy specifically designed for speech. Our qualitative-quantitative assessment of its reliability led us to suggest fundamental revisions of the structure of the taxonomy, striving to find a better balance between reliability and granularity. The resulting model articulates two independent levels of annotation (domains and functions) which, once combined, provide a robust tool for the analysis of DMs and relates them to more general functions of spoken language.
Crible, L., & Degand, L. (2019). Reliability vs. granularity in discourse annotation: What is the trade-off? Corpus Linguistics and Linguistic Theory, 15(1), 71-99. https://hdl.handle.net/2078.5/241731 (Original work published 2019)