Aims and theoretical background It has been well-established since Sweetser (1990), that because can be used to express causal relations in the content, epistemic and speech-act domains. In other words, the connective because is a multifunctional linguistic expression that is underspecified for its contexts of use, while other languages like Dutch, French or German have developed so-called “subjective” and “objective” connectives, which may function as strong indicators of the subjectivity level of the causal relation at hand (Stukker & Sanders 2012). The connective because being uninformative in this respect, we aim at investigating whether it is possible to anchor the different uses of because in context (or rather, cotext), examining a large number of syntactic, morphological and semantic cues with a minimal cost of manual annotation. Therefore, we propose an innovative method making use of information available from an English/Dutch parallel corpus to distinguish between different uses of because. We chose Dutch because the division of labour between omdat (objective) and want (subjective) is uncontroversial (but see Sanders & Spooren 2013). Data and procedure The data comes from the Europarl corpus, accessed via the OPUS interface (Tiedemann 2012). The data set consists of a random sample of English sentences (2,000 instances) that contain because as a causal connector. Every instance of because corresponds to either omdat or want “because” in the Dutch subcorpus of Europarl. The English sentences are then parsed with the help of the Stanford Parser, which helps us extract the syntactic dependencies (de Marneffe et al. 2006) and part-of-speech annotation of the main and subordinate clauses connected by because. On the basis of this information, every English sentence is coded semi-automatically for approximately 30 contextual variables, such as the part of speech, number, person, semantic class of the subject, and modality, tense and semantic class of the predicate of the main and subordinate clause. Finally, we use a number of machine learning techniques, such as logistic regression, random forests and conditional inference trees, to determine whether these contextual variables help predict which of the two causal connectors is used in the corresponding Dutch sentences, and which contextual features of the English sentences are associated with each of the functions of because. Results and interpretation The preliminary results based on a smaller sample show that the English sentences with such characteristics as the present tense, pronominal subject, active voice, presence of modal or mental verbs (e.g. think and believe), tend to correspond to the Dutch sentences with want, the more subjective connective. In contrast, the past tense, passive voice and lack of modal verbs increase the probability of observing omdat, the more objective connective. A multiple logistic regression model performs better than a baseline model with random prediction (C = 0.75). Although there is still room for improvement, our first results indicate that the more subjective and objective uses of because can indeed be anchored in immediate linguistic context with sufficient reliability.
Levshina, N., & Degand, E. (2015). Just because: In search of an objective approach to subjectivity. International Pragmatics Association (IprA) conference, Antwerpen. https://hdl.handle.net/2078.5/241763