Much research in learner language involves studying the degree to which frequencies of use of some linguistic element or its preferred co-occurrences with other elements by learners differ (i) from those of that same element by native speakers or (ii) over time as learners become more proficient. Much of the work involving co-occurrence (association) in particular involves association measures (AMs) that aim to quantify how much words are associated with other words (collocations) or with other, more schematic constructions (e.g. collostructions). Learner corpus studies often use measures such as MI and/or t score (Gablasova et al., 2017; Forsberg-Lundell, 2021); it is also not infrequent that they rely on cut-off points to identify important collocations (e.g. Durrant & Schmitt, 2009; Granger & Bestgen, 2014; Siyanova-Chanturia, 2015). In this paper, we will discuss several problems we see in much existing work. First, we will show that much work suffers from a very elementary – in a sense – validity problem because the association measures that are used really reflect co-occurrence frequency more than they do association proper. On the basis of hypothetical collostructional data, observed keyness statistics from the ClintonTrump corpora, and actual Adj-N collocations (involving four-speed adjectives in the BNC), we show that, in monofactorial tests, esp. the loglikelihood ratio and t are extremely predictable from just the co-occurrence frequency alone (R 2 GAM>0.94) and are hardly correlated with what one might consider a gold standard measure of association (the log odds ratio, R 2 GAM<0.06). In addition, we will exemplify how the characteristics of some measures (esp. MI) are not always properly understood or at least discussed in the literature. Second, if one tries to tease apart the contribution that co-occurrence frequency and association make together, i.e. in a multifactorial setting, one finds that the loglikelihood ratio and t can be nearly predicted perfectly from an interaction of logged frequency and association (the log odds ratio), but with frequency playing a much stronger role than the association. Crucially, that very high correlation of frequency and association also means that these AMs do not permit the user to properly address the separate contributions that association (i.e., contingency) and frequency (i.e. entrenchment) make cognitively or psycholinguistically (Gries & Ellis, 2015): if one's AM conflates contingency and frequency, one can be definition not separate high frequency from high association in one's theoretical explanations; we, therefore, recommend a two-dimensional representation of frequency and association for any kind of association measure scenario. The question that arises from both of these related issues is how to decouple frequency and association in our studies involving AMs. In other words, how do we make sure that our AMs really reflect association and only association so that we can test and/or develop our theories in such a way that frequency and association make measurable but separate components that our hopefully cognitively-informed SLA models can handle? In this paper, we will present an answer to this question by outlining a three-step procedure of how one can take any AM and completely decouple it from frequency. First, we quantify the association for a certain collocation in the data (let’s call this value obs). Second, we take the frequencies of the two co-occurring elements in question (i.e. the totals a+b and a+c) and the corpus size (i.e. a+b+c+d), hold them constant (which virtually eliminates any way in which co-occurrence frequency can unduly boost/lower the resulting association-only measure), and determine (i) the lowest and the highest possible associations given the values we are holding constant (let’s call these low and upp for lower and upper limit). Third and because these maximal-attraction and minimal-attraction/maximal-repulsion values will exhibit different ranges (due to the marginal totals), we then transform/min-max scale these three values (obs, low, and upp) such that they fit into the interval [0,1], and our new association-without-frequency measure becomes the value that corresponds to obs in that [0,1] interval. We exemplify this approach by applying it to one specific AM – conditional probability – for co-occurrences of speed adjectives with nouns in the BNC and show that it is indeed uncorrelated with frequency (R 2 with co-occurrence frequency < 0.01) but also supports the use of MI and the log odds ratio as true measures of association only. We conclude with some comments on what the results mean for the use of cut-off points for significant/interesting associations.
Gries, S. Th., & Paquot, M. (2022). Association measures in learner corpus research: problems and pointers for improvement. Learner Corpus Research 6, Padua, Italy. https://hdl.handle.net/2078.5/106323