Learner corpus research has seen major development since its inception some 25 years ago. Nevertheless, theoretical, methodological and empirical advances have been summarized in the literature only rarely and, in such cases, selectively rather than systematically. To the authors’ knowledge, in fact, there is no meta-analysis to date that summarizes and synthesizes the body of knowledge resulting from learner corpus research in a specific area of study (e.g. English as a Foreign Language learners’ use of collocations or tense, aspect and modality in learner writing). Equally concerning is that relatively little attention has been paid to the state or development of the field’s methodological practices, an unfortunate circumstance given the empirical rigor needed to reliably and accurately make use of corpus data and analyse frequencies of (co-)occurrence (Gries, 2013; Gries, forthcoming; Gries & Deshors, 2014). Progress in any discipline, however, crucially “depends on sound research methods, principled data analysis, and transparent reporting practices” (Plonsky & Gass, 2011: 325). This study thus aims to provide the first empirical assessment of quantitative research methods and study quality in learner corpus research. Study quality is defined rather broadly as "(a) adherence to standards of contextually appropriate methodological rigor in research practices and (b) transparent and complete reporting of such practices” (Plonsky, 2013: 657). Specifically, we systematically review all quantitative, primary studies referenced in the Learner Corpus Bibliography (LCB), a representative bibliography of learner corpus research maintained by the Learner Corpus Association (http://learnercorpusassociation.org) which currently contains approximately 1180 references. The techniques used to retrieve, code, and analyze this body of primary research are characteristic of research synthesis and meta-analysis. Following Plonsky (2013), however, this study differs from those traditions of synthetic research in that the focus here is almost exclusively methodological (i.e. the “how” of learner corpus research) rather than substantive (i.e. the “what”). Each reference in the LCB is surveyed using a coding scheme inspired from the protocol developed and first used by Plonsky & Gass (2011) to assess methodological quality in second language acquisition, and more particularly interaction research. The coding scheme is however revised and expanded to account for the methodological characteristics of corpus linguistics. Quantitative studies are coded for over 50 categories representing six dimensions: (a) publication type (i.e. conference paper, book chapter, journal article), (b) research focus (e.g. lexis, grammar), (c) methodological features (e.g. Contrastive Interlanguage Analysis, keyword analysis, error analysis, use of reference corpus), (d) statistical analyses (e.g. X², t-test, regression analysis), and (e) reporting practices (e.g. reliability coefficients, means). The 25-year span of research represented in the LCB provides a unique opportunity to examine the resulting data cumulatively and also permits analyses of changes taking place over time in the research and reporting practices of this domain. Preliminary results point to several systematic strengths as well as many flaws, such as the absence of research questions or hypotheses, incomplete and inconsistent reporting practices (e.g. means without standard deviations), and low statistical power (i.e. LCR studies generally overrely on tests of statistical significance such as the X² test, do not report effect sizes, rarely check or report whether statistical assumptions have been met, rarely use multivariate analyses). Improvements over time are however clearly noted and there are signs that, like other related disciplines, learner corpus research is slowly “undergoing a change to becoming much more empirical, much more rigorous, and much more quantitative/statistical” (Gries, 2013: 287) In addition to providing direction for future research and research practices, the study’s findings will also be discussed and contextualized within the research cultures of corpus linguistics, second language acquisition, and applied linguistics more generally. References Gries, S. (2013). Statistical tests for the analysis of learner corpus data. In Diaz-Negrillo A., Ballier N. & P. Thompson (eds). Automatic Treatment and Analysis of Learner Corpus Data. Amsterdam & Philadelphia: Benjamins. Gries, S. (forthcoming). Statistics for learner corpus research. In Granger S., G. Gilquin & F. Meunier (Eds). The Cambridge Handbook of Learner Corpus Research. Cambridge University Press. Gries, S., & Deshors, S. (2014). Using regressions to explore deviations between corpus data and a standard/target: two suggestion. Corpora, 9(1), 109–136. Plonsky, L. (2013). Study quality in SLA. An assessment of designs, analyses, and reporting practices in quantitative L2 research. Studies in Second Language Acquisition, 35, 655-687. Plonsky, L. & Gass, S. (2011). Quantitative research methods, study quality and outcomes: the case of interaction research. Language Learning 61(2): 325-366.
Paquot, M., & Plonsky, L. (2015). Quantitative research methods and study quality in learner corpus research. Learner Corpus Research 2015, Nijmegen, Netherlands. https://hdl.handle.net/2078.5/190163