Core Metadata Schema for L2 data

Frey, Jennifer - Carmen;König, Alexander;Stemle, Egon;Paquot, Magali
(2023) 32nd conference of the European Second Language Association (EUROSLA) — Location: University of Birmingham (30.August.2023)

Files

EuroSLAMetadataposter.pdf
  • Open Access
  • Adobe PDF
  • 358.13 KB

Details

Authors
  • Frey, Jennifer - Carmen
    Author
  • König, Alexander
    Author
  • Stemle, Egon
    Author
  • Author
Abstract
The main objective of this presentation is to introduce a core metadata schema for L2 production data, more particularly learner corpora, which is the result of extensive collaboration between learner corpus compilers at the Centre for English Corpus Linguistics (UCLouvain, Belgium) and EURAC Research (Bolzano, Italy), and a research data infrastructure expert and member of CLARIN's metadata taskforce. The project stems from the recognition that one area that would benefit significantly from standardization is L2 data description, which includes metadata at the level of the dataset as a whole and metadata used to describe the individual learners and task types/registers the corpus is meant to represent. There are a number of reasons why this is important. First, standardized and well-structured metadata increases the findability and usability of existing learner corpora. Second, it should enhance the comparability of datasets and comparability of L2 studies, provided researchers agree on a common set of definitions. Extensive metadata that follow - at best - a standardized vocabulary, and have a strong focus on findability, accessibility, interoperability, and reusability (FAIR) are an essential aspect of FAIR research data (Wilkinson et al. 2016). In continuation of Granger & Paquot (2017), our proposed metadata schema is divided into a number of different sections for Corpus metadata (itself divided into Administrative metadata (e.g. authors or license) and Corpus design metadata (e.g. date and place of collection or type of task)), Text metadata (fine-grained per-text information), Learner metadata (details about the learners, e.g. age, languages spoken), Annotation metadata (e.g. details about manual or automatic annotation), Annotator metadata (e.g. professional and language background), Transcriber metadata (e.g. native language or language repertoire) and Task metadata (e.g. instructions, time constraints). While basic information about learners (authors) and language samples (texts) are typically found as part of metadata associated with a learner corpus, other aspects such as those related to the annotation or transcription procedure or the specificities of a task are often found elsewhere (e.g. corpus manual) or are just absent from currently available learner corpora. Our proposal is to provide a systematic description of all these aspects as part of core metadata. A beta version of the core metadata schema was tested on a number of learner corpora representing a variety of learners and language samples. It was then presented at the LCR2022 conference (König et al., 2022), which opened a feedback collection phase. We are currently in the process of revising our initial proposal based on the comments received and will release a stable version of the core metadata schema in Spring 2023. References: Granger, S. & Paquot, M. (2017). Towards standardization of metadata for L2 corpora. Invited talk at the CLARIN workshop on Interoperability of Second Language Resources and Tools, 6-8 December 2017, University of Gothenburg, Sweden. https://sweclarin.se/sites/sweclarin.se/files/event_atachements/Granger_Paquot_Metadata_G%C3%B6teborg_final.pdf König, A., Frey J.-C., Stemle, E., Glaznieks, A. & Paquot, M. (2022). Towards standardizing LCR metadata. Paper presented at Learner Corpus Research 6, 22-24 September 2022, University of Padua, Italy. Wilkinson, M. D. et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific data, 3(1), 1-9.
Affiliations

Citations

Frey, J.-C., König, A., Stemle, E., & Paquot, M. (2023). Core Metadata Schema for L2 data. 32nd conference of the European Second Language Association (EUROSLA), University of Birmingham. https://hdl.handle.net/2078.5/100027