LLM-Generated Contexts to Practice Specialised Vocabulary: Corpus Presentation and Comparison

Nikolova, Iglika;Bibauw, Serge;Dumont, Amandine;Stas, Françoise;François, Thomas;et.al.
(2024) 31e Conférence sur le Traitement Automatique des Langues Naturelles (TALN 2024) — Location: Toulouse (8.July.2024)

Files

Nikolova-Stoupaketal2024atTALN.pdf
  • Open Access
  • Adobe PDF
  • 346.37 KB
  • https://creativecommons.org/licenses/by/4.0/

Details

Authors
Show more
Abstract
This project evaluates the potential of LLM and dynamic corpora to generate contexts ai- med at the practice and acquisition of specialised English vocabulary. We compared reference contexts—handpicked by expert teachers—for a specialised vocabulary list to contexts generated by three recent large language models (LLM) of different sizes (Mistral-7B-Instruct, Vicuna-13B, and Gemini 1.0 Pro) and to contexts extracted from articles web-crawled from specialised websites. The comparison uses a representative set of length-based, morphosyntactic, semantic, and discourse- related textual characteristics. We conclude that the LLM-based corpora can be combined effectively with a web-crawled one to form an academic corpus characterised by appropriate complexity and textual variety.
Affiliations

Citations

Nikolova, I., Bibauw, S., Dumont, A., Stas, F., Watrin, P., & François, T. (2024). LLM-Generated Contexts to Practice Specialised Vocabulary: Corpus Presentation and Comparison. In Mathieu Balaguer, Nihed Bendahman, Lydia-Mai Ho-Dac, Julie Mauclair, José G. Moreno, Julien Pinquier (ed.), Actes de la 31ème Conférence sur le Traitement Automatique des Langues Naturelles, volume 1 : articles longs et prises de position (pp. 472-498). Association pour le Traitement Automatique des Langues. https://hdl.handle.net/2078.5/274569