Using character ngrams for word-level language identification in trilingual code-mixed data (and even more)

(2023) Forum for Information Retrieval Evaluation (FIRE-WN 2023) — Location: Goa, India (15.December.2023)

Files

Camera_Ready_9UsingCharacterN-GramsforWord-Level.pdf
  • Open Access
  • Adobe PDF
  • 374.26 KB

Details

Authors
Abstract
This paper presents the solution proposed by the SATLab to classify all the words in short utterances into one of seven categories which include three languages, two of which are closely related Dravidian languages sparsely endowed with linguistic resources (Tulu and Kannada) and a category for tokens that mix several languages. This language-agnostic system uses only character ngrams as features and a classical supervised learning procedure. After optimizing a series of parameters, it ranked first in the CoLI-Tunglish challenge, with a Macro-F1 of 0.813, virtually on a par with the second-place system. Part of its effectiveness comes from taking into account the context in which each word to be categorized is used.
Affiliations

Citations

Bestgen, Y. (2023). Using character ngrams for word-level language identification in trilingual code-mixed data (and even more). In Ghosh, K., Mandl, T., Majumder, P. & Mitra M. (ed.), Working Notes of FIRE 2023 - Forum for Information Retrieval Evaluation (FIRE-WN 2023) (pp. 191-197). https://hdl.handle.net/2078.5/238454