This paper presents the solution proposed by the SATLab to classify all the words in short utterances into one of seven categories which include three languages, two of which are closely related Dravidian languages sparsely endowed with linguistic resources (Tulu and Kannada) and a category for tokens that mix several languages. This language-agnostic system uses only character ngrams as features and a classical supervised learning procedure. After optimizing a series of parameters, it ranked first in the CoLI-Tunglish challenge, with a Macro-F1 of 0.813, virtually on a par with the second-place system. Part of its effectiveness comes from taking into account the context in which each word to be categorized is used.
Bestgen, Y. (2023). Using character ngrams for word-level language identification in trilingual code-mixed data (and even more). In Ghosh, K., Mandl, T., Majumder, P. & Mitra M. (ed.), Working Notes of FIRE 2023 - Forum for Information Retrieval Evaluation (FIRE-WN 2023) (pp. 191-197). https://hdl.handle.net/2078.5/238454