Optimizing a supervised classifier for a difficult language identification problem.

(2021) Eigth Workshop on NLP for Similar Languages, Varieties and Dialects, — Location: Kiev (online)

Files

2021vardial-111L.pdf
  • Open Access
  • Adobe PDF
  • 428.11 KB

Details

Authors
Abstract
This paper describes the system developed by the Laboratoire d’analyse statistique des textes for the Dravidian Language Identifica- tion (DLI) shared task of VarDial 2021. This task is particularly difficult because the ma- terials consists of short YouTube comments, written in Roman script, from three closely re- lated Dravidian languages, and a fourth cate- gory consisting of several other languages in varying proportions, all mixed with English. The proposed system is made up of a logistic regression model which uses as only features n-grams of characters with a maximum length of 5. After its optimization both in terms of the feature weighting and the classifier parameters, it ranked first in the challenge. The additional analyses carried out underline the importance of optimization, especially when the measure of effectiveness is the Macro-F1.
Affiliations

Citations

Bestgen, Y. (2021). Optimizing a supervised classifier for a difficult language identification problem. Proceedings of the Eigth Workshop on NLP for Similar Languages, Varieties and Dialects, p. 96-101. https://hdl.handle.net/2078.5/223190