Using only character ngrams for hate speech and offensive content identification in five low-ressource languages

(2023) Forum for Information Retrieval Evaluation — Location: Goa, India (15.December.2023)

Files

51_final_crs.pdf
  • Open Access
  • Adobe PDF
  • 389.36 KB

Details

Authors
Abstract
This paper describes the system proposed by the SATLab for hate speech and offensive content identification in five low-ressource languages. This language-agnostic system applies a classical supervised learning to character n-grams, using no other data than the learning materials. After optimizing a series of parameters, it ranked first in the Bodo task and second in the Gujarati task, for which the learning material contained only 200 tweets. It also performed well in the Sinhala and Assamese task, but was outperformed by several systems in the Bengali task.
Affiliations

Citations

Bestgen, Y. (2023). Using only character ngrams for hate speech and offensive content identification in five low-ressource languages. In Ghosh, K., Mandl, T., Majumder, P. & Mitra M. (ed.), Working Notes of FIRE 2023 - Forum for Information Retrieval Evaluation (FIRE-WN 2023) (pp. 411-417). https://hdl.handle.net/2078.5/238453