Confirming the effectiveness of a simple language-agnostic yet very strong system for hate speech and offensive content identification

(2022) FIRE 2022, Forum for Information Retrieval Evaluation — Location: Indian Statistical Institute, Kolkata, India (9.December.2022)

Files

sample-1colSoumcopie.pdf
  • Open Access
  • Adobe PDF
  • 884.41 KB

Details

Authors
Abstract
At the 2021 edition of HASOC, the SATLab team proposed a very simple language-agnostic system for hate speech and offensive content identification. This system proved to be extremely effective for the two less resourced languages (e.g., Hindi and Marathi). The present paper describes the use of the same system for task 3 of the 2022 edition of HASOC on hate speech and offensive content identification in Marathi. It consists of a logistic regression applied to character n-grams. It ranked fifth on subtask 3A (macro-F1=0.937), quite close to the first ones, second on subtask 3B (macro-F1=0.915), very close to the first one, and first (macro-F1=0.961) with more than 16 Macro F1 points ahead of the second one in subtask 3C. These results confirm the effectiveness of the approach and suggest that studies evaluating different systems for this kind of problem should employ a character n-gram based approach as a baseline. They also show that the task is extremely simple since all macro-F1s are greater than or equal to 0.915.
Affiliations

Citations

Bestgen, Y. (2022). Confirming the effectiveness of a simple language-agnostic yet very strong system for hate speech and offensive content identification. Forum for Information Retrieval Evaluation (Working Notes)(FIRE), CEUR-WS. org. Published. FIRE 2022, Forum for Information Retrieval Evaluation, Indian Statistical Institute, Kolkata, India. https://hdl.handle.net/2078.5/224361