Can rules still beat neural networks? The case of automatic normalisation for 18th-century Portuguese texts

Zilio, Leonardo;Lazzari, Rafaela Radünz;Finatto, Maria José Bocorny
(2024) Proceedings of the 16th International Conference on Computational Processing of the Portuguese Language, PROPOR 2024 - vol. 2 — ISBN: [979-8-89176-062-2], p. 205, published

Files

2024propor-212.pdf
  • Open Access
  • Adobe PDF
  • 154.95 KB

Details

Authors
  • Author
  • Lazzari, Rafaela RadünzUFRGS
    Author
  • Finatto, Maria José BocornyUFRGS
    Author
Abstract
In this paper, we tested whether fine-tuned neural machine translation (NMT) models can produce better results than a rulebased method for the task of normalising historical medical documents written in 18th-century Portuguese. We used a replacement glossary as basis for the rulebased method, and tested three NMT models against it in an in-domain setting and in two out-of-domain scenarios. In-domain results showed that the rule-based method was better than off-the-shelf NMT models, and it still surpassed one of the in-domain fine-tuned models. The fine-tuned models showed their efficacy on out-of-domain settings, where only one NMT model did not surpass the rule-based method in one scenario.
Affiliations

Citations

Zilio, L., Lazzari, R. R., & Finatto, M. J. B. (2024). Can rules still beat neural networks? The case of automatic normalisation for 18th-century Portuguese texts. In Gamallo, Pablo ; et al. (ed.), Proceedings of the 16th International Conference on Computational Processing of the Portuguese Language, PROPOR 2024 - vol. 2 (p. p. 205). Association for Computational Linguistics. https://hdl.handle.net/2078.5/274803