Part-of-speech tagging of yoruba standard, language of niger-congo family

Adedjouma Sèmiyou;Aoga, John;Igue, M. Akanni
(2013) Research Journal of Computer and Information Technology Sciences — Vol. 1, n° 1, p. 2-5 (2013)

Files

2013-jan-1ISCA-RJCITS-2013-003.pdf
  • Open Access
  • Adobe PDF
  • 169.87 KB

Details

Authors
  • Adedjouma Sèmiyou
    Author
  • Aoga, JohnUCLouvain
    Author
  • Igue, M. Akanni
    Author
Abstract
The utilization of corpora is a critical phase of systems of Natural Language Processing (NLP) based on statistical methods. This point is crucial for less equipped and less computerized languages like African languages. This paper aimed to design a yoruba corpus. Yoruba is an African language of Niger-Congo family. It is spoken by more than thirty million people around the world and particularly in Nigeria and Benin. The main motivation of this work was to obtain training data for PoS taggers and to provide applications of Yoruba Language Processing (YLP) with a basic tool. The tagging was performed with SVMTool one of the Part-of-Speech taggers widely used. The preprocessing of the text general outline has been ensured by Perl scripts. The corpus with 312,562 words, formed from the Web, was annotated with an accuracy of 98.04%. This annotated corpus might be used in translation system.
Affiliations
  • Université d'Abomey CalaviLETIA

Citations

Adedjouma Sèmiyou, Aoga, J., & Igue, M. A. (2013). Part-of-speech tagging of yoruba standard, language of niger-congo family. Research Journal of Computer and Information Technology Sciences, 1(1), 2-5. https://hdl.handle.net/2078.5/125839 (Original work published 2013)