The utilization of corpora is a critical phase of systems of Natural Language Processing (NLP) based on statistical methods. This point is crucial for less equipped and less computerized languages like African languages. This paper aimed to design a yoruba corpus. Yoruba is an African language of Niger-Congo family. It is spoken by more than thirty million people around the world and particularly in Nigeria and Benin. The main motivation of this work was to obtain training data for PoS taggers and to provide applications of Yoruba Language Processing (YLP) with a basic tool. The tagging was performed with SVMTool one of the Part-of-Speech taggers widely used. The preprocessing of the text general outline has been ensured by Perl scripts. The corpus with 312,562 words, formed from the Web, was annotated with an accuracy of 98.04%. This annotated corpus might be used in translation system.
Affiliations
Université d'Abomey CalaviLETIA
Citations
APA
Chicago
FWB
Adedjouma Sèmiyou, Aoga, J., & Igue, M. A. (2013). Part-of-speech tagging of yoruba standard, language of niger-congo family. Research Journal of Computer and Information Technology Sciences, 1(1), 2-5. https://hdl.handle.net/2078.5/125839 (Original work published 2013)