Accès sémantique aux bases de données documentaires : techniques symboliques de traitement automatique du langage pour l'indexation thématique et l'extraction d'informations temporelles
(en) Within our information society and economy, access to the growing bulk of electronic documents is a major concern. Surprisingly, information retrieval tools often only work on a classical search scheme, based on words. But, as a matter of fact, the relation between words and concepts can not always be reduced to a single link. A concept can be named by several expressions (« unemployed person » and « job-seeker »), and conversely one word can sometimes depict more than one concept (« carrot » as a vegetable or in geology). To cope with this difficulty, it might be interesting to represent documents in a concepts space in place of a words space. In this thesis we are proposing 3 methods designed to complete the documents representation with semantic data, which should allow for an improved access to the information. First, a (semi) automatic classification method was set up to index documents, using a set of defined categories (supervised classification). These categories, generally assigned manually by human specialists having an expertise in a specific domain, add sense thanks to their well-defined semantics, given by a thesaurus or an ontology. This method allows to partly automate the work and, as a consequence, to lower the cost, improve scalability and consistency, and finally allow human validation. Moreover, specific processing of some information items, as time, allows to improve the semantic representation of the document. Using a set of manually built transducers, the recognition of temporal expressions (mainly adverbials) is followed by their interpretation which results into a unique and normalized value within a well-defined temporal space. The interest of such a system is broader than information retrieval as it can be useful for plenty of other NLP applications (information extraction, automatic translation, question answering systems, technology watch, etc.) in various domains (process analysis, planning, project management, etc.). Finally, document indexing can also be carried out in a more complex way. Some information items, like spatial or temporal information, can be viewed as complementary to others (« bank robbery in Brussels the 14/01/2011 »). This kind of complex information can be better indexed if these two particular dimensions are taken into account as metadata of the main information. We proposed a thematic and temporal indexing system that outlines interest and perspectives for applications within this multidimensional indexing approach.
Kevers, L. (2011). Accès sémantique aux bases de données documentaires : techniques symboliques de traitement automatique du langage pour l’indexation thématique et l’extraction d’informations temporelles. https://hdl.handle.net/2078.5/148434