In this paper, we propose a way of incorporating additional knowledge in probabilistic automata inference, by using typed automata. We compare two kinds of knowledge that are introduced into the learning algorithms. A statistical clustering algorithm and a part-of-speech tagger are used to label the data according to statistical or syntactic information automatically obtained from the data. The labeled data is then used to infer correctly typed automata. Compared to a previously proposed method which improved grammatical inference, our approach yields better automata. The inference of typed automata with statistically labeled data provides language models competitive with state-of-the art n-grams on the Air Travel Information System (ATIS) task.
Affiliations
Université de Montréal, CanadaDepartement IRO
Université Jean Monnet, Saint-Etienne, FranceEurise