Introduction This study investigates issues that arise when dealing with sparse data problems in developing language models. The practical problem under consideration is Spanish conversational speech recognition over the telephone. The primary data source, the Call Home Spanish corpus, is small in comparison with the English Switchboard corpus. With this problem in mind we explore alternative smoothing techniques. As decribed below, one of these techniques relies on a word-to-class mapping and an associated class bigram model [3]. Future extensions of this approach may allow for learning of more complex languages models, e.g. general stochastic regular grammars, at the class level or serve as constraints for language model adaptation within the maximum entropy framework. The Katz Model versus Absolute Discounting The Katz model, described in [1], may be summarized as follows. Given an initial maximum likelihood estimation of an N-gram, the model is smoothed by discounting a certain ...
Dupont, P. (1995). Interpolated Word and Class Bigram Models for Spanish Conversational Speech Recognition. IEEE Workshop on Automatic Speech Recognition, Snowbird, Utah, USA. https://hdl.handle.net/2078.5/253715