Word sense discrimination aims at automatically determining which instances of a n ambiguous word share the same sense. A fully unsupervised technique based on a vector representation of word senses w as proposed by Schütze [Schütze98]. While the original model was assumed to be Gaussian, practical results were only reported for an approximated model mak ing hard decisions between sense clusters. We show in the present study that a r eal Gaussian model provides a significant accuracy improvement while remaining f ully tractable. An alternative discrete naïve Bayes model was presented in [Manning99]. We propose here a description of both models in a unified statistical formalism in order to stress the similarities and diffe rences between both approaches. Several practical experiments are conducted on t he New York Times News 1997 corpus. They illustrate the respective advantages of various approaches trading off discrimination accuracy and computation time. We also show the interest of a global selection of content words to characterize the context of an ambiguous instance in the naï ve Bayes model.
de Marneffe, M.-C., & Dupont, P. (2004). Comparative study of statistical word sense discrimination,. Proceedings of the 7th International Conference on Textual Data Statistical Analysis, Louvain-la-Neuve, Belgique. https://hdl.handle.net/2078.5/254165