Do NLP and machine learning improve traditional readability formulas?

François, Thomas;Miltsakaki, Eleni
(2012) First Workshop on Predicting and improving text readability for target reader populations (PITR2012) — Location: Montréal (7.June.2012)

Files

No attached file found for this publication.

Details

Authors
Abstract
Readability formulas are methods used to match texts with the readers’ reading level. Several methodological paradigms have previously been investigated in the field. The most popular paradigm dates several decades back and gave rise to well known readability formulas such as the Flesch formula (among several others). This paper compares this approach (henceforth ”classic”) with an emerging paradigm which uses sophisticated NLP-enabled features and machine learning techniques. Our experiments, carried on a corpus of texts for French as a foreign language, yield four main results: (1) the new readability formula performed better than the “classic” formula; (2) “non-classic” features were slightly more informative than “classic” features; (3) modern machine learning algorithms did not improve the explanatory power of our readability model, but allowed to better classify new observations; and (4) combining “classic” and “non-classic” features resulted in a significant gain in performance.
Affiliations

Citations

François, T., & Miltsakaki, E. (2012). Do NLP and machine learning improve traditional readability formulas? In Williams, S. and Siddharthan, A. and Nenkova, A. (ed.), Proceedings of the First Workshop on Predicting and improving text readability for target reader populations (PITR2012). https://hdl.handle.net/2078.5/51690