A LightGBM model fed with target word lex- ical characteristics and features obtained from word frequency lists, psychometric data and bigram association measures has been opti- mized for the 2021 CMCL Shared Task on Eye-Tracking Data Prediction. It obtained the best performance of all teams on two of the five eye-tracking measures to predict, allowing it to rank first on the official challenge criterion and to outperform all deep-learning based sys- tems participating in the challenge.
Bestgen, Y. (2021). LAST at CMCL 2021 shared task: Predicting gaze data during reading with a gradient boosting decision tree approach. Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics. Published. Workshop on Cognitive Modeling and Computational Linguistics, Mexico (online). https://hdl.handle.net/2078.5/223337