This paper describes the system developed by the Laboratoire d’analyse statistique des textes (LAST) for the Lexical Complexity Prediction shared task at SemEval-2021. The proposed system is made up of a LightGBM model fed with features obtained from many word fre- quency lists, published lexical norms and psy- chometric data. For tackling the specificity of the multi-word task, it uses bigram associ- ation measures. Despite that the only contex- tual feature used was sentence length, the sys- tem achieved an honorable performance in the multi-word task, but poorer in the single word task. The bigram association measures were found useful, but to a limited extent.
Bestgen, Y. (2021). LAST at SemEval-2021 Task 1: Improving multi-word complexity prediction using bigram association measures. Proceedings of SemEval-2021. Published. SemEval-2021, (online). https://hdl.handle.net/2078.5/223260