A Question on the Explainability of Large Language Models and the Word-Level Univariate First-Order Plausibility Assumption

Bogaert, Jérémie;Standaert, François-Xavier
(2024) ReLM 2024 — Location: Vancouver, Canada

Files

AQuestionontheExplainabilityofLargeLanguageModelsandtheWord-LevelUnivariateFirst-OrderPlausibilityAssumpt.pdf
  • Open Access
  • Adobe PDF
  • 818.51 KB

Details

Authors
Abstract
The explanations of large language models have recently been shown to be sensitive to the randomness used for their training, creating a need to characterize this sensitivity. In this paper, we propose a characterization that questions the possi- bility to provide simple and informative explanations for such models. To this end, we give statistical definitions for the ex- planations’ signal, noise and signal-to-noise ratio. We high- light that, in a typical case study where word-level univariate explanations are analyzed with first-order statistical tools, the explanations of simple feature-based models carry more sig- nal and less noise than those of transformer ones. We then discuss the possibility to improve these results with alterna- tive definitions of signal and noise that would capture more complex explanations and analysis methods, while also ques- tioning the tradeoff with their plausibility for readers.
Affiliations

Citations

Bogaert, J., & Standaert, F.-X. (2024). A Question on the Explainability of Large Language Models and the Word-Level Univariate First-Order Plausibility Assumption. ReLM 2024, Vancouver, Canada. https://hdl.handle.net/2078.5/235628