A new paper argues that Surprisal Theory, often framed as a computational-level explanation, is not truly representation-agnostic, especially in the context of large language models (LLMs). The authors contend that the uncritical use of LLM-surprisal metrics obscures the underlying representational and algorithmic choices made by different models. Through three analyses, the paper demonstrates that algorithm and model architecture significantly influence language model probabilities, urging researchers to reconsider treating LLM probabilities as interchangeable when testing Surprisal Theory. AI
IMPACT Challenges the theoretical underpinnings of using LLM probabilities for research, potentially impacting how language model capabilities are evaluated.
RANK_REASON Academic paper published on arXiv discussing theoretical aspects of language models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →