Researchers have introduced a new metric called the Fatigue Index (FI) to measure and diagnose "cognitive fatigue" in autoregressive language models. This phenomenon, characterized by degraded performance during long-horizon generation, can lead to repetitive text and loss of instruction adherence. The FI, which is model-agnostic and lightweight, aggregates signals related to prompt attention decay, representational drift, and entropy miscalibration. Across nine models ranging from 1B to 13B parameters, the FI effectively predicts task degradation and repetition, revealing non-monotonic scaling behaviors and identifying factors that accelerate fatigue onset. AI
IMPACT Introduces a new diagnostic tool to monitor and potentially mitigate performance degradation in LLMs during long-form generation.
RANK_REASON The cluster contains an academic paper detailing a new measurement and formalization of a phenomenon in language models.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →