This paper investigates how autoregressive models trained on poker games represent information about an opponent's hand. Researchers found that while models show some predictive capability regarding opponent ranges, this information is largely explained by visible betting patterns rather than internal hidden states. The study introduces the concept of 'composition-bounded predictive support' to describe this phenomenon, suggesting that positive probes for belief tracking should be carefully interpreted against alternative explanations. AI
IMPACT Suggests that current AI models may not possess true 'belief tracking' capabilities, influencing how we interpret their behavior in complex strategic environments.
RANK_REASON The item is an academic paper published on arXiv detailing research findings on AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →