A new study published on arXiv investigates whether large language models (LLMs) possess a Theory of Mind (ToM), the ability to understand others' beliefs, intentions, and emotions. Researchers compared the performance of five LLMs, including GPT-4o, against human controls using the adapted Strange Stories Paradigm. While smaller models showed limitations, GPT-4o demonstrated human-comparable accuracy and robustness in inferring character states, raising questions about the nature of LLM understanding versus sophisticated pattern matching. AI
IMPACT This research probes the depth of LLM understanding, potentially influencing how we develop and interpret AI's social-cognitive abilities.
RANK_REASON The cluster contains an academic paper published on arXiv evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- Anna Babarczy
- arXiv
- DagsHub
- GPT-4o
- Hugging Face
- large-language models
- Strange Stories Paradigm
- Theory of Mind
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →