A new benchmark called "Humanity's Sixth Sense" has been introduced to evaluate intuitive visual reasoning, encompassing spatial, causal, and social understanding. The benchmark reveals a substantial performance gap between humans and AI models, with humans achieving a score of 93.1%. The leading AI model, GPT-6-astra, scored 53.6%, while the median model performance was significantly lower at 30.9%. This highlights the ongoing challenge in developing AI systems that can match human-level intuitive reasoning. AI
IMPACT Highlights a significant gap in AI's intuitive reasoning capabilities compared to humans, indicating areas for future research and development.
RANK_REASON The cluster introduces a new benchmark for AI evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →