A new benchmark called LaughBench has been introduced to evaluate AI models' ability to generate novel, funny jokes, which the creator posits is a strong indicator of general intelligence. While current frontier models like GPT 5.6 "Sol" and Fable can occasionally elicit a smile, none have yet succeeded in consistently producing laughter. The benchmark aims to measure "real intelligence" beyond tasks that can be mastered through memory or computation, distinguishing it from benchmarks like FrontierMath Open Problems and Arc Agi, which may be solvable with specialized skills rather than broad intelligence. AI
IMPACT Proposes a novel benchmark for assessing AI general intelligence through humor, potentially revealing limitations in current models.
RANK_REASON New benchmark proposed for AI general intelligence. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →