A new open-source benchmark called Nonobench has been released to evaluate the performance of 49 large language models on nonogram puzzles. The benchmark includes standard puzzles of varying sizes and more challenging random puzzles designed to test logic beyond simple pattern recognition. Results show a significant drop in solve rates as puzzle complexity increases, with GPT-6 Astra being the only model to solve all standard puzzles, while Claude Opus 5.5 performed best on the harder puzzles. AI
IMPACT Provides a new evaluation metric for LLM reasoning capabilities beyond traditional NLP tasks.
RANK_REASON The cluster describes a new open-source benchmark for evaluating LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →