PulseAugur
EN
LIVE 10:55:47

Nonobench benchmark tests 49 LLMs on nonogram puzzles

A new open-source benchmark called Nonobench has been released to evaluate the performance of 49 large language models on nonogram puzzles. The benchmark includes standard puzzles of varying sizes and more challenging random puzzles designed to test logic beyond simple pattern recognition. Results show a significant drop in solve rates as puzzle complexity increases, with GPT-6 Astra being the only model to solve all standard puzzles, while Claude Opus 5.5 performed best on the harder puzzles. AI

IMPACT Provides a new evaluation metric for LLM reasoning capabilities beyond traditional NLP tasks.

RANK_REASON The cluster describes a new open-source benchmark for evaluating LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/MachineLearning →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Nonobench benchmark tests 49 LLMs on nonogram puzzles

How we ranked this

Signal score
6 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster describes a new open-source benchmark for evaluating LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/mauricekleine ·

    Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1wxa2bs/nonobench_an_open_benchmark_of_49_llms_on/"> <img alt="Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]" src="https://preview.redd.it/qfn8kr1ppeth1.png?width…