Researchers have introduced BenchBench-Protocol, a new benchmark designed to evaluate the reasoning and modification capabilities of large language models in real-world wet-lab scientific protocols. This benchmark comprises 149 tasks derived from actual modifications made by scientists to published experimental procedures. Claude Opus 5 achieved the highest score at 59.2% on the normalized rubric, outperforming other evaluated models which ranged from 34.1% to 47.1%. The benchmark remains unsaturated, indicating potential for further improvement, and highlights the growing importance of assessing LLMs on practical scientific tasks. AI
影响 This benchmark could drive LLM development towards more practical scientific applications and improve their utility in life sciences research.
排序理由 The item is a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- BenchBench-Protocol
- CatalyzeX
- Claude Opus 5
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →