PulseAugur
EN
LIVE 06:27:14

New benchmark reveals LLMs struggle with enzyme classification due to knowledge gaps

Researchers have developed EC-Reason-Bench, a new benchmark designed to diagnose why large language models struggle with enzyme classification tasks. The benchmark focuses on four key areas: output structure, external knowledge, reasoning structure, and reasoning robustness. Experiments revealed that providing external knowledge significantly improves LLM performance, often narrowing the gap between different models. The study also found that in closed-book settings, the effectiveness of cascading and chain-of-thought reasoning depends on a model's tendency to abstain from answering. AI

IMPACT Highlights critical limitations in LLM's scientific reasoning and knowledge integration, suggesting a need for improved external knowledge access.

RANK_REASON The item describes a new benchmark for evaluating LLMs on a specific scientific task, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals LLMs struggle with enzyme classification due to knowledge gaps

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Linyu Li, Zhi Jin, Yichi Zhang, Dongming Jin, Yuanpeng He, Huanyao Zhang, Xuan Zhang, Gadeng Luosang, Nyima Tashi ·

    Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

    arXiv:2607.26397v1 Announce Type: new Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC …