Researchers have developed EC-Reason-Bench, a new benchmark designed to diagnose why large language models struggle with enzyme classification tasks. The benchmark focuses on four key areas: output structure, external knowledge, reasoning structure, and reasoning robustness. Experiments revealed that providing external knowledge significantly improves LLM performance, often narrowing the gap between different models. The study also found that in closed-book settings, the effectiveness of cascading and chain-of-thought reasoning depends on a model's tendency to abstain from answering. AI
IMPACT Highlights critical limitations in LLM's scientific reasoning and knowledge integration, suggesting a need for improved external knowledge access.
RANK_REASON The item describes a new benchmark for evaluating LLMs on a specific scientific task, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
- EC-Reason-Bench
- Enzyme classification by ligand binding
- European Community number
- Protein function classification via support vector machine approach
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →