Researchers have developed RxnOptBench, a new benchmark designed to evaluate the capabilities of large language models (LLMs) in optimizing chemical reaction conditions. Unlike previous benchmarks that focus on tasks like reaction classification or retrosynthesis, RxnOptBench specifically assesses LLMs' ability to analyze real-world reaction-screening tables and select optimal parameters such as catalyst, ligand, solvent, temperature, and time. The benchmark, which uses data from 2025 organic methodology papers, measures performance based on a combined score of yield and stereoselectivity metrics like enantiomeric excess (ee), diastereomeric ratio (dr), and regioisomeric ratio (rr). Initial testing across several frontier and chemistry-specialized LLMs revealed significant room for improvement, with even top-performing models struggling on multi-axis selection tasks. AI
IMPACT This benchmark could drive the development of more capable LLMs for scientific applications, particularly in chemistry.
RANK_REASON The item describes a new benchmark and evaluation of LLMs for a specific scientific task, fitting the research category. [lever_c_demoted from research: ic=1 ai=1.0]
- Atmosphere
- catalyst
- chemical reagent
- crop yield
- diastereomeric ratio (dr)
- enantiomeric excess (ee)
- ligand
- organic methodology
- regioisomeric ratio (rr)
- RxnOptBench
- solvent
- stereoselectivity
- temperature
- Time
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →