A new benchmark called BioEVAL has been developed to assess the experimental reasoning capabilities of large language and multimodal models in bioengineering. This initiative, involving 22 research groups, created a PhD-level benchmark with 608 items across 11 subfields, including multiple-choice questions, literature synthesis tasks, and multimodal problems. Cloud-scale models like ChatGPT, Gemini, and Grok achieved up to 90% accuracy on multiple-choice questions, but performance varied significantly across different bioengineering domains. AI
IMPACT This benchmark could guide future LLM development for specialized scientific applications and reveal performance gaps in complex reasoning tasks.
RANK_REASON The cluster describes a new academic benchmark for evaluating LLMs and multimodal models in a specific scientific domain. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- bioengineering
- ChatGPT
- Gemini
- Grok
- Hugging Face
- large language models
- multimodal models
- Vinny Chandran Suja
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →