A new benchmark, MedMCP-Calc, has been developed to evaluate Large Language Models (LLMs) in realistic medical calculator scenarios. The benchmark, which integrates the Model Context Protocol (MCP), includes 118 tasks across four clinical domains, simulating real-world adaptive processes like EHR data acquisition and multi-step computations. Evaluations revealed that even advanced models like Claude Opus 4.5 struggle with selecting appropriate calculators, performing SQL-based database interactions, and utilizing external tools for numerical tasks. To address these limitations, CalcMate, a fine-tuned model, was developed and demonstrated state-of-the-art performance among open-source models. AI
IMPACT Highlights critical gaps in LLM capabilities for complex, real-world clinical applications, driving development of specialized models.
RANK_REASON The cluster describes a new benchmark and research paper published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →