A new benchmark, GPAgentBench-2K, has been developed to evaluate large language model (LLM) agents in complex clinical decision-making scenarios. This benchmark utilizes Constrained Markov Decision Processes (CMDPs) based on real-world GP encounter records, incorporating a six-action clinical workflow and safety-informed abstention. Evaluations of 16 LLMs showed a significant drop in performance as the action space increased, with even top-performing models failing to meet safety constraints in over half of high-risk cases. While constrained reinforcement learning methods improved performance compared to unconstrained approaches, they still fell short of clinical safety standards. AI
IMPACT Highlights critical safety gaps in LLM agents for complex, real-world applications like healthcare.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating LLM agents in a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
- Constrained Group Relative Policy Optimization
- Constrained MDP
- GPAgentBench-2K
- large-language models
- Markov decision processes
- reinforcement learning
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →