A new benchmark called LogiMed-RoB has been developed to assess the logical consistency of large language models (LLMs) in medical risk-of-bias assessments. The benchmark, based on expert logic from Cochrane Risk of Bias 2.0, includes over 860 randomized controlled trials and 14,820 queries. Experiments revealed that while LLMs can achieve high atomic consistency, their overall logical consistency collapses significantly, with some open-weight models performing poorly. The study also identified a gap where models fail to deduce correct outcomes even when provided with relevant evidence, highlighting the need for rigorous logical verification before clinical deployment. AI
IMPACT Highlights critical reasoning flaws in LLMs, underscoring the need for white-box logical verification for clinical deployment.
RANK_REASON Academic paper introducing a new benchmark for LLM evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Cochrane Risk of Bias (RoB) 2.0
- Hierarchical Logical Consistency (HLC)
- Hugging Face
- large language models
- LogiMed-RoB
- open-weight architectures
- randomized controlled trial
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →