A new benchmark called InsClaimBench has been developed to evaluate the performance of large language models in insurance claim adjudication. The benchmark, which includes 3,780 cases across auto, property, and health insurance, assesses models' ability to connect evidence, rules, judgments, and payout calculations. Evaluations of six LLMs showed a decline in reliability along the decision chain, with payout-decision accuracy ranging from 74.23% to 80.19%, and joint decision-amount accuracy dropping to 47.54% to 73.15%. The study highlights that while models may perform well on individual rules, they struggle with consistent propagation of information across the entire adjudication process. AI
IMPACT Highlights limitations of current LLMs in complex, multi-step decision-making tasks, indicating a need for improved reasoning and consistency.
RANK_REASON The item is a research paper introducing a new benchmark for evaluating LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- InsClaimBench
- Litmaps
- ScienceCast
- Scite
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →