A new study evaluated six large language models on their ability to distinguish between actual pain signals and fabricated ones in clinical speech transcripts. While most models correctly abstained from predicting pain when no explicit cues were present, Gemini 2.5 Flash and Llama 3.1 8B demonstrated a tendency to confidently fabricate pain scores. The research highlights that models can be influenced by prompt framing, affecting their abstention rates, and that transcript-based predictions alone are insufficient for accurate pain assessment in such scenarios. AI
IMPACT Highlights potential for LLMs to confidently fabricate information, impacting reliability in sensitive applications.
RANK_REASON Academic paper detailing LLM evaluation on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →