PulseAugur
中
实时 18:46:08
English(EN) Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives

大型语言模型在通过文本诊断人格障碍方面优于心理健康专家

一项新研究评估了大型语言模型(特别是Gemini Pro)在根据自传体叙述诊断人格障碍方面与心理健康专业人士的对比情况。虽然大型语言模型在整体诊断评分上表现更高,尤其是在边缘性人格障碍方面,但它们在诊断自恋型人格障碍方面存在显著不足。模型提供了详细的、以模式为中心的理由,这与人类专家的更简洁、以患者为中心的方法形成对比,突显了大型语言模型临床评估中潜在的偏见和可靠性问题。 AI

影响 大型语言模型在临床叙事分析方面显示出潜力,但由于偏见和可靠性问题,需要仔细验证。

排序理由 学术论文,评估大型语言模型在特定临床任务上与人类专家的表现。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

大型语言模型在通过文本诊断人格障碍方面优于心理健康专家

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
学术论文,评估大型语言模型在特定临床任务上与人类专家的表现。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
163 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Karolina Dro\.zd\.z, Kacper Dudzic, Anna Sterna, Marcin Moskalewicz ·

    模式 vs. 患者:通过第一人称叙述评估大型语言模型与精神科医生在人格障碍诊断上的表现

    arXiv:2512.20298v2 Announce Type: replace Abstract: Growing reliance on LLMs for psychiatric self-assessment raises questions about their ability to interpret qualitative patient narratives. This depth-first case study provides the first direct comparison of state-of-the-art LLMs…