PulseAugur
实时 02:13:20
English(EN) When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding

LLM政治事件编码的准确性与可靠性之争

一项新的研究论文探讨了在社会科学研究中使用大型语言模型(LLMs)进行政治事件编码的挑战。虽然更清晰、对LLM友好的代码本能显著提高分类准确性,但这种预测性能并不总是能转化为行为可靠性。研究表明,用于编码的LLM系统不仅应根据准确性进行评估,还应根据其维持底层编码逻辑的能力进行评估。 AI

影响 强调了在社会科学等专业领域,对LLM进行超越简单准确性的稳健评估的必要性。

排序理由 关于LLM应用和评估的学术论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

LLM政治事件编码的准确性与可靠性之争

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Zixian He, Bharath Raahul Murugesan, Patrick Brandt, Yibo Hu ·

    当更好的代码本不足以应对:LLM政治事件编码中的预测性能与行为可靠性

    arXiv:2606.06781v1 Announce Type: new Abstract: High accuracy does not necessarily make an LLM a faithful coder. This issue matters because many social-science studies rely on expert-written codebooks to turn text into structured data. We study this problem in political event cod…

  2. arXiv cs.CL TIER_1 English(EN) · Yibo Hu ·

    当更好的代码本不足以应对:LLM政治事件编码的预测性能和行为可靠性

    High accuracy does not necessarily make an LLM a faithful coder. This issue matters because many social-science studies rely on expert-written codebooks to turn text into structured data. We study this problem in political event coding, a challenging source-target relation classi…