PulseAugur
实时 07:27:16
English(EN) HarmReduction: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

新基准测试揭示大型语言模型在减害信息方面存在安全风险

一项名为HarmReduction的新基准测试已被开发出来,用于评估大型语言模型(LLMs)在为药物使用者提供减害信息的准确性和安全性。该基准测试包含2,160个问答证据对,涵盖三个任务:评估安全边界、提供定量数据以及推断多物质使用风险。初步结果表明,当前最先进的大型语言模型在准确性方面存在困难,并且可能对寻求此类敏感信息的用户构成重大的安全风险。 AI

影响 强调了在公共卫生等敏感领域对大型语言模型进行专门评估以防止伤害的迫切需求。

排序理由 该集群包含一篇介绍用于评估大型语言模型的新基准测试的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准测试揭示大型语言模型在减害信息方面存在安全风险

本文如何被排名

Signal score
22 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇介绍用于评估大型语言模型的新基准测试的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Kaixuan Wang, Chenxin Diao, Jason T. Jacques, Zhongliang Guo, Shuai Zhao ·

    HarmReduction:在减少伤害信息提供方面对大型语言模型进行基准测试,以支持药物使用者

    arXiv:2507.21815v2 Announce Type: replace Abstract: Millions of individuals' well-being are challenged by the harms of substance use. Harm reduction as a public health strategy provides non-judgemental, evidence-based information intended to improve health outcomes and reduce ass…