PulseAugur
实时 17:18:55
English(EN) The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior

新论文揭示,LLM的拒绝行为在不同模型和设置下存在不一致性

两篇新研究论文探讨了大语言模型(LLM)拒绝行为的复杂性。第一篇论文《知识型和安全型拒绝的统一机制分析》提出,尽管知识型和安全型拒绝共享潜在机制,但在后续层级上存在分歧,安全型拒绝向知识型拒绝的迁移更强。第二篇论文《安全性的不稳定性:随机种子和温度如何暴露LLM不一致的拒绝行为》表明,由于随机种子和温度设置引入的不一致性,LLM的安全评估并不可靠,有相当比例的提示会表现出决策翻转。 AI

影响 强调需要更鲁棒的LLM安全评估协议,以应对模型行为中的随机变化。

排序理由 两篇发表在arXiv上的学术论文,分析了LLM的拒绝机制和稳定性。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新论文揭示,LLM的拒绝行为在不同模型和设置下存在不一致性

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两篇发表在arXiv上的学术论文,分析了LLM的拒绝机制和稳定性。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
16 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Yuri Son, Seunghee Kim, Hyuhng Joon Kim, Taeuk Kim ·

    知识与安全相关拒绝的统一机制分析

    arXiv:2609.00760v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly trained to decline queries that fall outside their knowledge (knowledge-based refusal, KR) or violate safety policies (safety-based refusal, SR). Although KR and SR result in superficial…

  2. arXiv cs.CL TIER_1 English(EN) · Erik Larsen ·

    安全性的不稳定性:随机种子和温度如何暴露 LLM 拒绝行为的不一致性

    arXiv:2512.12066v3 Announce Type: replace-cross Abstract: Current safety evaluations of large language models rely on single-shot testing, implicitly assuming that model responses are deterministic and representative of the model's safety alignment. We challenge this assumption b…