PulseAugur
实时 09:32:05
English(EN) The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior

研究发现大型语言模型安全拒绝在随机种子和温度下不稳定

一项新的研究论文强调了大型语言模型安全拒绝行为存在显著的不一致性。研究发现,在测试了不同的随机种子和温度设置后,18%到28%的有害提示导致了不同的拒绝决定。研究表明,增加温度参数会降低决策的稳定性,平均稳定性从温度0.0时的0.977下降到温度1.0时的0.942。研究结果表明,目前单次抽样安全评估方法不足,评估协议必须考虑到大型语言模型响应的随机性。 AI

影响 强调需要更稳健的大型语言模型安全评估方法,以考虑模型输出的随机变化。

排序理由 详细介绍研究结果的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现大型语言模型安全拒绝在随机种子和温度下不稳定

本文如何被排名

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍研究结果的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Erik Larsen ·

    安全性的不稳定性:随机种子和温度如何暴露 LLM 拒绝行为的不一致性

    arXiv:2512.12066v3 Announce Type: replace-cross Abstract: Current safety evaluations of large language models rely on single-shot testing, implicitly assuming that model responses are deterministic and representative of the model's safety alignment. We challenge this assumption b…