PulseAugur
实时 01:29:12
English(EN) Latent-space Attacks for Refusal Evasion in Language Models

大型语言模型拒绝策略研究:PsychoSafe提升支持,规避攻击破坏安全

研究人员开发了PsychoSafe框架,通过采用心理学启发的沟通策略来改进大型语言模型拒绝有害请求的方式。这种方法将拒绝重新定义为支持性互动,增强了外部资源推荐和心理基础。另外,另一项研究引入了用于规避拒绝的潜在空间攻击,该研究分析了如何通过操纵模型内部表征来抑制拒绝行为,从而绕过大型语言模型的安全机制。 AI

影响 大型语言模型拒绝策略和规避技术的发展凸显了人工智能安全和对齐方面持续存在的挑战。

排序理由 两篇关于大型语言模型安全和拒绝机制的研究论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

大型语言模型拒绝策略研究:PsychoSafe提升支持,规避攻击破坏安全

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两篇关于大型语言模型安全和拒绝机制的研究论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
safety, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
93 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Anne Lauscher ·

    PsychoSafe:在大型语言模型中引发心理学知情拒绝

    Large language models (LLMs) routinely face requests that should be refused, creating a trade-off between helpfulness and harm prevention. However, refusals themselves can be helpful. In high-risk interactions involving crisis, coercion, or escalating intent, blunt non-compliance…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    PsychoSafe: 在大型语言模型中引发心理学启发的拒绝

    A psychologically-informed refusal framework called PsychoSafe is developed for large language models to improve harmful request handling through structured supportive communication, showing enhanced refusal quality and resource referral while maintaining performance on non-refus…

  3. arXiv cs.AI TIER_1 English(EN) · Giorgio Piras, Raffaele Mura, Fabio Brau, Maura Pintor, Luca Oneto, Fabio Roli, Battista Biggio ·

    语言模型中的潜在空间攻击用于规避拒绝

    arXiv:2605.21706v2 Announce Type: replace Abstract: Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations. Existing methods do so by ablating a refusal direction from model activati…