PulseAugur
EN
LIVE 03:55:20

LLM refusal behavior inconsistent across models and settings, new papers reveal

Two new research papers explore the complexities of large language model (LLM) refusals. The first paper, "A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals," suggests that while knowledge-based and safety-based refusals share underlying mechanisms, they diverge in later layers, with safety refusals showing a stronger transfer to knowledge refusals. The second paper, "The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior," demonstrates that LLM safety evaluations are unreliable due to inconsistencies introduced by random seeds and temperature settings, with a significant percentage of prompts exhibiting decision flips. AI

IMPACT Highlights the need for more robust LLM safety evaluation protocols that account for stochastic variations in model behavior.

RANK_REASON Two academic papers published on arXiv analyzing LLM refusal mechanisms and stability.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

LLM refusal behavior inconsistent across models and settings, new papers reveal

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Two academic papers published on arXiv analyzing LLM refusal mechanisms and stability.
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
7 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Yuri Son, Seunghee Kim, Hyuhng Joon Kim, Taeuk Kim ·

    A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals

    arXiv:2609.00760v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly trained to decline queries that fall outside their knowledge (knowledge-based refusal, KR) or violate safety policies (safety-based refusal, SR). Although KR and SR result in superficial…

  2. arXiv cs.CL TIER_1 English(EN) · Erik Larsen ·

    The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior

    arXiv:2512.12066v3 Announce Type: replace-cross Abstract: Current safety evaluations of large language models rely on single-shot testing, implicitly assuming that model responses are deterministic and representative of the model's safety alignment. We challenge this assumption b…