PulseAugur
EN
LIVE 15:42:48
ENTITY Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

PulseAugur coverage of Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents — every cluster mentioning Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
1 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
1 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 1 TOTAL
  1. RESEARCH · CL_131843 ·

    New research explores LLM agent advancements in skill selection, autonomous driving, and compliance

    Multiple research papers released on arXiv explore advancements in Large Language Model (LLM) agents, focusing on improving their capabilities and reliability. One paper introduces Best Prefix Selection (BPS) for optima…