AgentHarm
PulseAugur coverage of AgentHarm — every cluster mentioning AgentHarm across labs, papers, and developer communities, ranked by signal.
-
Qwen 3.8 27B shows strong safety guardrails in AgentHarm benchmark
The Qwen 3.8 27B model demonstrated a strong refusal rate of 75% on harmful tasks within the AgentHarm benchmark. This performance indicates robust safety guardrails, contrasting with the Blackfrost package which achiev…
-
New prompt injection detection techniques leverage cross-domain methods
Researchers have developed seven novel techniques for detecting prompt injection attacks, moving beyond traditional pattern matching and fine-tuned transformer classifiers. These new methods draw inspiration from divers…
-
New framework guides LLM agents to revise plans, improving safety
Researchers have developed TRIAD, a new framework for LLM agents that integrates guardrails to improve safety and utility. Unlike traditional guardrails that simply block unsafe actions, TRIAD provides feedback to guide…
-
New framework LiSA enhances AI guardrails with sparse failure data
Researchers have developed LiSA (Lifelong Safety Adaptation), a new framework designed to improve AI guardrails by learning from sparse and noisy failure data. LiSA uses structured memory to generalize from individual i…