PulseAugur
EN
LIVE 20:17:34
ENTITY AgentHarm

AgentHarm

PulseAugur coverage of AgentHarm — every cluster mentioning AgentHarm across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
0
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
1 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 4 TOTAL
  1. TOOL · CL_216901 ·

    Qwen 3.8 27B shows strong safety guardrails in AgentHarm benchmark

    The Qwen 3.8 27B model demonstrated a strong refusal rate of 75% on harmful tasks within the AgentHarm benchmark. This performance indicates robust safety guardrails, contrasting with the Blackfrost package which achiev…

  2. TOOL · CL_169815 ·

    New prompt injection detection techniques leverage cross-domain methods

    Researchers have developed seven novel techniques for detecting prompt injection attacks, moving beyond traditional pattern matching and fine-tuned transformer classifiers. These new methods draw inspiration from divers…

  3. TOOL · CL_74391 ·

    New framework guides LLM agents to revise plans, improving safety

    Researchers have developed TRIAD, a new framework for LLM agents that integrates guardrails to improve safety and utility. Unlike traditional guardrails that simply block unsafe actions, TRIAD provides feedback to guide…

  4. TOOL · CL_32708 ·

    New framework LiSA enhances AI guardrails with sparse failure data

    Researchers have developed LiSA (Lifelong Safety Adaptation), a new framework designed to improve AI guardrails by learning from sparse and noisy failure data. LiSA uses structured memory to generalize from individual i…