AgentHarm
PulseAugur coverage of AgentHarm — every cluster mentioning AgentHarm across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New prompt injection detection techniques leverage cross-domain methods
Researchers have developed seven novel techniques for detecting prompt injection attacks, moving beyond traditional pattern matching and fine-tuned transformer classifiers. These new methods draw inspiration from divers…
-
New framework guides LLM agents to revise plans, improving safety
Researchers have developed TRIAD, a new framework for LLM agents that integrates guardrails to improve safety and utility. Unlike traditional guardrails that simply block unsafe actions, TRIAD provides feedback to guide…
-
New framework LiSA enhances AI guardrails with sparse failure data
Researchers have developed LiSA (Lifelong Safety Adaptation), a new framework designed to improve AI guardrails by learning from sparse and noisy failure data. LiSA uses structured memory to generalize from individual i…