InjecAgent
PulseAugur coverage of InjecAgent — every cluster mentioning InjecAgent across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New prompt injection detection techniques leverage cross-domain methods
Researchers have developed seven novel techniques for detecting prompt injection attacks, moving beyond traditional pattern matching and fine-tuned transformer classifiers. These new methods draw inspiration from divers…
-
New method enhances LLM agent safety by stratifying risk in tool calls
Researchers have developed a new method called role-stratified per-field conformal risk control to enhance the safety of language-model agents. This technique calibrates risk budgets separately for different semantic ro…
-
New RL framework PISmith tests and breaks prompt injection defenses
Researchers have developed PISmith, a novel reinforcement learning (RL) framework designed to rigorously test the effectiveness of prompt injection defenses in large language models (LLMs). The framework trains an attac…
-
Prompt optimization may weaken LLM adversarial robustness, new benchmark suggests
A new benchmark has been developed to investigate whether prompt optimization techniques for Large Language Models (LLMs) weaken their robustness against adversarial attacks, specifically prompt injection. Initial findi…
-
LLM attack benchmarks cover less than 25% of threat landscape
Researchers have developed a new framework to audit the coverage of benchmarks designed to test Large Language Model (LLM) attacks. This framework, based on a taxonomy of over 500 inference-time attacks, reveals that cu…
-
Arc Gate offers solution to OpenAI's 'unfixable' prompt injection vulnerability
OpenAI has stated that prompt injection in browser agents is an unfixable structural vulnerability at the model level. However, a new architectural solution called Arc Gate has demonstrated significant success in mitiga…
-
LLM attack benchmarks show significant gaps in security coverage
Researchers have developed a new framework to audit the coverage of LLM attack benchmarks, revealing significant gaps in current evaluations. Their analysis of six public benchmarks showed they collectively cover less t…