AgentDojo
PulseAugur coverage of AgentDojo — every cluster mentioning AgentDojo across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New defense probes detect and mitigate indirect prompt injection in LLMs
Researchers have developed a method to detect indirect prompt injection (IPI) attacks in agentic large language models (LLMs). By training simple linear probes on the models' internal states, they can predict IPI exposu…
-
New PIMiner system automates LLM prompt injection red-teaming
Researchers have developed PIMiner, an agentic system designed for automated prompt injection red-teaming of large language models. Unlike existing methods that often struggle with generalization, PIMiner builds a trans…
-
New prompt injection detection techniques leverage cross-domain methods
Researchers have developed seven novel techniques for detecting prompt injection attacks, moving beyond traditional pattern matching and fine-tuned transformer classifiers. These new methods draw inspiration from divers…
-
New method enhances LLM agent safety by stratifying risk in tool calls
Researchers have developed a new method called role-stratified per-field conformal risk control to enhance the safety of language-model agents. This technique calibrates risk budgets separately for different semantic ro…
-
New RL framework PISmith tests and breaks prompt injection defenses
Researchers have developed PISmith, a novel reinforcement learning (RL) framework designed to rigorously test the effectiveness of prompt injection defenses in large language models (LLMs). The framework trains an attac…
-
New methods compress LLM agent context for improved security and efficiency
Researchers have developed new methods for compressing context in large language model (LLM) agents to improve efficiency and security. One approach, "Twin Agent," separates agents into an "Explore Agent" for untrusted …
-
New pipeline hardens agentic AI apps against data leaks
Researchers have developed a new pre-deployment pipeline designed to prevent data leakage and tool misuse in agentic applications. This pipeline scans, hardens, and validates agentic systems by analyzing prompt template…
-
Prompt optimization may weaken LLM adversarial robustness, new benchmark suggests
A new benchmark has been developed to investigate whether prompt optimization techniques for Large Language Models (LLMs) weaken their robustness against adversarial attacks, specifically prompt injection. Initial findi…
-
LLM attack benchmarks cover less than 25% of threat landscape
Researchers have developed a new framework to audit the coverage of benchmarks designed to test Large Language Model (LLM) attacks. This framework, based on a taxonomy of over 500 inference-time attacks, reveals that cu…
-
New Protocol Enables LLMs to Safely Control Small Devices
Researchers have introduced the Device Context Protocol (DCP), a new architecture designed to enable large language models (LLMs) to safely control constrained devices. DCP is significantly more lightweight than existin…
-
Arc Gate offers solution to OpenAI's 'unfixable' prompt injection vulnerability
OpenAI has stated that prompt injection in browser agents is an unfixable structural vulnerability at the model level. However, a new architectural solution called Arc Gate has demonstrated significant success in mitiga…
-
LLM attack benchmarks show significant gaps in security coverage
Researchers have developed a new framework to audit the coverage of LLM attack benchmarks, revealing significant gaps in current evaluations. Their analysis of six public benchmarks showed they collectively cover less t…
-
New attack exploits LLM agent relays, bypassing alignment defenses
Researchers have identified a new vulnerability in LLM agent architectures that use Bring-Your-Own-Key (BYOK) systems. These architectures route LLM traffic through third-party relays, creating an integrity gap where a …
-
New benchmarks and safety methods emerge for advanced LLM agents
New research explores the development and evaluation of AI agents, focusing on their ability to navigate complex environments and adhere to policies. StarDojo benchmarks agent performance in open-ended simulations like …