AgentDojo
PulseAugur coverage of AgentDojo — every cluster mentioning AgentDojo across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New multi-agent AI security protocol combats hidden prompt injection attacks
Researchers are developing new methods to secure multi-agent AI systems against indirect prompt injection attacks, where malicious instructions are hidden within data processed by the agents. A new protocol called multi…
-
LLM safety evaluations flawed by distribution shift, research finds
A new research paper highlights significant flaws in how Large Language Model (LLM) safety routing evaluations are conducted. The study, "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift," dem…
-
New LLM frameworks tackle privacy-utility trade-offs in multi-user systems
Researchers have developed new frameworks and methods to address privacy concerns in large language model (LLM) interactions. One approach, AIM, introduces a privacy-aware memory system for multi-agent, multi-user LLM e…
-
New MOAE method optimizes LLM agents across multiple objectives
Researchers have introduced Multi-Objective Agent Evolution (MOAE), a novel approach to optimizing LLM-based agents across multiple criteria simultaneously. Unlike methods that collapse diverse metrics into a single sco…
-
New FGLGuard system enhances LLM multi-agent safety via federated graph learning
Researchers have developed FGLGuard, a novel system for enhancing the safety of multi-agent systems (MAS) powered by large language models (LLMs). This approach utilizes federated graph learning to train a graph neural …
-
New covert prompt injection attack targets LLM agents using tools
Researchers have identified a new threat to LLM agents that use tools, known as covert indirect prompt injection (ICoA). This attack allows malicious prompts to be executed without the user noticing, unlike overt inject…
-
Qwen 3.8 27B model vulnerable to prompt injection in AgentDojo tests
A recent test of the Qwen 3.8 27B model within the AgentDojo framework revealed vulnerabilities to prompt injection attacks. Approximately 11.7% of these attacks resulted in unauthorized actions, highlighting a structur…
-
StepGuard system enhances AI agent safety with step-level control
Researchers have developed StepGuard, a novel system designed to monitor and control the actions of AI agents at a step-by-step level, rather than just evaluating completed trajectories. This approach aims to prevent se…
-
New frameworks LongGuard and StepGuard enhance LLM safety guardrails
Researchers have developed two new frameworks, LongGuard and StepGuard, to address safety failures in large language models (LLMs). LongGuard focuses on analyzing and mitigating failures in long-context guardrails, prop…
-
New defense probes detect and mitigate indirect prompt injection in LLMs
Researchers have developed a method to detect indirect prompt injection (IPI) attacks in agentic large language models (LLMs). By training simple linear probes on the models' internal states, they can predict IPI exposu…
-
New PIMiner system automates LLM prompt injection red-teaming
Researchers have developed PIMiner, an agentic system designed for automated prompt injection red-teaming of large language models. Unlike existing methods that often struggle with generalization, PIMiner builds a trans…
-
New prompt injection detection techniques leverage cross-domain methods
Researchers have developed seven novel techniques for detecting prompt injection attacks, moving beyond traditional pattern matching and fine-tuned transformer classifiers. These new methods draw inspiration from divers…
-
New method enhances LLM agent safety by stratifying risk in tool calls
Researchers have developed a new method called role-stratified per-field conformal risk control to enhance the safety of language-model agents. This technique calibrates risk budgets separately for different semantic ro…
-
New RL framework PISmith tests and breaks prompt injection defenses
Researchers have developed PISmith, a novel reinforcement learning (RL) framework designed to rigorously test the effectiveness of prompt injection defenses in large language models (LLMs). The framework trains an attac…
-
New methods compress LLM agent context for improved security and efficiency
Researchers have developed new methods for compressing context in large language model (LLM) agents to improve efficiency and security. One approach, "Twin Agent," separates agents into an "Explore Agent" for untrusted …
-
New pipeline hardens agentic AI apps against data leaks
Researchers have developed a new pre-deployment pipeline designed to prevent data leakage and tool misuse in agentic applications. This pipeline scans, hardens, and validates agentic systems by analyzing prompt template…
-
Prompt optimization may weaken LLM adversarial robustness, new benchmark suggests
A new benchmark has been developed to investigate whether prompt optimization techniques for Large Language Models (LLMs) weaken their robustness against adversarial attacks, specifically prompt injection. Initial findi…
-
LLM attack benchmarks cover less than 25% of threat landscape
Researchers have developed a new framework to audit the coverage of benchmarks designed to test Large Language Model (LLM) attacks. This framework, based on a taxonomy of over 500 inference-time attacks, reveals that cu…
-
New Protocol Enables LLMs to Safely Control Small Devices
Researchers have introduced the Device Context Protocol (DCP), a new architecture designed to enable large language models (LLMs) to safely control constrained devices. DCP is significantly more lightweight than existin…
-
Arc Gate offers solution to OpenAI's 'unfixable' prompt injection vulnerability
OpenAI has stated that prompt injection in browser agents is an unfixable structural vulnerability at the model level. However, a new architectural solution called Arc Gate has demonstrated significant success in mitiga…