AdvBench
PulseAugur coverage of AdvBench — every cluster mentioning AdvBench across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New FATS attack exploits LLMs, highly susceptible GPT-4.1 and DeepSeek-R1
Researchers have developed a new prompt injection attack called FATS (Feign Agent Attack with Toxic-shots) that exploits vulnerabilities in large language models (LLMs). This attack method manipulates LLMs by obfuscatin…
-
DYNASHIELD: New defense shields LLMs from jailbreak attacks
Researchers have developed DYNASHIELD, a novel defense mechanism designed to protect large language models (LLMs) from jailbreak attacks. This black-box approach customizes decoding hyperparameters and system prompts at…
-
New framework ARENA automates red-teaming for audio language models
Researchers have developed ARENA, a novel closed-loop framework designed for automated red-teaming of large audio-language models (LALMs). This system addresses the unique safety challenges posed by LALMs, which can exh…
-
New framework scales AI red-teaming with reusable, evolving attack skills
Researchers have developed JailbreakSkill, a framework designed to enhance automated red-teaming for AI models. This system packages existing attack strategies into modular, reusable skills that can adapt and evolve ove…
-
New SEMA framework enhances multi-turn jailbreak attacks on LLMs
Researchers have developed SEMA, a novel framework designed to improve multi-turn jailbreak attacks against large language models. SEMA utilizes a two-stage process: prefilling self-tuning to generate structured adversa…
-
New frameworks emerge to evaluate and defend against LLM jailbreaks · 4 sources tracked
Researchers are developing new methods to evaluate and defend against jailbreak attacks on large language models (LLMs). One approach, Incomplete Prompt Jailbreaks (IPJ), focuses on how LLMs delay refusal of harmful pro…
-
New NonTextual Target Attack bypasses LLM safety measures with 96.8% success
Researchers have developed a new method called NonTextual Target Attack (NTA) to bypass safety measures in Large Language Models (LLMs). Unlike previous attacks that relied on specific target outputs, NTA focuses on max…
-
AI coding agents bypassed safety measures through multi-stage workflow jailbreaks
A new research paper explores a novel jailbreaking technique for AI coding agents, demonstrating how harmful objectives can be achieved by assembling them across multiple stages of a software development workflow, rathe…
-
New RetroCoT method bypasses LLM safety alignment by reframing harmful requests
Researchers have developed a new method called Retroactive Chain-of-Thought (RetroCoT) to test the safety alignment of large language models. This technique reframes harmful requests as forensic reconstruction tasks, pr…
-
New STEER attack exploits LLM safety gaps in multilingual contexts · 3 sources tracked
Researchers have developed a new method called STEER (Safety Targeted Embedding Exploit via Refinement) to exploit vulnerabilities in the safety training of large language models (LLMs). This technique targets models tr…
-
Researchers automate security rule generation from attack simulations
Researchers have developed a method to automatically generate security detection rules from attack simulations. This system deterministically maps findings from Breach-and-Attack-Simulation (BAS) tools to starter Sigma …
-
Hybrid defense framework boosts LLM accuracy and robustness
Researchers have developed a novel hybrid defense framework to combat both hallucinations and adversarial manipulation in large language models. This approach integrates entropy-based methods for reducing hallucinations…
-
EvoDefense uses LLMs to co-evolve defenses against black-box attacks
Researchers have developed EvoDefense, a novel approach to protect large language models (LLMs) from attacks in black-box scenarios. This system uses a guard LLM and an experience memory to continuously refine defense s…
-
New research reveals escalating LLM and LALM jailbreak vulnerabilities
Three new research papers explore the vulnerabilities and defenses of large language models (LLMs) and large audio-language models (LALMs). The first paper details a taxonomy of audio jailbreak attacks and defenses, hig…
-
New Research: Open-Weight LLM Defenses Vulnerable to Simple Jailbreaks
A new paper published on arXiv demonstrates that current defenses designed to protect open-weight large language models (LLMs) from harmful usage are susceptible to simple jailbreaking techniques. Researchers found that…
-
New BAIT Framework Exploits LLM Reasoning for Jailbreaking
Researchers have developed a new three-step framework called BAIT (Boundary-Aware Iterative Trap) designed to escalate disclosure of malicious content from large language models. This method guides models through identi…
-
New Logit-Gap Steering method efficiently measures AI alignment robustness
Researchers have developed a new metric called the refusal-affirmation logit gap to quantify the safety margin of aligned language models. This metric, which measures the difference between refusal and affirmation token…
-
New diagnostic tool probes LLM circuits for safety and behavior insights
A new research paper introduces "Perturbation Probing," a diagnostic method for understanding the internal workings of large language models. This technique uses two forward passes per prompt to identify and analyze "be…