PulseAugur
EN
LIVE 02:36:18
ENTITY Jailbreak Attacks

Jailbreak Attacks

PulseAugur coverage of Jailbreak Attacks — every cluster mentioning Jailbreak Attacks across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
7
7 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
7
7 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 7 TOTAL
  1. TOOL · CL_244924 ·

    New SAFEGuard framework detects advanced LLM jailbreak attacks

    Researchers have developed a new framework called SAFEGuard to detect optimization-based jailbreak attacks on large language models. This method combines fluency measurement, using cross-layer distribution distance and …

  2. TOOL · CL_239273 ·

    New AlcaTRAz defense targets LLM jailbreaks at prompt level

    Researchers have developed AlcaTRAz, a novel defense mechanism against jailbreak attacks on large language models. This defense operates at the prompt level, meaning it does not require access to the model's internal we…

  3. RESEARCH · CL_218978 ·

    New LLM safety techniques target neuron-level and jailbreak attacks · 4 sources tracked

    Researchers have developed new methods to enhance the safety alignment of Large Language Models (LLMs) against various attacks. NeuronFuzz utilizes internal safety neurons as continuous feedback for fuzzing, achieving h…

  4. TOOL · CL_135330 ·

    New framework uses computation graphs to diagnose LLM jailbreak vulnerabilities

    Researchers have developed a new framework to understand how large language models (LLMs) are vulnerable to adversarial prompts and jailbreak attacks. This method uses paired internal computation graphs to represent pro…

  5. RESEARCH · CL_117828 ·

    New research explores multimodal and sparse autoencoder methods to combat LLM jailbreaks

    Researchers are developing new methods to combat jailbreaking attacks on spoken language models (SLMs). One approach, JAMA, uses a joint multimodal optimization framework to simultaneously attack both audio and text mod…

  6. TOOL · CL_41189 ·

    New safeguard uses draft models to detect LLM jailbreaks

    Researchers have developed a new safeguard to improve the safety of large language models (LLMs) against jailbreak attacks. This system leverages the transferability of attacks from larger models to smaller "draft" mode…

  7. RESEARCH · CL_15872 ·

    New research tackles LLM jailbreaks with dynamic evaluation and robust defense strategies

    Multiple research papers explore advanced techniques for enhancing the safety and robustness of large language models (LLMs) against jailbreak attacks. These studies introduce novel frameworks and methods for evaluating…