PulseAugur
EN
LIVE 17:56:41
ENTITY WildJailbreak

WildJailbreak

PulseAugur coverage of WildJailbreak — every cluster mentioning WildJailbreak across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
3 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
3 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 3 TOTAL
  1. RESEARCH · CL_230863 ·

    BenchMIRT method reveals what LLM benchmarks truly measure · 2 sources tracked

    Researchers have introduced BenchMIRT, a novel methodology designed to dissect the performance of large language models (LLMs) on benchmarks by analyzing individual prompts. This approach, inspired by Item Response Theo…

  2. TOOL · CL_193471 ·

    LLM safety probes generalize across model families, study finds

    A new study reproduced and extended previous research on using latent-space safety probes to detect harmful prompts in Large Language Models. The researchers found that lightweight MLP probes, trained on activations fro…

  3. TOOL · CL_183233 ·

    New AI safety method allows models to generate and internalize own guidelines

    Researchers have developed a novel method called Self-Guided Adaptive Safety Alignment (SGASA) to enable reasoning models to generate and internalize their own safety guidelines. This approach involves the model creatin…