WildJailbreak
PulseAugur coverage of WildJailbreak — every cluster mentioning WildJailbreak across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
BenchMIRT method reveals what LLM benchmarks truly measure · 2 sources tracked
Researchers have introduced BenchMIRT, a novel methodology designed to dissect the performance of large language models (LLMs) on benchmarks by analyzing individual prompts. This approach, inspired by Item Response Theo…
-
LLM safety probes generalize across model families, study finds
A new study reproduced and extended previous research on using latent-space safety probes to detect harmful prompts in Large Language Models. The researchers found that lightweight MLP probes, trained on activations fro…
-
New AI safety method allows models to generate and internalize own guidelines
Researchers have developed a novel method called Self-Guided Adaptive Safety Alignment (SGASA) to enable reasoning models to generate and internalize their own safety guidelines. This approach involves the model creatin…