WildguardMix
PulseAugur coverage of WildguardMix — every cluster mentioning WildguardMix across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
AI Safety Guard Models Vulnerable to "Refusal-Cue Shortcut"
Researchers have identified a significant vulnerability in AI safety guard models, termed the "refusal-cue shortcut." This shortcut allows harmful AI responses to be misclassified as safe by simply including a refusal p…
-
New 184M-parameter safety classifier Semalith v1.4 outperforms Llama-Guard-3-8B on prompt injection
Researchers have introduced Semalith v1.4, a new safety classifier designed for large language models. This 184M-parameter model, built on DeBERTa-v3-base, excels at detecting prompt injection attacks and ensuring regul…
-
New D^2-Monitor system enhances safety for diffusion LLMs
Researchers have introduced $D^2$-Monitor, a novel safety monitoring system designed for diffusion large language models (D-LLMs). This system addresses the unique challenges of monitoring D-LLMs, which generate text th…