gpt-oss-safeguard-20b
PulseAugur coverage of gpt-oss-safeguard-20b — every cluster mentioning gpt-oss-safeguard-20b across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
LLM safety judges vulnerable to content-invariant wrappers, study finds
Researchers have discovered that automatic safety judges for large language models can be easily manipulated by altering the tone or framing of a response without changing its content. By adding "content-invariant style…
-
Mistral AI releases Shieldstral 1.0 3B, an adaptable open-weights safety classifier
Mistral AI has launched Shieldstral 1.0 3B, an open-weights safety classifier designed for policy adaptability. Unlike traditional models that rely on fixed harm categories, Shieldstral uses natural language questions t…
-
Mistral AI unveils policy-adaptive multimodal safety model Shieldstral
Mistral AI has released Shieldstral 1.0 3B, a new multimodal safety classifier designed for efficient content moderation. Unlike traditional models that predict fixed categories, Shieldstral adapts to natural language s…
-
AI agents discover advanced LLM attack methods, revealing non-monotonic safety gains
AI agents are capable of discovering novel adversarial attack algorithms that outperform existing methods against large language models. One study demonstrated that these AI-discovered attacks achieved up to 80% success…
-
New framework identifies elderly-specific risks in AI chatbots
Researchers have developed GrandGuard, a new framework to address safety concerns specific to elderly users interacting with AI chatbots. The framework includes a taxonomy of 50 risk types across mental well-being, fina…