ToxicChat
PulseAugur coverage of ToxicChat — every cluster mentioning ToxicChat across labs, papers, and developer communities, ranked by signal.
-
Mistral releases self-hostable 3B moderation model Shieldstral 1.0
Mistral has released Shieldstral 1.0, a 3-billion-parameter model designed for self-hosted text and image content moderation. Available under an Apache 2.0 license with full weights, the model can run on a single 16GB G…
-
New AI moderation methods balance usefulness and safety
Researchers have developed a new method for evaluating content moderation in AI systems, focusing on end-to-end trade-offs rather than isolated classifier accuracy. The study introduces two key metrics: Usefulness, whic…
-
New 184M-parameter safety classifier Semalith v1.4 outperforms Llama-Guard-3-8B on prompt injection
Researchers have introduced Semalith v1.4, a new safety classifier designed for large language models. This 184M-parameter model, built on DeBERTa-v3-base, excels at detecting prompt injection attacks and ensuring regul…
-
New OS Kernel Primitive Enhances LLM Safety Checks
A new kernel-level operation called ProbeLogits has been developed for AI-native operating systems, allowing them to directly read an LLM's logit distribution before token generation. This primitive enables the OS to cl…
-
Gnosys improves AI classifiers with sparse labels using autonomous engineering
Gnosys, an autonomous model engineer, has developed a method to improve AI classifiers when labeled data is scarce. Their approach, tested on the ToxicChat safety benchmark, demonstrated an improvement in harm detection…
-
New D^2-Monitor system enhances safety for diffusion LLMs
Researchers have introduced $D^2$-Monitor, a novel safety monitoring system designed for diffusion large language models (D-LLMs). This system addresses the unique challenges of monitoring D-LLMs, which generate text th…