ToxicChat
PulseAugur coverage of ToxicChat — every cluster mentioning ToxicChat across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New AI moderation methods balance usefulness and safety
Researchers have developed a new method for evaluating content moderation in AI systems, focusing on end-to-end trade-offs rather than isolated classifier accuracy. The study introduces two key metrics: Usefulness, whic…
-
New 184M-parameter safety classifier Semalith v1.4 outperforms Llama-Guard-3-8B on prompt injection
Researchers have introduced Semalith v1.4, a new safety classifier designed for large language models. This 184M-parameter model, built on DeBERTa-v3-base, excels at detecting prompt injection attacks and ensuring regul…
-
New OS Kernel Primitive Enhances LLM Safety Checks
A new kernel-level operation called ProbeLogits has been developed for AI-native operating systems, allowing them to directly read an LLM's logit distribution before token generation. This primitive enables the OS to cl…
-
Gnosys improves AI classifiers with sparse labels using autonomous engineering
Gnosys, an autonomous model engineer, has developed a method to improve AI classifiers when labeled data is scarce. Their approach, tested on the ToxicChat safety benchmark, demonstrated an improvement in harm detection…
-
New D^2-Monitor system enhances safety for diffusion LLMs
Researchers have introduced $D^2$-Monitor, a novel safety monitoring system designed for diffusion large language models (D-LLMs). This system addresses the unique challenges of monitoring D-LLMs, which generate text th…