XSTest
PulseAugur coverage of XSTest — every cluster mentioning XSTest across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
PL-Guard architecture separates LLM grounding from reasoning for enhanced safety
Researchers have introduced PL-Guard, a novel neurosymbolic architecture designed to enhance the safety of large language models (LLMs) by separating semantic grounding from policy reasoning. This approach uses a symbol…
-
LLM Safety Benchmarks Underestimate Risks Due to Prompt Sensitivity
A new research paper highlights that standard benchmarks may underestimate the safety risks of large language models (LLMs) by relying on single, canonical prompts. The study found that varying the surface form of promp…
-
New AI Safety Research: Activation Probes Detect Harmful Requests
A new research paper titled "The Entanglement Wall" proposes using activation-space probes as a method to detect potentially harmful AI requests. These probes demonstrated a high success rate in blocking compliant attac…
-
New J-Space Protocol Assesses AI Model Safety Internally
Researchers have introduced JADR, a new protocol for evaluating the internal safety mechanisms of AI models. This method analyzes a model's Jacobian space (J-space) before response generation, offering a more direct ass…
-
New OS Kernel Primitive Enhances LLM Safety Checks
A new kernel-level operation called ProbeLogits has been developed for AI-native operating systems, allowing them to directly read an LLM's logit distribution before token generation. This primitive enables the OS to cl…