PulseAugur
EN
LIVE 18:11:35
ENTITY XSTest

XSTest

PulseAugur coverage of XSTest — every cluster mentioning XSTest across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
8
8 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
8
8 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 8 TOTAL
  1. RESEARCH · CL_239380 ·

    New LLM safety method refines refusal boundaries, reduces over-refusal

    Researchers have developed a new framework called Boundary-Aware Self-Distillation to improve the safety and usability of large language models. This method focuses on creating narrow refusal boundaries, allowing models…

  2. RESEARCH · CL_230863 ·

    BenchMIRT method reveals what LLM benchmarks truly measure · 2 sources tracked

    Researchers have introduced BenchMIRT, a novel methodology designed to dissect the performance of large language models (LLMs) on benchmarks by analyzing individual prompts. This approach, inspired by Item Response Theo…

  3. TOOL · CL_226812 ·

    New LMSM Framework Enhances LLM Security with Modular Design

    Researchers have introduced Language Model Security Modules (LMSM), a novel framework designed to enhance the security of large language model (LLM) deployments. Inspired by Linux Security Modules, LMSM separates the pr…

  4. TOOL · CL_206119 ·

    PL-Guard architecture separates LLM grounding from reasoning for enhanced safety

    Researchers have introduced PL-Guard, a novel neurosymbolic architecture designed to enhance the safety of large language models (LLMs) by separating semantic grounding from policy reasoning. This approach uses a symbol…

  5. TOOL · CL_183121 ·

    LLM Safety Benchmarks Underestimate Risks Due to Prompt Sensitivity

    A new research paper highlights that standard benchmarks may underestimate the safety risks of large language models (LLMs) by relying on single, canonical prompts. The study found that varying the surface form of promp…

  6. TOOL · CL_145799 ·

    New AI Safety Research: Activation Probes Detect Harmful Requests

    A new research paper titled "The Entanglement Wall" proposes using activation-space probes as a method to detect potentially harmful AI requests. These probes demonstrated a high success rate in blocking compliant attac…

  7. RESEARCH · CL_143640 ·

    New J-Space Protocol Assesses AI Model Safety Internally

    Researchers have introduced JADR, a new protocol for evaluating the internal safety mechanisms of AI models. This method analyzes a model's Jacobian space (J-space) before response generation, offering a more direct ass…

  8. TOOL · CL_129350 ·

    New OS Kernel Primitive Enhances LLM Safety Checks

    A new kernel-level operation called ProbeLogits has been developed for AI-native operating systems, allowing them to directly read an LLM's logit distribution before token generation. This primitive enables the OS to cl…