PulseAugur
EN
LIVE 06:33:37
ENTITY Harmful Detection Heads

Harmful Detection Heads

PulseAugur coverage of Harmful Detection Heads — every cluster mentioning Harmful Detection Heads across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
1 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
1 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 1 TOTAL
  1. TOOL · CL_231337 ·

    New research reveals LLM safety circuit, improving refusal rates

    Researchers have identified a multi-stage safety circuit within Large Language Models (LLMs) that governs their ability to refuse harmful content. This circuit comprises Harmful Detection Heads, Safety Neurons, and Refu…