Researchers have identified "Harmfulness Propagation Dynamics" (HPD), a phenomenon where the representation of harmful intent in large language models increases with the model's depth. This suggests that harmfulness is a progressively resolved semantic property, with its full intent consolidating later in the model's layers. To address this, a new lightweight input moderator called "Herald" has been developed. Herald extracts features from the cross-layer projection sequence to classify harmful prompts, achieving high accuracy and providing interpretable audit trails of how harmfulness emerges within the model. AI
IMPACT Provides a new method for understanding and mitigating harmful outputs in LLMs, potentially improving model safety and interpretability.
RANK_REASON The cluster contains an academic paper detailing a new research finding and a proposed method for analyzing LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Harmfulness Propagation Dynamics
- Herald
- Hugging Face
- multilayer perceptron
- Noor S. Mohammad
- OLMo2-7B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →