Researchers have developed a new method called TRACE (Trajectory Return Attribution and Contrastive Erasure) to improve the safety of large language models (LLMs). This technique addresses the issue where LLMs can be tricked into generating harmful content by spreading the request across multiple turns, even if they refuse single-turn harmful prompts. TRACE works by assigning a token-level objective that weights tokens based on the discounted return of a refusal-attributable advantage, allowing earlier tokens to be credited for later refusal evidence. The method has demonstrated a significant reduction in attack success rates across various models and attacks, with minimal impact on model utility. AI
IMPACT Enhances LLM safety by providing a more robust defense against multi-turn harmful prompts, potentially leading to safer AI deployments.
RANK_REASON Academic paper detailing a new method for LLM safety. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- HellaSwag
- Hugging Face
- large language models
- Massive Multitask Language Understanding
- ScienceCast
- TRACE
- Trajectory Return Attribution and Contrastive Erasure
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →