Researchers have introduced PersonaUnlearnBench, a new benchmark designed to evaluate the effectiveness of unlearning specific personas from large language models (LLMs). The benchmark, which covers six LLMs and five distinct personas, reveals that current unlearning methods struggle to reliably remove target personas without negatively impacting general utility. To address this, the paper proposes PaCE, a novel method that identifies an internal behavior direction and trains the model to suppress the target persona while preserving desirable responses and overall functionality. AI
IMPACT This research could lead to more controllable and safer LLMs by enabling the removal of undesirable behaviors.
RANK_REASON The cluster contains a research paper detailing a new benchmark and method for LLM persona unlearning. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- LLM Persona Unlearning
- LLMs
- PaCE
- PersonaUnlearnBench
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →