Researchers have developed the EPITOME framework to analyze supportive empathy in large language models (LLMs) by decomposing it into emotional reactions, interpretations, and explorations. Their study on three instruction-tuned LLMs revealed that while contrastive activation addition can steer empathy scores, the recovered directions are not fully separable, with interventions causing off-target shifts. Furthermore, persona prompts significantly alter empathy scores, but the activation shifts captured only a small fraction of the persona-induced changes, indicating that controlling persona-conditioned empathy requires targeting structures beyond individual mechanism directions. AI
IMPACT Provides a new method for analyzing and potentially controlling nuanced aspects of LLM behavior like empathy.
RANK_REASON Academic paper detailing a new framework and findings on LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →