PulseAugur
实时 10:14:26

新框架揭示LLM中的共情是可控但多轴的

研究人员开发了EPITOME框架,通过将支持性共情分解为情感反应、解释和探索,来分析大型语言模型(LLMs)中的支持性共情。他们对三个指令微调的LLM进行的研究表明,虽然对比激活加法可以引导共情得分,但恢复的方向并非完全可分离,干预会导致非目标偏移。此外,角色提示会显著改变共情得分,但捕获的激活偏移仅占角色提示引起的变化的一小部分,这表明控制角色条件共情需要针对个体机制方向之外的结构。 AI

影响 提供了一种分析和潜在控制LLM行为细微方面(如共情)的新方法。

排序理由 学术论文,详细介绍了LLM行为的新框架和发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新框架揭示LLM中的共情是可控但多轴的

本文如何被排名

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了LLM行为的新框架和发现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · JuHeon Ha, Byounghan Lee, Yunseo Choi, Kyung-Ah Sohn ·

    Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs

    arXiv:2609.15654v1 Announce Type: new Abstract: Activation steering has been used to control traits such as honesty, refusal, and sycophancy, yet supportive empathy is evaluated along multiple dimensions that need not correspond to independently controllable activation directions…