PulseAugur
实时 09:44:37
English(EN) Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

研究发现Qwen3大型语言模型的时间偏好决策可被引导

研究人员识别并操纵了Qwen3-32B大型语言模型中的时间偏好。通过训练对比线性探针,他们发现了模型残差流中代表短期与长期决策的方向。使用这些方向进行对比激活加法引导,显著改变了模型在时间选择任务中的选择,包括货币跨期选择和规划能力基准测试。这项工作表明,大型语言模型的跨期偏好是可衡量的和可引导的,这对提供延迟成本和收益建议的AI系统以及关于长期规划的AI安全具有启示意义。 AI

影响 证明了大型语言模型跨期偏好的可引导性,影响了提供延迟成本/收益建议的AI系统以及关于长期规划的AI安全。

排序理由 学术论文,详细介绍了一种引导大型语言模型行为的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现Qwen3大型语言模型的时间偏好决策可被引导

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Michal Mr\'az, Justin Shenk ·

    Qwen3 通过对比激活加法实现跨期偏好引导

    arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-…