PulseAugur
实时 09:10:40
English(EN) Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief Systems

大型语言模型(LLM)代理的动机可预测,但信念系统仍不透明

研究人员使用 Llama-3.1-8B 代理进行了一项大规模实验,以了解从其行为中可以推断出代理动机和信念系统的程度。研究发现存在显著的不对称性,动机高度可预测(准确率 98-100%),而信念系统则难以辨别,即使是先进的 Transformer 模型也只能达到 34.0% 的准确率。这种推断信念系统(尤其是中性对齐)的困难,表明仅凭可观察到的行为来理解 LLM 代理的价值观存在局限性。 AI

影响 凸显了推断 LLM 代理价值观的局限性,影响了对齐研究和代理可解释性。

排序理由 学术论文,详细介绍了对 LLM 代理行为进行的受控实验。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

大型语言模型(LLM)代理的动机可预测,但信念系统仍不透明

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Jason Starace, Terence Soule ·

    大规模行为推断:动机与信念系统之间的根本不对称性

    arXiv:2509.05624v3 Announce Type: replace-cross Abstract: How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach that infers agent properties from action sequences, yet remains empirically open…