PulseAugur
实时 07:10:54
Nederlands(NL) Interpretable GOHR Agents via Sparse Autoencoders

稀疏自编码器揭示GOHR游戏中可解释的AI智能体

研究人员为在隐藏规则游戏(GOHR)中运行的Transformer智能体开发了可解释的实验。通过在智能体的决策令牌嵌入上训练稀疏自编码器,他们能够恢复游戏的底层结构。这些自编码器维度被证明对诸如选择的形状或存储桶等特定概念具有选择性,并且还对应于诸如假设探测和负反馈后切换等可解释的策略。 AI

影响 为理解AI智能体内部决策过程提供了新方法。

排序理由 该集群包含一篇详细介绍AI智能体可解释性实验的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

稀疏自编码器揭示GOHR游戏中可解释的AI智能体

报道来源 [1]

  1. arXiv cs.LG TIER_1 Nederlands(NL) · Shiwei Tan, Yusong Zhao, Weiyi Qin, Wentian Wang, Jacob Feldman, Lazaros K. Gallos, Paul B. Kantor, Vladimir Menkov, Hao Wang ·

    通过稀疏自编码器实现可解释的 GOHR 代理

    arXiv:2607.25132v2 Announce Type: replace Abstract: A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help explain their behavior. We report interpretability experiments for a tokenized…