PulseAugur
实时 11:49:10
English(EN) Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

新PRISM方法将LLM组件映射到人类大脑功能

研究人员开发了一种名为PRISM(基于扰动的区域可解释性通过减法映射)的新方法来分析大型语言模型的内部工作机制。该技术借鉴了人类神经影像学的方法,以识别Transformer模型中的专业化组件。通过比较扰动后的LLaVA-1.6-Vicuna-13B模型的错误模式与中风后失语症患者的病灶模式,PRISM旨在为检验LLM的功能专业化主张提供一种可证伪的方式。 AI

影响 通过与人类认知研究建立联系,为理解LLM内部机制提供了一种新颖的方法。

排序理由 该条目是一篇学术论文,详细介绍了一种新的LLM可解释性方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新PRISM方法将LLM组件映射到人类大脑功能

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Xiang Guan, Roger D. Newman-Norlund, Yong Yang, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Srihari Nelakuditi, Chris Rorden, Leonardo Bonilha, Julius Fridriksson ·

    基于扰动的区域可解释性通过减法映射(PRISM):语言模型中的命名错误分离与中风后失语症

    arXiv:2608.12717v1 Announce Type: cross Abstract: Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. We adapt subtraction analysis, the standa…