PulseAugur
实时 10:35:19

Google 的 Gemma-4-E4B-it 模型内部科学表征可读且可控

研究人员开发了读取和控制开源语言模型 google/gemma-4-E4B-it 中材料科学机制内部表征的方法。该研究表明,概念在单个隐藏状态中是可辨识的,构成方向通过状态转换传达,并且特定的内部表征可以对工程答案产生因果影响。通过使用雅可比读出和因果干预,该团队能够识别机制家族,甚至操纵模型的输出来符合物理定律。 AI

影响 这项研究为理解和控制科学领域中 LLM 的行为提供了新技术,有望提高可靠性和可解释性。

排序理由 该集群包含一篇学术论文,详细介绍了对开源语言模型内部工作原理的新研究。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Google 的 Gemma-4-E4B-it 模型内部科学表征可读且可控

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Markus J. Buehler ·

    开放权重语言模型中材料科学机制的表征阅读与引导

    arXiv:2607.20058v1 Announce Type: new Abstract: Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight goo…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Open-Weight语言模型中材料科学机制的读取与引导表示

    Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentall…