PulseAugur
中
实时 06:23:59
English(EN) ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

开源ZGCM-1模型在数学和智能体搜索方面实现高效率

研究人员推出ZGCM-1,一个为数学推理和智能体搜索设计的7B参数基础模型。该模型利用了一种高效的训练方法,结合了交错注意力等架构创新、稳定的FP8 Muon优化器以及渐进式课程学习。ZGCM-1在特定任务上展现出与更大规模前沿模型相媲美的性能,并在训练时间上提供了显著的效率提升。 AI

影响 这个开源模型的发布可能会加速在高效LLM训练和智能体系统方面的研究,从而可能降低开发先进AI能力的门槛。

排序理由 该集群描述了一个新发布的基础模型,包含其权重、训练代码和数据,符合OSS模型发布的科研类别。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

开源ZGCM-1模型在数学和智能体搜索方面实现高效率

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一个新发布的基础模型,包含其权重、训练代码和数据,符合OSS模型发布的科研类别。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
27 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Y… ·

    ZGCM-1:一个完全开放且极其高效的数学与智能体搜索基础模型

    arXiv:2609.13356v1 Announce Type: new Abstract: In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the op…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    ZGCM-1:一个完全开放且极其高效的数学与智能体搜索基础模型

    ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency.