PulseAugur
中
实时 17:42:05
English(EN) Distilled DeepSeek into Gemma 4 26B-A4B vs 12B. Not very useful, but I learned a lot.

用户将 DeepSeek V4 Pro 蒸馏到 Gemma 26B MoE 和 12B 密集模型

一位用户详细介绍了将 DeepSeek V4 Pro 模型蒸馏成两个 Gemma 版本的过程:一个 26B 参数的 MoE 模型和一个 12B 参数的密集模型。蒸馏过程包括重新填充 Natural Questions QA 对,DeepSeek API 调用成本为 0.36 美元。用户在 Unsloth Studio 中遇到了 bug,但最终成功在一台配备两块 RTX 3090 GPU 的服务器上训练了模型。26B 模型消耗了更多的 VRAM 并实现了更低的训练损失,表明知识吸收更好,尽管两个模型之间的评估差距很小。12B 模型速度更快,每 GPU 吞吐量更高。 AI

影响 展示了模型蒸馏和微调的实用技术,可能为未来的开源模型开发提供信息。

排序理由 用户主导的模型蒸馏和微调研究与实验。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

用户将 DeepSeek V4 Pro 蒸馏到 Gemma 26B MoE 和 12B 密集模型

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
用户主导的模型蒸馏和微调研究与实验。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
91 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Paramecium_caudatum_ ·

    将 DeepSeek 蒸馏到 Gemma 4 26B-A4B 对比 12B。用处不大,但我学到了很多。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1ur1i1a/distilled_deepseek_into_gemma_4_26ba4b_vs_12b_not/"> <img alt="Distilled DeepSeek into Gemma 4 26B-A4B vs 12B. Not very useful, but I learned a lot." src="https://preview.redd.it/irn879iku1ch1.png?widt…