PulseAugur
实时 17:04:43
English(EN) Squeezing a 744B Model Onto Two Tenstorrent Cards... Kinda

开发者在双 Tenstorrent 卡上运行了 744B GLM-5.2 模型

Tenstorrent 的一位开发者成功地在双卡设置上运行了拥有 7440 亿参数的 GLM-5.2 模型,实现了每秒 0.35 个 token 的性能。这是通过借鉴 Colibri 项目的方法实现的,该项目通过按需流式传输专家子网络来优化内存使用。开发者在关注性能之前,仔细地将模型实现的每个组件与参考版本进行了核对,以确保高准确性。 AI

影响 展示了在更易于获得的硬件配置上运行大型模型的潜力。

排序理由 在特定硬件上运行大型模型的演示,改编自现有项目。[lever_c_demoted from research: ic=1 ai=1.0]

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

开发者在双 Tenstorrent 卡上运行了 744B GLM-5.2 模型

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Eric Zietlow ·

    将一个744B模型塞进两张Tenstorrent卡里……勉强算吧

    <p>I did a mad science thing recently and I need to tell you about it. This one didn't touch the whole home lab, just one small piece of it: the master node from the swarm cluster I described back in my first post (the box I called the swarm host there), running just two p150a ca…