PulseAugur
EN
LIVE 17:04:18

Developer runs 744B GLM-5.2 model on dual Tenstorrent cards

A developer at Tenstorrent successfully ran the 744 billion parameter GLM-5.2 model on a dual-card setup, achieving a performance of 0.35 tokens per second. This was accomplished by adapting the Colibri project's approach, which optimizes memory usage by streaming expert sub-networks as needed. The developer meticulously verified each component of the model implementation against reference versions, ensuring high accuracy before focusing on performance. AI

IMPACT Shows potential for running large models on more accessible hardware configurations.

RANK_REASON Demonstration of running a large model on specific hardware, adapting an existing project. [lever_c_demoted from research: ic=1 ai=1.0]

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Developer runs 744B GLM-5.2 model on dual Tenstorrent cards

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Eric Zietlow ·

    Squeezing a 744B Model Onto Two Tenstorrent Cards... Kinda

    <p>I did a mad science thing recently and I need to tell you about it. This one didn't touch the whole home lab, just one small piece of it: the master node from the swarm cluster I described back in my first post (the box I called the swarm host there), running just two p150a ca…