PulseAugur
EN
LIVE 15:45:31

GLM-5.3-flash achieves 206 tok/s on DGX Station GB300

A user on Reddit shared their experience running the GLM-5.3-flash model on a DGX Station GB300, achieving a speed of approximately 206 tokens per second in a single stream. The setup involved using NVFP4 for its compatibility with Blackwell architecture and HBM3e memory. The user provided detailed Docker commands and configuration parameters for replicating the setup, noting a bug in the provided image that requires pre-downloading the model weights. AI

IMPACT Demonstrates performance capabilities of GLM-5.3-flash on high-end hardware, offering insights for users aiming for similar setups.

RANK_REASON User-level report on running a specific model on specific hardware, not an official release or benchmark from the model's creators.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

GLM-5.3-flash achieves 206 tok/s on DGX Station GB300

How we ranked this

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
User-level report on running a specific model on specific hardware, not an official release or benchmark from the model's creators.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/funding__secured ·

    GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context

    <!-- SC_OFF --><div class="md"><p>Hey all!</p> <p>I'm finally doing some cool stuff with my &quot;thinking heater&quot; (h/t <a href="/u/-TV-Stand-">u/-TV-Stand-</a>). I'm still experimenting with GLM-5.2 (in anticipation of 5.3 coming tomorrow, I hope!) and things are very cool …