PulseAugur
EN
LIVE 15:18:33

GLM-5.3-Flash unveiled, serving 100T tokens/day on Chinese chips

A new model, Ox Alpha, has been unveiled as GLM-5.3-Flash, capable of processing 100 trillion tokens per day. Notably, this immense capacity is reportedly served using Chinese chips, achieving hardware efficiency and per-token costs comparable to Nvidia GPUs. This development challenges the dominance of established hardware providers and suggests a shift in the compute landscape for large-scale AI. AI

IMPACT Challenges Nvidia's CUDA moat and suggests a shift in AI compute infrastructure towards Chinese hardware.

RANK_REASON The cluster announces a new model (GLM-5.3-Flash) and its capabilities, with a focus on the underlying infrastructure (Chinese chips) and its competitive implications.

Read on X — SemiAnalysis →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

GLM-5.3-Flash unveiled, serving 100T tokens/day on Chinese chips

How we ranked this

Signal score
47 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Frontier Release
The cluster announces a new model (GLM-5.3-Flash) and its capabilities, with a focus on the underlying infrastructure (Chinese chips) and its competitive implications.
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [3]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    But ALL traffic was served on Chinese chips, attaining hardware efficiency and per-token cost comparable to Nvidia GPUs. The cuda moat is being tested once agai

    But ALL traffic was served on Chinese chips, attaining hardware efficiency and per-token cost comparable to Nvidia GPUs. The cuda moat is being tested once again after Jalapeño's announcement yesterday. (3/3) https://t.co/hZDrwPqbi1

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    100T tokens per day free tokens and people were saying only frontier labs has this amount of compute. (2/3)

    100T tokens per day free tokens and people were saying only frontier labs has this amount of compute. (2/3) https://t.co/7DyX5G8bpx

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Ox Alpha has been unveiled as GLM-5.3-Flash, but what's shocking is that the 100T tokens per day is served on Chinese chip. (1/3)🧵 https://t.co/kjl18yxqRm

    Ox Alpha has been unveiled as GLM-5.3-Flash, but what's shocking is that the 100T tokens per day is served on Chinese chip. (1/3)🧵 https://t.co/kjl18yxqRm