PulseAugur
EN
LIVE 16:03:40
中文(ZH) 顶流里最快!智谱,你是在「喷」代码吧

Zhipu AI launches GLM-5.1-highspeed API at 400 tokens/s

Zhipu AI has released GLM-5.1-highspeed, a new API for its GLM-5.1 model that achieves an inference speed of 400 tokens per second. This new offering is positioned as the fastest among leading global LLM providers and has demonstrated impressive performance in real-world tests, including rapid code generation and content summarization. The speed enhancement is attributed to significant system engineering optimizations in the inference engine, scheduling system, and underlying infrastructure, aiming to improve the user experience for AI agents by reducing wait times and increasing feedback frequency. AI

IMPACT Accelerates AI agent responsiveness and real-time interaction capabilities across various applications.

RANK_REASON Model release from a frontier lab with a new speed benchmark. [lever_c_demoted from frontier_release: ic=2 ai=1.0]

Read on 量子位 (QbitAI) →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

Zhipu AI launches GLM-5.1-highspeed API at 400 tokens/s

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
Model release from a frontier lab with a new speed benchmark. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
127 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [3]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 十三 ·

    The fastest among the top streams! Zhipu, are you 'spraying' code?

    400 tokens/s

  2. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    Zhipu AI Launches GLM-5.1 High-Speed API: 400 Tokens/s Sets New Global Benchmark

    Zhipu AI has launched GLM-5.1-highspeed, an API variant of its GLM-5.1 model delivering 400 tokens per second — reportedly the fastest inference speed among major global LLM providers.

  3. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Zhipu AI has launched GLM-5.1-highspeed, a high-speed API variant of its GLM-5.1 large language model, delivering 400 tokens per second and reportedly setting a

    Zhipu AI has launched GLM-5.1-highspeed, a high-speed API variant of its GLM-5.1 large language model, delivering 400 tokens per second and reportedly setting a new global benchmark for inference speed among major LLM providers. The API targets enterprise applications requiring r…