A new model, Ox Alpha, has been unveiled as GLM-5.3-Flash, capable of processing 100 trillion tokens per day. Notably, this immense capacity is reportedly served using Chinese chips, achieving hardware efficiency and per-token costs comparable to Nvidia GPUs. This development challenges the dominance of established hardware providers and suggests a shift in the compute landscape for large-scale AI. AI
IMPACT Challenges Nvidia's CUDA moat and suggests a shift in AI compute infrastructure towards Chinese hardware.
RANK_REASON The cluster announces a new model (GLM-5.3-Flash) and its capabilities, with a focus on the underlying infrastructure (Chinese chips) and its competitive implications.
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →