PulseAugur
EN
LIVE 00:01:32

GLM-5.3 challenges GPT-5.6 Sol and Claude Fable 5 on coding tasks

Together AI has benchmarked its GLM-5.3 model against both OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 on the DeepSWE software engineering tasks. GLM-5.3 demonstrates competitive performance, narrowly trailing GPT-5.6 Sol in single-shot accuracy but surpassing it with multiple attempts and at a significantly lower cost. Against Claude Fable 5, GLM-5.3 matches accuracy while being over five times cheaper, with both models exhibiting similar failure modes and high task agreement. AI

IMPACT GLM-5.3's performance suggests open-weight models are closing the gap with frontier models in specialized tasks like software engineering, offering significant cost advantages.

RANK_REASON Comparison benchmark of multiple LLMs on a specific task.

Read on X — Together (inference / OSS) →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

GLM-5.3 challenges GPT-5.6 Sol and Claude Fable 5 on coding tasks

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Comparison benchmark of multiple LLMs on a specific task.
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
27 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [3]

  1. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    We analyzed GLM-5.3 and Claude Fable 5 on DeepSWE.

    We analyzed GLM-5.3 and Claude Fable 5 on DeepSWE. GLM-5.3 MATCHED Fable 5 on single-shot solve rate at about one-fifth the cost, then pulled ahead with multiple attempts. The economics get interesting fast when retries are cheap 👇

  2. Together AI blog TIER_1 English(EN) ·

    GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

  3. Together AI blog TIER_1 English(EN) ·

    GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.