PulseAugur
EN
LIVE 22:04:32

GLM-5.3 challenges GPT-5.6 Sol and Claude Fable 5 on coding tasks

Together AI has benchmarked its GLM-5.3 model against both OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 on the DeepSWE software engineering tasks. GLM-5.3 demonstrates competitive performance, narrowly trailing GPT-5.6 Sol in single-shot accuracy but surpassing it with multiple attempts and at a significantly lower cost. Against Claude Fable 5, GLM-5.3 matches accuracy while being over five times cheaper, with both models exhibiting similar failure modes and high task agreement. AI

IMPACT GLM-5.3's performance suggests open-weight models are closing the gap with frontier models in specialized tasks like software engineering, offering significant cost advantages.

RANK_REASON Comparison benchmark of multiple LLMs on a specific task.

Read on Together AI blog →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

GLM-5.3 challenges GPT-5.6 Sol and Claude Fable 5 on coding tasks

COVERAGE [2]

  1. Together AI blog TIER_1 English(EN) ·

    GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

  2. Together AI blog TIER_1 English(EN) ·

    GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.