PulseAugur
EN
LIVE 04:24:44

Benzi coding agent outperforms Claude Code on SWE-bench benchmark

Benzi, a new coding harness and agent, has demonstrated superior performance compared to Claude Code on the SWE-bench Verified benchmark. Benzi utilizes a novel approach of compiling the entire codebase into a resolved map before answering queries, rather than simply dumping files into a context window. This method allows Benzi to trace calls, data flow, and references with greater accuracy, leading to a higher resolution rate and lower cost per instance, especially when run on models like DeepSeek v4-flash. AI

IMPACT This new coding agent's approach of compiling codebases into a resolved map could set a new standard for AI-assisted software development.

RANK_REASON This is a demonstration of a new coding harness/agent, not a release from a frontier AI lab.

Read on HN — claude cli stories →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Benzi coding agent outperforms Claude Code on SWE-bench benchmark

COVERAGE [1]

  1. HN — claude cli stories TIER_1 English(EN) · showhz ·

    Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on Sonnet