Benzi, a new coding harness and agent, has demonstrated superior performance compared to Claude Code on the SWE-bench Verified benchmark. Benzi utilizes a novel approach of compiling the entire codebase into a resolved map before answering queries, rather than simply dumping files into a context window. This method allows Benzi to trace calls, data flow, and references with greater accuracy, leading to a higher resolution rate and lower cost per instance, especially when run on models like DeepSeek v4-flash. AI
IMPACT This new coding agent's approach of compiling codebases into a resolved map could set a new standard for AI-assisted software development.
RANK_REASON This is a demonstration of a new coding harness/agent, not a release from a frontier AI lab.
Read on HN — claude cli stories →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →