Databricks has conducted an internal benchmark of coding agents using its own diverse codebase, revealing that the open-source landscape now rivals proprietary models in performance. A key finding is that GLM-5.2 demonstrates strong capabilities even on codebases significantly different from typical benchmarks. Furthermore, the study highlights that the 'harness' or framework surrounding the AI model has a substantial impact on cost and performance, with a minimal open-source harness matching proprietary ones at a fraction of the expense by optimizing token usage. AI
IMPACT Suggests that optimizing the surrounding harness, not just the model, is key to cost-effective AI agent deployment.
RANK_REASON Internal benchmark results published by a company on coding agents.
Read on Mastodon — mastodon.social →
- Databricks
- Mastodon
- Anthropic
- OpenAI
- GLM-5.2
- Java
- Jsonnet
- Matei Zaharia
- Protocol Buffers
- Rust
- Scala
- SWE-bench
- Terminal-Bench
- TypeScript
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →