PulseAugur
EN
LIVE 09:21:18

Databricks benchmark finds open-source coding agents competitive, harness impacts cost

Databricks has conducted an internal benchmark of coding agents using its own diverse codebase, revealing that the open-source landscape now rivals proprietary models in performance. A key finding is that GLM-5.2 demonstrates strong capabilities even on codebases significantly different from typical benchmarks. Furthermore, the study highlights that the 'harness' or framework surrounding the AI model has a substantial impact on cost and performance, with a minimal open-source harness matching proprietary ones at a fraction of the expense by optimizing token usage. AI

IMPACT Suggests that optimizing the surrounding harness, not just the model, is key to cost-effective AI agent deployment.

RANK_REASON Internal benchmark results published by a company on coding agents.

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

Databricks benchmark finds open-source coding agents competitive, harness impacts cost

COVERAGE [4]

  1. Towards AI TIER_1 English(EN) · Samarth Banodia ·

    Databricks Benchmarked Coding Agents on Its Own Codebase. The Results Should Change How You Buy

    <h4><em>Four findings from Matei Zaharia’s team: open source caught up, GLM-5.2 is the real deal outside benchmark-land, a minimal harness matched vendor harnesses at half the cost — and cheaper per-token can mean pricier per-task.</em></h4><figure><img alt="" src="https://cdn-im…

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🤖 Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase The main conclusions from analysis were: The Pareto frontier for coding tasks (i.e. best

    🤖 Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase The main conclusions from analysis were: The Pareto frontier for coding tasks (i.e. best quality for a given cost) includes models from OpenAI, Anthropic, and open source. This means today, only a ... 📰 Sourc…

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Benchmarking coding agents on Databricks' multi-million line codebase https://www. databricks.com/blog/benchmarki ng-coding-agents-databricks-multi-million-line

    Benchmarking coding agents on Databricks' multi-million line codebase https://www. databricks.com/blog/benchmarki ng-coding-agents-databricks-multi-million-line-codebase Comments: https:// news.ycombinator.com/item?id=4 8837696 # HackerNews # benchmarking # coding # agents # Data…

  4. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Which matters more for your coding-agent bill: the model, or the harness around it? Databricks benchmarked both on its own merged pull requests and found the ha

    Which matters more for your coding-agent bill: the model, or the harness around it? Databricks benchmarked both on its own merged pull requests and found the harness swings cost as much as the model. A minimal open-source harness called Pi matched the native ones on quality for u…