PulseAugur
EN
LIVE 02:01:18

Atlassian's Rovo Dev AI cuts PR times; new Real-SWE benchmark launched · 3 sources tracked

Atlassian has reported significant improvements in its Rovo Dev AI reviewer, which reduced pull request cycle times by up to 45% internally and 32% for external contributions. Separately, Specific Labs has launched Real-SWE, an enterprise-focused code benchmark that tests frontier models against private codebases. Initial findings from Real-SWE indicate that models like Claude Code and Codex CLI do not perform as well as expected on this specialized benchmark. AI

IMPACT AI code review tools show measurable efficiency gains, while new benchmarks highlight the need for specialized evaluation of models on private enterprise code.

RANK_REASON The cluster discusses a product feature improvement (Rovo Dev AI) and the launch of a new benchmark tool (Real-SWE).

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

Atlassian's Rovo Dev AI cuts PR times; new Real-SWE benchmark launched · 3 sources tracked

How we ranked this

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster discusses a product feature improvement (Rovo Dev AI) and the launch of a new benchmark tool (Real-SWE).
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
product, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [3]

  1. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Atlassian reports its Rovo Dev AI reviewer cut PR cycle time by up to 45% internally and 32% for... # codereview # ai # engineeringmanagement # software # codin

    Atlassian reports its Rovo Dev AI reviewer cut PR cycle time by up to 45% internally and 32% for... # codereview # ai # engineeringmanagement # software # coding # development # engineering # inclusive # community Cutting PR review time is really changing where review happens

  2. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI"

    Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name. Same harness, two different brains: GPT-6 Astra on Codex CLI: 33.8% resolution GPT-5.6 Sol on Codex CL…

  3. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Real-SWE ran frontier models against licensed, private enterprise codebases (billing, tax, multi-service work) and one number jumped out at me: rollout duration

    Real-SWE ran frontier models against licensed, private enterprise codebases (billing, tax, multi-service work) and one number jumped out at me: rollout duration barely moves resolution. 71.4% of rollouts that finished in under 10 minutes FAILED. 73.4% of rollouts that ran 10 minu…