Atlassian has reported significant improvements in its Rovo Dev AI reviewer, which reduced pull request cycle times by up to 45% internally and 32% for external contributions. Separately, Specific Labs has launched Real-SWE, an enterprise-focused code benchmark that tests frontier models against private codebases. Initial findings from Real-SWE indicate that models like Claude Code and Codex CLI do not perform as well as expected on this specialized benchmark. AI
IMPACT AI code review tools show measurable efficiency gains, while new benchmarks highlight the need for specialized evaluation of models on private enterprise code.
RANK_REASON The cluster discusses a product feature improvement (Rovo Dev AI) and the launch of a new benchmark tool (Real-SWE).
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →