ENTITY
SWEbench Pro
SWEbench Pro
PulseAugur coverage of SWEbench Pro — every cluster mentioning SWEbench Pro across labs, papers, and developer communities, ranked by signal.
Total · 30d
0
1 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
0 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
-
Kimi K3 self-hosting costs 20% more but boosts task resolution by 24%
Self-hosting the Kimi K3 model requires approximately 20% more hardware cost compared to GLM-5.2, utilizing an 8xB300 node instead of an 8xB200 node. While Kimi K3 exhibits lower token throughput and longer task resolut…
-
DeepSWE benchmark shows GPT-5.5 outperforming Claude Opus
A new benchmark called DeepSWE, designed to more realistically assess AI coding capabilities, has revealed that GPT-5.5 outperforms Anthropic's Claude Opus. The DeepSWE benchmark is noted for its contamination-free task…