SWE-Marathon
PulseAugur coverage of SWE-Marathon — every cluster mentioning SWE-Marathon across labs, papers, and developer communities, ranked by signal.
-
xAI unveils Grok 4.5, targeting programmers and office tasks
xAI has launched Grok 4.5, its most advanced model to date, specifically designed for programming, agentic tasks, and knowledge work. Trained with Cursor on NVIDIA GB300 GPUs, Grok 4.5 shows strong performance in coding…
-
AI models struggle with marathon tasks, revealing benchmark limitations · 1 source tracked
New benchmarks reveal a significant gap between AI models' performance on short, single-session tasks and their ability to handle long, multi-hour operations. While models like GLM-5.2 and GPT-5.5 excel on benchmarks li…
-
Zhipu AI releases GLM-5.2 with 1M context window, challenging top proprietary models
Zhipu AI has released GLM-5.2, a 744B-parameter Mixture-of-Experts model featuring a 1 million token context window and MIT-licensed weights. This model achieves a high ranking on the BenchLM leaderboard and demonstrate…
-
Z.ai releases GLM-5.2, setting new open-source benchmark for long-context AI
Z.ai has released GLM-5.2, an open-source language model with a 1 million token context window, positioning it as a strong contender in long-horizon tasks and coding benchmarks. The model features an improved architectu…