Codeforces
PulseAugur coverage of Codeforces — every cluster mentioning Codeforces across labs, papers, and developer communities, ranked by signal.
-
SpecCoder framework enhances Code LLMs with formal specifications
Researchers have developed SpecCoder, a new framework designed to enhance the reasoning capabilities of Code LLMs by utilizing intermediate formal specifications. Unlike natural language, these executable specifications…
-
Research probes how language agents effectively use feedback for improvement
A new research paper investigates the effectiveness of feedback in improving language agent performance. The study introduces a controlled student-teacher protocol across multiple benchmarks, comparing external feedback…
-
AI benchmark scores predictable from just two factors, study finds
A new research paper proposes a method called BenchPress that can predict a frontier model's performance across numerous benchmarks using only two key scores. The study analyzed 84 models and 133 benchmarks, finding tha…
-
AI agents struggle to autoformalize code specs despite Gemini 3.1 Pro success
Researchers have introduced Verus-SpecGym, an agentic environment and benchmark designed to evaluate the ability of AI models to translate informal programming problems into formal specifications. The system tests gener…
-
New self-distillation methods boost LLM performance on reasoning tasks
Researchers have developed new self-distillation techniques for large language models to improve their performance without relying on external feedback. AVSD (Adaptive-View Self-Distillation) balances consensus signals …
-
DeepSeek V4's architecture slashes costs, impresses with high coding benchmark performance
DeepSeek V4 was released on April 24, 2026, with its architecture featuring five key tricks that contribute to its cost-effectiveness. The model achieved a Codeforces rating of 3,206, marking it as the highest rating ev…
-
Google upgrades Gemini 3 Deep Think for science and engineering
Google has released an upgraded version of Gemini 3 Deep Think, a specialized reasoning mode designed for complex scientific, research, and engineering challenges. This new iteration is available to Google AI Ultra subs…
-
new Gemini 3 Deep Think, Anthropic $30B @ $380B, GPT-5.3-Codex Spark, MiniMax M2.5
Google DeepMind has released Gemini 3 Deep Think V2, a new reasoning mode for Google AI Ultra subscribers and available via API early access. This model achieves new state-of-the-art results on benchmarks like ARC-AGI-2…
-
OpenAI's o1 model shows advanced reasoning, while Google and Apple explore new LLM training methods.
OpenAI has released an early version of its new model, OpenAI o1-preview, which demonstrates significant improvements in reasoning capabilities compared to GPT-4o. The model excels in competitive programming, advanced m…