PulseAugur
EN
LIVE 15:15:49
ENTITY Codeforces

Codeforces

PulseAugur coverage of Codeforces — every cluster mentioning Codeforces across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
5
8 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
3
5 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/1 · 15 TOTAL
  1. TOOL · CL_256056 ·

    Google's Stellar Colosseum tackles long-horizon math proofs with Gemini models

    Google Research has developed a novel multi-agent system called Stellar Colosseum, designed to tackle long-horizon tasks, particularly in mathematical proofs. This system operates in stages, generating candidate solutio…

  2. RESEARCH · CL_252313 ·

    DeepSeek's V4.1 Flash model offers speed and low cost but struggles with market share

    DeepSeek has released its V4.1 Flash model, a 552B parameter Mixture-of-Experts model that boasts impressive speed and a significantly reduced KV cache size, making it one of the cheapest frontier-class models available…

  3. TOOL · CL_244747 ·

    AI reasoning traces found to be easily decryptable, exposing data

    Researchers have discovered that the "encrypted" reasoning traces used by major AI providers like OpenAI, Anthropic, and Google are not truly secure. These traces, intended to protect proprietary reasoning and sensitive…

  4. TOOL · CL_231333 ·

    New benchmark reveals LLMs struggle with interactive programming problems

    Researchers have introduced InteractBench, a new benchmark designed to evaluate the algorithmic reasoning capabilities of large language models (LLMs) on interactive problems. These problems, common in competitive progr…

  5. TOOL · CL_228949 ·

    New paper frames LLM post-training as 'brownfield maintenance'

    A new paper from Amir M. Ebrahimi on arXiv proposes an industrial perspective on dataware engineering for post-training large language models. The research frames this process as "brownfield maintenance," where improvem…

  6. TOOL · CL_197371 ·

    AI Reasoning Traces Vulnerable to Cross-Model Decryption Attacks

    Researchers have discovered a vulnerability in the API ecosystems of Anthropic, OpenAI, and Google that allows for the extraction of supposedly hidden reasoning traces. By replaying encrypted reasoning blocks into weake…

  7. RESEARCH · CL_128531 ·

    SpecCoder framework enhances Code LLMs with formal specifications

    Researchers have developed SpecCoder, a new framework designed to enhance the reasoning capabilities of Code LLMs by utilizing intermediate formal specifications. Unlike natural language, these executable specifications…

  8. TOOL · CL_119403 ·

    Research probes how language agents effectively use feedback for improvement

    A new research paper investigates the effectiveness of feedback in improving language agent performance. The study introduces a controlled student-teacher protocol across multiple benchmarks, comparing external feedback…

  9. RESEARCH · CL_107855 ·

    AI benchmark scores predictable from just two factors, study finds

    A new research paper proposes a method called BenchPress that can predict a frontier model's performance across numerous benchmarks using only two key scores. The study analyzed 84 models and 133 benchmarks, finding tha…

  10. RESEARCH · CL_53697 ·

    AI agents struggle to autoformalize code specs despite Gemini 3.1 Pro success

    Researchers have introduced Verus-SpecGym, an agentic environment and benchmark designed to evaluate the ability of AI models to translate informal programming problems into formal specifications. The system tests gener…

  11. RESEARCH · CL_40825 ·

    New self-distillation methods boost LLM performance on reasoning tasks

    Researchers have developed new self-distillation techniques for large language models to improve their performance without relying on external feedback. AVSD (Adaptive-View Self-Distillation) balances consensus signals …

  12. SIGNIFICANT · CL_18151 ·

    DeepSeek V4's architecture slashes costs, impresses with high coding benchmark performance

    DeepSeek V4 was released on April 24, 2026, with its architecture featuring five key tricks that contribute to its cost-effectiveness. The model achieved a Codeforces rating of 3,206, marking it as the highest rating ev…

  13. SIGNIFICANT · CL_97397 ·

    Google upgrades Gemini 3 Deep Think for science and engineering

    Google has released an upgraded version of Gemini 3 Deep Think, a specialized reasoning mode designed for complex scientific, research, and engineering challenges. This new iteration is available to Google AI Ultra subs…

  14. FRONTIER RELEASE · CL_01763 ·

    new Gemini 3 Deep Think, Anthropic $30B @ $380B, GPT-5.3-Codex Spark, MiniMax M2.5

    Google DeepMind has released Gemini 3 Deep Think V2, a new reasoning mode for Google AI Ultra subscribers and available via API early access. This model achieves new state-of-the-art results on benchmarks like ARC-AGI-2…

  15. FRONTIER RELEASE · CL_01020 ·

    OpenAI's o1 model shows advanced reasoning, while Google and Apple explore new LLM training methods.

    OpenAI has released an early version of its new model, OpenAI o1-preview, which demonstrates significant improvements in reasoning capabilities compared to GPT-4o. The model excels in competitive programming, advanced m…