PulseAugur
EN
LIVE 19:09:51
ENTITY LiveCodeBench V6

LiveCodeBench V6

PulseAugur coverage of LiveCodeBench V6 — every cluster mentioning LiveCodeBench V6 across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
6 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 11 TOTAL
  1. TOOL · CL_228949 ·

    New paper frames LLM post-training as 'brownfield maintenance'

    A new paper from Amir M. Ebrahimi on arXiv proposes an industrial perspective on dataware engineering for post-training large language models. The research frames this process as "brownfield maintenance," where improvem…

  2. TOOL · CL_226646 ·

    New Agentic Coding Index ranks LLMs by intelligence density

    A Reddit user has developed a new metric called the Agentic Coding Index (ACI) to evaluate Large Language Models (LLMs) on coding tasks. The ACI aggregates scores from several coding benchmarks, including SWE-bench Pro,…

  3. SIGNIFICANT · CL_210098 ·

    Qwen3.8-27B open-source model tops leaderboards with advanced agency

    Alibaba's Qwen team has open-sourced Qwen3.8-27B, a 27-billion parameter model that has achieved top rankings on several benchmarks, including SWE Bench Pro and OSWorld. This model demonstrates significant advancements …

  4. SIGNIFICANT · CL_182151 ·

    Google DeepMind's DiffusionGemma achieves 1500 tokens/sec via discrete diffusion

    Google DeepMind has released DiffusionGemma, an open-weight language model that utilizes discrete diffusion for text generation, offering significantly faster output speeds compared to traditional autoregressive models.…

  5. TOOL · CL_161625 ·

    Bad Theory Labs releases BTL-3 agent model for coding tasks

    Bad Theory Labs has released BTL-3, a 27-billion parameter open-weight model designed for agentic coding and structured tool use. This model is a post-trained version of Qwen3.6-27B, offering strong performance on codin…

  6. TOOL · CL_145035 ·

    Agents-A1-4B model shows strong performance in long-horizon search

    Agents-A1-4B, a new model developed by InternScience, demonstrates strong performance across various benchmarks, particularly in long-horizon search and agentic tasks. The model, which is based on Qwen3.7-4B, significan…

  7. RESEARCH · CL_94915 ·

    New 3B model VibeThinker matches frontier math & coding performance

    Researchers have developed VibeThinker-3B, a compact 3-billion parameter model that achieves performance comparable to much larger models in mathematics and coding tasks. This model, built upon Qwen2.5-Coder-3B and util…

  8. RESEARCH · CL_53559 ·

    New CPPO method boosts code generation by exploring multiple strategies

    Researchers have introduced Coordinated Pass@K Policy Optimization (CPPO), a novel method to enhance code generation by exploring multiple distinct algorithmic strategies simultaneously. Unlike standard approaches that …

  9. RESEARCH · CL_40825 ·

    New self-distillation methods boost LLM performance on reasoning tasks

    Researchers have developed new self-distillation techniques for large language models to improve their performance without relying on external feedback. AVSD (Adaptive-View Self-Distillation) balances consensus signals …

  10. RESEARCH · CL_02960 ·

    Process Supervision via Verbal Critique Improves Reasoning in Large Language Models

    Researchers have developed a new framework called Verbal Process Supervision (VPS) that enhances the reasoning capabilities of large language models without requiring gradient updates. This method utilizes structured na…

  11. FRONTIER RELEASE · CL_01735 ·

    Google DeepMind launches Deep Think for Gemini Ultra subscribers

    Google DeepMind has released a new AI capability called Deep Think, now available to Google AI Ultra subscribers via the Gemini app. This feature utilizes parallel thinking techniques, allowing the model to explore mult…