PulseAugur
EN
LIVE 07:20:34
ENTITY BigCodeBench

BigCodeBench

PulseAugur coverage of BigCodeBench — every cluster mentioning BigCodeBench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
7
7 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
7
7 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 8 TOTAL
  1. TOOL · CL_254740 ·

    Thesis: LLM hidden states can predict code correctness

    A new thesis explores the use of Introspective Uncertainty Estimation (IUE) to gauge the correctness of code generated by Large Language Models (LLMs). The research indicates that LLM hidden states can effectively signa…

  2. RESEARCH · CL_235465 ·

    LLMs Over-Edit Code, New Research Finds

    A new research paper explores the issue of "over-editing" in large language models (LLMs) when they are used to repair code. The study found that even advanced models like GPT-5.5 tend to make larger edits than necessar…

  3. TOOL · CL_231493 ·

    New memory system APEX-EM boosts LLM agent performance

    Researchers have developed APEX-EM, a novel non-parametric memory system designed to enhance the capabilities of large language model agents. This system stores complete procedural-episodic traces within a structured kn…

  4. TOOL · CL_218870 ·

    New framework STEP-KTODER optimizes code generation with function-level feedback

    Researchers have introduced STEP-KTODER, a novel framework designed to enhance code generation models through function-level process supervision. This method defines 'steps' as module-level functions within decomposed p…

  5. TOOL · CL_205902 ·

    Self-correction methods fail to improve LLM code generation without verification

    A new study on arXiv investigates the effectiveness of self-correction methods for large language models (LLMs) in code generation. Researchers found that while some uncertainty estimation techniques correlate weakly wi…

  6. RESEARCH · CL_77299 ·

    New metrics and benchmarks advance AI code quality evaluation

    Researchers have developed FASE, a new metric for evaluating code quality in multi-agent AI systems. FASE approximates functional correctness by analyzing code dissimilarity, offering a significant speed improvement ove…

  7. RESEARCH · CL_68146 ·

    FLARE framework improves LLM code generation with fine-grained bug detection

    Researchers have developed FLARE, a new framework designed to improve the accuracy of code generated by large language models. FLARE utilizes a lightweight diagnostic model to pinpoint specific lines of code that are li…

  8. TOOL · CL_18865 ·

    ReCode framework enhances AI code generation by rewarding reasoning processes

    Researchers have developed ReCode, a novel reinforcement learning framework designed to improve code generation by focusing on the reasoning process. This framework uses Contrastive Reasoning-Process Reward Learning (CR…