BigCodeBench
PulseAugur coverage of BigCodeBench — every cluster mentioning BigCodeBench across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Thesis: LLM hidden states can predict code correctness
A new thesis explores the use of Introspective Uncertainty Estimation (IUE) to gauge the correctness of code generated by Large Language Models (LLMs). The research indicates that LLM hidden states can effectively signa…
-
LLMs Over-Edit Code, New Research Finds
A new research paper explores the issue of "over-editing" in large language models (LLMs) when they are used to repair code. The study found that even advanced models like GPT-5.5 tend to make larger edits than necessar…
-
New memory system APEX-EM boosts LLM agent performance
Researchers have developed APEX-EM, a novel non-parametric memory system designed to enhance the capabilities of large language model agents. This system stores complete procedural-episodic traces within a structured kn…
-
New framework STEP-KTODER optimizes code generation with function-level feedback
Researchers have introduced STEP-KTODER, a novel framework designed to enhance code generation models through function-level process supervision. This method defines 'steps' as module-level functions within decomposed p…
-
Self-correction methods fail to improve LLM code generation without verification
A new study on arXiv investigates the effectiveness of self-correction methods for large language models (LLMs) in code generation. Researchers found that while some uncertainty estimation techniques correlate weakly wi…
-
New metrics and benchmarks advance AI code quality evaluation
Researchers have developed FASE, a new metric for evaluating code quality in multi-agent AI systems. FASE approximates functional correctness by analyzing code dissimilarity, offering a significant speed improvement ove…
-
FLARE framework improves LLM code generation with fine-grained bug detection
Researchers have developed FLARE, a new framework designed to improve the accuracy of code generated by large language models. FLARE utilizes a lightweight diagnostic model to pinpoint specific lines of code that are li…
-
ReCode framework enhances AI code generation by rewarding reasoning processes
Researchers have developed ReCode, a novel reinforcement learning framework designed to improve code generation by focusing on the reasoning process. This framework uses Contrastive Reasoning-Process Reward Learning (CR…