PulseAugur
EN
LIVE 18:31:46

New UOJ-Bench evaluates LLMs on code repair and error detection

A new benchmark called UOJ-Bench has been developed to evaluate Large Language Models (LLMs) on code generation, hacking, and repair tasks, moving beyond simple problem-solving. Initial tests show that even top-tier models struggle with identifying errors in human-written code, with success rates below 50% in one-shot evaluations. While test-time scaling improves performance significantly, it incurs substantial computational costs, limiting practical deployment. However, the best models can still identify errors in a small percentage of full-score submissions, suggesting potential for LLMs to offer complementary insights to existing judging systems. AI

IMPACT New benchmarks like UOJ-Bench and FrontierCode are pushing LLM evaluations beyond simple problem-solving to assess more nuanced capabilities like code repair and maintainability, highlighting current limitations.

RANK_REASON The cluster focuses on new benchmarks and evaluations for LLMs in code-related tasks, rather than a new model release or significant industry event.

Read on Latent Space (swyx) →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

New UOJ-Bench evaluates LLMs on code repair and error detection

COVERAGE [4]

  1. arXiv cs.AI TIER_1 English(EN) · Tingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu, Kaifeng Lyu ·

    Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming

    arXiv:2606.12864v1 Announce Type: cross Abstract: Despite strong performance in competitive programming, the role of Large Language Models (LLMs) in supporting human learning in the same setting remains largely unexplored. In this work, we introduce UOJ-Bench, a benchmark designe…

  2. Latent Space (swyx) TIER_1 English(EN) ·

    [AINews] FrontierCode: Benchmarking for Code Quality over Slop

    We made a thing!

  3. HN — claude cli stories TIER_1 English(EN) · bugvader ·

    Claude Fable 5: mid-tier results on coding tasks

  4. r/singularity TIER_2 English(EN) · /u/acoolrandomusername ·

    FrontierCode: a coding eval that raises the bar for difficulty & quality.

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1u0k192/frontiercode_a_coding_eval_that_raises_the_bar/"> <img alt="FrontierCode: a coding eval that raises the bar for difficulty &amp; quality." src="https://preview.redd.it/ihk4ib8nd46h1.png?width=640&amp;…