PulseAugur
EN
LIVE 13:35:01

LLM judges unreliable for table recognition regeneration, study finds

A new research paper challenges the reliability of using Large Language Models (LLMs) as judges for evaluating and selecting outputs in closed-loop regeneration tasks, particularly in table recognition. The study found that LLM judge signals were weak, often resulting in tied scores and irreproducible rankings. While iterative refinement improved candidate outputs, the LLM judges failed to consistently identify the better ones. The research suggests that LLMs' evaluation ability does not directly translate to optimization utility, and that deterministic verification signals are necessary for effective iterative refinement. AI

IMPACT Suggests current LLM evaluation methods may not be sufficient for optimizing regeneration tasks, potentially impacting development of AI systems that rely on such feedback loops.

RANK_REASON Research paper published on arXiv detailing findings about LLM evaluation capabilities.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

LLM judges unreliable for table recognition regeneration, study finds

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Donghwan Kim ·

    Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition

    arXiv:2607.13347v1 Announce Type: cross Abstract: LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a co…

  2. arXiv cs.AI TIER_1 English(EN) · Donghwan Kim ·

    Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition

    LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, using FinTabNet and OmniDocBench…