PulseAugur
EN
LIVE 10:29:41
ENTITY LLM judges

LLM judges

PulseAugur coverage of LLM judges — every cluster mentioning LLM judges across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
5
15 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
5
15 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/1 · 15 TOTAL
  1. TOOL · CL_198024 ·

    New Graph-Structured Rubrics improve LLM evaluation accuracy

    Researchers have developed Graph-Structured Rubrics (GSR), a novel method for compiling rubrics into typed evaluation graphs before response observation. This approach allows for explicit criterion composition and type …

  2. TOOL · CL_193755 ·

    New FanarGuard filter enhances Arabic LLMs with cultural awareness

    Researchers have developed FanarGuard, a new moderation filter designed to address safety and cultural alignment in Arabic language models. This filter was trained on a dataset of over 468,000 prompt and response pairs,…

  3. TOOL · CL_193611 ·

    New framework detects hidden behavioral entanglement in LLMs

    Researchers have developed a new statistical framework to detect and quantify behavioral entanglement among large language models (LLMs). This framework uses information-theoretic metrics, specifically a Difficulty-Weig…

  4. TOOL · CL_177085 ·

    Researchers probe MUD benchmark for AI evaluation flaws

    Researchers are investigating MUD, a benchmark environment, as a method for evaluating AI systems. Their study reveals that Large Language Model (LLM) judges can exhibit biases that are not detected by traditional aggre…

  5. RESEARCH · CL_143664 ·

    LLM judges may over-credit incorrect answers in reference-free evaluations

    A new research paper highlights a significant issue with using Large Language Model (LLM) judges for evaluating open-ended responses, particularly in scenarios lacking a ground truth answer. The study found that these L…

  6. TOOL · CL_141539 ·

    New protocol probes LLM judges for bias in RAG systems

    A new meta-evaluation protocol called Eval-Pair Matrix has been developed to assess the reliability of Large Language Models (LLMs) when used as judges in retrieval-augmented generation (RAG) systems. This method aims t…

  7. RESEARCH · CL_128710 ·

    LLM-as-a-Tutor framework enhances reinforcement learning for instruction following

    Researchers have developed a new framework called LLM-as-a-Tutor to improve reinforcement learning for instruction following. This system dynamically adjusts the difficulty of training prompts by having a single LLM act…

  8. TOOL · CL_123204 ·

    New framework improves LLM judges by accounting for bias

    A new research paper introduces a bias-aware Bayesian active learning framework designed to improve the accuracy of large language models (LLMs) when used as judges for ranking tasks. The framework explicitly models jud…

  9. TOOL · CL_117472 ·

    Specialized Clinical AI Outperforms Frontier Models in Real-World Tests

    A new study evaluated the performance of leading AI models, including Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5, against a specialized clinical tool called OpenEvidence. The evaluation used 620 real-world clinical qu…

  10. TOOL · CL_62713 ·

    New framework audits LLM judge rubrics for reliability and robustness

    Researchers have developed PReMISE, a framework designed to evaluate the effectiveness of rubrics used by Large Language Model (LLM) judges. The framework treats rubrics as measurement specifications, analyzing their st…

  11. TOOL · CL_53666 ·

    New BITE framework exploits LLM judge biases to inflate scores

    Researchers have developed a novel black-box adversarial framework called BITE that exploits stylistic biases in LLM judges to artificially inflate their scores. By framing the selection of stylistic edits as a contextu…

  12. TOOL · CL_51221 ·

    LLM judges show rationalization bias, new framework reveals

    Researchers have developed a causal framework to analyze rationalization bias in large language models (LLMs) when they act as judges for text evaluation. The study introduces new metrics and cue interventions to test i…

  13. TOOL · CL_51073 ·

    New framework tackles preference cycles in AI feedback

    Researchers have developed a new framework called Topological Consensus Rewards (TCR) to improve the stability of Reinforcement Learning from AI Feedback (RLAIF). This method addresses the issue of preference cycles, wh…

  14. TOOL · CL_40852 ·

    New benchmark reveals LLM judges unreliable for research agents

    Researchers have developed a new benchmark called REFLECT to evaluate the reliability of Large Language Models (LLMs) when used as judges for deep research agents. These agents automate complex information-seeking tasks…

  15. TOOL · CL_21933 ·

    LLM judges evaluate agentic stock predictors, improving accuracy via reinforcement learning

    Researchers have developed a novel framework for evaluating agentic stock prediction systems by utilizing large language models as judges. This system breaks down performance into six specific dimensions, including regi…