PulseAugur
EN
LIVE 06:18:43
ENTITY MLLM-as-a-Judge

MLLM-as-a-Judge

PulseAugur coverage of MLLM-as-a-Judge — every cluster mentioning MLLM-as-a-Judge across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
3 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
3 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 3 TOTAL
  1. TOOL · CL_239312 ·

    New PRISM-Bench evaluates audio in text-to-video generation

    Researchers have introduced PRISM-Bench, a new benchmark designed to specifically evaluate the audio generation capabilities of text-to-audio-video (T2AV) systems. Unlike previous benchmarks that treated audio as second…

  2. RESEARCH · CL_197790 ·

    New multimodal benchmark and agents aim to boost AI business ideation

    Researchers have developed MBA-Bench, a novel multimodal benchmark designed to train and evaluate AI agents for real-world business ideation. This benchmark includes 30,000 samples across six domains, incorporating visu…

  3. RESEARCH · CL_148028 ·

    New benchmark MultiRef-Compass evaluates multi-reference audio-video generation

    Researchers have introduced MultiRef-Compass, a new benchmark designed to evaluate multi-reference-to-audio-video (MR2AV) generation. This benchmark addresses the limitations of existing methods by focusing on the compl…