MLLM-as-a-Judge
PulseAugur coverage of MLLM-as-a-Judge — every cluster mentioning MLLM-as-a-Judge across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New PRISM-Bench evaluates audio in text-to-video generation
Researchers have introduced PRISM-Bench, a new benchmark designed to specifically evaluate the audio generation capabilities of text-to-audio-video (T2AV) systems. Unlike previous benchmarks that treated audio as second…
-
New multimodal benchmark and agents aim to boost AI business ideation
Researchers have developed MBA-Bench, a novel multimodal benchmark designed to train and evaluate AI agents for real-world business ideation. This benchmark includes 30,000 samples across six domains, incorporating visu…
-
New benchmark MultiRef-Compass evaluates multi-reference audio-video generation
Researchers have introduced MultiRef-Compass, a new benchmark designed to evaluate multi-reference-to-audio-video (MR2AV) generation. This benchmark addresses the limitations of existing methods by focusing on the compl…