ENTITY
MLLM-as-Judge
MLLM-as-Judge
PulseAugur coverage of MLLM-as-Judge — every cluster mentioning MLLM-as-Judge across labs, papers, and developer communities, ranked by signal.
Total · 30d
1
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D
1 day(s) with sentiment data
RECENT · PAGE 1/1 · 2 TOTAL
-
AI models struggle with visual metaphor generation, new benchmarks show
Two research papers explore the generation of visual metaphors using AI. The first paper introduces VMetaphor-Bench, a benchmark for evaluating text-to-image models on their ability to create visual metaphors, finding t…
-
New Sci-VBench benchmark reveals gap in AI video generation for science
Researchers have introduced Sci-VBench, a new benchmark designed to evaluate the capabilities of AI models in generating videos for scientific domains. This benchmark includes over 1,200 expert-annotated examples across…