Researchers have introduced MultiRef-Compass, a new benchmark designed to evaluate multi-reference-to-audio-video (MR2AV) generation. This benchmark addresses the limitations of existing methods by focusing on the complex task of generating coherent audio-video content from multiple references and textual instructions. MultiRef-Compass includes 350 curated samples and a four-dimensional evaluation protocol, incorporating automatic metrics and an MLLM-as-a-Judge framework to assess perceptual fidelity and reference-conditioned composition. AI
IMPACT Provides a standardized method for evaluating complex audio-video generation models, driving progress in the field.
RANK_REASON The cluster describes a new academic benchmark for evaluating a specific type of AI generation task.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →