Researchers have introduced AVENUE, a new benchmark and evaluation framework designed to improve audio-video editing models. The AVENUE benchmark includes 1,291 source clips and 7,957 editing instructions, curated from the VGGSound dataset, covering audio-targeted, video-targeted, and coupled edits. The accompanying evaluation framework is modality-aware and sample-specific, addressing limitations in existing systems that often overlook unintended cross-modal changes. Initial evaluations using AVENUE reveal that current models frequently introduce unwanted modifications to one modality when editing the other, regardless of the editing paradigm used. AI
IMPACT This benchmark aims to drive progress in controllable audio-video editing, potentially leading to more sophisticated multimedia manipulation tools.
RANK_REASON The cluster describes a new benchmark and evaluation framework for audio-video editing models, presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →