Researchers have developed Rigel, a new metric for evaluating image and video captioning systems that aims to better align with human judgments than existing methods. Rigel utilizes a self-distilled score adaptation approach, where an evaluation-specific scoring head is distilled from a large language model and then refined with human judgment data. This method avoids the limitations of large-vocabulary language models by focusing on task-aligned scoring. The effectiveness of Rigel was demonstrated using the newly constructed Vid-Lepus dataset, showing significant improvements over current state-of-the-art metrics on benchmarks like ActivityNet-Fact. AI
IMPACT This new metric could lead to more accurate benchmarking of multimodal AI systems, driving progress in image and video captioning.
RANK_REASON The cluster describes a new research paper introducing a novel metric for AI evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →