A new paper titled "The FID Lottery" investigates the reproducibility of the Fréchet Inception Distance (FID) metric used in generative model evaluation. The study reveals that retraining a model with a different seed impacts FID 3.2 times more than simply resampling from a fixed network. This variance is attributed to factors like random initialization, data ordering, and the flow-matching loss. The research suggests a revised FID evaluation protocol that includes per-cell optimal guidance and error bars over multiple training seeds to account for this inherent randomness. AI
IMPACT Highlights the need for more robust evaluation protocols in generative AI, potentially impacting how model performance is reported and compared.
RANK_REASON The cluster contains an academic paper detailing a new research finding about a standard evaluation metric in generative AI.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →