Researchers have explored the potential of vision-language models (VLMs) for assessing the quality of Olympic diving performances. A proposed framework leverages VLMs' semantic reasoning and phase-level sub-scores, combined with TF-IDF vectorization and ensemble learning, to predict final competition scores. While standalone VLMs showed limited correlation, the ensemble approach achieved a Spearman correlation of 0.67, indicating that VLM-generated explanations are valuable for sports performance evaluation. AI
影响 VLMs show potential as assistive tools for explainable and semi-automated sports performance evaluation.
排序理由 Academic paper detailing a new methodology for action quality assessment using VLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →