PulseAugur
实时 20:28:01

新的Apple-PI基准测试视频模型对物理定律的理解能力

研究人员推出了一种名为Apple-PI的新型基准测试,旨在评估视频生成模型对物理定律的理解能力。与以往仅评估输出合理性的方法不同,Apple-PI scrutinizes reasoning process itself。该基准测试包含一个名为Orchard的数据集,其中有400个关于经典力学的视频,一个使用链式帧提示的三阶段协议(感知、形成、推导),以及一个结合主观和客观度量的混合评估套件。对11个模型的初步测试显示,当前的视频模型在基于定律的物理智能方面存在显著不足,最好的模型得分仅为0.473,凸显了在形成和推导阶段的瓶颈。 AI

影响 为评估视频生成模型的物理推理能力树立了新标准。

排序理由 该集群描述了一个用于评估AI模型的新学术基准和数据集,该基准和数据集在一篇研究论文中提出。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的Apple-PI基准测试视频模型对物理定律的理解能力

报道来源 [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    Apple-π:通过视频进行思维基准测试,迈向基于法律的物理智能

    Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithfu…

  2. arXiv cs.CV TIER_1 English(EN) · Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu ·

    Apple-$\pi$:通过视频对标思考,实现以法律为基础的物理智能

    arXiv:2607.16401v1 Announce Type: new Abstract: Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying w…