实体
PhysBench: A Benchmark Framework for Remote Physiological Sensing with New Dataset and Baseline
PhysBench: A Benchmark Framework for Remote Physiological Sensing with New Dataset and Baseline
PulseAugur coverage of PhysBench: A Benchmark Framework for Remote Physiological Sensing with New Dataset and Baseline — every cluster mentioning PhysBench: A Benchmark Framework for Remote Physiological Sensing with New Dataset and Baseline across labs, papers, and developer communities, ranked by signal.
总计 · 30天
1
90 天内 2
发布 · 30天
0
90 天内 0
论文 · 30天
1
90 天内 2
层级分布 · 90 天
主题
情绪 · 30 天
1 天有情绪数据
最近 · 第 1/1 页 · 共 2 条
-
视觉语言模型(VLMs)物理测试失败,依赖模式匹配而非理解
根据最新的基准测试,视觉语言模型(VLMs)经常能正确回答与物理学相关的问题,但原因却不正确。像PhysBench和IntPhys 2这样的研究表明,即使是像GPT-4o这样先进的模型,在需要理解物理属性和动力学的任务上的得分也只有40-50%,而人类的得分接近满分。这表明VLMs在很大程度上依赖于训练数据中的模式匹配,而不是真正理解物理定律。
-
PhysBrain 1.0 从视频中提取物理常识用于机器人学习
研究人员推出 PhysBrain 1.0,这是一种通过从大规模人类自我中心视频中提取物理常识来增强机器人学习的新方法。该方法将视频数据转换为结构化的问答监督,然后用于训练视觉-语言-动作 (VLA) 模型。PhysBrain 1.0 在各种多模态 QA 和具身控制基准测试中表现出最先进的性能,尤其显示出强大的域外泛化能力。