PulseAugur
实时 01:28:08
English(EN) Velox: Learning Representations of 4D Geometry and Appearance

新研究通过新颖的框架探索 4D 几何和动态场景理解

研究人员推出了几个新的框架和数据集,用于从视觉数据推进 4D(三个空间维度加上时间)理解和重建。其中包括 4DThinker,它通过在连续隐藏空间中模拟场景演变,使视觉语言模型能够“用 4D 进行思考”;以及 Ground4D,一个用于在非结构化环境中进行无姿态 4D 重建的空间锚定框架。此外,Velox 提供了一种从动态点云中学习 4D 几何和外观潜在表示的方法,而 Syn4D 为动态场景重建和跟踪提供了合成数据集。Flux4D 提出了一种可扩展的无监督方法,用于大规模动态场景的 4D 重建,ISExplore 通过选择信息丰富的短参考视频片段,为个性化 3D 说话人脸生成提供了一种有效的策略。 AI

影响 这些在 4D 理解和重建方面的进步可以显著改善机器人技术、自动驾驶和逼真的虚拟环境生成。

排序理由 arXiv 上发表了多篇研究论文,详细介绍了用于 4D 重建和理解的新框架和数据集。

在 Apple Machine Learning Research 阅读 →

AI 生成摘要 · Google Gemini · 来自 12 个来源。 我们如何撰写摘要 →

新研究通过新颖的框架探索 4D 几何和动态场景理解

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
arXiv 上发表了多篇研究论文,详细介绍了用于 4D 重建和理解的新框架和数据集。
Source corroboration
12 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
134 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [12]

  1. Apple Machine Learning Research TIER_1 English(EN) ·

    Velox:学习四维几何和外观的表示

    We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input, i.e., an unstructured dynamic point cloud,…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Velox:学习四维几何和外观的表示

    We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input, i.e., an unstructured dynamic point cloud,…

  3. arXiv cs.CV TIER_1 English(EN) · Zhangquan Chen, Manyuan Zhang, Xinlei Yu, Xiang An, Bo Li, Xin Xie, ZiDong Wang, Mingze Sun, Shuang Chen, Hongyu Li, Xiaobin Hu, Ruqi Huang ·

    4DThinker:通过四维图像进行思考以实现动态空间理解

    arXiv:2605.05997v1 Announce Type: new Abstract: Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vision-language models (VLMs). Prior approaches either verbalize spatial-temporal reaso…

  4. arXiv cs.CV TIER_1 English(EN) · Ruqi Huang ·

    4DThinker:通过四维图像进行思考以实现动态空间理解

    Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vision-language models (VLMs). Prior approaches either verbalize spatial-temporal reasoning entirely as text, which is inherently verbo…

  5. arXiv cs.CV TIER_1 English(EN) · Anagh Malik, Dorian Chan, Xiaoming Zhao, David B. Lindell, Oncel Tuzel, Jen-Hao Rick Chang ·

    Velox:学习四维几何和外观的表示

    arXiv:2605.04527v1 Announce Type: new Abstract: We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal i…

  6. arXiv cs.CV TIER_1 English(EN) · Shuo Wang, Jilin Mei, Fuyang Liu, Wenfei Guan, Fanjie Kong, Zhihua Zhao, Shuai Wang, Chen Min, Yu Hu ·

    Ground4D:非结构化越野场景的空间地面前馈4D重建

    arXiv:2605.04435v1 Announce Type: new Abstract: Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-road scenes, its performance degrades due to high-frequency geometry, ego-motion …

  7. arXiv cs.CV TIER_1 English(EN) · Zeren Jiang, Yushi Lan, Yihang Luo, Yufan Deng, Zihang Lai, Edgar Sucar, Christian Rupprecht, Iro Laina, Diane Larlus, Chuanxia Zheng, Andrea Vedaldi ·

    Syn4D:一个多视角合成4D数据集

    arXiv:2605.05207v1 Announce Type: new Abstract: Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by the scarcity of high-quality datasets with dense, comp…

  8. arXiv cs.CV TIER_1 English(EN) · Andrea Vedaldi ·

    Syn4D:一个多视角合成4D数据集

    Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by the scarcity of high-quality datasets with dense, complete, and accurate geometric annotations. To add…

  9. arXiv cs.CV TIER_1 English(EN) · Jen-Hao Rick Chang ·

    Velox:学习四维几何和外观的表示

    We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input, i.e., an unstructured dynamic point cloud,…

  10. arXiv cs.CV TIER_1 English(EN) · Yihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan, Chen Change Loy ·

    4RC:随时随地通过条件查询进行四维重建

    arXiv:2602.10094v2 Announce Type: replace Abstract: We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple motion from geometry or produce limited 4D attributes such as sparse trajectories o…

  11. arXiv cs.CV TIER_1 English(EN) · Jingkang Wang, Henry Che, Yun Chen, Ze Yang, Lily Goli, Sivabalan Manivasagam, Raquel Urtasun ·

    Flux4D:基于流的无监督4D重建

    arXiv:2512.03210v2 Announce Type: replace Abstract: Reconstructing large-scale dynamic scenes from visual observations is a fundamental challenge in computer vision, with critical implications for robotics and autonomous systems. While recent differentiable rendering methods such…

  12. arXiv cs.CV TIER_1 English(EN) · Rui-Qing Sun, Ang Li, Zhijing Wu, Tian Lan, Qianyu Lu, Xingshan Yao, Chen Xu, Xian-Ling Mao ·

    ISExplore:信息片段选择用于高效个性化3D谈话人脸生成

    arXiv:2511.07940v2 Announce Type: replace Abstract: Talking Face Generation (TFG) methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have recently achieved impressive progress in personalized talking head synthesis. However, existing methods typically…