PulseAugur
实时 10:11:45

新研究应对单目深度估计挑战

研究人员正在开发新的方法来改进单目深度估计,这是一种允许单个摄像头推断3D场景几何形状的技术。一种方法,不变深度约束(ID-Constraint),通过在训练期间引入几何和空间推理任务来解决由相机滚动角度引起的不鲁棒性问题。另一种方法,GIFT(几何不变微调),专门针对镜子和玻璃等非朗伯表面,防止模型在反射性或透明内容上产生深度幻觉。此外,还创建了一个新的人类中心基准来评估深度估计模型,不仅评估其物理准确性,还评估其预测与人类感知的匹配程度,强调高准确性并不总是等同于类人行为。 AI

影响 单目深度估计的进步可以通过从单个摄像头输入实现更准确的3D场景理解,从而改进机器人、自动驾驶和增强现实应用。

排序理由 多篇arXiv论文介绍了单目深度估计的新方法和基准。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新研究应对单目深度估计挑战

报道来源 [4]

  1. arXiv cs.CV TIER_1 English(EN) · Kaihua Tang, Ziqing Xia, Xiaoxu Zheng, Xiaoxue Zhang, Michael Bi Mi, Zhan Xu, Dave Zhenyu Chen ·

    Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation

    arXiv:2608.00678v1 Announce Type: new Abstract: Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can result in substantial degradation in depth estimation…

  2. arXiv cs.CV TIER_1 English(EN) · Xianghui Fan, Zhaoyu Chen, Bingqian Wu, Dayu Li, Xin Zeng, Huanran Cui, Guangzhen Xu, Xiangru Huang, Hang Yang ·

    GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation

    arXiv:2608.02068v1 Announce Type: new Abstract: Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirro…

  3. arXiv cs.CV TIER_1 English(EN) · Yuki Kubota, Taiki Fukiage ·

    Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

    arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks largely evaluate physical accuracy rather than behavioral alignment with humans. We…

  4. arXiv cs.CV TIER_1 English(EN) · Ying Zang, Xuanyi Liu, Yidong Han, Deyi Ji, Chaotao Ding, Yuanqi Hu, Qi Zhu, Xuanfu Li, Jin Ma, Lingyun Sun, Tianrun Chen, Lanyun Zhu ·

    4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation

    arXiv:2605.12027v2 Announce Type: replace Abstract: Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. T…