PulseAugur
中
实时 19:17:08
English(EN) An active-learning framework for real-time depth perception from monocular vision streams

新数据集和轻量级模型推动单目深度估计发展

研究人员正在开发用于单目深度估计的新方法和数据集,这项技术对于增强现实和虚拟现实等应用至关重要。新的数据集(如MODEST)正在被创建,以提供高分辨率的真实世界图像,捕捉复杂的视觉效果,从而解决当前训练数据的局限性。同时,在轻量级神经网络架构和专为资源受限设备设计的主动学习框架方面也取得了进展,旨在提高在动态环境中的适应性和性能。 AI

影响 单目深度估计的进步可以为机器人、自动驾驶和沉浸式技术领域带来更复杂的人工智能应用。

排序理由 多篇在arXiv上发表的研究论文,详细介绍了用于单目深度估计的新数据集、模型和技术。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

新数据集和轻量级模型推动单目深度估计发展

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇在arXiv上发表的研究论文,详细介绍了用于单目深度估计的新数据集、模型和技术。
Source corroboration
8 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
65 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [8]

  1. arXiv cs.AI TIER_1 English(EN) · Nisarg K. Trivedi, Vinayaka A. Belludi, Li-Yun Wang ·

    MODEST: 多光学景深立体数据集

    arXiv:2511.20853v4 Announce Type: replace-cross Abstract: Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus deblurring remain constrained by a persistent lack of large-scale, full-frame, high fid…

  2. arXiv cs.CV TIER_1 English(EN) · Xiaorong Zeng, Weiqiang Chen, Peng Shi, Liang Su, Zirui Wang, Xuewu Ji, Shuiwen Shen ·

    用于单目视觉流实时深度感知的激活学习框架

    arXiv:2608.04917v1 Announce Type: new Abstract: Biological visual systems can perceive depth from monocular vision flow, continuously integrating temporal visual cues while maintaining a balance between stability and plasticity in dynamic environments. In contrast, artificial per…

  3. arXiv cs.CV TIER_1 English(EN) · Elena Izzo, Riccardo Toniolo, Lamberto Ballan ·

    XiDepth:一种轻量级高效的自监督单目深度估计网络

    arXiv:2608.03666v1 Announce Type: new Abstract: Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computationally constrained devices due to its reduced reliance on expensive depth sensors.…

  4. arXiv cs.CV TIER_1 English(EN) · Ziyang Chen, Yansong Qu, You Shen, Xuan Cheng, Liujuan Cao ·

    StereoVGGT:一种无需训练的立体视觉几何变换器

    arXiv:2603.29368v2 Announce Type: replace Abstract: Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research frontier. Contemporary stereo vision backbones typically rely on either Monocular D…

  5. arXiv cs.CV TIER_1 English(EN) · Kaihua Tang, Ziqing Xia, Xiaoxu Zheng, Xiaoxue Zhang, Michael Bi Mi, Zhan Xu, Dave Zhenyu Chen ·

    打破水平先验:从长尾方向偏差到滚动鲁棒单目深度估计

    arXiv:2608.00678v1 Announce Type: new Abstract: Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can result in substantial degradation in depth estimation…

  6. arXiv cs.CV TIER_1 English(EN) · Xianghui Fan, Zhaoyu Chen, Bingqian Wu, Dayu Li, Xin Zeng, Huanran Cui, Guangzhen Xu, Xiangru Huang, Hang Yang ·

    GIFT:非朗伯反射单目深度估计的几何不变微调

    arXiv:2608.02068v1 Announce Type: new Abstract: Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirro…

  7. arXiv cs.CV TIER_1 English(EN) · Yuki Kubota, Taiki Fukiage ·

    准确性不保证类人:单目深度估计中的跨域以人为中心的基准

    arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks largely evaluate physical accuracy rather than behavioral alignment with humans. We…

  8. arXiv cs.CV TIER_1 English(EN) · Ying Zang, Xuanyi Liu, Yidong Han, Deyi Ji, Chaotao Ding, Yuanqi Hu, Qi Zhu, Xuanfu Li, Jin Ma, Lingyun Sun, Tianrun Chen, Lanyun Zhu ·

    4DVGGT-D:具有改进动态深度估计的4D视觉几何Transformer

    arXiv:2605.12027v2 Announce Type: replace Abstract: Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. T…