English(EN)An active-learning framework for real-time depth perception from monocular vision streams
新数据集和轻量级模型推动单目深度估计发展
作者PulseAugur 编辑部·[8 个来源]·
研究人员正在开发用于单目深度估计的新方法和数据集,这项技术对于增强现实和虚拟现实等应用至关重要。新的数据集(如MODEST)正在被创建,以提供高分辨率的真实世界图像,捕捉复杂的视觉效果,从而解决当前训练数据的局限性。同时,在轻量级神经网络架构和专为资源受限设备设计的主动学习框架方面也取得了进展,旨在提高在动态环境中的适应性和性能。
AI
arXiv:2511.20853v4 Announce Type: replace-cross Abstract: Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus deblurring remain constrained by a persistent lack of large-scale, full-frame, high fid…
arXiv cs.CV
TIER_1English(EN)·Xiaorong Zeng, Weiqiang Chen, Peng Shi, Liang Su, Zirui Wang, Xuewu Ji, Shuiwen Shen·
arXiv:2608.04917v1 Announce Type: new Abstract: Biological visual systems can perceive depth from monocular vision flow, continuously integrating temporal visual cues while maintaining a balance between stability and plasticity in dynamic environments. In contrast, artificial per…
arXiv:2608.03666v1 Announce Type: new Abstract: Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computationally constrained devices due to its reduced reliance on expensive depth sensors.…
arXiv cs.CV
TIER_1English(EN)·Ziyang Chen, Yansong Qu, You Shen, Xuan Cheng, Liujuan Cao·
arXiv:2603.29368v2 Announce Type: replace Abstract: Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research frontier. Contemporary stereo vision backbones typically rely on either Monocular D…
arXiv cs.CV
TIER_1English(EN)·Kaihua Tang, Ziqing Xia, Xiaoxu Zheng, Xiaoxue Zhang, Michael Bi Mi, Zhan Xu, Dave Zhenyu Chen·
arXiv:2608.00678v1 Announce Type: new Abstract: Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can result in substantial degradation in depth estimation…
arXiv:2608.02068v1 Announce Type: new Abstract: Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirro…
arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks largely evaluate physical accuracy rather than behavioral alignment with humans. We…
arXiv:2605.12027v2 Announce Type: replace Abstract: Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. T…