New datasets and lightweight models advance monocular depth estimation
ByPulseAugur Editorial·[8 sources]·
Researchers are developing new methods and datasets for monocular depth estimation, a technique crucial for applications like augmented and virtual reality. New datasets such as MODEST are being created to provide high-resolution, real-world images that capture complex optical effects, addressing limitations in current training data. Concurrently, advancements are being made in lightweight neural network architectures and active learning frameworks designed for resource-constrained devices, aiming to improve adaptability and performance in dynamic environments.
AI
IMPACT
Advances in monocular depth estimation could enable more sophisticated AI applications in robotics, autonomous driving, and immersive technologies.
RANK_REASON
Multiple research papers published on arXiv detailing new datasets, models, and techniques for monocular depth estimation.
arXiv:2511.20853v4 Announce Type: replace-cross Abstract: Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus deblurring remain constrained by a persistent lack of large-scale, full-frame, high fid…
arXiv cs.CV
TIER_1English(EN)·Xiaorong Zeng, Weiqiang Chen, Peng Shi, Liang Su, Zirui Wang, Xuewu Ji, Shuiwen Shen·
arXiv:2608.04917v1 Announce Type: new Abstract: Biological visual systems can perceive depth from monocular vision flow, continuously integrating temporal visual cues while maintaining a balance between stability and plasticity in dynamic environments. In contrast, artificial per…
arXiv:2608.03666v1 Announce Type: new Abstract: Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computationally constrained devices due to its reduced reliance on expensive depth sensors.…
arXiv cs.CV
TIER_1English(EN)·Ziyang Chen, Yansong Qu, You Shen, Xuan Cheng, Liujuan Cao·
arXiv:2603.29368v2 Announce Type: replace Abstract: Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research frontier. Contemporary stereo vision backbones typically rely on either Monocular D…
arXiv cs.CV
TIER_1English(EN)·Kaihua Tang, Ziqing Xia, Xiaoxu Zheng, Xiaoxue Zhang, Michael Bi Mi, Zhan Xu, Dave Zhenyu Chen·
arXiv:2608.00678v1 Announce Type: new Abstract: Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can result in substantial degradation in depth estimation…
arXiv:2608.02068v1 Announce Type: new Abstract: Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirro…
arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks largely evaluate physical accuracy rather than behavioral alignment with humans. We…
arXiv:2605.12027v2 Announce Type: replace Abstract: Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. T…