PulseAugur
EN
LIVE 09:11:58

New HOPformer model and EPIC-Contact dataset advance egocentric 3D hand-object pose estimation

Researchers have introduced HOPformer, a novel end-to-end transformer model designed for egocentric 3D hand-object pose estimation. This model jointly predicts hand and object poses in a single pass, improving robustness through cross-attention that conditions object features on hand priors. To facilitate this research, the team also released EPIC-Contact, a new dataset featuring 2.3K clips of in-the-wild egocentric scenes with detailed 3D hand-object contact correspondences and posed meshes. HOPformer demonstrates significant performance gains on both existing and the new dataset, nearly doubling success rates on EPIC-Contact. AI

IMPACT Advances egocentric 3D computer vision capabilities, potentially improving robotics and augmented reality applications.

RANK_REASON The cluster describes a new research paper detailing a novel model and dataset for a specific computer vision task.

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

New HOPformer model and EPIC-Contact dataset advance egocentric 3D hand-object pose estimation

COVERAGE [4]

  1. arXiv cs.CV TIER_1 English(EN) · Hui Yang, Wei Sun, Jian Liu, Jin Zheng, Jian Xiao, Ajmal Mian ·

    Occlusion-Aware 3D Hand-Object Pose Estimation with Masked AutoEncoders

    arXiv:2506.10816v2 Announce Type: replace Abstract: Hand-object pose estimation from monocular RGB images remains a significant challenge mainly due to the severe occlusions inherent in hand-object interactions. Existing methods do not sufficiently explore global structural perce…

  2. arXiv cs.CV TIER_1 English(EN) · Hui Yang, Wei Sun, Jian Liu, Jian Xiao, Tao Xie, Hossein Rahmani, Ajmal Saeed Mian, Nicu Sebe, Gim Hee Lee ·

    GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness

    arXiv:2603.19013v4 Announce Type: replace Abstract: Generalized 3D hand-object pose estimation from a single RGB image remains challenging due to the large variations in object appearances and interaction patterns, especially under heavy occlusion. We propose GenHOI, a framework …

  3. arXiv cs.CV TIER_1 English(EN) · Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Jiahe Zhao, Michael J. Black, Dima Damen ·

    Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

    arXiv:2606.30598v1 Announce Type: new Abstract: Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often struggle to generalise to in-the-wild scenes and are l…

  4. arXiv cs.CV TIER_1 English(EN) · Dima Damen ·

    Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

    Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often struggle to generalise to in-the-wild scenes and are limited by the scarcity of supervision. We addres…