PulseAugur
实时 09:50:01

New HOPformer model and EPIC-Contact dataset advance egocentric 3D hand-object pose estimation

研究人员推出了一种新颖的端到端 Transformer 模型 HOPformer,用于自主式 3D 手部-物体姿态估计。该模型通过一次性预测手部和物体姿态,并利用交叉注意力将物体特征条件化于手部先验信息,从而提高了鲁棒性。为了促进这项研究,该团队还发布了 EPIC-Contact,这是一个包含 2.3K 个自主式真实场景片段的新数据集,具有详细的 3D 手部-物体接触对应关系和姿态网格。HOPformer 在现有数据集和新数据集上均展现出显著的性能提升,在 EPIC-Contact 上的成功率几乎翻倍。 AI

影响 推动了自主式 3D 计算机视觉能力的发展,可能改进机器人和增强现实应用。

排序理由 该集群描述了一篇关于针对特定计算机视觉任务的新模型和数据集的最新研究论文。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

New HOPformer model and EPIC-Contact dataset advance egocentric 3D hand-object pose estimation

报道来源 [4]

  1. arXiv cs.CV TIER_1 English(EN) · Hui Yang, Wei Sun, Jian Liu, Jin Zheng, Jian Xiao, Ajmal Mian ·

    具有遮挡感知的掩码自编码器三维手部-物体姿态估计

    arXiv:2506.10816v2 Announce Type: replace Abstract: Hand-object pose estimation from monocular RGB images remains a significant challenge mainly due to the severe occlusions inherent in hand-object interactions. Existing methods do not sufficiently explore global structural perce…

  2. arXiv cs.CV TIER_1 English(EN) · Hui Yang, Wei Sun, Jian Liu, Jian Xiao, Tao Xie, Hossein Rahmani, Ajmal Saeed Mian, Nicu Sebe, Gim Hee Lee ·

    GenHOI:具有遮挡感知的通用手部-物体姿态估计

    arXiv:2603.19013v4 Announce Type: replace Abstract: Generalized 3D hand-object pose estimation from a single RGB image remains challenging due to the large variations in object appearances and interaction patterns, especially under heavy occlusion. We propose GenHOI, a framework …

  3. arXiv cs.CV TIER_1 English(EN) · Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Jiahe Zhao, Michael J. Black, Dima Damen ·

    迈向野外以自我为中心的3D手部-物体姿态估计

    arXiv:2606.30598v1 Announce Type: new Abstract: Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often struggle to generalise to in-the-wild scenes and are l…

  4. arXiv cs.CV TIER_1 English(EN) · Dima Damen ·

    迈向野外以自我为中心的3D手部-物体姿态估计

    Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often struggle to generalise to in-the-wild scenes and are limited by the scarcity of supervision. We addres…