PulseAugur
EN
LIVE 10:03:05

ProxyPose uses video-to-video translation for 6-DoF pose tracking

Researchers have developed ProxyPose, a novel method for tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video. This approach reframes the problem as a video-to-video translation task, utilizing a fine-tuned video diffusion model to generate a synthetic proxy video. By analyzing the motion of a known object within this proxy video, ProxyPose can accurately recover the 6-DoF trajectory of the original surface. This technique bypasses the need for additional inputs like 3D models or depth maps, demonstrating state-of-the-art performance and extending to applications such as face tracking and camera pose estimation. AI

IMPACT This research could advance computer vision applications by enabling more robust and input-agnostic object and surface pose tracking.

RANK_REASON The cluster describes a new research paper published on arXiv detailing a novel method for pose tracking.

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

ProxyPose uses video-to-video translation for 6-DoF pose tracking

COVERAGE [2]

  1. arXiv cs.CV TIER_1 English(EN) · Ruihang Zhang, Felix Taubner, Pooja Ravi, Kiriakos N. Kutulakos, David B. Lindell ·

    ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

    arXiv:2607.06555v1 Announce Type: new Abstract: Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself-such as 3D m…

  2. arXiv cs.CV TIER_1 English(EN) · David B. Lindell ·

    ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

    Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself-such as 3D models, depth maps, object masks, or task-specifi…