PulseAugur
EN
LIVE 11:39:50

New DeVA model enhances robot policy learning with decoupled video-action approach · 2 sources tracked

Researchers have developed DeVA, a new Decoupled Video-Action model designed to improve robot policy learning. DeVA separates video and action prediction into specialized experts, allowing for richer information exchange and more tractable policy adaptation. The model incorporates physically salient guidance, such as affordance and depth, to supervise intermediate video features and the action stream. Experiments show that DeVA achieves strong performance with limited data, converges faster than unified architectures, and demonstrates clear benefits from its physical guidance approach. AI

IMPACT Enhances robot manipulation capabilities by improving policy learning with visual and physical dynamics.

RANK_REASON The cluster describes a new research paper detailing a novel model for robot policy learning.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New DeVA model enhances robot policy learning with decoupled video-action approach · 2 sources tracked

COVERAGE [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

    Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limite…

  2. arXiv cs.CV TIER_1 English(EN) · Mengqi Zhang, Sahil Khose, Simar Kareer, Yuchen Song, Unnat Jain, Judy Hoffman ·

    DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

    arXiv:2607.24159v1 Announce Type: cross Abstract: Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predomin…