Researchers have developed DeVA, a new Decoupled Video-Action model designed to improve robot policy learning. DeVA separates video and action prediction into specialized experts, allowing for richer information exchange and more tractable policy adaptation. The model incorporates physically salient guidance, such as affordance and depth, to supervise intermediate video features and the action stream. Experiments show that DeVA achieves strong performance with limited data, converges faster than unified architectures, and demonstrates clear benefits from its physical guidance approach. AI
IMPACT Enhances robot manipulation capabilities by improving policy learning with visual and physical dynamics.
RANK_REASON The cluster describes a new research paper detailing a novel model for robot policy learning.
Read on Hugging Face Daily Papers →
- arXiv
- DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
- Hugging Face
- Vision-Language-Action model
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →