PulseAugur
EN
LIVE 14:47:04

Robots learn and improve policies using visual and tactile feedback · 5 sources tracked

Researchers have developed new frameworks for improving robot policy performance through inference-time steering and self-improvement. VERITAS, a generator-verifier framework, uses a pre-trained policy and a visual verifier to steer actions without additional training, achieving performance gains comparable to expert demonstrations. ViTaL enhances this by incorporating tactile feedback alongside visual data for contact-rich manipulation tasks, significantly improving success rates. Additionally, Visual-OPSD and ViGOS explore on-policy self-distillation techniques for multimodal large language models, decoupling perception and reasoning to improve grounded behavior and reduce inference costs. AI

IMPACT These advancements could lead to more adaptable and efficient AI systems in robotics and multimodal reasoning, reducing reliance on human intervention and computational costs.

RANK_REASON Multiple arXiv papers detailing new research frameworks for AI and robotics.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 8 sources. How we write summaries →

Robots learn and improve policies using visual and tactile feedback · 5 sources tracked

COVERAGE [8]

  1. arXiv cs.LG TIER_1 English(EN) · Sihan Wang, Xiyao Liu, Lianqing Liu, Zhi Han ·

    Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

    arXiv:2606.19120v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target. This works well for LLM reasoning, but a direct extension to multim…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

    Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference cost from multi-step diffusion. We find this cost yields limited direct benefit. On ThinkMorph, removi…

  3. arXiv cs.AI TIER_1 English(EN) · Mingtong Zhang, Dhruv Shah ·

    Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement

    arXiv:2606.18247v1 Announce Type: cross Abstract: Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism of practicing and learning from feedback. In this paper, we propose VERITAS, a generator-verifier framework for …

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

    ViGOS is a visually grounded on-policy self-distillation framework for multimodal large language models that improves image-grounded behavior by using specialized teachers for different stages of reasoning and handling invalid rollouts.

  5. arXiv cs.AI TIER_1 English(EN) · Dhruv Shah ·

    Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement

    Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism of practicing and learning from feedback. In this paper, we propose VERITAS, a generator-verifier framework for generalist robot policies for inference-time polic…

  6. arXiv cs.AI TIER_1 English(EN) · Yilin Wu, Zilin Si, Zeynep Temel, Oliver Kroemer, Andrea Bajcsy ·

    Inference-time Policy Steering via Vision and Touch

    arXiv:2606.14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution. While prior methods typically perform this verification only with visual observations, vision …

  7. arXiv cs.CV TIER_1 English(EN) · Zhi Han ·

    Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

    On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target. This works well for LLM reasoning, but a direct extension to multimodal large language models (MLLMs) can create a …

  8. arXiv cs.CV TIER_1 English(EN) · Jun Liu ·

    Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

    Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference cost from multi-step diffusion. We find this cost yields limited direct benefit. On ThinkMorph, removi…