Vision-Language-Action (VLA)
PulseAugur coverage of Vision-Language-Action (VLA) — every cluster mentioning Vision-Language-Action (VLA) across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New framework aligns VLA driving supervision with policy optimization
Researchers have developed a new framework to improve Vision-Language-Action (VLA) driving methods by aligning multi-trajectory imitation learning with policy optimization. The proposed method addresses issues where hig…
-
New VLA models enhance autonomous driving with multi-expert reasoning and multi-modality interaction
Two new research papers explore advanced Vision-Language-Action (VLA) models for autonomous driving. The first paper, CoWorld-VLA, introduces a multi-expert world reasoning framework that uses specialized tokens to cond…
-
New autonomous driving model learns from past failures with memory augmentation
Researchers have introduced DriveVLA-M0, a novel Vision-Language-Action (VLA) model designed to improve autonomous driving systems by learning from past failures. This model incorporates a failure-aware latent memory th…
-
New WAM-Diff2 framework boosts autonomous driving VLA model efficiency
Researchers have developed WAM-Diff2, a new framework designed to improve the efficiency of Vision-Language-Action (VLA) models for autonomous driving. This framework uses a hierarchical distillation strategy to convert…
-
New LIFT framework enhances VLA policies with reactive force injection
Researchers have developed LIFT (Late Reactive Injection of Force for VLA Post-Training), a new framework designed to enhance the performance of vision-language-action (VLA) policies, particularly in contact-rich manipu…
-
New VLA models enhance robot manipulation with foresight and large-scale data
Researchers have developed new methods to improve the performance of Vision-Language-Action (VLA) models in robotics, particularly for complex, long-horizon tasks. One approach, Foresight Residual RL, enhances credit as…
-
Lift3D-VLA enhances robotic manipulation with 3D geometry and temporal action modeling
Researchers have introduced Lift3D-VLA, a novel framework designed to enhance Vision-Language-Action (VLA) models for robotic manipulation by integrating explicit 3D geometric reasoning and temporal action modeling. The…
-
New frameworks enhance embodied agents for complex manipulation tasks · 2 sources tracked
Two new research papers introduce frameworks for embodied agents to perform long-horizon manipulation tasks. Cortex utilizes a bidirectionally aligned embodied agent framework with a customized planning interface to con…
-
New framework trains AI action models using unlabeled human videos
Researchers have developed a new framework for training Vision-Language-Action (VLA) models using unlabeled human videos. The system, called Motion-Focused Latent Action, employs a Hybrid Disentangled VQ-VAE to separate…
-
MM-Nav: Multi-View VLA Model Enhances Visual Navigation Capabilities
Researchers have developed MM-Nav, a novel multi-view Vision-Language-Action (VLA) model designed for robust visual navigation. This model leverages pretrained large language and visual foundation models, trained in a t…
-
InSight framework enables VLA models to autonomously acquire new manipulation skills
Researchers have developed InSight, a novel framework designed to enhance the skill acquisition capabilities of Vision-Language-Action (VLA) models. This system enables VLAs to learn new manipulation skills autonomously…
-
Pose6DAug framework enhances robot data augmentation for VLA policies · 2 sources tracked
Researchers have developed Pose6DAug, a novel data augmentation framework designed to improve the performance of Vision-Language-Action (VLA) policies in robotics. This method leverages successful robot manipulation epi…
-
New AI models enhance robot manipulation with advanced memory systems · 4 sources tracked
Researchers have introduced two new methods for improving robot manipulation through enhanced memory systems. Mem-World, a memory-augmented multi-view action-conditioned world model, addresses challenges in persistent w…
-
ROVE framework improves humanoid manipulation with imperfect human interventions
Researchers have introduced ROVE, a reinforcement learning framework designed to improve humanoid manipulation by effectively utilizing imperfect human interventions. The system addresses challenges in collecting high-q…
-
New GAM framework enhances embodied AI generalization
Researchers have developed a new framework called Generalized Action Manifold (GAM) to improve generalization in embodied intelligence tasks. GAM enforces general covariance by decoupling spatial path geometry from temp…