Vision-language-action model
PulseAugur coverage of Vision-language-action model — every cluster mentioning Vision-language-action model across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
New paper categorizes Generative Physical AI approaches for robotics
A new arXiv paper provides a comprehensive review of Generative Physical Artificial Intelligence (GPAI), a field that integrates large foundation models with physical robots. The paper categorizes GPAI systems into five…
-
New research enhances robot manipulation with temporal context and visual foresight · 4 sources tracked
Researchers are developing new methods to improve robot manipulation by incorporating temporal context and visual foresight. PACT-WAM uses compact temporal encoding to predict action trajectories and visual outcomes, ac…
-
New frameworks enhance robot policies with flow matching and safety constraints · 4 sources tracked
Researchers have developed several new frameworks for improving flow-based policies in reinforcement learning, particularly for robotics. These methods aim to address challenges like multimodal action distributions and …
-
MobileVLA-R1 2.0 enhances robot instruction following with reasoning and RL
Researchers have introduced MobileVLA-R1 2.0, a new framework designed to enhance the ability of mobile robots to follow complex, long-horizon instructions. This system integrates structured reasoning with reinforcement…
-
New EGR method boosts robot policy robustness against sensor issues
Researchers have developed a new training objective called Evidence-Gated Regularization (EGR) to improve the robustness of Vision-Language-Action (VLA) policies in robotics. This method addresses the issue of modality …
-
New RoboSPA benchmark tests VLA models on complex robotic reasoning
Researchers have introduced RoboSPA, a new dataset and benchmark designed to evaluate the embodied reasoning capabilities of Vision-Language-Action (VLA) models in robotics. RoboSPA focuses on fine-grained spatial reaso…
-
New RedLight-VLA model enhances AI driving policies at intersections
Researchers have developed RedLight-VLA, a novel training objective designed to improve the performance of Vision-Language-Action (VLA) driving policies, particularly in complex scenarios like signalized intersections. …
-
New latent-space reasoning model enhances end-to-end autonomous driving
Researchers have introduced Latent Chain-of-Thought-Drive (LCDrive), a novel approach for end-to-end autonomous driving that utilizes a latent language for reasoning instead of natural language. This model interleaves a…
-
New research explores parallel drafting for speculative decoding in LLMs
Two new research papers explore advancements in speculative decoding for large language models, focusing on improving efficiency and coherence in parallel drafting. The first paper surveys the applicability of block-par…
-
New methods enhance VLA policy efficiency and robustness in robotics · 4 sources tracked
Researchers have developed new methods to improve the efficiency and robustness of vision-language-action (VLA) policies in robotics. One approach, EXIMO, uses a vision-language model (VLM) as a planner to break down co…
-
SpecVLA framework enhances VLA model efficiency for embodied AI
Researchers have developed SpecVLA, a novel framework for co-designing algorithms and hardware architectures to improve the efficiency of Vision-Language-Action (VLA) models in embodied AI. This approach leverages the o…
-
G0.5 model integrates robot reasoning and action in single stream
Researchers have introduced G0.5, a novel autoregressive Vision-Language-Action (VLA) model that integrates reasoning and action generation within a single Transformer decoder. This approach allows the VLM to act as a d…
-
New benchmark and VLA model advance cooperative autonomous driving research
Researchers have introduced CMU-Drive, a new benchmark for evaluating cooperative autonomous driving among multiple connected vehicles. Alongside this benchmark, they propose V2V-VLA, a vision-language-action model desi…
-
New adaptive VLA framework enhances embodied intelligence with environment-aware model selection
Researchers have developed a new framework called Environment-aware Model Selection (EMS) for embodied intelligence, which adaptively switches between two distinct Vision-Language-Action (VLA) systems. This approach dec…
-
New SkillMemo framework enhances robotic manipulation generalization
Researchers have developed SkillMemo, a novel framework designed to improve the compositional generalization of embodied visuomotor models in robotics. This framework addresses the limitations of current models, which a…
-
New RESample framework improves robotic manipulation with failure recovery data
Researchers have developed RESample, a novel data augmentation framework designed to improve the performance of Vision-Language-Action (VLA) models in robotic manipulation tasks. This framework addresses the issue of di…
-
DyPES-VLA model enhances robot manipulation across diverse embodiments
Researchers have introduced DyPES-VLA, a novel Vision-Language-Action (VLA) model designed to improve robot manipulation across different embodiments. The model addresses limitations in current VLA approaches by learnin…
-
New VLM techniques enhance autonomous driving reasoning and efficiency
Researchers are developing new methods for vision-language models (VLMs) used in autonomous driving to improve reasoning and reduce hallucinations. One approach, DEFT-RLVR, addresses trajectory anchoring bias by making …
-
New VLAGuard framework enhances robot defense against physical attention hijacking
Researchers have developed VLAGuard, a framework designed to protect Vision-Language-Action (VLA) robots operating as mobile edge nodes in wireless sensor networks from physical adversarial attacks. The framework includ…
-
Google DeepMind unveils Gemini Robotics 2 for adaptable robots
Google DeepMind has introduced Gemini Robotics 2, an intelligence layer designed to power adaptable robots. This new system includes an embodied reasoning model, Gemini Robotics ER 2, which acts as the robot's high-leve…