Vision-language-action model
PulseAugur coverage of Vision-language-action model — every cluster mentioning Vision-language-action model across labs, papers, and developer communities, ranked by signal.
12 day(s) with sentiment data
-
New benchmark and VLA model advance cooperative autonomous driving research
Researchers have introduced CMU-Drive, a new benchmark for evaluating cooperative autonomous driving among multiple connected vehicles. Alongside this benchmark, they propose V2V-VLA, a vision-language-action model desi…
-
New adaptive VLA framework enhances embodied intelligence with environment-aware model selection
Researchers have developed a new framework called Environment-aware Model Selection (EMS) for embodied intelligence, which adaptively switches between two distinct Vision-Language-Action (VLA) systems. This approach dec…
-
New SkillMemo framework enhances robotic manipulation generalization
Researchers have developed SkillMemo, a novel framework designed to improve the compositional generalization of embodied visuomotor models in robotics. This framework addresses the limitations of current models, which a…
-
New RESample framework improves robotic manipulation with failure recovery data
Researchers have developed RESample, a novel data augmentation framework designed to improve the performance of Vision-Language-Action (VLA) models in robotic manipulation tasks. This framework addresses the issue of di…
-
DyPES-VLA model enhances robot manipulation across diverse embodiments
Researchers have introduced DyPES-VLA, a novel Vision-Language-Action (VLA) model designed to improve robot manipulation across different embodiments. The model addresses limitations in current VLA approaches by learnin…
-
New VLM techniques enhance autonomous driving reasoning and efficiency
Researchers are developing new methods for vision-language models (VLMs) used in autonomous driving to improve reasoning and reduce hallucinations. One approach, DEFT-RLVR, addresses trajectory anchoring bias by making …
-
New VLAGuard framework enhances robot defense against physical attention hijacking
Researchers have developed VLAGuard, a framework designed to protect Vision-Language-Action (VLA) robots operating as mobile edge nodes in wireless sensor networks from physical adversarial attacks. The framework includ…
-
Google DeepMind unveils Gemini Robotics 2 for adaptable robots
Google DeepMind has introduced Gemini Robotics 2, an intelligence layer designed to power adaptable robots. This new system includes an embodied reasoning model, Gemini Robotics ER 2, which acts as the robot's high-leve…
-
New DeVA model enhances robot policy learning with decoupled video-action approach · 2 sources tracked
Researchers have developed DeVA, a new Decoupled Video-Action model designed to improve robot policy learning. DeVA separates video and action prediction into specialized experts, allowing for richer information exchang…
-
New benchmark MulRobBench tests safe and secure decision-making for UAV agents
Researchers have introduced MulRobBench, a new benchmark designed to evaluate multimodal Uncrewed Aerial Vehicle (UAV) agents in smart-city environments. This benchmark focuses on decision-making, integrating real-world…
-
Action QFormer enhances VLA models by shaping multimodal representations
Researchers have introduced Action QFormer, a novel approach to enhance vision-language-action (VLA) models by treating action supervision as a representation shaping mechanism rather than a mere downstream objective. T…
-
RoboTTT scales robot policy context to 8K timesteps for enhanced imitation and long-horizon tasks · 3 sources tracked
Researchers have developed RoboTTT, a novel training methodology and robot model that significantly expands the visuomotor context window to 8,000 timesteps. This advancement enables robots to perform one-shot imitation…
-
Robotics VLA models improved with semantic anchoring technique · 2 sources tracked
Researchers have developed a new method called "Semantic Anchoring" to improve the performance of Vision-Language-Action (VLA) models in robotics. These models often lose their rich semantic understanding when fine-tune…
-
New models RxBrain and GTA-VLA advance embodied AI reasoning
Researchers have introduced RxBrain, a novel foundation model for embodied cognition that integrates language and visual reasoning for planning. Unlike existing models that focus on scene understanding or future state p…
-
AI Agents: Defining the Future of Robotics with Vision-Language-Action Models
An AI agent is defined by its ability to perceive, decide, act, and pursue a goal, distinguishing it from a simple chatbot through the use of tools and a feedback loop. This framework is crucial for developing robots ca…
-
Nomagic deploys AI 'brain' for warehouse robots, halving intervention rates
Nomagic has successfully deployed its first vision-language-action (VLA) model to paying customers, marking a significant step in applying AI to real-world robotic tasks. This AI 'brain' for warehouse robots, led by for…
-
SIEVE method enhances VLA imitation learning with structure-aware data selection
Researchers have introduced SIEVE, a novel method for selecting data in vision-language-action (VLA) imitation learning. SIEVE identifies reusable visuo-motor primitives and transition interfaces within robot demonstrat…
-
New CamVLA Model Adapts to Unseen Camera Views Without Calibration
Researchers have developed a new Vision-Language-Action (VLA) model called CamVLA that can adapt to varying camera positions without explicit calibration. This model decouples manipulation controls from camera geometry …
-
New VLA Model Achieves Calibration-Free Robot Control
Researchers have developed a new Vision-Language-Action (VLA) model called Camera-Centric VLA (CamVLA) that can operate without explicit camera calibration. This model predicts camera-centric actions and a hand-eye matr…
-
New models enhance robot manipulation by integrating vision and state
Researchers have developed several new methods to improve robot manipulation capabilities by better integrating visual information with the robot's state and actions. GeoProp, for instance, is a lightweight adapter that…