RGB-D Visual Simultaneous Localization and Mapping (SLAM) Application
PulseAugur coverage of RGB-D Visual Simultaneous Localization and Mapping (SLAM) Application — every cluster mentioning RGB-D Visual Simultaneous Localization and Mapping (SLAM) Application across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
New perception pipeline aims to automate mining rock-breakers
Researchers have developed a real-time RGB-D perception pipeline to automate the operation of hydraulic impact hammers, commonly known as rock-breakers, in mining. This system integrates image-based instance segmentatio…
-
Seg2Grasp pipeline enhances robotic bin picking with modular approach
Researchers have developed Seg2Grasp, a novel modular pipeline for robust suction grasping in bin picking tasks. This system employs a three-step process: segmentation using a Transformer-based model to create object ma…
-
AlayaWorld advances interactive video world modeling with 720p generation
Researchers have introduced AlayaWorld, an interactive video world model capable of generating 24-fps video at 540p and 720p resolutions. This model utilizes a 15B video diffusion transformer and incorporates several me…
-
DA-Fusion Transformer enhances unseen object segmentation for logistics
Researchers have developed DA-Fusion, a novel Transformer model that uses deformable attention to fuse RGB and depth data for improved unseen object instance segmentation. This advancement is particularly beneficial for…
-
Depth data boosts surgical vision foundation models, study finds
A new study explored the impact of incorporating depth information into vision foundation models for surgical applications. The research found that models pre-trained with RGB-D data, such as MultiMAE, significantly out…
-
New TACTIC controller enhances whole-arm manipulation with tactile and vision data
Researchers have developed TACTIC, a new controller designed for whole-arm manipulation tasks that involve complex contact dynamics. This system integrates RGB-D vision, distributed tactile sensing, and a proximity repr…
-
GenVid2Robot framework translates generated video motion into executable robot trajectories
Researchers have developed GenVid2Robot, a framework that translates generated video motion into executable robot manipulation trajectories. This system addresses the limitations of using generated videos directly for r…
-
SplatCtrl framework enables reactive robot control in dynamic environments
Researchers have developed SplatCtrl, a new framework that integrates real-time scene reconstruction with reactive robot motion generation. This system uses 3D Gaussian Splatting to efficiently build 3D models of dynami…
-
Utonia: Unified 3D Point Cloud Encoder Advances Perception and Reasoning
Researchers have introduced Utonia, a novel self-supervised point transformer encoder designed to process diverse 3D point cloud data from various domains. This unified approach aims to create a single model capable of …
-
Image2Sim framework generates realistic 3D environments for AI navigation training
Researchers have developed Image2Sim, a novel framework for creating realistic and interactive 3D environments for embodied navigation training. This system leverages decoupled 3D spatial anchoring and photorealistic re…
-
New LINet architecture enables continuous cross-modal learning in RGB-D scene classification
Researchers have introduced LINet, a novel Multi-Stream Neural Network (MSNN) designed for RGB-D scene classification. Unlike existing architectures that fuse features discretely, LINet employs a continuous integration …
-
VCS-SLAM enhances semantic 3D Gaussian SLAM with geometry validation
Researchers have developed VCS-SLAM, a novel framework designed to enhance the accuracy and consistency of semantic 3D Gaussian SLAM systems. This new approach addresses limitations in current methods that often fuse 2D…
-
New SWAM model enables efficient embodied navigation with single-pass RGB input
Researchers have developed SWAM (Spatial-perceiving World Action Model), a novel framework for embodied navigation that jointly generates intermediate visual sequences and action trajectories in a single pass. Unlike pr…
-
New AISPO framework boosts robotic depth reliability for challenging objects
Researchers have developed AISPO, a novel depth completion framework designed to enhance depth reliability for robotic manipulation, particularly with challenging non-Lambertian objects like transparent or specular surf…
-
UniRED framework unifies RGB-D video interpolation with event guidance
Researchers have developed UniRED, a novel framework for interpolating RGB-D videos by integrating RGB appearance, depth geometry, and event-based temporal cues. This approach addresses limitations in existing methods t…
-
New AI Methods Enhance Point Cloud Registration for Robotics and Surgery
Two new research papers explore advanced techniques for point cloud registration. The first, Generalized-CVO, uses Riemannian optimization to achieve up to a 10x speedup over previous methods for LiDAR and RGB-D data, s…
-
New GPS Representation Enhances Robotic Manipulation with VR Data
Researchers have introduced a new geometric representation called Geometric Primary Structure (GPS) for improving robotic manipulation by perceiving articulated parts. This method utilizes virtual reality for efficient …
-
Con-DSO improves RGB-D odometry with learned consistency priors
Researchers have developed Con-DSO, a novel RGB-D direct sparse odometry framework designed to improve accuracy in challenging environments. This system learns to predict pixel-level uncertainty in photometric and depth…
-
VLM pipeline enables viewpoint-agnostic grasping for robots with partial observations
Researchers have developed a new end-to-end pipeline for language-guided grasping that enhances the robustness of mobile manipulators in cluttered environments. This system uses visual-language models (VLMs) and partial…
-
CLAMP framework uses contrastive learning for 3D robotic manipulation pretraining
Researchers have developed CLAMP, a new pre-training framework for robotic manipulation that leverages 3D multi-view image data and robot actions. CLAMP uses contrastive learning on simulated trajectories to associate g…