ScanRefer
PulseAugur coverage of ScanRefer — every cluster mentioning ScanRefer across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New framework GUIDE enhances MLLMs with progressive geometric integration
Researchers have developed GUIDE (Geometric Unrolling Inside MLLM Early-layers), a novel framework designed to enhance Multimodal Large Language Models (MLLMs) in understanding physical space and 3D scenes. Unlike previ…
-
New frameworks enhance 3D visual grounding with LLMs and VLMs · 2 sources tracked
Researchers have developed two new frameworks, TDVR and GuideGround, to improve zero-shot 3D visual grounding. TDVR addresses challenges of ambiguous text and missing viewpoints by using LLMs for text disambiguation and…
-
OpenGround framework enhances 3D visual grounding with planning and online perception
Researchers have introduced OpenGround, a novel framework designed for open-world 3D visual grounding. This system addresses limitations in current methods by integrating Task-Chain Planning to break down complex querie…
-
PVCap enhances 3D dense captioning with new data augmentation and network architecture
Researchers have introduced PVCap, a novel approach to enhance 3D dense captioning, a task focused on generating descriptions for objects within 3D scenes. The method addresses limitations in existing techniques by inco…
-
PruneGround framework enhances 3D visual grounding with spatial pruning
Researchers have introduced PruneGround, a novel framework designed to improve 3D Visual Grounding by focusing on language-relevant regions within 3D scenes. This approach utilizes Language-Guided Spatial Pruning (LGSP)…
-
New multi-agent framework boosts zero-shot 3D understanding · 2 sources tracked
Researchers have introduced a novel collaborative multi-agent framework for zero-shot 3D understanding, addressing limitations in existing video-based methods. The system employs a Planning Agent to strategically select…
-
AgentGrounder enables zero-shot 3D visual grounding on point clouds
Researchers have introduced AgentGrounder, a novel framework for zero-shot 3D visual grounding that operates directly on colored point clouds. This approach bypasses the need for task-specific 3D training by employing a…
-
SceneGraphGrounder uses 3D scene graphs for zero-shot visual grounding
Researchers have introduced SceneGraphGrounder, a novel framework designed for zero-shot 3D visual grounding. This approach tackles the challenge of locating objects in unstructured environments using natural language b…
-
New frameworks MCM-VG and DEGround advance zero-shot 3D visual grounding
Researchers have developed two new frameworks, DEGround and MCM-VG, to improve ego-centric 3D visual grounding, a key task for embodied intelligence. DEGround utilizes a homogeneous pipeline that shares object represent…
-
Chat-Scene++ advances 3D LLM scene understanding with context-rich object identification
Researchers have introduced Chat-Scene++, a novel framework designed to enhance multi-modal large language models (MLLMs) for 3D scene understanding. This approach structures 3D scenes as sequences of objects, incorpora…