ScanQA
PulseAugur coverage of ScanQA — every cluster mentioning ScanQA across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
SparseTalk cuts 3D Gaussian language field costs for VQA
Researchers have developed SparseTalk, a method to significantly reduce the storage and computational costs associated with 3D Gaussian language fields used in 3D visual question answering (VQA). By systematically spars…
-
New research probes LLM spatial reasoning with 3D scene graphs and 2D benchmarks
Two new research papers explore spatial reasoning capabilities in large language models. The first, "GraFT," introduces a training-free framework that uses 3D scene graphs to enhance multimodal LLMs' geometric understan…
-
ViewMind3D framework enables training-free 3D question answering
Researchers have introduced ViewMind3D, a novel framework designed for training-free 3D question answering using multi-view observations. This modular system bypasses the need for costly 3D-specific training by breaking…
-
New framework enhances LLMs' 3D spatial reasoning from sparse inputs
Researchers have developed SpaR3D-MoE, a novel framework designed to enhance the 3D spatial reasoning capabilities of Multimodal Large Language Models (MLLMs) using only sparse RGB inputs. The system employs an adaptive…
-
New method prunes tokens for efficient 3D question answering
Researchers have developed a novel online token-pruning method designed to enhance the efficiency of multi-modal large language models (MLLMs) in 3D question answering tasks. This approach projects input frames into a s…
-
Chat-Scene++ advances 3D LLM scene understanding with context-rich object identification
Researchers have introduced Chat-Scene++, a novel framework designed to enhance multi-modal large language models (MLLMs) for 3D scene understanding. This approach structures 3D scenes as sequences of objects, incorpora…