PulseAugur
EN
LIVE 23:16:04
ENTITY VSI-Bench

VSI-Bench

PulseAugur coverage of VSI-Bench — every cluster mentioning VSI-Bench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
3
10 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
9 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

3 day(s) with sentiment data

RECENT · PAGE 1/1 · 18 TOTAL
  1. TOOL · CL_242828 ·

    Tsinghua, Tencent, NTU unveil Spatial-TTT for AI spatial memory

    Researchers from Tsinghua University, Tencent Hunyuan, and Nanyang Technological University have developed Spatial-TTT, a novel approach to endow AI models with "streaming spatial memory." This method addresses the limi…

  2. FRONTIER RELEASE · CL_230690 ·

    World Labs unveils Atlas, an omni world model for spatial intelligence · 8 sources tracked

    World Labs has introduced Atlas, a new AI model designed for spatial intelligence that can generate, reconstruct, and simulate 3D worlds from various inputs including images, video, text, and depth data. Unlike speciali…

  3. RESEARCH · CL_228744 ·

    New research probes LLM spatial reasoning with 3D scene graphs and 2D benchmarks

    Two new research papers explore spatial reasoning capabilities in large language models. The first, "GraFT," introduces a training-free framework that uses 3D scene graphs to enhance multimodal LLMs' geometric understan…

  4. RESEARCH · CL_216149 ·

    OraRL framework boosts video MLLM training efficiency

    Researchers have introduced OraRL, a novel reinforcement learning framework designed to enhance the training of video multimodal large language models (MLLMs). This method improves sample efficiency and scalability by t…

  5. RESEARCH · CL_206606 ·

    New frameworks tackle video reasoning challenges in large vision-language models

    Researchers have developed new frameworks to address the challenges in video reasoning for large vision-language models (LVLMs). One approach, the Chain of Evidence (CoE), decouples grounding and reasoning to improve ef…

  6. TOOL · CL_200301 ·

    New Spa3R framework boosts 3D spatial reasoning in vision-language models

    Researchers have developed Spa3R, a novel self-supervised framework designed to enhance 3D spatial reasoning in vision-language models. Unlike existing methods that rely on explicit 3D data or partial geometric priors, …

  7. TOOL · CL_194152 ·

    New framework GUIDE enhances MLLMs with progressive geometric integration

    Researchers have developed GUIDE (Geometric Unrolling Inside MLLM Early-layers), a novel framework designed to enhance Multimodal Large Language Models (MLLMs) in understanding physical space and 3D scenes. Unlike previ…

  8. RESEARCH · CL_154642 ·

    ConsiSpace framework boosts video spatial reasoning in LLMs

    Researchers have introduced ConsiSpace, a novel framework designed to enhance video spatial reasoning capabilities in multimodal large language models (MLLMs). This framework addresses the current semantic-centric limit…

  9. TOOL · CL_153707 ·

    RynnBrain 1.1 embodied models outperform existing systems on benchmarks

    Researchers have introduced RynnBrain 1.1, a new family of embodied foundation models available in 2B, 9B, and 122B-A10B scales. This model family is designed for robots, supporting perception, spatial reasoning, locali…

  10. TOOL · CL_133506 ·

    New framework enhances LLMs' 3D spatial reasoning from sparse inputs

    Researchers have developed SpaR3D-MoE, a novel framework designed to enhance the 3D spatial reasoning capabilities of Multimodal Large Language Models (MLLMs) using only sparse RGB inputs. The system employs an adaptive…

  11. RESEARCH · CL_105024 ·

    New framework DR-MV3D enhances 3D visual question answering with dense rewards

    Researchers have introduced DR-MV3D, a novel framework designed to enhance multi-view 3D visual question answering (MV3D-VQA). This approach utilizes dense, verifiable rewards to supervise the reasoning process, moving …

  12. RESEARCH · CL_97820 ·

    OneCanvas simplifies 3D scene understanding for VLMs with panoramic reprojection

    Researchers have developed OneCanvas, a novel approach to 3D scene understanding for vision-language models (VLMs). Instead of complex geometry encoders or extensive training, OneCanvas projects patch features onto a si…

  13. TOOL · CL_79746 ·

    New framework AlloSpatial boosts foundation model spatial reasoning

    Researchers have introduced AlloSpatial, a new framework designed to enhance the spatial reasoning capabilities of foundation models. This framework converts egocentric observations into structured allocentric represent…

  14. TOOL · CL_92090 ·

    New AlloSpatial Framework Boosts AI Spatial Reasoning

    Researchers have developed AlloSpatial, a new framework designed to improve the spatial reasoning capabilities of foundation models. This framework addresses the limitation of current models by converting egocentric obs…

  15. RESEARCH · CL_44057 ·

    Cambrian-P video model uses camera pose for improved spatial reasoning

    Researchers have introduced Cambrian-P, a novel video multimodal large language model (MLLM) that incorporates camera pose information. This approach treats video frames not as isolated images but as part of a continuou…

  16. RESEARCH · CL_14362 ·

    GeoThinker framework actively integrates geometry for advanced spatial reasoning

    Researchers have developed GeoThinker, a novel framework that enhances spatial reasoning in multimodal large language models (MLLMs) by actively integrating geometric information. Unlike previous passive fusion methods,…

  17. RESEARCH · CL_06186 ·

    VLMs tackle visual illusions, spatial reasoning, and evaluation benchmarks

    Researchers are developing new methods to improve the robustness and reasoning capabilities of Vision-Language Models (VLMs). One approach, Structured Qualitative Inference (SQI), aims to mitigate visual illusions by en…

  18. RESEARCH · CL_02944 ·

    New frameworks enhance VLM spatial reasoning with world models and multi-agent systems

    Researchers have developed World2VLM, a novel training framework that distills spatial reasoning capabilities from generative world models into vision-language models (VLMs). This approach synthesizes future views to pr…