PulseAugur
EN
LIVE 19:54:49

New GST-Bench benchmark reveals VLM struggles with global spatial awareness in video

Researchers have introduced GST-Bench, a new benchmark designed to evaluate the global spatial awareness of Vision-Language Models (VLMs) using video data. The benchmark, which includes questions derived from over 6,790 minutes of synthetic video, requires models to infer spatial relationships from novel viewpoints and map egocentric observations to global top-down perspectives. Evaluations of 22 state-of-the-art VLMs revealed a significant performance gap compared to humans, with the best-performing model achieving only 42.68% accuracy, far below the human score of 79.08%. Further analysis indicated that while models possess strong local spatial understanding, they struggle to consolidate long-horizon observations into a consistent global scene representation. AI

IMPACT Highlights a critical gap in VLM capabilities, potentially guiding future research towards improved spatial reasoning and scene representation.

RANK_REASON The cluster describes a new academic benchmark and associated dataset for evaluating AI models.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New GST-Bench benchmark reveals VLM struggles with global spatial awareness in video

COVERAGE [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

    Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-…

  2. arXiv cs.CV TIER_1 English(EN) · Qifeng Zhang, Kaixiang Huang, Heng Dong, Huang Fang, Junting Chen, Junjie Zhu, Yonghang Chen, Zhiyu Zhang, Wei Li ·

    GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

    arXiv:2608.05747v1 Announce Type: new Abstract: Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To a…