Researchers have developed Hi-Token, a novel method for generative visual grounding that improves the accuracy of bounding-box predictions by tokenizing coordinates hierarchically. This approach encodes digits for hundreds, tens, and ones, adding structure and promoting token reuse within vision-language models. Complementary to Hi-Token, Hi-GAR uses a geometry-based reward during training to further refine localization accuracy. Experiments across multiple vision-language model backbones and benchmarks demonstrate consistent performance gains, with Hi-R1 achieving superior results compared to existing specialist methods. AI
IMPACT Improves localization accuracy in AI models for visual grounding tasks, potentially enhancing applications that require precise object identification from text descriptions.
RANK_REASON The cluster contains two arXiv papers detailing new research and surveys in the field of visual grounding.
- arXiv
- Generative Visual Grounding
- Group Relative Policy Optimization
- GRPO
- Hi-GAR
- Hi-R1
- Hi-Token
- Linhui Xiao
- RefCOCO+
- vision-language model
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →