Researchers have introduced ReRef-3D, a new benchmark designed to evaluate language-guided object placement in 3D environments. The benchmark comprises over 33,000 instructions across nearly 1,000 scenes, focusing on various referencing complexities. Initial evaluations show that LLaVA-3D achieved the highest performance, correctly placing objects in 68.3% of cases, significantly outperforming 3D-LLM and PlaceIt3D. The study also found that relational difficulties, such as 'nearest' or 'between,' pose the greatest challenge for current models. AI
IMPACT Establishes a new standard for evaluating spatial reasoning and object manipulation in AI models within 3D environments.
RANK_REASON Publication of a new benchmark and research paper on arXiv.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →