This paper introduces a new pipeline for recognizing objects indicated by human pointing gestures in RGB images. The system combines object detection, pose estimation, depth estimation, and vision-language models to identify targets of non-verbal communication. Experiments show that incorporating 3D spatial information from a single image enhances target identification, particularly in cluttered scenes, and that image captioning can help correct classification errors. AI
IMPACT This research could lead to more intuitive human-robot interaction by enabling robots to understand pointing gestures.
RANK_REASON This is a research paper detailing a new technical approach. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →