Researchers have developed a novel active learning framework to improve the efficiency of collecting annotated data for referring image segmentation and grounding tasks. This method addresses the bottleneck of annotators needing to write descriptive text by focusing on images with ambiguous regions. The framework utilizes foundation models to generate auxiliary text and introduces a new acquisition function, Referred Region Ambiguity, to identify informative samples. Experiments on RIS and REC benchmarks demonstrate its superiority over existing active learning baselines, and a user study indicated a 1.6x speedup in description labeling. AI
IMPACT This research could significantly reduce the cost and time associated with data annotation for visual grounding tasks, potentially accelerating the development of AI systems that understand and interact with images.
RANK_REASON The cluster contains a research paper detailing a new method for active learning in computer vision. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →