Researchers have developed a new label-free test-time training method for GUI grounding, a crucial step for autonomous agents to interpret natural language commands. The approach, called Confidence-Anchored Negative Learning (CANL), leverages confidence patterns in coordinate tokens and prioritizes learning from negative samples over potentially noisy positive ones. CANL-7B demonstrated significant improvements, achieving 92.1% on the ScreenSpot-V2 benchmark and a 8.9% absolute improvement on the more challenging ScreenSpot-Pro dataset. AI
IMPACT This new method could significantly reduce the annotation costs for training autonomous agents, enabling more scalable development of GUI-grounding capabilities.
RANK_REASON Academic paper detailing a new method for GUI grounding. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CANL-7B
- Confidence-Anchored Learning
- Confidence-Anchored Negative Learning
- ScreenSpot-Pro
- ScreenSpot-V2
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →