Researchers have explored methods for small GUI grounding models to effectively receive action types, comparing five different approaches. Fine-tuning the Qwen2-VL-2B model using LoRA on the Android in the Wild dataset, they found that auxiliary loss, additive learned embeddings, and prompt-based action words significantly improved performance over a baseline. However, hard routing and prepended tokens showed no discernible benefit, and a preprocessing choice to clamp off-screen touch points negatively impacted grounding accuracy. AI
IMPACT This research offers insights into improving the efficiency and accuracy of small AI models for GUI interaction, potentially impacting the development of more capable assistive technologies.
RANK_REASON Academic paper detailing a novel approach to model training and evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →