ScreenSpot-Pro
PulseAugur coverage of ScreenSpot-Pro — every cluster mentioning ScreenSpot-Pro across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New CA-OPD framework improves vision-language models with confidence-aware distillation
Researchers have developed a new framework called Confidence-Aware On-Policy Distillation (CA-OPD) to improve autoregressive vision-language models. This method addresses compounding errors by using teacher confidence t…
-
New benchmarks probe VLM spatial reasoning, revealing localization and relation understanding gaps · 5 sources tracked
Researchers are developing new benchmarks and methodologies to better understand and diagnose spatial reasoning failures in vision-language models (VLMs). One approach, GUI-Primitives, uses contrastive instruction pairs…
-
New AI methods enhance GUI grounding with self-evolution and reflection · 4 sources tracked
Researchers are developing advanced methods for GUI visual grounding, enabling AI agents to better interact with graphical user interfaces. One approach, Test-Time Self-Evolving GUI Visual Grounding, uses a closed-loop …
-
New GUI-AIMA framework enhances multimodal LLM grounding capabilities
Researchers have developed GUI-AIMA, a novel framework for improving graphical user interface (GUI) grounding in multimodal large language models (MLLMs). This attention-based approach aligns intrinsic multimodal attent…
-
New methods enhance VLM accuracy for GUI grounding tasks · 2 papers
Two new research papers introduce novel methods for improving the accuracy and reliability of vision-language models (VLMs) in GUI grounding tasks. The first paper, "Trust the Right Teacher," proposes quality-aware self…
-
New methods BAMI and AutoFocus improve GUI grounding for AI agents
Researchers have developed two new training-free methods, BAMI and AutoFocus, to improve the accuracy of GUI grounding for AI agents. BAMI addresses precision and ambiguity biases by using coarse-to-fine focus and candi…
-
New method corrects MLLM coordinate prediction bias from positional encoding failures
Researchers have developed a new method called Vision-PE Shuffle Guidance (VPSG) to address inaccuracies in coordinate prediction within multimodal large language models (MLLMs). These models often struggle with precise…