GUI agents
PulseAugur coverage of GUI agents — every cluster mentioning GUI agents across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New Gated Hindsight Distillation enhances GUI agent training
Researchers have developed a new training technique called Gated Hindsight Distillation (GHD) to improve the performance of GUI agents. GHD utilizes future screenshots as privileged information during training, allowing…
-
New Android GUI Agent Vulnerability Exploits Multimodal Model Weaknesses
Researchers have identified a novel security vulnerability in Android GUI agents powered by large multimodal models. These agents, designed to perceive screen content and inject inputs, are susceptible to "Action Rebind…
-
New red-teaming method exploits GUI agent vulnerabilities
Researchers have developed a new black-box red-teaming method called Semantic-level UI Element Injection to test the robustness of GUI agents. This technique overlays harmless UI elements onto screenshots to misdirect a…
-
New metric reveals GUI agents prioritize structure over pixels
Researchers have developed a new metric called the Perception-Fusion Gap to diagnose how multimodal GUI agents form beliefs about their interface state. This metric measures the extent to which an agent relies on visual…
-
New GUIDE framework reduces domain bias in GUI agents using video retrieval
Researchers have developed GUIDE, a novel framework designed to mitigate domain bias in GUI agents. This plug-and-play system leverages real-time web video retrieval and an automated annotation pipeline to equip agents …
-
New GAIA system trains critic models to improve GUI agent performance
Researchers have developed GAIA, a data flywheel system designed to improve the performance of GUI agents by training an Intuitive Critic Model (ICM). This ICM evaluates the correctness of an agent's actions, selecting …
-
VisCritic framework enhances GUI agents with visual state comparison
Researchers have introduced VisCritic, a novel visual process reward framework designed to enhance the performance of GUI agents. Unlike previous methods that rely solely on textual reasoning, VisCritic directly compare…
-
New EVA framework evolves semantic attacks on GUI agents
Researchers have developed EVA, an evolutionary framework designed to identify semantic vulnerabilities in GUI agents powered by multimodal large language models (MLLMs). This method focuses on manipulating the semantic…
-
StainFlow improves GUI agent training with novel reward model
Researchers have introduced StainFlow, a novel process reward model designed to enhance the training of GUI agents. This method addresses the sparsity of feedback in reinforcement learning by providing finer-grained tra…
-
New DragOn dataset boosts GUI agent drag-and-drop capabilities
Researchers have introduced DragOn, a new benchmark and dataset designed to improve the performance of GUI agents in handling drag-based interactions. The dataset includes 286,000 training screenshots and 3.5 million tr…
-
New benchmark tests AI agents on dynamic short-video platforms
Researchers have introduced "LivingScreen," a new benchmark designed to evaluate GUI agents on dynamic short-video platforms. Unlike previous benchmarks that assume static screens, LivingScreen accounts for continuously…
-
New benchmark and data synthesis boost GUI agent error recovery
Researchers have developed a new benchmark and data synthesis framework to improve the error recovery capabilities of GUI agents. The benchmark, GUI-RobustEval, includes over 1,200 test cases to systematically measure h…
-
MaskClaw system offers edge-side privacy for GUI agents
Researchers have developed MaskClaw, a novel edge-side privacy arbitrator designed for GUI agents. This system aims to protect sensitive information within screenshots by making privacy decisions locally, before data is…
-
New method GUI-CIDER boosts GUI agent knowledge
Researchers have developed GUI-CIDER, a novel mid-training method designed to enhance the world knowledge of GUI agents built with multimodal large language models. This approach explicitly internalizes GUI operational …
-
New frameworks and benchmarks advance mobile GUI agent capabilities
Researchers have developed several new frameworks and benchmarks to advance the capabilities of mobile GUI agents. STAMP introduces explicit memory training for agents in virtual environments, improving task resilience.…
-
New CutVerse benchmark reveals GUI agents struggle with media editing tasks
Researchers have introduced CutVerse, a new benchmark designed to assess the capabilities of GUI agents in media post-production tasks. The benchmark features over 180 complex tasks across seven professional application…
-
New AQuaUI method slashes GUI agent visual tokens
Researchers have developed AQuaUI, a novel method to reduce the number of visual tokens processed by Large Multimodal Models (LMMs) when interacting with graphical user interfaces (GUIs). This training-free technique co…
-
DocOS benchmark tests GUI agents' ability to use online docs
Researchers have introduced DocOS, a new benchmark designed to evaluate GUI agents' ability to proactively use online documentation for task completion. Current GUI agents struggle with tasks requiring procedural knowle…
-
Mobile GUI agents guided by new world models trained on code and text
Researchers have developed a novel approach to enhance mobile GUI agents by training world models across four modalities: delta text, full text, diffusion-based images, and renderable code. These models achieved state-o…