PulseAugur
中
实时 20:08:13
English(EN) One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

InnerZoom框架在单次前向传播中实现SOTA GUI基础定位 · 跟踪3个来源

研究人员开发了InnerZoom,一个新颖的框架,用于在单次前向传播中实现准确高效的GUI基础定位。该方法通过在解码器层之间保留目标区域感知来解决现有多模态大语言模型(MLLM)方法的局限性,这对于GUI交互中精确坐标的生成至关重要。InnerZoom在多个基准测试中取得了最先进的性能,在提高精度的同时降低了计算成本和延迟。 AI

影响 这种新方法可以提高AI代理与图形用户界面交互的效率和准确性。

排序理由 该集群报道了一篇详细介绍一种新GUI基础定位方法的最新研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

InnerZoom框架在单次前向传播中实现SOTA GUI基础定位 · 跟踪3个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群报道了一篇详细介绍一种新GUI基础定位方法的最新研究论文。
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
101 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [4]

  1. arXiv cs.AI TIER_1 English(EN) · Yangyue Wang, Harshvardhan Sikka, Yash Mathur, Tony Zhou, Jinu Nyachhyon, Pranav Guruprasad ·

    GUI-Perturbed:领域随机化揭示GUI基础模型系统性脆弱性

    arXiv:2604.14262v2 Announce Type: replace-cross Abstract: GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial reasoning rather than direct element naming. Current benchmarks miss this because the…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    一当二:InnerZoom实现精确高效的GUI定位

    InnerZoom addresses GUI grounding challenges by preserving target-region awareness across decoder layers through a single-forward pass that bridges cross-layer evidence, achieving state-of-the-art performance with reduced computational cost.

  3. arXiv cs.CV TIER_1 English(EN) · Chen Liu, Ling Chen, Hanzhang Zhou, Liangyu Chen, Chenglin Cai, Xin Yu, Steven Hoi, Yue Wang ·

    一当二:InnerZoom实现精准高效的GUI定位

    arXiv:2606.30084v1 Announce Type: new Abstract: MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enabling models to leverage the strong instruction-following and semantic understanding capabilities of MLLMs. However,…

  4. arXiv cs.CV TIER_1 English(EN) · Yue Wang ·

    一当二:InnerZoom实现精准高效的GUI基础定位

    MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enabling models to leverage the strong instruction-following and semantic understanding capabilities of MLLMs. However, this formulation requires the model to retain r…