PulseAugur
实时 22:23:15
English(EN) Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining

Video2GUI 从无标签视频生成1200万条GUI轨迹

研究人员开发了Video2GUI,一个旨在为GUI代理训练生成大规模交互轨迹的自动化框架。该系统从无标签的互联网视频中提取数据,通过过滤过程将其转换为结构化的代理轨迹。由此产生的WildGUI数据集包含1500多个应用程序的1200万条轨迹,显著改进了Qwen2.5-VL和Mimo-VL等模型的预训练。 AI

影响 能够为GUI代理创建大规模数据集,可能提高其在各种应用程序中的泛化能力和性能。

排序理由 介绍GUI代理预训练新方法和数据集的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Video2GUI 从无标签视频生成1200万条GUI轨迹

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
介绍GUI代理预训练新方法和数据集的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
124 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Hao Tian ·

    Video2GUI:为通用GUI代理预训练合成大规模交互轨迹

    Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization remains constrained by the scarcity of large-scale training data spanning diverse real-world applications. Existing datasets rely he…