PulseAugur
实时 18:42:33
English(EN) Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms

新的基准测试在动态短视频平台上测试AI智能体

研究人员推出了“LivingScreen”,这是一个旨在评估动态短视频平台上GUI智能体的新基准。与假设屏幕静态的先前基准不同,LivingScreen考虑了连续播放的内容,要求智能体就观察和交互做出实时决策。对当前前沿模型的评估显示,在成本准确性方面均未达到人类水平,常见故障包括观察时间不当,凸显了未来GUI智能体在观察控制方面需要改进。 AI

影响 突出了当前GUI智能体在动态环境中的能力差距,可能指导未来在观察控制方面的研究。

排序理由 该集群包含一篇介绍AI智能体新基准的学术论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的基准测试在动态短视频平台上测试AI智能体

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Jiashu Yao, Heyan Huang, Daiqing Wu, Wangke Chen, Huaxi Ai, Haoyu Wen, Zeming Liu, Yuhang Guo ·

    在短视频平台上对原生动态屏幕GUI代理进行基准测试

    arXiv:2606.04701v1 Announce Type: cross Abstract: GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications violate this assumption, as their content keeps playing, and a competent user must d…

  2. arXiv cs.CL TIER_1 English(EN) · Yuhang Guo ·

    在短视频平台上对原生动态屏幕GUI代理进行基准测试

    GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications violate this assumption, as their content keeps playing, and a competent user must decide what to watch and for how long. We formalize…