PulseAugur
中
实时 22:57:36
English(EN) Best AI agents fail 50% of visual tool tasks, Apple benchmark shows A new open benchmark with 500+ tools reveals even frontier AI models can't reliably read an

AI代理在视觉任务中挣扎;EvoCUA-1.5在OSWorld基准测试中达到63.2%

一项名为OSWorld的新基准测试显示,AI代理EvoCUA-1.5通过在模拟环境中进行试错学习,取得了63.2%的分数。另外,Apple Inc.开发的一项新的开放基准测试表明,即使是先进的AI模型在视觉工具任务上也面临困难,由于图像识别和解释方面的挑战,失败率超过50%。 AI

影响 突出了当前AI代理在视觉任务方面的能力局限性,并展示了通过模拟环境学习的进展。

排序理由 该集群讨论了新的基准测试和AI代理在特定任务上的表现,属于研究类别。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

AI代理在视觉任务中挣扎;EvoCUA-1.5在OSWorld基准测试中达到63.2%

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群讨论了新的基准测试和AI代理在特定任务上的表现,属于研究类别。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
89 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    EvoCUA-1.5 在 OSWorld 上达到 63.2%,采用在线强化学习训练 EvoCUA-1.5 通过在沙盒环境中进行试错来教会 AI 代理使用计算机,达到

    EvoCUA-1.5 hits 63.2% on OSWorld with online RL training EvoCUA-1.5 teaches AI agents to use computers through trial-and-error in sandbox environments, reaching 63.2% on OSWorld-Verified. https://www. notatechguy.com/evocua-1-5-hit s-63-2-on-osworld-with-online-rl-training/ # Not…

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    最佳AI代理在视觉工具任务中失败率达50%,Apple基准测试显示:拥有500多个工具的新开放基准测试表明,即使是前沿AI模型也无法可靠地读取

    Best AI agents fail 50% of visual tool tasks, Apple benchmark shows A new open benchmark with 500+ tools reveals even frontier AI models can't reliably read an image and act on it, with most failures traced to seeing, not https://www. notatechguy.com/best-ai-agents -fail-50-of-vi…