PulseAugur
EN
LIVE 10:43:31
中文(ZH) 华尔街实测8款全球主流Agent:千问办公综合排名第一,超Claude Cowork、Codex

Alibaba's Qwen Office tops AI agent real-world office task test

A recent real-world office task test conducted by Wall Street investment bank Jefferies evaluated eight leading AI agents. Alibaba's Qwen Office emerged as the top performer, surpassing competitors like Claude Cowork and Codex. The evaluation highlighted Qwen Office's balanced performance across complex tasks, web browser control, and multimodal content generation, being the only agent to score above 90 in all tested dimensions. The report also emphasized the importance of the 'Harness' component—the engineering mechanisms surrounding the model—in translating AI intelligence into practical results, where Qwen Office also led. AI

IMPACT Sets a new benchmark for AI agent performance in real-world office tasks, highlighting the importance of engineering alongside core models.

RANK_REASON The cluster reports on the results of an independent benchmark test of AI agents, which is a form of research milestone. [lever_c_demoted from research: ic=1 ai=1.0]

Read on 雷峰网 (Leiphone) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Alibaba's Qwen Office tops AI agent real-world office task test

COVERAGE [1]

  1. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Wall Street tests 8 global mainstream Agents: Qwen Office ranks first overall, surpassing Claude Cowork and Codex

    <p>8月19日消息,华尔街投行杰富瑞分析师近日对八款全球主流AI Agent 进行了真实办公任务实测,结果显示,通过模型和Harness的协同,阿里千问办公综合得分排名第一,超过美国的Claude Cowork和Codex等Agent工具。</p><p>&nbsp;</p><p>此次测试共设置5项真实办公任务,包括基于多份文件完成公司年报摘要、联网查找并比较公司经营数据、操作真实桌面浏览器完成信息检索和文档生成、根据数据制作英文PPT,以及基于参考图片生成营销海报。测试结果显示,千问办公整体表现较为均衡,尤其在复杂办公任务、网页浏览器控制以及多模态内容…