PulseAugur
EN
LIVE 07:20:35
中文(ZH) 华尔街实测8款全球主流Agent:千问办公综合排名第一,超Claude Cowork、Codex

Alibaba's Qwen Office tops AI agent benchmark, outperforming Claude Cowork and Codex

A recent test by Jefferies analysts evaluated eight leading AI agents on real-world office tasks, with Alibaba's Qwen Office ranking first. The agent demonstrated strong performance across complex tasks, including web browsing and multimodal content generation, achieving over 90 points in all tested dimensions. The report highlighted that the synergy between Qwen Office and its underlying Qwen 3.8 Max model, combined with its engineering mechanisms (Harness), contributed to its superior performance and cost-effectiveness compared to competitors like Claude Cowork and Codex. The analysis also emphasized the growing importance of cost per task and enterprise-level features like workflow integration and ecosystem collaboration in the competitive AI agent market. AI

IMPACT This benchmark suggests that integrated AI agent solutions, focusing on both model performance and engineering efficiency, are becoming key differentiators in the enterprise market.

RANK_REASON A benchmark evaluation of multiple AI agents on real-world tasks, highlighting a specific product's performance and cost advantages.

Read on 雷峰网 (Leiphone) →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Alibaba's Qwen Office tops AI agent benchmark, outperforming Claude Cowork and Codex

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
A benchmark evaluation of multiple AI agents on real-world tasks, highlighting a specific product's performance and cost advantages.
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
15 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Wall Street tests 8 global mainstream Agents: Qwen Office ranks first overall

    成本也正成为Agent商业化需要考虑的重要因素

  2. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Wall Street tests 8 global mainstream Agents: Qwen Office ranks first overall, surpassing Claude Cowork and Codex

    <p>8月19日消息,华尔街投行杰富瑞分析师近日对八款全球主流AI Agent 进行了真实办公任务实测,结果显示,通过模型和Harness的协同,阿里千问办公综合得分排名第一,超过美国的Claude Cowork和Codex等Agent工具。</p><p>&nbsp;</p><p>此次测试共设置5项真实办公任务,包括基于多份文件完成公司年报摘要、联网查找并比较公司经营数据、操作真实桌面浏览器完成信息检索和文档生成、根据数据制作英文PPT,以及基于参考图片生成营销海报。测试结果显示,千问办公整体表现较为均衡,尤其在复杂办公任务、网页浏览器控制以及多模态内容…