A recent real-world office task test conducted by Wall Street investment bank Jefferies evaluated eight leading AI agents. Alibaba's Qwen Office emerged as the top performer, surpassing competitors like Claude Cowork and Codex. The evaluation highlighted Qwen Office's balanced performance across complex tasks, web browser control, and multimodal content generation, being the only agent to score above 90 in all tested dimensions. The report also emphasized the importance of the 'Harness' component—the engineering mechanisms surrounding the model—in translating AI intelligence into practical results, where Qwen Office also led. AI
IMPACT Sets a new benchmark for AI agent performance in real-world office tasks, highlighting the importance of engineering alongside core models.
RANK_REASON The cluster reports on the results of an independent benchmark test of AI agents, which is a form of research milestone. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →