A recent test by Jefferies analysts evaluated eight leading AI agents on real-world office tasks, with Alibaba's Qwen Office ranking first. The agent demonstrated strong performance across complex tasks, including web browsing and multimodal content generation, achieving over 90 points in all tested dimensions. The report highlighted that the synergy between Qwen Office and its underlying Qwen 3.8 Max model, combined with its engineering mechanisms (Harness), contributed to its superior performance and cost-effectiveness compared to competitors like Claude Cowork and Codex. The analysis also emphasized the growing importance of cost per task and enterprise-level features like workflow integration and ecosystem collaboration in the competitive AI agent market. AI
IMPACT This benchmark suggests that integrated AI agent solutions, focusing on both model performance and engineering efficiency, are becoming key differentiators in the enterprise market.
RANK_REASON A benchmark evaluation of multiple AI agents on real-world tasks, highlighting a specific product's performance and cost advantages.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →