PulseAugur
EN
LIVE 11:58:25
中文(ZH) 百度文心助手任务Agent登顶国际权威榜单,超越Claude、GPT拿下全球智能体冠军

Baidu's Ernie Bot Agent Tops Global Intelligence Benchmark, Outperforming Claude and GPT

Baidu's Ernie Bot Task Agent has achieved the top position on the international PinchBench v2 leaderboard, surpassing models like Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.6-luna. The agent scored 94.6% highest and 94.4% average, marking it as the first Chinese agent system to secure the number one spot on this benchmark. PinchBench evaluates an agent's ability to complete complex, real-world tasks with verifiable results, rather than just knowledge recall, highlighting the importance of system architecture and engineering alongside model capabilities. AI

IMPACT Sets a new benchmark for AI agent capabilities in complex task completion, potentially driving competition in agent framework development.

RANK_REASON The cluster reports a top ranking on a significant AI agent benchmark by a major tech company's product, surpassing established competitors. [lever_c_demoted from frontier_release: ic=2 ai=1.0]

Read on 量子位 (QbitAI) →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Baidu's Ernie Bot Agent Tops Global Intelligence Benchmark, Outperforming Claude and GPT

COVERAGE [2]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Baidu Wenxin Assistant Task Agent Tops International Authoritative List, Surpassing Claude and GPT to Become Global Intelligent Agent Champion

  2. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Baidu Wenxin Assistant Task Agent Tops International Authoritative List, Surpassing Claude and GPT to Become Global Intelligent Agent Champion

    <p>2026 年 7 月 17 日,百度文心助手任务 Agent,以最高分 94.6%、平均分 94.4% 的成绩,登顶全球工程向 AI 智能体评测榜单 PinchBench v2。成为首个以正式产品身份获得 PinchBench 总榜第一的国产智能体系统。</p><p>&nbsp;</p><p>在 59 个参评模型中,文心助手任务 Agent 排名第一,领先 Anthropic Claude Opus 4.8-fast(93.5%)、阿里通义千问 Qwen3.7-max(92.5%)、Anthropic Claude Opus 4.8(90.5%)、…