Baidu's Ernie Bot Task Agent has achieved the top position on the international PinchBench v2 leaderboard, surpassing models like Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.6-luna. The agent scored 94.6% highest and 94.4% average, marking it as the first Chinese agent system to secure the number one spot on this benchmark. PinchBench evaluates an agent's ability to complete complex, real-world tasks with verifiable results, rather than just knowledge recall, highlighting the importance of system architecture and engineering alongside model capabilities. AI
IMPACT Sets a new benchmark for AI agent capabilities in complex task completion, potentially driving competition in agent framework development.
RANK_REASON The cluster reports a top ranking on a significant AI agent benchmark by a major tech company's product, surpassing established competitors. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- Anthropic
- Baidu
- Claude Opus 4.8
- Claude Opus 4.8-fast
- GPT-5.6-luna
- Kilo AI
- OpenAI
- OpenClaw
- PinchBench v2
- 阿里通义千问 Qwen3.7-max
- Wenxin Assistant Task Agent
- Anthropic Claude Opus 4.8
- Anthropic Claude Opus 4.8-fast
- OpenAI GPT-5.6-luna
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →