A new benchmark called mcpbench has been developed to evaluate AI models' ability to construct MCP (Model Context Protocol) clients and servers. MCP is a standard created by Anthropic to enable AI systems to access external services and data, with its latest version, 2026-07-28, introducing statelessness. The benchmark tests various AI models, including GPT-5.6 variants and Claude Opus, on their capability to generate MCP clients and servers according to different MCP versions and documentation inputs. Results indicate that GPT-5.6 Sol performed best in creating MCP clients for version 2025-11-25, while Claude Opus 4.8 and Kimi K2.7 Code achieved 100% in server creation for the same version. However, performance dropped significantly for the newer 2026-07-28 version, suggesting a lack of knowledge about the updated specification among the tested models. AI
IMPACT Highlights potential gaps in LLM knowledge of evolving external communication standards like MCP.
RANK_REASON New benchmark for evaluating AI models' ability to interact with external protocols. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
- Anthropic
- Claude Opus 4.8
- Claude Sonnet 5
- Cloudflare
- GPT-5.6 Luna
- GPT-5.6 Sol
- GPT-5.6 Terra
- Kimi K2.7 Code
- Matt Carey
- MCP
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →