PulseAugur
EN
LIVE 01:01:51
日本語(JA) GPT-5.6やClaude Opusの性能を「MCPサーバーを構築できるか?」という観点で測定したベンチマークテスト「mcpbench」 https:// fed.brid.gy/r/https://gigazine .net/news/20260729-mcpbench/

New mcpbench tests AI models on MCP client/server construction

A new benchmark called mcpbench has been developed to evaluate AI models' ability to construct MCP (Model Context Protocol) clients and servers. MCP is a standard created by Anthropic to enable AI systems to access external services and data, with its latest version, 2026-07-28, introducing statelessness. The benchmark tests various AI models, including GPT-5.6 variants and Claude Opus, on their capability to generate MCP clients and servers according to different MCP versions and documentation inputs. Results indicate that GPT-5.6 Sol performed best in creating MCP clients for version 2025-11-25, while Claude Opus 4.8 and Kimi K2.7 Code achieved 100% in server creation for the same version. However, performance dropped significantly for the newer 2026-07-28 version, suggesting a lack of knowledge about the updated specification among the tested models. AI

IMPACT Highlights potential gaps in LLM knowledge of evolving external communication standards like MCP.

RANK_REASON New benchmark for evaluating AI models' ability to interact with external protocols. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New mcpbench tests AI models on MCP client/server construction

COVERAGE [1]

  1. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    Benchmark test "mcpbench" that measures the performance of GPT-5.6 and Claude Opus from the perspective of "Can it build an MCP server?" https://fed.brid.gy/r/https://gigazine.net/news/20260729-mcpbench/

    GPT-5.6やClaude Opusの性能を「MCPサーバーを構築できるか?」という観点で測定したベンチマークテスト「mcpbench」 https:// fed.brid.gy/r/https://gigazine .net/news/20260729-mcpbench/