PulseAugur
中
实时 17:37:11
English(EN) MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

MCP-Atlas基准测试使用真实服务器评估LLM的工具使用能力

研究人员推出了MCP-Atlas,这是一个旨在评估大型语言模型工具使用能力的新基准测试。该基准测试包含36个真实的MCP服务器和220个工具,有1000个任务需要多步工作流和多工具调用编排。对先进模型的初步评估显示,尽管顶级模型的通过率超过50%,但常见的失败源于工具使用和任务理解方面的问题。 AI

影响 为评估LLM的工具使用能力建立了新的标准,有望推动智能体能力和现实世界应用集成的改进。

排序理由 引入了一个新的基准数据集来评估LLM的工具使用能力。 [lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

MCP-Atlas基准测试使用真实服务器评估LLM的工具使用能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
引入了一个新的基准数据集来评估LLM的工具使用能力。 [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
155 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Chaithanya Bandi, Ben Hertzberg, Geobio Boo, Tejas Polakam, Jeff Da, Sami Hassaan, Manasi Sharma, Andrew Park, Ernesto Hernandez, Dan Rambado, Ivan Salazar, Rafael Cruz, Chetan Rane, Ben Levin, Brad Kenstler, Bing Liu ·

    MCP-Atlas:具有真实MCP服务器的工具使用能力的大规模基准测试

    arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external tools. However, existing evaluations often fail to capture the complexity of real…