跳到正文
@omarsar0· @omarsar0 · X·· 2026-08-30AI 评分45
AI 导读

Apple 的 Agent Seer 能从单份 MCP 规范出发,在无示例、无实时工具访问、无领域微调的情况下合成多轮智能体测试场景,并生成带评分标准的对话评测。研究在 7 份跨领域 MCP 规范上验证,中小型规范实现完整工具覆盖;参数 schema 复杂度比工具套件规模更能预测质量差异,参数值准确率是主要失败模式,而粗粒度名称匹配的工具调用指标完全看不到这一点。

正文

Banger paper from Apple.

If you build MCP servers, this can help you turn your specification into an evaluation suite.

(bookmark it)

It's actually a really neat idea that's easy to implement. And it showcases the awesomeness of MCP.

Agent Seer starts from a single MCP spec and synthesizes multi-turn agent test scenarios with no examples, no live tool access, and no domain-specific tuning.

Function names, natural-language descriptions and typed parameter schemas already carry enough semantics to generate graded scenarios with synthetic tool outputs, which then expand into mock-data-grounded dialogues.

Hand-built agent benchmarks demand deep domain expertise, do not scale across tool ecosystems, and go stale as soon as an API changes. Generating them from the live spec keeps pace with the ecosystem instead.

They ran it on seven MCP specifications spanning different domains and suite sizes, with complete tool coverage on small and medium specs.

Parameter schema complexity predicts quality variation far better than tool-suite size does. And argument value accuracy is the dominant failure mode, a sub-dimension that coarse name-match tool-calling metrics cannot see at all.

Paper: https://t.co/ByU0tYn39y

Chat with Paper: https://t.co/w3YaQ2RXRe

来源:@omarsar0 · x.com