跳到正文

评测基准

模型到底谁强:Benchmark 成绩、评测方法论争议与排行榜变化的持续记录。

当前仅显示精选新闻

最新精选

第 21–38 条 · 共 38 条
9月2日周三
  1. @testingcatalog71

    Anthropic 宣布推出 Claude Fable 5.1 和 Claude Mythos 5.1,并称其面向编码与知识工作。据 Testing Catalog 引述,Fable 5.1 在 Terminal-Bench-Science 0.1 上得分 52.6%,是 Fable 5 的两倍多;在 Terminal-Bench 4.0 上得分 55.8%,Fable 5 为 42.0%。两款模型目前正在 Claude 上推送。

    引用@claudeai@claudeai

    We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3

    推荐理由:两款新模型在 Terminal-Bench 系列基准上的分数对比,可直观看出相对前代的提升幅度与推送状态。

8月27日周四
8月19日周三
  1. @kimmonismus65

    Anthropic 测试 Claude 从零自主设计蛋白质结合体,基于一份人类专家撰写的蛋白设计提示词,针对 15 个靶点中的 14 个完成设计,命中率约 27%。Adaptyv Bio 与 Twist Bioscience 独立构建并测试后确认 1320 个设计中有 354 个成功结合。不同设置下命中率在 22.6% 到 35.1% 之间,排名最高的设计在 49% 的实验轮次中成功结合。

    引用@AnthropicAI@AnthropicAI

    Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work. We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets. We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.

    推荐理由:材料给出 Claude 自主蛋白质结合体设计的命中率数字,可与该领域人类主导流程的典型水平直接对比。

8月16日周日
  1. LangChain Blog71

    LangChain 实测 Switchyard 路由:仅 7% 的 agent 调用需要前沿模型

    LangChain 用自家 145 个多步任务的 Deep Agents 评测套件测试 NVIDIA 开源路由库 Switchyard,只有 7% 的模型调用被送到 Claude Opus 4.8,其余 93% 由 30B 的 Nemotron 3.5 Lightning 处理。

    推荐理由:用同一套 agent 评测套件对比单模型与路由方案,并给出判断是否值得上路由的成本公式。

8月14日周五
  1. 量子位 · 微信公众号78

    深度体验 DeepSeek Harness:开源编码 Agent 的插件化架构与实测对比

    DeepSeek Harness 正式发布并开源,作者内测半个月后把自己的 Vibe Coding 项目从 Codex 迁移到 DSH。DSH 采用一切皆插件的 Cordis 架构,内置 100 多个插件,提供标准、PTC、极简、创造四类 Agent 预设,并有可直接查看原始事件的轨迹回放视图。

    推荐理由:作者用半个月内测体验对比 Codex,呈现 DSH 的插件化架构与轨迹回放等设计,便于判断它与现有编码 Agent 的差异。

8月12日周三
  1. OpenRouter Announcements65

    OpenRouter 上线实时网络搜索基准:比较引擎、深度与模型的取舍

    OpenRouter 发布实时网络搜索基准榜单,覆盖 BrowseComp、DeepSearchQA、WideSearch 和 HLE 四套评测,比较模型、引擎(Exa、Parallel、Perplexity 及 OpenAI、Anthropic、Google 原生搜索)、搜索方法与预算的组合。

    推荐理由:原文用四套实时榜单量化搜索预算、引擎与模型对质量和成本的影响,还给出可直接套用的参数配置方法。

7月7日周二
  1. OpenRouter Announcements71

    OpenRouter 实测图像 detail 参数:推理模型用 low 更贵更差

    OpenRouter 在 MMMU-Pro Vision(1,730 题)上基准测试 OpenAI 和 Google 五款模型的图像 detail 设置,发现 gpt-5.5 用 low 比 auto 低 13.8 分(65.2% vs 79.0%)且每题更贵(5.1¢ vs 4.5¢),原因是模型为看清热压缩图像多花 1.6x 推理 token。

    推荐理由:基于 MMMU-Pro Vision 实测五款模型,给出推理模型保持 auto 细节、改调 reasoning effort 省钱的可操作结论。

6月24日周三
  1. Hugging Face Blog63

    Hugging Face 与 Treble Technologies 发布 FFASR 远场 ASR 评测榜单

    Treble Technologies 与 Hugging Face 推出 FFASR 榜单,称其为首个开放的远场 ASR 基准,在 14 个仿真房间和不同信噪比条件下评测语音识别模型。榜单同时报告 WER 与 RTFx,现有提交显示低 SNR 下的远场 WER 普遍是近场 WER 的数倍。用户可在 Submit 页粘贴 Hugging Face 模型 ID,评测在服务端对留出数据集运行。

    推荐理由:新榜单将近场与远场语音识别的性能差距量化并公开,读者可据此判断模型在真实部署声学环境下的鲁棒性。

6月18日周四
6月10日周三
  1. @swyx70

    karpathy 称 Claude Fable 5 与 Mythos 是同一底层模型,只是增加了防护措施,在各项基准上都以一定优势达到 SOTA,并认为这是值得大版本号跃迁的阶跃式进步,尤其擅长在极难问题上的长时间解题会话。swyx 转发表示自己重跑了历史图表上的 FC Diamond,认为官方表格和图表都没有体现出这种起飞幅度,因为 Fable 属于不同级别的模型。

    引用Andrej Karpathy (@karpathy)@karpathy

    This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward (imo of the same order as Claude 4.5 was in November), peaking especially for long problem-solving sessions on very difficult problems. You can give it a lot more ambitious tasks than what you're used to, the model "gets it" and it will just go, and it's never felt this tempting to stop looking at the code at all (but don't do this in prod!). The model still has quirks that people will run into and the safeguards are configured to be a little too trigger happy for launch, which can hopefully be tuned over time. I feel a lot of things changing as working software increasingly comes out on a tap. The Jevon's paradox kicks in and I feel my own demand for software growing substantially. You can ask for anything - explainers, visualizers, dashboards, bespoke single-use apps (e.g. a full wandb that is hyper-specific just for your project), you can 10X your test suite, auto-optimize code, run giant research projects with custom HTML for the results, anything! "Free your mind" (Matrix ref). Really looking forward to all the things people build!

    推荐理由:转发的评测者认为官方榜单未体现 Fable 5 的进步幅度,可作为判断这次模型跃迁的定性参考。

5月28日周四
  1. Hugging Face Blog68

    Artificial Analysis 与 IBM 发布 ITBench-AA 基准,前沿模型在智能体企业 IT 任务上得分不足 50%

    Artificial Analysis 与 IBM 软件创新实验室发布 ITBench-AA,这是首个面向智能体企业 IT 任务的基准系列,首批聚焦 SRE 场景,前沿模型得分均低于 50%。

    推荐理由:前沿模型在这套企业IT智能体基准上全部低于50%,并给出开源权重模型的每任务成本对照,可作为能力与成本参照。

5月18日周一
  1. Hugging Face Blog75

    IBM Research 发布 Open Agent Leaderboard 与 Exgentic 评测框架

    IBM Research 发布 Open Agent Leaderboard,把完整智能体系统而非单个模型作为评测单位,同时报告成功率和每任务成本,并开源配套的 Exgentic 评测框架与论文。

    推荐理由:榜单把完整智能体系统而非单个模型作为评测单位,同时给出质量与成本,便于判断不同实现的部署取舍。

5月4日周一
5月2日周六
  1. Sierra Blog64

    Sierra 发布 𝜏-voice 基准,在真实语音条件下评测实时语音 Agent

    Sierra 团队推出 𝜏-voice 基准,将 𝜏-bench 的 278 个客服任务与全双工实时语音、可控真实音频结合,任务、工具和评估器与文本版逐字节一致,语音成绩可与文本 Agent 直接对比。

    推荐理由:语音与文本能力可在同一套任务上直接对比,还给出了噪音与口音下各厂商的具体失败模式。

4月17日周五
  1. AI as Normal Technology74

    Sayash Kapoor 等人提出开放世界评测并发布 CRUX 项目,AI 智能体成功上架 iOS 应用

    Sayash Kapoor 与 Arvind Narayanan 等 17 位研究者发布 8000 字论文,定义开放世界评测这一新兴评测类型,并介绍定期开展此类评测的 CRUX 项目。

    推荐理由:文章界定了开放世界评测这一新兴评测类型,并以 CRUX 首个实验给出智能体发布 iOS 应用的具体过程与成本细节。

2月24日周二
  1. AI as Normal Technology70

    AI as Normal Technology 团队发布论文《Towards a Science of AI Agent Reliability》

    Sayash Kapoor、Arvind Narayanan 与 Stephan Rabanser 等人发布 66 页论文《Towards a Science of AI Agent Reliability》,将智能体可靠性拆解为一致性、鲁棒性、可预测性、安全性四维共 12 项指标。

    推荐理由:论文把智能体可靠性拆成12个维度并实测14个模型,显示18个月内准确率大涨而可靠性仅小幅提升。