推荐理由:评测给出 Claude Fable 5.1 的智能指数分数与单任务成本,可对照它在 Fable 5 和 Opus 5 之间的性能与价格取舍。
评测基准
全部主题模型到底谁强:Benchmark 成绩、评测方法论争议与排行榜变化的持续记录。
当前仅显示精选新闻最新精选
第 21–38 条 · 共 38 条@ArtificialAnlys@ArtificialAnlys精选AI 评分7373 
@testingcatalog@testingcatalog精选AI 评分7171 
引用@claudeai@claudeaiWe’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3
推荐理由:两款新模型在 Terminal-Bench 系列基准上的分数对比,可直观看出相对前代的提升幅度与推送状态。
@ArtificialAnlys@ArtificialAnlys精选AI 评分6666 Z AI 发布 GLM-5.3-Flash,320B 总参数、18B 激活参数,在 Artificial Analysis 智能指数以最高推理档位得 57 分,单任务成本 0.09 美元。

推荐理由:GLM-5.3-Flash 以约七分之一的单任务成本把智能指数做到与 GPT-5.6 Terra 持平,可作为成本敏感选型的比较样本。
@kimmonismus@kimmonismus精选AI 评分6565
引用@AnthropicAI@AnthropicAIMany drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work. We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets. We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
推荐理由:材料给出 Claude 自主蛋白质结合体设计的命中率数字,可与该领域人类主导流程的典型水平直接对比。
LangChain Blog精选AI 评分7171 LangChain 实测 Switchyard 路由:仅 7% 的 agent 调用需要前沿模型
LangChain 用自家 145 个多步任务的 Deep Agents 评测套件测试 NVIDIA 开源路由库 Switchyard,只有 7% 的模型调用被送到 Claude Opus 4.8,其余 93% 由 30B 的 Nemotron 3.5 Lightning 处理。
推荐理由:用同一套 agent 评测套件对比单模型与路由方案,并给出判断是否值得上路由的成本公式。
量子位 · 微信公众号精选AI 评分7878 深度体验 DeepSeek Harness:开源编码 Agent 的插件化架构与实测对比
DeepSeek Harness 正式发布并开源,作者内测半个月后把自己的 Vibe Coding 项目从 Codex 迁移到 DSH。DSH 采用一切皆插件的 Cordis 架构,内置 100 多个插件,提供标准、PTC、极简、创造四类 Agent 预设,并有可直接查看原始事件的轨迹回放视图。
推荐理由:作者用半个月内测体验对比 Codex,呈现 DSH 的插件化架构与轨迹回放等设计,便于判断它与现有编码 Agent 的差异。
OpenRouter Announcements精选AI 评分6565 OpenRouter 上线实时网络搜索基准:比较引擎、深度与模型的取舍
OpenRouter 发布实时网络搜索基准榜单,覆盖 BrowseComp、DeepSearchQA、WideSearch 和 HLE 四套评测,比较模型、引擎(Exa、Parallel、Perplexity 及 OpenAI、Anthropic、Google 原生搜索)、搜索方法与预算的组合。
推荐理由:原文用四套实时榜单量化搜索预算、引擎与模型对质量和成本的影响,还给出可直接套用的参数配置方法。
OpenRouter Announcements精选AI 评分7171 OpenRouter 实测图像 detail 参数:推理模型用 low 更贵更差
OpenRouter 在 MMMU-Pro Vision(1,730 题)上基准测试 OpenAI 和 Google 五款模型的图像 detail 设置,发现 gpt-5.5 用 low 比 auto 低 13.8 分(65.2% vs 79.0%)且每题更贵(5.1¢ vs 4.5¢),原因是模型为看清热压缩图像多花 1.6x 推理 token。
推荐理由:基于 MMMU-Pro Vision 实测五款模型,给出推理模型保持 auto 细节、改调 reasoning effort 省钱的可操作结论。
Hugging Face Blog精选AI 评分6363 Hugging Face 与 Treble Technologies 发布 FFASR 远场 ASR 评测榜单
Treble Technologies 与 Hugging Face 推出 FFASR 榜单,称其为首个开放的远场 ASR 基准,在 14 个仿真房间和不同信噪比条件下评测语音识别模型。榜单同时报告 WER 与 RTFx,现有提交显示低 SNR 下的远场 WER 普遍是近场 WER 的数倍。用户可在 Submit 页粘贴 Hugging Face 模型 ID,评测在服务端对留出数据集运行。
推荐理由:新榜单将近场与远场语音识别的性能差距量化并公开,读者可据此判断模型在真实部署声学环境下的鲁棒性。
Hugging Face Blog精选AI 评分6969 Hugging Face 发布 agent-eval,评测开放模型在自有工具上的智能体表现
Hugging Face 发布开源评测工具 agent-eval,用于衡量开放模型驱动智能体使用某个库完成任务时的过程成本,而不只看最终答案。
推荐理由:这篇评测以 transformers 为样本,把智能体完成任务的过程成本拆成可量化指标,方法可迁移到其他库。
Simon Willison精选AI 评分7979 Simon Willison 实测 Claude Fable 5 初体验
Simon Willison 用约 5.5 小时实测 Anthropic 新发布的 Claude Fable 5,认为它慢、贵但知识储备丰富。
推荐理由:作者投入数小时实测,并用同一提示对比 Opus 4.8 与 GPT-5.5,呈现新模型在知识召回上的差异。
@swyx@swyx精选AI 评分7070
引用Andrej Karpathy (@karpathy)@karpathyThis is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward (imo of the same order as Claude 4.5 was in November), peaking especially for long problem-solving sessions on very difficult problems. You can give it a lot more ambitious tasks than what you're used to, the model "gets it" and it will just go, and it's never felt this tempting to stop looking at the code at all (but don't do this in prod!). The model still has quirks that people will run into and the safeguards are configured to be a little too trigger happy for launch, which can hopefully be tuned over time. I feel a lot of things changing as working software increasingly comes out on a tap. The Jevon's paradox kicks in and I feel my own demand for software growing substantially. You can ask for anything - explainers, visualizers, dashboards, bespoke single-use apps (e.g. a full wandb that is hyper-specific just for your project), you can 10X your test suite, auto-optimize code, run giant research projects with custom HTML for the results, anything! "Free your mind" (Matrix ref). Really looking forward to all the things people build!
推荐理由:转发的评测者认为官方榜单未体现 Fable 5 的进步幅度,可作为判断这次模型跃迁的定性参考。
Hugging Face Blog精选AI 评分6868 Artificial Analysis 与 IBM 发布 ITBench-AA 基准,前沿模型在智能体企业 IT 任务上得分不足 50%
Artificial Analysis 与 IBM 软件创新实验室发布 ITBench-AA,这是首个面向智能体企业 IT 任务的基准系列,首批聚焦 SRE 场景,前沿模型得分均低于 50%。
推荐理由:前沿模型在这套企业IT智能体基准上全部低于50%,并给出开源权重模型的每任务成本对照,可作为能力与成本参照。
Hugging Face Blog精选AI 评分7575 IBM Research 发布 Open Agent Leaderboard 与 Exgentic 评测框架
IBM Research 发布 Open Agent Leaderboard,把完整智能体系统而非单个模型作为评测单位,同时报告成功率和每任务成本,并开源配套的 Exgentic 评测框架与论文。
推荐理由:榜单把完整智能体系统而非单个模型作为评测单位,同时给出质量与成本,便于判断不同实现的部署取舍。
OpenRouter Announcements精选AI 评分6565 OpenAI 将 GPT-5.5 每 token 价格翻倍,实际成本究竟如何
OpenAI 在 GPT-5.5 上将每 token 价格翻倍,但该模型输出更为简洁。OpenRouter 基于真实用量数据测算这一变化带来的净成本影响。
推荐理由:OpenRouter 基于真实用量测算 GPT-5.5 单价翻倍后的净成本变化,读者可据此判断实际开销是否更贵。
Sierra Blog精选AI 评分6464 Sierra 发布 𝜏-voice 基准,在真实语音条件下评测实时语音 Agent
Sierra 团队推出 𝜏-voice 基准,将 𝜏-bench 的 278 个客服任务与全双工实时语音、可控真实音频结合,任务、工具和评估器与文本版逐字节一致,语音成绩可与文本 Agent 直接对比。
推荐理由:语音与文本能力可在同一套任务上直接对比,还给出了噪音与口音下各厂商的具体失败模式。
AI as Normal Technology精选AI 评分7474 Sayash Kapoor 等人提出开放世界评测并发布 CRUX 项目,AI 智能体成功上架 iOS 应用
Sayash Kapoor 与 Arvind Narayanan 等 17 位研究者发布 8000 字论文,定义开放世界评测这一新兴评测类型,并介绍定期开展此类评测的 CRUX 项目。
推荐理由:文章界定了开放世界评测这一新兴评测类型,并以 CRUX 首个实验给出智能体发布 iOS 应用的具体过程与成本细节。
AI as Normal Technology精选AI 评分7070 AI as Normal Technology 团队发布论文《Towards a Science of AI Agent Reliability》
Sayash Kapoor、Arvind Narayanan 与 Stephan Rabanser 等人发布 66 页论文《Towards a Science of AI Agent Reliability》,将智能体可靠性拆解为一致性、鲁棒性、可预测性、安全性四维共 12 项指标。
推荐理由:论文把智能体可靠性拆成12个维度并实测14个模型,显示18个月内准确率大涨而可靠性仅小幅提升。