跳到正文

OpenRouter

OpenRouter 模型路由平台动态:新模型上架、用量排行与开发者选型的真实市场信号。

当前仅显示精选新闻

最新精选

第 61–80 条 · 共 85 条
7月22日周三
7月21日周二
  1. OpenRouter Announcements64

    OpenRouter 解析提示词缓存与粘性路由如何降低 Agent 多轮成本

    OpenRouter 发布教程讲解提示词缓存与粘性路由如何降低 Agent 多轮成本。缓存读取为输入价格的 0.1x 到 0.5x,Anthropic 缓存写入为 1.25x(5 分钟 TTL)或 2.0x(1 小时 TTL),未被复用的单次写入比不缓存更贵。

    推荐理由:原文给出各厂商缓存读写倍率、session_id 粘性路由用法和缓存未命中排查方法,读者可以直接照做降低 Agent 多轮成本。

7月7日周二
  1. OpenRouter Announcements71

    OpenRouter 实测图像 detail 参数:推理模型用 low 更贵更差

    OpenRouter 在 MMMU-Pro Vision(1,730 题)上基准测试 OpenAI 和 Google 五款模型的图像 detail 设置,发现 gpt-5.5 用 low 比 auto 低 13.8 分(65.2% vs 79.0%)且每题更贵(5.1¢ vs 4.5¢),原因是模型为看清热压缩图像多花 1.6x 推理 token。

    推荐理由:基于 MMMU-Pro Vision 实测五款模型,给出推理模型保持 auto 细节、改调 reasoning effort 省钱的可操作结论。

6月30日周二
  1. OpenRouter Announcements74

    OpenRouter 数据分析:DeepSeek V4 上线后 token 份额从 9% 翻倍至 18%

    OpenRouter 基于 2026 年 1 月 1 日至 6 月 14 日超 450 万亿 token 的日志分析,DeepSeek V4 于 4 月 24 日发布后,DeepSeek token 份额从年初约 9% 升至 6 月初近 20%,并自 5 月中旬起成为 OpenRouter 上用量第一的模型,智能体工作负载贡献了主要增量。

    推荐理由:基于 OpenRouter 450 万亿 token 日志分析 V4 上线后的份额变化,并区分智能体与人类流量,数据口径和价格对比都有可参考价值。

6月27日周六
  1. OpenRouter Announcements76

    OpenRouter 评出 2026 年 6 月最值得关注的四款开放权重模型

    OpenRouter 认为 2026 年 6 月最值得关注的四款开放权重模型是 DeepSeek V4 Flash、GLM 5.2、MiniMax M3 和 NVIDIA Nemotron 3 Ultra,并指出开放权重模型与美国前沿实验室的能力差距已连续 18 个月保持在 3-6 个月且未在扩大。

    推荐理由:OpenRouter 用自家价格与吞吐数据逐一点评四款开放权重模型的适用场景,读者可以按成本、质量和模态对号入座选型。

6月10日周三
  1. @OpenRouter67

    Anthropic 的 Claude Fable 5 已在 OpenRouter 上线,被称为 Anthropic 最强的编码模型。该模型面向长时程、模糊任务,包括 legacy migrations、生产环境疑难 bug 以及运行数小时到数天的异步会话,在几乎所有测试基准上达到 SOTA。

    推荐理由:Anthropic 这款主打长时程异步编码任务的模型已接入 OpenRouter,读者可据此了解它的能力定位与使用入口。

6月3日周三
  1. @OpenRouter67

    OpenRouter 推出 Pareto Code,一个免费、实验性的编程路由。开发者在请求中设置 min_coding_score,即可路由到满足该门槛的最便宜代码能力模型,排名由 Artificial Analysis 提供,并可实时看到 Pareto 前沿的移动。作者本人的推文仅表示将提供该路由的更多信息。

    引用OpenRouter (@OpenRouter)@OpenRouter

    Introducing Pareto Code: a new, free, experimental coding router Set `min_coding_score` in your request and route to the cheapest code-capable model that clears your bar, ranked by @ArtificialAnlys. See the Pareto frontier shifting in real time👇

    推荐理由:请求可按代码能力门槛自动筛选最便宜可用模型,为控制编码模型成本提供了一种可配置的路由思路。

6月1日周一
  1. @MiniMax_AI71

    MiniMax 发布 M3 模型,并在发布当天同步上线 OpenRouter。该模型具备 1M token 上下文窗口、前沿编码与智能体能力,以及原生多模态,首周提供 50% 折扣。OpenRouter 称其为前沿级开源权重模型,原生多模态覆盖图像与视频。

    引用OpenRouter (@OpenRouter)@OpenRouter

    MiniMax-M3 is live on OpenRouter! A frontier-class open-weight model that combines a 1M-token context window, frontier coding and agentic performance, and native multimodality (image & video) in one model.

    推荐理由:MiniMax-M3 上线 OpenRouter 并给出首周折扣,可了解其 1M 上下文与多模态能力的组合方式。

5月29日周五
  1. @OpenRouter69

    阶跃星辰发布 Step 3.7 Flash,主打 agent 效率,采用 198B 稀疏 MoE、约 11B 激活、256K 上下文和 3 级推理,以 Apache 2.0 开放权重。该模型在 ClawEval-1.1 得 67.1、SimpleVQA Search 得 79.2 均列第一,SWE-PRO 得 56.3,并支持 Claude Code、MCP 等工具调用,可在 Mac Studio M4 Max、DGX Spark 等设备本地运行。

    引用StepFun (@StepFun_ai)@StepFun_ai

    ⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SWE-PRO (56.3), 95.3 on V* Python. Open weights under Apache 2.0. Built for agentic, coding, search, and multimodal workflows — balancing speed, cost, and reliable execution. - 400 TPS. 198B sparse MoE, ~11B active. 256K context, 3 reasoning levels. - Understands UIs, charts, docs, images — then writes code or calls tools to act on what it sees. - Web + visual search reaches further: more sources, deeper follow-up. - Reliable tool use — less drift, fewer broken toolcalls. 98%+ on τ²-bench across all difficulty levels. - Works with Claude Code, KiloCode, Hermes Agent, OpenClaw, and protocols like MCP. - Runs locally on Mac Studio M4 Max, DGX Spark, AMD AI Max+ 395. GitHub: github.com/stepfun-ai/Step-3… HuggingFace: huggingface.co/stepfun-ai/St… GGUF: huggingface.co/stepfun-ai/St… ModelScope: modelscope.cn/models/stepfun… API: platform.stepfun.ai Blog: static.stepfun.com/blog/step…

    推荐理由:从开放权重和速度成本平衡切入 agent 场景,读者可对照其编程与搜索评测及本地部署支持。

5月27日周三
  1. @berryxia67

    OpenRouter 宣布完成1.13亿美元B轮融资,由CapitalG领投,a16z、Menlo Ventures、NVIDIA的NVentures、ServiceNow、MongoDB、Snowflake、Databricks等跟投;过去半年其每周token处理量从5T增至25T。作者介绍,其统一API可切换500多个模型,其中50多个免费,并提供私有聊天与可探索数据。

    引用OpenRouter (@OpenRouter)@OpenRouter

    Today we’re announcing our $113M Series B led by @CapitalGVC. Over the last 6 months, weekly volume on OpenRouter grew from 5T to 25T tokens as AI rapidly shifts from experimentation into production. We’re excited for what comes next.

    推荐理由:OpenRouter 的 B 轮融资与 token 增长数据,为观察多模型基础设施在生产环境中的需求变化提供参照。

5月26日周二
5月22日周五
  1. @alibaba_cloud73

    阿里云 Qwen3.7-Max 已在 OpenRouter 上线。OpenRouter 称其为 Qwen3.7 系列旗舰,面向以智能体为核心的工作,涵盖编码、办公与生产力任务以及长程自主执行,在编码和智能体基准上较 Qwen3.6 有明显提升,并支持显式 prompt caching。

    引用OpenRouter (@OpenRouter)@OpenRouter

    The new Qwen3.7-Max from @Alibaba_Qwen is live on OpenRouter. The flagship of the Qwen3.7 series, built for agent-centric work: coding, office and productivity tasks, and long-horizon autonomous execution. Big jumps in coding and agent benchmarks over Qwen3.6, with explicit prompt caching for repeated context.

    推荐理由:Qwen3.7-Max 已在 OpenRouter 上线,面向编码与办公的智能体场景,读者可了解这一旗舰版本的定位。

  2. @OpenRouter69

    阿里发布最新旗舰模型 Qwen3.7-Max,定位为面向 Agent 场景的基础模型。官方称其可完成前端原型、多文件重构与调试等端到端编码任务,通过 MCP 集成和多智能体编排充当办公助手,并曾在内核优化任务上连续运行 35 小时、完成 1000+ 次工具调用。模型支持 Claude Code、OpenClaw、Qwen Code 等不同脚手架,API 已在 Alibaba Model Studio 上线,也可在 Qwen Studio 试用。

    引用Qwen (@Alibaba_Qwen)@Alibaba_Qwen

    📣Meet Qwen3.7-Max — our latest flagship, made for the Agent Era. A versatile foundation for agents that actually get things done: 🧑‍💻 Coding agent, end to end. Frontend prototypes, multi-file refactors, real debugging — nails it. 🗂️ A reliable office and productivity assistant. Get your work done through MCP integrations and multi-agent orchestration. ⏱️ Long-horizon autonomy. 35 hours straight on a kernel optimization task — 1,000+ tool calls, zero hand-holding. 🔌 Scaffold-agnostic. Claude Code, OpenClaw, Qwen Code, or your own stack. Consistent reliability everywhere. API's up on Alibaba Model Studio. You can also take it for a spin on Qwen Studio. Go build something wild!🏃🏃‍♂️ 📖 Blog: qwen.ai/blog?id=qwen3.7 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.7… ⚡️ API:modelstudio.console.alibabac…

    推荐理由:原文列出端到端编码、MCP 集成与长时间自主运行等能力,可用以判断该旗舰模型在 Agent 场景中的定位。

5月20日周三
  1. @OpenRouter79

    Google DeepMind 发布 Gemini 3.5 模型家族,称其将前沿智能与现实世界行动结合,3.5 Flash 是该家族发布的第一款模型,官方称其为面向智能体与编码的最强模型。OpenRouter 转发该消息,并附上阅读模型详情的链接。

    引用Google DeepMind (@GoogleDeepMind)@GoogleDeepMind

    Introducing Gemini 3.5: our newest family of models combining frontier intelligence with real-world action. The first release is 3.5 Flash, our strongest model yet for agents and coding 🧵

    推荐理由:Google DeepMind 公布 Gemini 3.5 家族,3.5 Flash 主打智能体与编码,可看到新模型的能力侧重。