Not Diamond 发布其模型路由方法论,可在达到 Opus xhigh 质量的同时将智能体成本降低 20–80%。该路由将模型选择视为序列决策问题,每步基于会话状态、消息与 token 数、任务复杂度、KV cache 状态及中间奖励信号,预测各模型与推理档位的未来收益和成本。其基准测试覆盖模拟用户轮次、任务复杂度变化与可能导致 KV cache 失效的响应延迟。
Not Diamond just released the methodology behind their model routing which gets Opus xhigh quality while cutting agent costs 20–80%.
Model routing becomes much harder decision for long-running coding agents. So Not Diamond is benchmarking the router across simulated user turns, changing task complexity, and response delays that can expire the KV cache.
And harder than it seems, switch models too aggressively and you can throw away the KV cache. Pick one model from the opening prompt and you are assuming task complexity stays fixed for the rest of the session.
So Not Diamond is treating routing as a sequential decision problem instead.
At each step, its router predicts future reward and cost for a particular model and reasoning effort, using current and previous session state, message and token counts, task complexity, KV-cache state, and intermediate reward signals.
Read more detail on their technical report.
来源:@rohanpaul_ai · x.com