跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-23AI 评分52
AI 导读

在 OpenRouter 上,智能体消耗 token 的速度接近人类的 5 倍,自 2 月以来用量增长约 14 倍,并已超过人类 token 总量。作者指出,智能体执行长任务时会反复发送同一段背景指令形成缓存命中,模型供应商把该缓存块存在内存中并按很小比例计费,因此 OpenRouter 上超过 85% 的智能体 token 属于这类廉价复用。

正文

Agents are consuming tokens at nearly 5x the human rate, while their usage has exploded ~14X since February.

Once agents became the majority of tokens on OpenRouter, the router's ability to play suppliers off each other may be much less.

OpenRouter sits between developers and the companies that run the models. It makes money by price shopping, sending each query to the cheapest provider at the time.

When the next question is a single question, this is fine; there is no need for any continuity between the questions.

But, the behavior of an agent doing a long task is different: it resends the same long block of background instructions at each stage (i.e. cache hit).

A cache hit only lives on the machine still holding the warm prefix, and rerouting mid-task will mean paying the full pre-fill over again.

i.e. the model provider stores that cache-hit block in memory and charges only a small fraction to reuse it.

Which is why more than 85% of agent tokens on OpenRouter are these cheap reuses (cached prompt) rather than fresh ones.

The copy is stored on the servers of a single company . If the job is transferred to a cheaper competitor halfway through , the stored copy is discarded and the full block is repaid .

That will mean the agent remains with the provider they started the task with until the task is finished and the threat from the router to take their business elsewhere is eliminated.

So looks like the discounts routers can squeeze out of model providers many shrink on agent traffic well before they shrink anywhere else.

来源:@rohanpaul_ai · x.com