跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-04AI 评分44
AI 导读

同一问题不同模型可能消耗 5K 到 50K+ token,叠加不同分词器、隐藏推理 token、工具调用与重试后,"每百万 token 价格"对实际工作负载已失去意义。智能体场景下 token 低效会逐级累积,多耗 2 倍 token 可能带来远超 2 倍的上下文、延迟与后续推理成本。更合理的经济指标是"每成功完成任务成本":模型+推理+工具+重试总成本除以成功完成任务数。

正文

A token is no longer a comparable unit of work.

2 models can solve roughly the same problem, yet 1 may need 5K tokens while another needs 50K+.

Add different tokenizers, hidden reasoning tokens, tool calls and retries, and "$ per 1M tokens" has no meaning for the workload.

There is an even bigger implication for agents. Token inefficiency compounds. A verbose output from step 1 often becomes input for step 2, then gets carried into step 3, step 4 and beyond.

So a model using 2x more tokens does not necessarily create only 2x more expense. It can also increase context size, generation latency, tool-call overhead and the cost of every later reasoning step. Token efficiency becomes much more valuable as workflows get longer.

The better economic measure is probably cost per successful task at a required quality level:

total model + reasoning + tool + retry cost ÷ successful completed tasks.

Token economy may be moving toward something similar to a semantic efficiency metric: how much useful work, intelligence or task completion you get from each dollar

来源:@rohanpaul_ai · x.com