Tinker 将长上下文 prefill 与采样的价格降至与短上下文相同,整体降价最高 70%,使长上下文 RL rollout 和评测成本大幅降低。GLM-5.3-Flash 和 DeepSeek-v4.1-Flash 也已上线,用于低成本长上下文任务。
Bullish on this trend of making post-training more accessible.
A new post-training era is upon us.
If you work on agentic RL, long-context tasks (a big focus today) are expensive, inefficient, and don't scale well.
I've been diving into RL envs and evals for long-context tasks, and I can see this being useful.
In agent RL, rollouts use most of the tokens. Every turn re-reads the whole growing context, including tool outputs, files, and earlier turns.
Tinker just cut the price of those tokens. Long-context prefill and sampling now cost the same as short context.
This means that evaluating your trained models on long inputs also gets cheaper. Huge win here.
I believe RL will keep unlocking specialized models that slash the cost of critical agent operations. Cheaper long rollouts make them more practical to build.
Own your intelligence stack!
Tinkerers have been busy scaling up long-context RL! We’ve made significant improvements to Tinker’s efficiency to support those, and are passing these on with price cuts up to 70%. GLM-5.3-Flash and DeepSeek-v4.1-Flash are also live for cost-efficient long-context work.在 X 查看被引用的帖子
来源:elvis · x.com