跳到正文
elvis· @omarsar0 · X·· 3 小时前AI 评分46
AI 导读

Tinker 将长上下文 prefill 与采样的价格降至与短上下文相同,整体降价最高 70%,使长上下文 RL rollout 和评测成本大幅降低。GLM-5.3-Flash 和 DeepSeek-v4.1-Flash 也已上线,用于低成本长上下文任务。

正文

Bullish on this trend of making post-training more accessible.

A new post-training era is upon us.

If you work on agentic RL, long-context tasks (a big focus today) are expensive, inefficient, and don't scale well.

I've been diving into RL envs and evals for long-context tasks, and I can see this being useful.

In agent RL, rollouts use most of the tokens. Every turn re-reads the whole growing context, including tool outputs, files, and earlier turns.

Tinker just cut the price of those tokens. Long-context prefill and sampling now cost the same as short context.

This means that evaluating your trained models on long inputs also gets cheaper. Huge win here.

I believe RL will keep unlocking specialized models that slash the cost of critical agent operations. Cheaper long rollouts make them more practical to build.

Own your intelligence stack!

引用Tinker@tinkerapi
Tinkerers have been busy scaling up long-context RL! We’ve made significant improvements to Tinker’s efficiency to support those, and are passing these on with price cuts up to 70%. GLM-5.3-Flash and DeepSeek-v4.1-Flash are also live for cost-efficient long-context work.
在 X 查看被引用的帖子

来源:elvis · x.com