AI 导读
本周早些时候,我们展示了 GPT-5.6 和 Claude 5 在长上下文 TTFT 扩展上表现出非常不同的模式,这与它们在长上下文长度下不同的定价结构相对应。 OpenAI 最近发布了 GPT-6 Astra,它表现出同样的模式:超过 272k 输入 token 后 API 定价上升。我们收集了额外的延迟测量数据,显示出与 GPT-5.6 模型相似的曲率。
正文
Earlier this week, we showed that GPT-5.6 and Claude 5 exhibit very different long-context TTFT scaling, which corresponds to their different pricing structures at long context lengths.
OpenAI recently released GPT-6 Astra, which shows the same pattern of increased API pricing beyond 272k input tokens. We collected additional latency measurements, which show a similar curvature to that of GPT-5.6 models.
OpenAI GPT-5.6 models and Anthropic Claude 5 models have different pricing structures at long context lengths. GPT model costs increase in price past 272k input tokens, while Claude model costs remain fixed. Does this reflect an underlying difference in the architecture of these models? Our measurements of serving latency suggest so. We studied time to first token (TTFT) on these models and how it scales with increasing context length. We found a significant difference in how they scale, with GPT showing a noticeable quadratic component, while Claude models remain closer to linear.在 X 查看被引用的帖子
来源:@EpochAIResearch · x.com