跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 14 天前AI 评分57
AI 导读

OpenAI 为 GPT-6 的提示词缓存增加更高命中率、诊断、断点、预热与保持缓存的推理改动。长时运行的智能体常重复发送指令、工具定义和早前上下文,缓存可避免重复计算这些前缀,最多降低 90% 的缓存输入 token 成本,新仪表盘展示缓存命中率与缓存、未缓存 token 量。原文也提醒,高缓存命中率仍部分取决于应用侧设计。

正文

OpenAI is making GPT-6 prompt caching more reliable and controllable, adding higher hit rates, diagnostics, breakpoints, prewarming, and cache-preserving reasoning changes.

Long-running agents often resend instructions, tool definitions, and earlier context, so caching avoids recomputing those prefixes across successive API calls.

That reuse can cut cached-input token costs by up to 90%, while a new dashboard exposes cache-hit rates and cached versus uncached token volume.

ofcourse, a high cache hit rate is still remains partly an application-design problem. because GPT-6’s improved caching alone doesn’t guarantee good economics; developers still need to structure long-running agents so stable instructions and tools remain reusable.

来源:@rohanpaul_ai · x.com