OpenAI 为 GPT-6 的提示词缓存增加更高命中率、诊断、断点、预热与保持缓存的推理改动。长时运行的智能体常重复发送指令、工具定义和早前上下文,缓存可避免重复计算这些前缀,最多降低 90% 的缓存输入 token 成本,新仪表盘展示缓存命中率与缓存、未缓存 token 量。原文也提醒,高缓存命中率仍部分取决于应用侧设计。
OpenAI is making GPT-6 prompt caching more reliable and controllable, adding higher hit rates, diagnostics, breakpoints, prewarming, and cache-preserving reasoning changes.
Long-running agents often resend instructions, tool definitions, and earlier context, so caching avoids recomputing those prefixes across successive API calls.
That reuse can cut cached-input token costs by up to 90%, while a new dashboard exposes cache-hit rates and cached versus uncached token volume.
ofcourse, a high cache hit rate is still remains partly an application-design problem. because GPT-6’s improved caching alone doesn’t guarantee good economics; developers still need to structure long-running agents so stable instructions and tools remain reusable.
来源:@rohanpaul_ai · x.com