OpenAI has updated prompt caching in GPT-6 — the reuse of the same context between requests: the system now finds reusable context more often by default, and developers can monitor cache performance and investigate requests for which the cache has no ready response.
The company now offers discounts on cached input tokens for shared request prefixes reused within 30 minutes. Explicit cache breakpoints let developers choose which prompt prefixes to reuse.
GPT-6 models can change reasoning effort between responses without breaking the cache if they append a configuration_update and leave request-level reasoning effort unchanged.
Claim check:
- OpenAI has updated prompt caching in GPT-6: the system now finds reusable context more often by default, and developers can monitor cache performance and investigate requests for which the cache has no ready response. (confirmed by the publication itself: evidence; «With the GPT‑6 family, we launched an improved prompt caching system that delivers higher cache hit rates by default. We now give cache discounts for eligible shared prefixes reused within a 30-minute window. We’re also introducing new tools to help developers monitor cache performance, diagnose misses, and choose how much of a prompt to cache.»)
- The company now offers discounts on cached input tokens for shared request prefixes reused within 30 minutes. (confirmed by the publication itself: evidence; «We now give cache discounts for eligible shared prefixes reused within a 30-minute window.»)
- Explicit cache breakpoints let developers choose which prompt prefixes to reuse. (confirmed by the publication itself: evidence; «Explicit cache breakpoints let you choose which prompt prefixes to reuse.»)
- GPT-6 models can change reasoning effort between responses without breaking the cache if they append a configuration_update and leave request-level reasoning effort unchanged. (confirmed by the publication itself: evidence; «On GPT‑6 models, you can now change reasoning effort between responses without breaking cache. Raise effort for a harder task or lower it for a routine follow-up by appending a configuration_update while leaving request-level reasoning effort unchanged.»)
Primary sources:
score 64.6 out of 100 · kind: guide