Better prompt caching for GPT-6: higher hit rates, diagnostics, and agent-friendly controls
- Published
- 2026-09-22
- Type
- Product
- Organization
- OpenAI
- Source
- OpenAI
Why it matters
OpenAI published Better prompt caching for GPT-6 (primary: https://openai.com/index/better-prompt-caching-for-gpt-6 ; RSS pubDate 2026-09-22T21:00:00Z). For persistent multi-turn agents, shared instructions, tool definitions, and context can be cached across requests—cutting latency and offering up to ~90% discount on cached input tokens. GPT-6 family raises default hit rates; eligible shared prefixes reused within a ~30-minute window get cache discounts. New Prompt Caching Dashboard and diagnostics explain misses (e.g. tools_changed). Optional controls: explicit cache breakpoints; change reasoning effort without breaking cache; keep tool defs/schemas/order stable (prefer allowed_tools / tool_choice=none over removing defs); append developer messages; optional prewarming. Customer quotes on the page are not independently verified here. Editorial: infra/observability for agent economics, not a new falsifiable AGI capability milestone → home_latest (not home_week). Do not confuse with e/669 Sol & Luna model launch.
Related developments
Aggregated by AGI Pulse. Titles and links only; no full text is republished.