AGI Pulse
ProductOpenAI · OpenAI

Better prompt caching for GPT-6: higher hit rates, diagnostics, and agent-friendly controls

Published
2026-09-22
Type
Product
Organization
OpenAI
Source
OpenAI

Why it matters

OpenAI published Better prompt caching for GPT-6 (primary: https://openai.com/index/better-prompt-caching-for-gpt-6 ; RSS pubDate 2026-09-22T21:00:00Z). For persistent multi-turn agents, shared instructions, tool definitions, and context can be cached across requests—cutting latency and offering up to ~90% discount on cached input tokens. GPT-6 family raises default hit rates; eligible shared prefixes reused within a ~30-minute window get cache discounts. New Prompt Caching Dashboard and diagnostics explain misses (e.g. tools_changed). Optional controls: explicit cache breakpoints; change reasoning effort without breaking cache; keep tool defs/schemas/order stable (prefer allowed_tools / tool_choice=none over removing defs); append developer messages; optional prewarming. Customer quotes on the page are not independently verified here. Editorial: infra/observability for agent economics, not a new falsifiable AGI capability milestone → home_latest (not home_week). Do not confuse with e/669 Sol & Luna model launch.

Read original →

Related developments

Aggregated by AGI Pulse. Titles and links only; no full text is republished.