English

Product LaunchesOpenAIGPT-6

OpenAI Enhances Prompt Caching for GPT-6 to Boost Agent Performance and Reduce Costs

OpenAI has introduced an improved prompt caching system designed to enhance the performance and cost-efficiency of persistent agents. These agents, which perform complex tasks like refactoring codebases or producing research-based documents, often reuse significant context across multiple API requests. By caching this shared context, OpenAI reduces response times and provides developers with discounts of up to 90% on cached input tokens.

The new system for the GPT-6 family offers higher cache hit rates by default and introduces several tools for developers. A new Prompt Caching Dashboard allows users to monitor cache performance, track hit rates over time, and use an input composition chart to compare cached versus uncached tokens. Additionally, a prompt caching diagnostics tool helps developers understand the causes of unexpected cache misses.

To optimize integration, developers can now use explicit cache breakpoints to choose exactly which prompt prefixes to reuse. For GPT-6 models, users can also adjust reasoning effort between responses using a configuration_update without breaking the existing cache, allowing for more flexible task management. Furthermore, a "prewarming" feature is available, which prepares known context ahead of time to reduce latency for the first response.

Sources

  1. Better prompt caching for GPT-6 (OpenAI News, 2026-09-22)
  2. prompt caching guide