Prompt Cache

Prompt Cache — an inter-request cache at the API level: identical prefixes of different requests reuse the already computed KV Cache. Reading from the cache is orders of magnitude cheaper (approximately 10x for Anthropic and DeepSeek).

Provider differences:

  • Anthropic — requires an explicit cache_control point;
  • OpenAI — caching is automatic.

Cache economics — not a post-factum optimization, but an architectural limitation: the order of prompt elements is determined by the boundaries of the cache.

Related: KV Cache, Three rules of KV Cache friendliness