Prompt Caching, End to End

A four-part series on LLM prompt caching — the mechanism, what the major clouds actually charge, why caches silently miss, and a measured threshold that disagrees with the vendor documentation.

4 parts

  1. 1

    Prompt Caching (Part 1): What Is Actually Cached

    Prompt caching saves the server's prefill compute, not the cost of re-sending tokens. Where the KV cache comes from, why a prefix must match token for token, and why caching can cost more.

    11 min
  2. 2

    Prompt Caching (Part 2): Five Vendors, Three Clouds, One Bill

    AWS documents a 4,096-token minimum for Sonnet 4.5 on Bedrock; I measured 1,024. Multipliers, TTL tiers and cross-region behaviour for five vendors — measured kept apart from documented.

    16 min
  3. 3

    Prompt Caching (Part 3): It Is On and Not Hitting — Now What

    Cache invalidation almost never errors: inference succeeds, the logs are clean, only the bill grows. Four culprits by frequency, plus a cross-vendor hit-rate self-check you can run today.

    14 min
  4. 4

    Prompt Caching (Part 4): AWS Documents 4,096 — I Measured 1,024

    AWS gives Claude Sonnet 4.5 a 4,096-token minimum cacheable prefix. Bisecting on Bedrock put the real threshold at 1,024. Three control models matched their docs; only this row was off.

    12 min