Lesson 12 / 25

Prompt Caching

Reuse a stable prompt prefix across turns to cut cost and latency.

Pay less for the part that does not change

The start of every request (system prompt, tool definitions, early history) is identical turn after turn. Prompt caching lets the provider remember that prefix, so cache reads are billed at a lower rate and are faster. With the Claude API you mark a block with cache_control. Put stable content first and changing content last, because only an unchanged prefix can be reused.

Marking a cacheable block

The marker goes on the last block of the stable prefix. Minimum cacheable sizes, lifetimes and prices depend on the model, so check current documentation.

system = [{
    "type": "text",
    "text": LONG_INSTRUCTIONS,
    "cache_control": {"type": "ephemeral"},   # cache everything up to here
}]
reply = client.messages.create(
    model=MODEL, max_tokens=1024, system=system, tools=TOOLS, messages=messages)

Do not edit the prefix

Inserting a timestamp or changing the order of tools at the top of the prompt breaks the cache for every following turn. Keep the early part byte-for-byte stable.

Quick check: Where should changing content go for caching to work well?

  • At the very start of the prompt
  • After the stable prefix
  • In the API key
  • Nowhere
Answer

After the stable prefix — Only an identical prefix can be reused, so volatile parts must come after it.