A huge part of managing an agent’s LLM costs is understanding how caching works because inference is incredibly compute intensive. So, if we can cache some of that intermediate steps, we can both save costs and speed up responses.
It turns out this is possible in the KV cache layer. When you are having a conversation with an agent, even through it seems you are posting one message at a time, the entire conversation is reprocessed from the very start. The good news is that new messages are normally added to the end of the conversation. This means a growing portion at the start of the conversation stays the same and can be cached.
So, in general, it is pretty simple: if you continually append new messages, you will pay about 5%-10% the costs for things you have already sent before.
Induction is continually working to optimize token efficiency, so we have looked deeply into how caching works (and what invalidates it) in order to make the right decisions for our customers’ trajectories.
Explicit and implicit caching
Anthropic has had explicit caching for a while now, but OpenAI has only recently added it.
In both, you tell it where there are “breakpoints” and it caches up until that point. If it sees the same content again exactly before that breakpoint, it will be a cache hit. If anything is different, it will be a miss and the cache will be invalidated back until there is a matching breakpoint.
Both services also have implicit (or automatic) caching. One way to think about this is as automatically setting breakpoints for you. Anthropic and gpt-5.6+ models do it at the end of every piece of content you send in a request. OpenAI gpt-5.5 models and earlier did it at 128-token blocks. This can be less than one message or more than one message depending on the size.
In both cases, there is now a cache-write premium of 1.25x when there is new content after a breakpoint. This gives us all even more incentive to not invalidate the cache unnecessarily.
Cache invalidation
We measured how each change to a request affects the prompt cache on OpenAI and Anthropic. The tests and raw data are available in on GitHub.
The tool below is a simulation based on the results. It lets you build a thread, change it in many different ways, and see how each change affects the cache.
Some things have no effect. For example, you can change the output_tokens or toggle on a strict tool setting. You can also generally reorder things like the JSON keys of a tool definition or exactly where you specify the system prompt.
Generally, the system prompt and tools are counted as one unit. So don’t go adding the current timestamp to the system prompt as it will invalidate the entire cache.
If you change the reasoning level or model in most models, it will invalidate the entire cache. The new exception is gpt-6 family, who allows this transition with the configuration_update input item.
All of this is to say: there are many nuances. To try and be exhaustive, we have included all the cases we could think of below and what happens.
Simulator
- cache read
- billed full price
- changed
A conversation that has grown normally for 3 turns
What do you want to do next?
Predicted from rules each measured one change at a time, then checked on 50 live conversations of random changes: 435 of 441 requests read within 1.5% of the prediction. After a switch to a sibling model the predictions keep the first model’s rules. gpt-5.5’s reads stop at fixed token positions, and its rule is approximate: it reads one block (1,024 tokens) less than predicted in five of the cases we measured, and about 5% of its requests miss the cache entirely at random.
Every change we measured, on every API
| Change | OpenAI gpt-5.5 | OpenAI gpt-5.6 | OpenAI gpt-6 | Anthropic, automatic | Anthropic, breakpoints |
|---|---|---|---|---|---|
| In a real conversation | |||||
| Append the next turn (a conversation that only grows) | caches up to a fixed block point (3584 tokens, of ~4134 in the previous request) | caches all of the previous request | caches all of the previous request | caches all of the previous request | caches all of the previous request |
| Append a word to user message 4 of 6 | caches up to a fixed block point (1536 tokens, of ~3451 unchanged), short of the previous turn | caches through user message 3; reply 3 onward re-billed | caches through user message 3; reply 3 onward re-billed | caches through user message 3; reply 3 onward re-billed | caches through user message 3; reply 3 onward re-billed |
| Append a word to reply 4 of 5 | caches up to a fixed block point (3584 tokens), past the previous turn | caches through user message 4; reply 4 onward re-billed | caches through user message 4; reply 4 onward re-billed | caches through user message 4; reply 4 onward re-billed | caches through user message 4; reply 4 onward re-billed |
| Replace user message 4 with a different one (branch) | caches up to a fixed block point (1536 tokens), short of the previous turn | caches through user message 3; reply 3 onward re-billed | caches through user message 3; reply 3 onward re-billed | caches through user message 3; reply 3 onward re-billed | caches through user message 3; reply 3 onward re-billed |
| Drop everything after reply 4 and ask something new | caches up to a fixed block point (3584 tokens), past the previous turn | caches through user message 4; reply 4 onward re-billed | caches through user message 4; reply 4 onward re-billed | caches through user message 4; reply 4 onward re-billed | caches through user message 4; reply 4 onward re-billed |
| The same request again | |||||
| Resend the request unchanged | fully cached | fully cached | fully cached | fully cached | fully cached |
| Editing a request whose history was never sent turn by turn | |||||
| Edit a tool description (first character, or one word appended) | nothing cached | nothing cached | nothing cached | nothing cached | nothing cached |
| Edit the system prompt: change its first character | nothing cached | nothing cached | nothing cached | nothing cached | caches tools |
| Edit the system prompt: append one word to its end | caches up to a fixed block point (2560 tokens) | nothing cached | nothing cached | nothing cached | caches tools |
| Edit the first history message: change its first character | caches up to a fixed block point (2560 tokens) | caches tools and system prompt | caches tools and system prompt | nothing cached | caches tools and system prompt |
| Edit the first history message: append one word to its end | caches up to a fixed block point (2560 tokens) | caches tools and system prompt | caches tools and system prompt | nothing cached | caches tools and system prompt |
| Edit the third history message: change its first character | caches up to a fixed block point (2560 tokens) | caches tools and system prompt | caches tools and system prompt | nothing cached | caches tools, system prompt and first exchange |
| Edit the third history message: append one word to its end | caches up to a fixed block point (4608 tokens) | caches tools and system prompt | caches tools and system prompt | nothing cached | caches tools, system prompt and first exchange |
| Append one word to the final user message | fully cached | caches tools and system prompt | caches tools and system prompt | nothing cached | fully cached |
| Same, after the previous turn was sent once | fully cached | caches all but the last reply and final message | caches all but the last reply and final message | caches all but the last reply and final message | fully cached |
| Edit the third message after a request that diverged there was sent | caches up to a fixed block point (4608 tokens) | caches tools and system prompt | caches tools and system prompt | nothing cached | caches tools, system prompt and first exchange |
| Messages | |||||
| Append a reply and a new user turn (normal conversation growth) | fully cached | fully cached | fully cached | fully cached | fully cached |
| Add a 1×1 image to the final user message | nothing cached | caches tools and system prompt | caches tools and system prompt | nothing cached | fully cached |
| Send a message as text parts instead of a string (same text) | fully cached | fully cached | fully cached | fully cached | fully cached |
| Tools | |||||
| Add a tool, remove one, reorder them, or send none | nothing cached | nothing cached | nothing cached | nothing cached | nothing cached |
Toggle strict on every tool (OpenAI true → false; Anthropic unset → true) | fully cached | fully cached | fully cached | fully cached | fully cached |
| Reorder JSON keys inside a tool’s schema (same meaning) | fully cached | fully cached | fully cached | fully cached | fully cached |
tool_choice: auto (default) → must call a tool (required / any) | caches all but a small tail at the end | caches all but a small tail at the end | caches all but a small tail at the end | nothing cached | caches tools and system prompt |
tool_choice: auto (default) → a named tool | caches all but a small tail at the end | caches all but a small tail at the end | caches all but a small tail at the end | nothing cached | caches tools and system prompt |
tool_choice: auto (default) → none | caches all but a small tail at the end | caches all but a small tail at the end | caches all but a small tail at the end | fully cached | fully cached |
| Parallel tool calls: allowed (default) → off | nothing cached | nothing cached | nothing cached | fully cached | fully cached |
| System prompt | |||||
| Where the system prompt lives: top-level field → a leading message (OpenAI) / string → text blocks (Anthropic) | fully cached | fully cached | fully cached | fully cached | fully cached |
Append a system message mid-conversation (OpenAI developer; Anthropic system) | fully cached | fully cached | fully cached | fully cached | fully cached |
| Model | |||||
| Switch to the sibling model | nothing cached | nothing cached | nothing cached | nothing cached | nothing cached |
| Reasoning | |||||
Change reasoning effort, any level to any other (on gpt-6, as a configuration_update item) | nothing cached | nothing cached | fully cached | nothing cached | nothing cached |
| Omit effort after warming at the default (medium on OpenAI, high on Anthropic) | fully cached | fully cached | fully cached | fully cached | fully cached |
| Change effort, then change it back (low → medium → low) | fully cached | fully cached | fully cached | fully cached | fully cached |
| Ask for reasoning summaries (none, the default → summarized) | fully cached | fully cached | fully cached | fully cached | fully cached |
| Thinking: adaptive (default) → disabled | n/a | n/a | n/a | nothing cached | nothing cached |
| Per-message effort change (beta), or just its beta header | n/a | n/a | n/a | nothing cached | nothing cached |
| Output | |||||
| Max output tokens: 256 → 512 | fully cached | fully cached | fully cached | fully cached | fully cached |
| Output format: plain text (default) → a JSON schema | nothing cached | nothing cached | nothing cached | nothing cached | caches tools |
text.verbosity: medium (default) → high | nothing cached | nothing cached | nothing cached | n/a | n/a |
| Add stop sequences (none by default) | n/a | n/a | n/a | fully cached | fully cached |
| Routing, account and cache settings | |||||
Service tier: default → another (OpenAI priority / flex; Anthropic standard_only) | nothing cached | nothing cached | nothing cached | fully cached | fully cached |
| Longer cache retention (OpenAI in-memory → 24h; Anthropic 5 min → 1 h) | fully cached | fully cached | fully cached | fully cached | fully cached |
Add request metadata (none by default) | fully cached | fully cached | fully cached | n/a | n/a |
Add an end-user id (OpenAI safety_identifier; Anthropic metadata.user_id) | fully cached | fully cached | fully cached | fully cached | fully cached |
prompt_cache_key: this trial’s key → a different key, or none | nothing cached | nothing cached | nothing cached | n/a | n/a |
store: false → true | fully cached | fully cached | fully cached | n/a | n/a |
truncation: disabled (default) → auto | fully cached | fully cached | fully cached | n/a | n/a |
inference_geo: default → us | n/a | n/a | n/a | nothing cached | nothing cached |
| Automatic caching ↔ explicit breakpoints on the same content | n/a | n/a | n/a | fully cached | fully cached |