Induction

Cache invalidation simulator

A huge part of managing an agent’s LLM costs is understanding how caching works because inference is incredibly compute intensive. So, if we can cache some of that intermediate steps, we can both save costs and speed up responses.

It turns out this is possible in the KV cache layer. When you are having a conversation with an agent, even through it seems you are posting one message at a time, the entire conversation is reprocessed from the very start. The good news is that new messages are normally added to the end of the conversation. This means a growing portion at the start of the conversation stays the same and can be cached.

So, in general, it is pretty simple: if you continually append new messages, you will pay about 5%-10% the costs for things you have already sent before.

Induction is continually working to optimize token efficiency, so we have looked deeply into how caching works (and what invalidates it) in order to make the right decisions for our customers’ trajectories.

Explicit and implicit caching

Anthropic has had explicit caching for a while now, but OpenAI has only recently added it.

In both, you tell it where there are “breakpoints” and it caches up until that point. If it sees the same content again exactly before that breakpoint, it will be a cache hit. If anything is different, it will be a miss and the cache will be invalidated back until there is a matching breakpoint.

Both services also have implicit (or automatic) caching. One way to think about this is as automatically setting breakpoints for you. Anthropic and gpt-5.6+ models do it at the end of every piece of content you send in a request. OpenAI gpt-5.5 models and earlier did it at 128-token blocks. This can be less than one message or more than one message depending on the size.

In both cases, there is now a cache-write premium of 1.25x when there is new content after a breakpoint. This gives us all even more incentive to not invalidate the cache unnecessarily.

Cache invalidation

We measured how each change to a request affects the prompt cache on OpenAI and Anthropic. The tests and raw data are available in on GitHub.

The tool below is a simulation based on the results. It lets you build a thread, change it in many different ways, and see how each change affects the cache.

Some things have no effect. For example, you can change the output_tokens or toggle on a strict tool setting. You can also generally reorder things like the JSON keys of a tool definition or exactly where you specify the system prompt.

Generally, the system prompt and tools are counted as one unit. So don’t go adding the current timestamp to the system prompt as it will invalidate the entire cache.

If you change the reasoning level or model in most models, it will invalidate the entire cache. The new exception is gpt-6 family, who allows this transition with the configuration_update input item.

All of this is to say: there are many nuances. To try and be exhaustive, we have included all the cases we could think of below and what happens.

Simulator

Cache invalidation simulator
  • cache read
  • billed full price
  • changed

A conversation that has grown normally for 3 turns

request 1request 2request 3request 3, as each API bills itgpt-5.5, approximatecaches 1536 tokens, partway into reply 1 · 25% of the conversation so far read from cachegpt-5.6caches through user 2 · 56% of the conversation so far read from cachegpt-6caches through user 2 · 56% of the conversation so far read from cacheclaude, automaticcaches through user 2 · 56% of the conversation so far read from cacheclaude, breakpointscaches through user 2 · 56% of the conversation so far read from cache

What do you want to do next?

change the conversation
change the prompt
change a setting

Predicted from rules each measured one change at a time, then checked on 50 live conversations of random changes: 435 of 441 requests read within 1.5% of the prediction. After a switch to a sibling model the predictions keep the first model’s rules. gpt-5.5’s reads stop at fixed token positions, and its rule is approximate: it reads one block (1,024 tokens) less than predicted in five of the cases we measured, and about 5% of its requests miss the cache entirely at random.

Every change we measured, on every API
ChangeOpenAI gpt-5.5OpenAI gpt-5.6OpenAI gpt-6Anthropic, automaticAnthropic, breakpoints
In a real conversation
Append the next turn (a conversation that only grows)caches up to a fixed block point (3584 tokens, of ~4134 in the previous request)caches all of the previous requestcaches all of the previous requestcaches all of the previous requestcaches all of the previous request
Append a word to user message 4 of 6caches up to a fixed block point (1536 tokens, of ~3451 unchanged), short of the previous turncaches through user message 3; reply 3 onward re-billedcaches through user message 3; reply 3 onward re-billedcaches through user message 3; reply 3 onward re-billedcaches through user message 3; reply 3 onward re-billed
Append a word to reply 4 of 5caches up to a fixed block point (3584 tokens), past the previous turncaches through user message 4; reply 4 onward re-billedcaches through user message 4; reply 4 onward re-billedcaches through user message 4; reply 4 onward re-billedcaches through user message 4; reply 4 onward re-billed
Replace user message 4 with a different one (branch)caches up to a fixed block point (1536 tokens), short of the previous turncaches through user message 3; reply 3 onward re-billedcaches through user message 3; reply 3 onward re-billedcaches through user message 3; reply 3 onward re-billedcaches through user message 3; reply 3 onward re-billed
Drop everything after reply 4 and ask something newcaches up to a fixed block point (3584 tokens), past the previous turncaches through user message 4; reply 4 onward re-billedcaches through user message 4; reply 4 onward re-billedcaches through user message 4; reply 4 onward re-billedcaches through user message 4; reply 4 onward re-billed
The same request again
Resend the request unchangedfully cachedfully cachedfully cachedfully cachedfully cached
Editing a request whose history was never sent turn by turn
Edit a tool description (first character, or one word appended)nothing cachednothing cachednothing cachednothing cachednothing cached
Edit the system prompt: change its first characternothing cachednothing cachednothing cachednothing cachedcaches tools
Edit the system prompt: append one word to its endcaches up to a fixed block point (2560 tokens)nothing cachednothing cachednothing cachedcaches tools
Edit the first history message: change its first charactercaches up to a fixed block point (2560 tokens)caches tools and system promptcaches tools and system promptnothing cachedcaches tools and system prompt
Edit the first history message: append one word to its endcaches up to a fixed block point (2560 tokens)caches tools and system promptcaches tools and system promptnothing cachedcaches tools and system prompt
Edit the third history message: change its first charactercaches up to a fixed block point (2560 tokens)caches tools and system promptcaches tools and system promptnothing cachedcaches tools, system prompt and first exchange
Edit the third history message: append one word to its endcaches up to a fixed block point (4608 tokens)caches tools and system promptcaches tools and system promptnothing cachedcaches tools, system prompt and first exchange
Append one word to the final user messagefully cachedcaches tools and system promptcaches tools and system promptnothing cachedfully cached
Same, after the previous turn was sent oncefully cachedcaches all but the last reply and final messagecaches all but the last reply and final messagecaches all but the last reply and final messagefully cached
Edit the third message after a request that diverged there was sentcaches up to a fixed block point (4608 tokens)caches tools and system promptcaches tools and system promptnothing cachedcaches tools, system prompt and first exchange
Messages
Append a reply and a new user turn (normal conversation growth)fully cachedfully cachedfully cachedfully cachedfully cached
Add a 1×1 image to the final user messagenothing cachedcaches tools and system promptcaches tools and system promptnothing cachedfully cached
Send a message as text parts instead of a string (same text)fully cachedfully cachedfully cachedfully cachedfully cached
Tools
Add a tool, remove one, reorder them, or send nonenothing cachednothing cachednothing cachednothing cachednothing cached
Toggle strict on every tool (OpenAI true → false; Anthropic unset → true)fully cachedfully cachedfully cachedfully cachedfully cached
Reorder JSON keys inside a tool’s schema (same meaning)fully cachedfully cachedfully cachedfully cachedfully cached
tool_choice: auto (default) → must call a tool (required / any)caches all but a small tail at the endcaches all but a small tail at the endcaches all but a small tail at the endnothing cachedcaches tools and system prompt
tool_choice: auto (default) → a named toolcaches all but a small tail at the endcaches all but a small tail at the endcaches all but a small tail at the endnothing cachedcaches tools and system prompt
tool_choice: auto (default) → nonecaches all but a small tail at the endcaches all but a small tail at the endcaches all but a small tail at the endfully cachedfully cached
Parallel tool calls: allowed (default) → offnothing cachednothing cachednothing cachedfully cachedfully cached
System prompt
Where the system prompt lives: top-level field → a leading message (OpenAI) / string → text blocks (Anthropic)fully cachedfully cachedfully cachedfully cachedfully cached
Append a system message mid-conversation (OpenAI developer; Anthropic system)fully cachedfully cachedfully cachedfully cachedfully cached
Model
Switch to the sibling modelnothing cachednothing cachednothing cachednothing cachednothing cached
Reasoning
Change reasoning effort, any level to any other (on gpt-6, as a configuration_update item)nothing cachednothing cachedfully cachednothing cachednothing cached
Omit effort after warming at the default (medium on OpenAI, high on Anthropic)fully cachedfully cachedfully cachedfully cachedfully cached
Change effort, then change it back (low → medium → low)fully cachedfully cachedfully cachedfully cachedfully cached
Ask for reasoning summaries (none, the default → summarized)fully cachedfully cachedfully cachedfully cachedfully cached
Thinking: adaptive (default) → disabledn/an/an/anothing cachednothing cached
Per-message effort change (beta), or just its beta headern/an/an/anothing cachednothing cached
Output
Max output tokens: 256 → 512fully cachedfully cachedfully cachedfully cachedfully cached
Output format: plain text (default) → a JSON schemanothing cachednothing cachednothing cachednothing cachedcaches tools
text.verbosity: medium (default) → highnothing cachednothing cachednothing cachedn/an/a
Add stop sequences (none by default)n/an/an/afully cachedfully cached
Routing, account and cache settings
Service tier: default → another (OpenAI priority / flex; Anthropic standard_only)nothing cachednothing cachednothing cachedfully cachedfully cached
Longer cache retention (OpenAI in-memory → 24h; Anthropic 5 min → 1 h)fully cachedfully cachedfully cachedfully cachedfully cached
Add request metadata (none by default)fully cachedfully cachedfully cachedn/an/a
Add an end-user id (OpenAI safety_identifier; Anthropic metadata.user_id)fully cachedfully cachedfully cachedfully cachedfully cached
prompt_cache_key: this trial’s key → a different key, or nonenothing cachednothing cachednothing cachedn/an/a
store: false → truefully cachedfully cachedfully cachedn/an/a
truncation: disabled (default) → autofully cachedfully cachedfully cachedn/an/a
inference_geo: default → usn/an/an/anothing cachednothing cached
Automatic caching ↔ explicit breakpoints on the same contentn/an/an/afully cachedfully cached

Request access

Leave your details and we’ll be in touch soon.