How Prompt Caching Works in Claude Code
Each Claude Code turn sends the accumulated context to the model again. Prompt caching lets the unchanged beginning of that context, such as system instructions, tool definitions, CLAUDE.md, and earlier conversation, be reused from a cache instead of processed from scratch, which costs less and responds faster. Changing early content, long pauses, or starting new sessions reduces cache hits.
Why caching matters in an agent
A coding agent's conversation grows with every turn: the instructions, tool definitions, and project memory stay the same, and the history of messages and tool results accumulates behind them. Each new turn resends all of it. Without caching, the model would reprocess the same content repeatedly, paying for it every time. Prompt caching lets the provider recognise an unchanged prefix and reuse its processed form.
What gets cached
Caching applies to the beginning of the request that matches a recent request exactly. In a Claude Code session that typically includes the system prompt, tool definitions including MCP servers, CLAUDE.md, and the earlier part of the conversation. Each turn extends the cacheable prefix, so in a steady session most of the context is served from cache and only the newest messages are processed fresh. Claude Code manages caching automatically; you do not need to configure it.
What breaks the cache
- Inactivity. Cached entries expire after a period without use, so returning to a session after a long break processes the context again.
- Changing early content. Editing CLAUDE.md, adding or removing MCP servers, or other changes near the start of the context alter the prefix, so everything after the change must be reprocessed.
- New sessions. A fresh session starts a new cache.
- Compaction. Summarising the conversation replaces the history, creating a new prefix.
None of these are reasons to avoid fresh sessions or compaction when they help; they are trade-offs to be aware of.
Habits that help
- Keep CLAUDE.md and the MCP server set stable during a working session; make configuration changes between sessions.
- Work in focused bursts on a task rather than leaving a session idle for long periods mid-task.
- Keep the stable prefix lean, since even cached tokens have a cost.
- Watch usage reports to see how much input is served from cache.
Reading the usage numbers
Usage reports separate input tokens processed fresh, tokens written to the cache, and tokens read from the cache. In a healthy steady session, cache reads dominate input and fresh input stays small. A spike in fresh or cache-write tokens points to an event that reset the prefix: a pause long enough for expiry, a configuration change, a compaction, or a new session. Matching those spikes to what you did is the fastest way to find habits worth changing, without guesswork.
Caching versus retrieval
Caching makes repeated context cheaper. It does not reduce how much context there is, and a large context still dilutes the model's attention. Retrieval addresses the other half: instead of loading whole documents or many files, the agent fetches the few relevant passages. The two combine well. A lean, stable prefix gets cached, and retrieved snippets keep each turn's new content small. For repeated questions across sessions, a retrieval layer can also return stored answers instead of regenerating them, which caching alone cannot do.
Frequently asked questions
- Does Claude Code use prompt caching automatically?
- Yes. Claude Code structures requests so the stable parts of the context, such as system instructions, tool definitions, project memory, and earlier conversation, can be served from the cache on later turns. No configuration is needed, though habits such as keeping configuration stable during a session improve cache hits.
- Why did my Claude Code session suddenly cost more?
- Common causes are a cache miss after a long pause, changes to CLAUDE.md or MCP servers mid-session, a compaction that created a new prefix, or simply a much larger context from reading big files or verbose outputs. Usage reports showing cached versus uncached input help identify which.
- How long does the prompt cache last?
- Cache entries expire after a period of inactivity, typically minutes rather than hours, and the timer resets when the cached content is used. Check Anthropic's current documentation for exact durations and options, since they can change. In practice, steady work within a session keeps the cache warm.
- Is prompt caching the same as semantic caching?
- No. Prompt caching reuses the processed form of an identical prompt prefix within a model provider. Semantic caching stores answers and returns them for new questions with similar meaning, avoiding a model call entirely. They solve different parts of the cost problem and can be used together.
- Does prompt caching affect output quality?
- No. Caching changes how the provider processes an identical prompt prefix, not what the model receives or how it responds. The model sees the same context whether or not parts of it come from the cache. The benefits are lower cost and faster responses for the repeated portion.
- Should I avoid compaction to keep the cache?
- Not when compaction helps. Compaction creates a new prefix, so the next turn reprocesses the summary, but it also shrinks the context, which lowers cost on every following turn and can improve focus. Compact when the session has accumulated a lot of stale material; the short-term cache miss is usually worth it.
- Does a longer CLAUDE.md cost more with caching?
- Less than without caching, but still something. Cached tokens are billed at a reduced rate, not free, and they still occupy the context window, diluting attention. Keep CLAUDE.md concise, and put detailed background in documentation the agent retrieves only when relevant, rather than in the prefix of every request.