Input, output and cache tokens in Claude Code, explained
Why Claude Code's token totals look so huge, what input, output and cache tokens are, and how to read the numbers on your receipt.
2 min read · updated 2026-10-02
Open your Claude Code usage for the first time and the number is usually startling: hundreds of millions of tokens for a few weeks of work. Almost all of that is not what you typed or what the model wrote. Here is what the numbers mean.
What is a token?
A token is the unit a language model reads and writes: a chunk of text, usually a few characters long. A short English word is often one token, a long word is several, and code tends to use more tokens per line than prose because of symbols and indentation. Models are limited and billed in tokens, not in words or lines.
Input and output
- Input tokens are everything sent to the model for a reply: your message, the files it has read, tool results, and the conversation so far.
- Output tokens are what the model writes back: text, plans and code edits.
An agent like Claude Code does not send just your last message. On every step it sends the whole working context again, so a long session sends the same early material over and over. That repetition is why input dwarfs output.
What the cache changes
Sending the same context repeatedly would be slow and expensive, so the API supports prompt caching: a large, unchanging prefix is stored after the first request and reused on the following ones. Two numbers describe it:
- Cache write (
cache_creation_input_tokens): tokens stored in the cache the first time they are seen. - Cache read (
cache_read_input_tokens): tokens served from the cache on later requests. Cache reads are cheaper and faster than fresh input, which is the whole point.
In a long Claude Code session the same context is read from the cache again and again, so cache reads usually make up the large majority of all tokens counted. This is normal. It does not mean you wrote or received that much text.
How a receipt counts them
On a showmytokens receipt the headline total adds everything up: input, output, cache reads and cache writes. The card also splits it into two parts so the big number is easy to interpret:
| Part | What it is |
|---|---|
| Input + output | The fresh tokens: what you sent and what the model wrote. |
| Cache | Cache reads and writes: context reused across steps. |
Both are shown per day and per model. If you want a number closer to "how much did the model actually produce", look at output tokens. If you want "how much context did Claude Code move around", look at the total.
Tokens are not dollars
A token total is not a bill. Fresh input, output, cache writes and cache reads are all priced differently, and plans may meter usage in their own way. Treat the receipt as a measure of activity, and check your plan or API dashboard for cost.
See your own numbers
To count your own tokens from the logs, see how to check your Claude Code token usage. To put them on a card, get your receipt.