Where the tokens go: why one AI job re-read 78 million tokens
Tokens, calls, context, and compaction, explained with one real run of my listing pipeline, and the three ways I plan to cut what it reads.
Part three put a price on one run of my listing pipeline: a median of $42 a batch at API list prices. That made me ask what I was paying for.
The answer isn’t thinking. It’s re-reading.
This part shows where the tokens go, using one real run, and explains the handful of terms you need to follow it. If you’re paying for any AI automation, or about to, these are the terms that decide the bill.
You don’t need a technical background for this one. If you can read a phone bill, you can read this.
Which steps use the AI
Here’s the pipeline from part one again, with the AI’s steps picked out. The AI agent runs the whole thing, so every step uses some tokens. But most of its calls happen in these three, where it’s reading photos, researching cards, and working a website:
The pipeline, start to finish
- MeA person decides
- CodeA fixed program decides, the same way every time
- AIThe AI uses its eyes or its judgment
Before the run
Step 1: Sort and photographMe
The run, at 2:00 a.m.
Step 2: Check the folderCode
Step 3: Give every card an addressCode
Step 4: Publish the photosCode
Step 5: Identify each cardAI
Step 6: Check the AI’s workCode
Step 7: Research the market priceAI
Step 8: Build the eBay fileCode
Step 9: Upload, then read every draft backAI
After the run
Step 10: Review and publishMe
Token Talk
AI is billed by the token. A token is a piece of a word, about three-quarters of one on average.
There are two kinds, and they’re priced differently:
- Input is everything the AI reads.
- Output is what it writes.
On the model my worker uses now, list price is $2 per million input tokens and $10 per million output tokens. Output costs five times more, so it’s natural to think the output is the expensive part.
It isn’t, because there is a third price to consider: cached input. That’s input the AI was already sent a moment ago. It costs $0.10 per million, a twentieth of the normal rate. Super cheap, right? Hold on to that one. It’s where most of the money goes, and I’ll explain it in this example.
A job is hundreds of calls
An AI agent doesn’t do a job in one go. It works in calls. On each call it reads what it’s sent, takes one step, such as running a command or clicking a button, and stops. Then the next call starts.
The part that changes everything: the AI remembers nothing between calls.
Picture playing fetch with a dog that forgets the game every time it drops the ball. Before each throw, you have to explain all of it again: what fetch is, where the yard ends, and every throw so far. Then you get one throw.
It’s a very good dog. It just starts from zero every time.
So every call is sent the whole story again. Its instructions. Every document it opened. Every step it took and every result that came back. All of that together is the context.
Step 1 sends a little. Step 300 sends everything from steps 1 through 299, and then does one more thing.
Where the dollars go
A batch is one stack of fifty cards. A batch run is one run of the pipeline working through a batch. Most batches take one run. A few take more.
Over 34 days, the pipeline made 14 batch runs. I added up what they would have cost at the provider’s pay-as-you-go prices, then sorted every dollar by the kind of token it paid for:
What the money paid for
- 73% Re-read (cached) input
- 19% New input
- 8% Output
Nearly three-quarters of the cost is re-reading. It’s cheap per token and enormous in volume.
Watch one run
This is one run of the worker, on September 27, 2026. Each point along the bottom is a call. The height is how many tokens that call had to re-read before it could do anything new.
One long session
- Tokens re-read on each call
- Compaction
- Startup reading
Read it left to right:
- It starts at 39,000 tokens, before any work is done. Nearly all of that is Codex itself: the app’s built-in instructions, plus a description of every tool it’s allowed to use, such as running a command or working a browser. My nightly note is in there too, but it’s only about 1,200 tokens. The procedure and my other documents haven’t been opened yet. (This isn’t the harness from part one. That’s my code. This is the app’s own rulebook, and every session starts with it.)
- Loading. The nightly note tells it to read about a dozen of my documents before acting. That’s roughly 50,000 more tokens, and it carries them on every call after.
- The climb. Every step adds its result to the context. The typical call adds under 1,000 tokens. The average is 2,200, because a few steps bring back a whole web page or a long file.
- The drop. Somewhere past 200,000 tokens, Codex steps in. More on that below.
The run made 540 calls, and the average call re-read 145,000 tokens. Multiply those and you get the headline: 78 million input tokens for one run.
What one average call sends today
- 39K The starting load
- 106K Everything read and done so far in the run
Compaction: a reset you don’t control
A model can only read so much at once. When the context gets close to that limit, the tool replaces it with a summary and carries on. That’s compaction, and it’s each dashed red line in the chart.
Compaction keeps a long job alive. But it has two problems:
- It’s late. By the time it happens, hundreds of calls have already paid to re-read a huge context.
- It’s blunt. It happens wherever the job happens to be, and a summary can drop a detail the next step needed.
That run was compacted five times. I didn’t choose any of them.
Three ways to make the AI read less
Reading is 92% of what a run costs. So the way to a cheaper run is to make the AI read less. There’s a second payoff too: the less it carries, the less often Codex has to step in and compact.
How much the AI reads comes down to two numbers: how many calls it makes, and how much it re-reads on each one. That leaves three things to change.
1. Reset sooner, on purpose. Split the job into short sessions. Each one does a clean piece of work, writes the result to a file, and ends. That’s a handoff. The next session starts small and reads the file, not the whole history.
2. Climb slower. Have each step bring back only what’s needed. A step that returns one price, not a whole page, adds a few hundred tokens where it used to add thousands.
3. Make fewer calls. Give the agent saved routines, so one call does what used to take five or ten.
There’s a fourth, and it’s the cheapest to try: pick the model with the re-read price in mind. Cached input is $0.10 per million on the model I use now. On the most expensive one I tried, it’s $1.00. Same job, ten times the price for the biggest slice.
The plan: thirteen short sessions
A batch of fifty cards already has natural seams. So the plan is to cut along them:
| Session | Work | How many |
|---|---|---|
| Open | Check the photos, claim the batch, and publish the images | 1 |
| Identify and price | Research ten cards, then hand off | 5 |
| Build | Merge the research, build the draft file, and upload it | 1 |
| Audit drafts | Read ten saved drafts back, then hand off | 5 |
| Close | Check every receipt and close the batch | 1 |
Here’s what that should look like, drawn to the same scale as the real run:
The same work as thirteen short sessions
- Tokens re-read on each call (projected)
- Handoff
What one average call would send in a short session
- 39K The starting load
- 5K A one-page instruction card
- 46K This chunk so far
The same work, about 48 million input tokens where there were 78 million.
This is a projection, not a result. It assumes the same number of calls and the same growth per call as the real run, and a one-page instruction card for each session in place of the dozen documents. Until it’s built and measured, treat it as a plan.
What turned out not to matter
- How hard the AI thinks. Output was 8% of the cost. Turning the reasoning down wouldn’t move much.
- Quiet nights. A night with nothing to list is about ten calls. It’s the long runs that cost.
Five questions to ask about any AI automation
- How many calls does one job take?
- What does each call re-read?
- Where does the context reset, and who decided that?
- Which model runs it, and what does it charge for cached input?
- Can I see what one run cost?
If the person building it can’t answer those, the bill will.
The terms, in one place
If you’re paying for AI by the run
This is the kind of thing I look at in AI consulting: what a job costs now, where it’s wasting effort, and what to change first. Or tell me what you’re running and what it costs.