Matt Murphy
All posts
No. 05
  • AI
  • Corridor Cards

Where the tokens go: why one AI job re-read 78 million tokens

Tokens, calls, context, and compaction, explained with one real run of my listing pipeline, and the three ways I plan to cut what it reads.

Part three put a price on one run of my listing pipeline: a median of $42 a batch at API list prices. That made me ask what I was paying for.

The answer isn’t thinking. It’s re-reading.

This part shows where the tokens go, using one real run, and explains the handful of terms you need to follow it. If you’re paying for any AI automation, or about to, these are the terms that decide the bill.

You don’t need a technical background for this one. If you can read a phone bill, you can read this.

Which steps use the AI

Here’s the pipeline from part one again, with the AI’s steps picked out. The AI agent runs the whole thing, so every step uses some tokens. But most of its calls happen in these three, where it’s reading photos, researching cards, and working a website:

The pipeline, start to finish

  • MeA person decides
  • CodeA fixed program decides, the same way every time
  • AIThe AI uses its eyes or its judgment

Before the run

  1. Step 1: Sort and photographMe

The run, at 2:00 a.m.

  1. Step 2: Check the folderCode

  2. Step 3: Give every card an addressCode

  3. Step 4: Publish the photosCode

  4. Step 5: Identify each cardAI

  5. Step 6: Check the AI’s workCode

  6. Step 7: Research the market priceAI

  7. Step 8: Build the eBay fileCode

  8. Step 9: Upload, then read every draft backAI

After the run

  1. Step 10: Review and publishMe

Token Talk

AI is billed by the token. A token is a piece of a word, about three-quarters of one on average.

There are two kinds, and they’re priced differently:

  • Input is everything the AI reads.
  • Output is what it writes.

On the model my worker uses now, list price is $2 per million input tokens and $10 per million output tokens. Output costs five times more, so it’s natural to think the output is the expensive part.

It isn’t, because there is a third price to consider: cached input. That’s input the AI was already sent a moment ago. It costs $0.10 per million, a twentieth of the normal rate. Super cheap, right? Hold on to that one. It’s where most of the money goes, and I’ll explain it in this example.

A job is hundreds of calls

An AI agent doesn’t do a job in one go. It works in calls. On each call it reads what it’s sent, takes one step, such as running a command or clicking a button, and stops. Then the next call starts.

The part that changes everything: the AI remembers nothing between calls.

Picture playing fetch with a dog that forgets the game every time it drops the ball. Before each throw, you have to explain all of it again: what fetch is, where the yard ends, and every throw so far. Then you get one throw.

It’s a very good dog. It just starts from zero every time.

So every call is sent the whole story again. Its instructions. Every document it opened. Every step it took and every result that came back. All of that together is the context.

Step 1 sends a little. Step 300 sends everything from steps 1 through 299, and then does one more thing.

Where the dollars go

A batch is one stack of fifty cards. A batch run is one run of the pipeline working through a batch. Most batches take one run. A few take more.

Over 34 days, the pipeline made 14 batch runs. I added up what they would have cost at the provider’s pay-as-you-go prices, then sorted every dollar by the kind of token it paid for:

What the money paid for

  • 73% Re-read (cached) input
  • 19% New input
  • 8% Output
The combined cost of 14 batch runs, September 2 to October 5, 2026, at API list prices. Each part is one kind of token’s share of that cost.

Nearly three-quarters of the cost is re-reading. It’s cheap per token and enormous in volume.

Watch one run

This is one run of the worker, on September 27, 2026. Each point along the bottom is a call. The height is how many tokens that call had to re-read before it could do anything new.

One long session

  • Tokens re-read on each call
  • Compaction
  • Startup reading
Measured: one run of the worker on September 27, 2026. 540 calls, 78 million input tokens, 145,000 tokens re-read on the average call.

Read it left to right:

  1. It starts at 39,000 tokens, before any work is done. Nearly all of that is Codex itself: the app’s built-in instructions, plus a description of every tool it’s allowed to use, such as running a command or working a browser. My nightly note is in there too, but it’s only about 1,200 tokens. The procedure and my other documents haven’t been opened yet. (This isn’t the harness from part one. That’s my code. This is the app’s own rulebook, and every session starts with it.)
  2. Loading. The nightly note tells it to read about a dozen of my documents before acting. That’s roughly 50,000 more tokens, and it carries them on every call after.
  3. The climb. Every step adds its result to the context. The typical call adds under 1,000 tokens. The average is 2,200, because a few steps bring back a whole web page or a long file.
  4. The drop. Somewhere past 200,000 tokens, Codex steps in. More on that below.

The run made 540 calls, and the average call re-read 145,000 tokens. Multiply those and you get the headline: 78 million input tokens for one run.

What one average call sends today

  • 39K The starting load
  • 106K Everything read and done so far in the run
145,000 tokens. The starting load is fixed: Codex’s built-in instructions and tool list, plus my short nightly note. Everything else is the run so far.

Compaction: a reset you don’t control

A model can only read so much at once. When the context gets close to that limit, the tool replaces it with a summary and carries on. That’s compaction, and it’s each dashed red line in the chart.

Compaction keeps a long job alive. But it has two problems:

  • It’s late. By the time it happens, hundreds of calls have already paid to re-read a huge context.
  • It’s blunt. It happens wherever the job happens to be, and a summary can drop a detail the next step needed.

That run was compacted five times. I didn’t choose any of them.

Three ways to make the AI read less

Reading is 92% of what a run costs. So the way to a cheaper run is to make the AI read less. There’s a second payoff too: the less it carries, the less often Codex has to step in and compact.

How much the AI reads comes down to two numbers: how many calls it makes, and how much it re-reads on each one. That leaves three things to change.

1. Reset sooner, on purpose. Split the job into short sessions. Each one does a clean piece of work, writes the result to a file, and ends. That’s a handoff. The next session starts small and reads the file, not the whole history.

2. Climb slower. Have each step bring back only what’s needed. A step that returns one price, not a whole page, adds a few hundred tokens where it used to add thousands.

3. Make fewer calls. Give the agent saved routines, so one call does what used to take five or ten.

There’s a fourth, and it’s the cheapest to try: pick the model with the re-read price in mind. Cached input is $0.10 per million on the model I use now. On the most expensive one I tried, it’s $1.00. Same job, ten times the price for the biggest slice.

The plan: thirteen short sessions

A batch of fifty cards already has natural seams. So the plan is to cut along them:

SessionWorkHow many
OpenCheck the photos, claim the batch, and publish the images1
Identify and priceResearch ten cards, then hand off5
BuildMerge the research, build the draft file, and upload it1
Audit draftsRead ten saved drafts back, then hand off5
CloseCheck every receipt and close the batch1

Here’s what that should look like, drawn to the same scale as the real run:

The same work as thirteen short sessions

  • Tokens re-read on each call (projected)
  • Handoff
Projected, not measured: the same number of calls and the same growth per call as the run above, split into thirteen sessions (open, five to identify and price, build, five to audit drafts, close) that each start with a one-page instruction card. About 48 million input tokens, 39% less.

What one average call would send in a short session

  • 39K The starting load
  • 5K A one-page instruction card
  • 46K This chunk so far
About 90,000 tokens, drawn to the same scale as the bar above. A projection, not a measurement.

The same work, about 48 million input tokens where there were 78 million.

This is a projection, not a result. It assumes the same number of calls and the same growth per call as the real run, and a one-page instruction card for each session in place of the dozen documents. Until it’s built and measured, treat it as a plan.

What turned out not to matter

  • How hard the AI thinks. Output was 8% of the cost. Turning the reasoning down wouldn’t move much.
  • Quiet nights. A night with nothing to list is about ten calls. It’s the long runs that cost.

Five questions to ask about any AI automation

  1. How many calls does one job take?
  2. What does each call re-read?
  3. Where does the context reset, and who decided that?
  4. Which model runs it, and what does it charge for cached input?
  5. Can I see what one run cost?

If the person building it can’t answer those, the bill will.

The terms, in one place

If you’re paying for AI by the run

This is the kind of thing I look at in AI consulting: what a job costs now, where it’s wasting effort, and what to change first. Or tell me what you’re running and what it costs.