Skip to content
The Money August 25, 2026

AI Agents Use Five Times the AI a Person Does. Most of It Is Re-Reading.

OpenRouter data shows agent AI use up fourteenfold since February. Most of that is the same context being read again, and the cheapest part of your bill is the part that's growing.

By The State of AI Marketing newsroom
Share
Editorial illustration for: AI Agents Use Five Times the AI a Person Does. Most of It Is Re-Reading.
Credit: JAC Growth Marketing

Agent traffic on OpenRouter went from 0.51 trillion tokens in February to 7.3 trillion. Human traffic grew 2.8 times across the same stretch. Agent traffic grew fourteenfold. The figures come from a16z’s Charts of the Week newsletter, published August 21, drawing on data from OpenRouter, the service that routes a company’s AI requests to whichever model it picks.

Tokens are the unit your AI invoice counts, roughly a fragment of a word going in and a fragment coming back out. Per task, an agent uses close to five times what a person typing into the same model uses. And more than 85% of that is cached prompts rather than new output, which is to say the agent spends most of its budget re-reading context it already has.

Three days later, Jon Morra, Chief AI Officer at the brand-safety firm Zefr, wrote a column for AdExchanger about what that does to a marketing stack. His diagnosis was that the industry built its systems on a pair of assumptions, both of which have quietly stopped holding:

“assuming both tokens would stay cheap forever and the number of tokens needed to achieve an outcome would stay flat.”

Neither is true now. The price per token keeps falling. The bill keeps climbing anyway.

Why the agent costs more: it starts over on every loop

An agent doesn’t do five times the thinking. It re-reads its own briefing, constantly, and you pay for every pass.

PPC Land’s write-up of the OpenRouter figures puts the mechanism plainly: “An agent is built to iterate towards a goal. The opening prompt carries a token-intensive pre-fill containing the core context around that goal.” Every loop the agent runs, it re-sends that context. Your brand guidelines, your product list, your last four campaign reports, the instructions you wrote in January. All of it, again, on every pass.

That re-sent context is where the 85% goes. The machine reads its own notes far more than it writes anything new.

For a marketing team, this changes what the AI line item is actually measuring. A person using a chat window generates cost roughly in proportion to the work they ask for. An agent generates cost in proportion to how much context it has to carry and how many times it loops. Those are different meters. A team can double its output and quadruple its bill, or hold output flat and still watch the number climb because someone added 40 pages to the reference file. That upkeep is the same maintenance tax we found sitting under agent deployments, showing up on the invoice instead of the calendar.

The honest counter-argument

A serious objection runs against everything above, and it deserves stating at full strength.

Cached tokens are billed at a steep discount. PPC Land says so directly: “Cached tokens cost far less than the pre-fill, which improves the unit economics of running agents.” The Decoder covered the same dataset on August 23. It puts the cached share near 70% and concludes that “actual costs aren’t rising as fast as the raw numbers suggest.”

Note that the two outlets don’t agree on the share. One says more than 85%, the other says nearly 70%, and both are reading the same a16z chart. That gap is worth holding onto, because it’s the difference between most of your agent spend being cheap and a quarter of it being expensive.

Either way, the direction is the same: token counts are a bad proxy for cost, and anyone forecasting an AI budget off volume alone will get the number wrong. The counter-argument doesn’t rescue the forecast. It just means the error runs in both directions.

What the number is really telling you

Morra’s prescription is a question, and it’s the sharpest thing written about this all month:

“The question every marketer should be asking their AI vendors this quarter isn’t ‘Which model do you use?’ It’s ‘Why does every task get the same model?’”

That is the finding buried in the OpenRouter data. Volume grew fourteenfold and almost nobody changed what they route it to. Morra’s own breakdown of the fix is unglamorous: “Lightweight models excel on simpler tasks, larger models on knowledge-intensive ones and reasoning models earn their keep only when the problem is complex enough.”

Some teams have already stopped waiting for vendors to sort this out. The agency PMG caps its people at $50 of AI per user per day through an internal governance layer. The number is designed to keep access open when campaign pressure spikes, not to ration it. Dollar Shave Club’s Laura Higgins moved her team off open-ended experimentation across four different assistants. What replaced it was a rough triage system that reserves the heavy models for heavy jobs.

Neither of those is a procurement policy. They’re both someone noticing that the meter runs differently now and putting a number on it before finance does.

Three things to check this quarter

  1. Ask your AI vendor which model runs which task. If the answer is one model for everything, you’re paying reasoning-model rates to reformat a spreadsheet.
  2. Find out what your agents re-read. The pre-fill is the cost driver. A bloated reference file is a recurring charge, not a one-time setup.
  3. Stop forecasting off token counts. Cached and uncached tokens bill at rates that differ by roughly an order of magnitude. A volume projection without that split is a guess wearing a decimal point.

The budget ceiling we covered on August 21 was about how much finance will approve. This is about what happens underneath it. The ceiling is holding. The thing pressing against it changed shape.

Quoted in this story

  • Jon Morra, Chief AI Officer, Zefr (source)

Want your perspective in coverage like this? Get quoted.

Sources

This story is part of our running coverage: the full picture →

Get Net Effect.

The net effect of AI on your marketing: the stories that matter, twice a week, in five minutes.

More from The Money