Why AI Agents Consume More Tokens Than You Expect
Quiz Recap: If a ReAct agent runs through 5 iterative steps to solve a user's question, what happens to token consumption across the full task?
Correct Answer: C — Token usage grows as the entire context must be re-sent per LLM call
When developers build their first autonomous AI agent, one of the biggest surprises is how quickly token usage and API costs increase.
A common assumption is:
If an agent takes 5 steps and generates 400 tokens per step, the total should be around 2,000 tokens.
In reality, it can be significantly higher.
The main reason is simple: LLM APIs are stateless.
The Stateless Problem
An LLM doesn't automatically remember what happened in your previous API call.
In an agentic workflow, the model might follow a loop like:
Think → Act → Observe → Think → Act → Observe
For the model to make a decision in Step 5, it needs context from the previous steps—what the user asked, which tools were called, and what those tools returned.
Because the API doesn't retain this state, your application typically sends the relevant conversation history again with every request.
And that's where token usage starts compounding.
How Token Usage Compounds
Consider a simple 5-step agent:
Step 1: System prompt + user request
Step 2: Previous context + Step 1 output
Step 3: Previous context + Step 2 output
Step 4: Previous context + Step 3 output
Step 5: Previous context + Step 4 output
The context sent to the model keeps getting larger.
So you're not simply paying for:
5 × tokens generated
You're also repeatedly paying for the growing input context processed at every iteration.
This becomes even more significant when tools return large responses.
The Observation Explosion
Imagine your agent performs a database query and receives a 2,000-token result.
If the agent needs four more iterations to analyze that result, those 2,000 tokens may become part of the context for each subsequent request.
A single verbose tool response can therefore create thousands of additional input tokens.
This is why returning raw database dumps, full HTML pages, or massive JSON responses from tools is usually a bad idea.
How Production Agents Reduce Token Costs
1. Keep Tool Outputs Small
Filter and transform tool responses before sending them back to the LLM.
Instead of:
Database → 5,000 rows → LLM
use:
Database → Filter → Relevant fields → LLM
Give the model only what it needs.
2. Summarize Long Histories
When an agent's context becomes too large, summarize older interactions into important facts, decisions, and results.
This keeps the context useful without carrying the entire history forever.
3. Use Prompt Caching
Many modern LLM providers support prompt caching.
If large portions of your prompt remain unchanged between iterations, caching can reduce the cost and latency of repeatedly processing that context.
4. Set Iteration Limits
Always enforce a maximum number of agent iterations.
Without a limit, a stuck agent can continuously call tools and consume tokens.
MAX_ITERATIONS = 10
This is both a cost-control and reliability mechanism.
The Key Takeaway
Building an efficient AI agent isn't just about choosing the right model.
It's about managing context intelligently.
The goal isn't to give the LLM more information.
It's to give it the right information, at the right time, with as little unnecessary context as possible.
That difference can have a major impact on the cost, latency, and reliability of production AI agents.