Why AI Agents Consume More Tokens Than You Expect

•
8 views
•
22 days ago

Quiz Recap: If a ReAct agent runs through 5 iterative steps to solve a user's question, what happens to token consumption across the full task?
Correct Answer: C — Token usage grows as the entire context must be re-sent per LLM call

When developers build their first autonomous AI agent, one of the biggest surprises is how quickly token usage and API costs increase.

A common assumption is:

If an agent takes 5 steps and generates 400 tokens per step, the total should be around 2,000 tokens.

In reality, it can be significantly higher.

The main reason is simple: LLM APIs are stateless.

The Stateless Problem

An LLM doesn't automatically remember what happened in your previous API call.

In an agentic workflow, the model might follow a loop like:

Think → Act → Observe → Think → Act → Observe

For the model to make a decision in Step 5, it needs context from the previous steps—what the user asked, which tools were called, and what those tools returned.

Because the API doesn't retain this state, your application typically sends the relevant conversation history again with every request.

And that's where token usage starts compounding.

How Token Usage Compounds

Consider a simple 5-step agent:

  • Step 1: System prompt + user request

  • Step 2: Previous context + Step 1 output

  • Step 3: Previous context + Step 2 output

  • Step 4: Previous context + Step 3 output

  • Step 5: Previous context + Step 4 output

The context sent to the model keeps getting larger.

So you're not simply paying for:

5 × tokens generated

You're also repeatedly paying for the growing input context processed at every iteration.

This becomes even more significant when tools return large responses.

The Observation Explosion

Imagine your agent performs a database query and receives a 2,000-token result.

If the agent needs four more iterations to analyze that result, those 2,000 tokens may become part of the context for each subsequent request.

A single verbose tool response can therefore create thousands of additional input tokens.

This is why returning raw database dumps, full HTML pages, or massive JSON responses from tools is usually a bad idea.

How Production Agents Reduce Token Costs

1. Keep Tool Outputs Small

Filter and transform tool responses before sending them back to the LLM.

Instead of:

Database → 5,000 rows → LLM

use:

Database → Filter → Relevant fields → LLM

Give the model only what it needs.

2. Summarize Long Histories

When an agent's context becomes too large, summarize older interactions into important facts, decisions, and results.

This keeps the context useful without carrying the entire history forever.

3. Use Prompt Caching

Many modern LLM providers support prompt caching.

If large portions of your prompt remain unchanged between iterations, caching can reduce the cost and latency of repeatedly processing that context.

4. Set Iteration Limits

Always enforce a maximum number of agent iterations.

Without a limit, a stuck agent can continuously call tools and consume tokens.

MAX_ITERATIONS = 10

This is both a cost-control and reliability mechanism.

The Key Takeaway

Building an efficient AI agent isn't just about choosing the right model.

It's about managing context intelligently.

The goal isn't to give the LLM more information.

It's to give it the right information, at the right time, with as little unnecessary context as possible.

That difference can have a major impact on the cost, latency, and reliability of production AI agents.

0

Discussion (0)

Markdown, bold, quotes & code blocks supported
No comments yet. Start the conversation!