The Stopping Condition Problem: How Agents Know When to Finish
The Infinite Loop Vulnerability
In traditional programming, functions usually have a predictable termination path. A function executes its logic and eventually reaches a return statement or the end of its execution block.
Autonomous AI agents are different.
When an agent follows the ReAct (Reasoning + Acting) pattern, it operates inside an open-ended loop:
Thought → Action → Observation → Thought → Action → Observation → ...
This creates an important engineering challenge: the stopping condition problem.
If an agent cannot determine whether its objective has been completed, it may continue reasoning and calling tools indefinitely. A search tool might return ambiguous information, a website might provide incomplete results, or the model might repeatedly attempt slightly different versions of the same action. Without proper safeguards, a simple task can turn into hundreds of unnecessary LLM calls, consuming tokens, increasing latency, and potentially generating significant API costs.
How Does an Agent Actually Stop?
A common misconception is that the LLM itself somehow "shuts down" the agent when it believes the task is complete. In reality, the runtime orchestrator controls the execution loop.
One common approach is to define an explicit terminal signal in the agent's output format.
For example, the model might produce:
Thought: I now have enough verified information.
Final Answer: The elevation difference is 822 meters.
The runtime parser examines the model's output. If it detects a tool call such as Action: search, it executes that tool and feeds the resulting observation back into the agent.
However, when it detects Final Answer:, the runtime recognizes that the agent has finished and breaks the execution loop.
But relying entirely on the model to stop is dangerous.
Defensive Guardrails
Production agents should always have deterministic safety mechanisms.
The most important is a maximum iteration limit:
max_iterations = 10
If the agent reaches this limit without producing a final answer, the runtime terminates execution instead of allowing the loop to continue indefinitely.
Other useful guardrails include token or context limits, which prevent the agent from exceeding its available context budget, and duplicate-action detection, which can identify when an agent repeatedly performs the same tool call without making progress.
The Key Principle
The most important distinction is:
The LLM signals when to stop; the runtime controls whether execution actually stops.
A well-designed agent should therefore never rely solely on the model's judgment. Always implement deterministic termination conditions and meaningful fallback behavior. If an agent reaches its limit, it should return its best available findings or clearly explain that it could not complete the task.
In agentic systems, knowing when to stop is just as important as knowing what to do next.