Who Actually Executes the Action in a ReAct Agent?

•
8 views
•
22 days ago

Quiz Recap: In a production implementation of the ReAct pattern, who executes the Action and generates the Observation?
Correct Answer: B — The external runtime / host application executes the tool and injects the output back into the prompt context.

The Common Misconception: "The LLM Runs the Tool"

When developers first encounter AI agents querying databases, scraping websites, or running Python scripts, it is easy to assume that the Large Language Model (LLM) itself has gained runtime execution capabilities.

In reality, an LLM is strictly a statistical text-prediction engine. It cannot open a TCP socket, execute bash commands, or make HTTP requests.

To build an agent, we rely on a clear separation of concerns between the Model and the Host Runtime.

The Architecture: Brain vs. Body

  1. The LLM (The Brain): Decides what needs to be done and generates structured text declaring its intent.

  2. The Host Runtime (The Hands & Eyes): Parses that text, executes the corresponding code in the real world, and feeds the resulting data back to the LLM.

How the 4-Stage Execution Cycle Works

Whenever a ReAct agent runs, the host runtime follows this strict state machine:

Step 1: Prompting & Output Generation

The runtime sends the conversation history and tool schemas to the model. The model emits formatted text:

Thought: I need to check the stock price of Apple.
Action: get_stock_price(ticker="AAPL")

Step 2: Output Parsing

The runtime uses regex or JSON parsers to intercept the output before the user sees it. It extracts:

  • Tool Name: get_stock_price

  • Arguments: {"ticker": "AAPL"}

Step 3: Tool Execution (In the Host Environment)

The host runtime looks up get_stock_price in its local Python tool registry and invokes the actual function:

# Executed locally by the Python runtime, NOT by the LLM
result = get_stock_price(ticker="AAPL")  # Returns: "$182.50"

Step 4: Observation Injection

The runtime wraps the return value in an Observation block and appends it to the conversation history:

Observation: $182.50

This updated transcript is then sent right back to the LLM for the next turn.

Key Takeaways for Developers

  • LLMs emit intent, not execution: The model only outputs the name and arguments of a tool.

  • Security lives in the runtime: Because your runtime executes the tools, all security sandboxing, rate limiting, and permission checks must happen at the runtime level before executing the tool.

  • Framework Agnostic: Whether you use LangChain's AgentExecutor, LlamaIndex, or raw OpenAI Function Calling, they all adhere to this exact runtime orchestrator pattern.

0

Discussion (0)

Markdown, bold, quotes & code blocks supported
No comments yet. Start the conversation!