The ReAct Loop
An agent built on a language model runs the same perceive-reason-act loop as anything else, with one detail that gives the pattern its name: the model reasons in words between the actions it takes. Yao et al. (2022) called this ReAct, for reasoning and acting. They found that interleaving short reasoning traces with tool use handles many tasks more reliably than planning everything first or acting with no reasoning step.
A tool is code your harness can run for the model: a named function with a description, an input schema, and a return value. A tool call is the model's structured request for that code. The harness parses the request, runs the function, appends the result, and calls the model again.
This loop is the foundation that most general-purpose agent frameworks
are built on, LangChain's createReactAgent
among them, and the rest of this module works up to it.
The one field that tells the loop what to do next
Every chat-completions response carries a finish_reason
field that tells you why the model stopped generating. Inside an agent
loop, read it alongside the response content and tool requests to decide what to do next:
run a tool, return the answer to the user, or handle an error. Before reading
what its values mean, make two calls and compare their replies.
The first call offers no tools. In the second, the model is offered a tool and asked
for the current time. It still chooses whether to request the clock tool. Compare the displayed questions, replies, and
finish_reason. If both calls return "stop", inspect the
second reply: did it admit it cannot check the time, or give an unsupported answer?
Change the model or question and rerun. A tool request names work for your code;
this preview does not execute it.
"stop"means the model has written its final answer and the loop can returncontentto the user."tool_calls"means the model is asking your code to run a named function and append the result before calling back."length"means the token budget ran out before the answer finished."content_filter"means a safety classifier stopped the output.
In modern ReAct, a short router reads that field. If the model asked for a tool, the router runs your function, appends the result, and calls the model again. If the model stopped, the router returns the answer.
The model can revise its plan after reading a tool result. Whenever a tool returns something it did not expect, that result joins the next round of reasoning, and the model can reconsider where it stands before deciding what to do next.
Map the model, memory, tools, and router?
The pieces an LLM agent is built from
Zoom in on the agent and you find the decomposition Lilian Weng drew in 2023. A language model sits at the center. Around it sit memory, planning, tools, and the actions that reach the world.
In the tool loop below, these responsibilities are divided as follows:
- Memory is the
messages[]array. - Planning happens when the model uses the supplied context to choose its next response or tool request.
- Tools are the functions your harness exposes.
- Action is the tool call the harness executes.
The harness maintains the messages and executes the tools. Its router reads
finish_reason to decide whether to process tool requests or return an answer.
When this loop fails, inspect the values passed between its components. Check for:
- the model returning the wrong
finish_reason, - a missing turn in the messages array,
- the tool executor handing back an unexpected value,
- or the router misreading the signal.
Building the loop from scratch with one tool
As an exercise, we can try to create a simple loop from scratch.
All we need is a single tool named get_current_time plus one system prompt
and one while loop. Run it with "What time is it in UTC?" to start, and see what else you can ask:
When the model requests the clock, the loop adds these turns to
state.messages: - the system prompt (turn 0),
- the user question (turn 1),
- an assistant message with a
tool_callsfield (turn 2, the request), - and a
tool-role message holding the function's return value (turn 3).
Need context-rot evidence and mitigations?
Context rot: why a longer loop reads worse
Context rot is the tendency for contradictions, stale observations, and poor conventions to accumulate in the prompt history until the model's next decision gets worse. The effect also has a positional component: Liu et al. (2023) showed that models can underuse information buried in the middle of a long context. The further an important signal drifts from the current turn, the easier it is for the loop to ignore or contradict it.
Common mitigations keep each decision close to the evidence it needs:
- plan before the loop accumulates observations,
- delegate sub-tasks so each sub-agent works on a short context,
- or summarize earlier turns before they become a liability.
Applying the Same Loop to the Course
The loop above hard-wires one tiny tool. LangChain's createReactAgent
generalizes the same machinery to support more tools, standard interfaces, and observability integrations. The agent below uses a single tool,
read_course_page, which hands the agent any page of this course as markdown. The model decides which pages to read for your question, so give it a try.
Run the cell to open a chat panel, then ask “How does the clock loop decide when to stop?” You can leave the page selection empty; the agent will choose which course pages to read. To choose the context yourself, open Model and context options and select the page title under Pages before your first message. The title's tooltip shows its source ID, such as 01b-react. If you change the selected pages later, start a New chat. Expand a source chip to compare the answer with what the agent actually read.
Before you continue
Return to the clock-loop canvas and edit its inputs for these comparisons.
- Force the model to call the tool even when it does not need to.
Try a question the tool genuinely cannot answer ("capital of France?")
and watch
finish_reasonon step 1. Was it"stop"or"tool_calls"? Now edit the system prompt to require a tool call before answering ("You must always call get_current_time once, then answer."). Does the model comply? What does that tell you about how strongly the system prompt steers a model that has training-data confidence in another direction? - Count the steps on a genuinely tool-needing question.
"How many minutes until 17:00 UTC?" needs a clock reading and a calculation.
Run it a few times. Did the model call
get_current_timeonce, twice, or three times? Why would a model re-call a tool whose output it already has, and how would you discourage that in the system prompt?
What should I learn from these comparisons?
The first comparison separates prompt instructions from enforced control flow. Compare the requested tool and finish_reason before deciding whether the instruction was followed. A request alone is not proof that the clock ran.
The second comparison asks whether each extra tool call supplies useful evidence. Inspect the clock results and their order in messages. If the sequence is unclear, return to the clock-loop canvas and expand its trace before changing the prompt. Your result may differ from another run; explain it from your own trace.
Trace these mechanisms to primary sources?
References
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022). The original alternating thought / action / observation pattern; the figure above is redrawn from this paper.
- Wei et al., Chain-of-Thought Prompting (2022). The "reason" half of ReAct; what the model is doing inside
reasoning_contentwhen it deliberates before acting. - Lilian Weng, LLM Powered Autonomous Agents (2023). The decomposition the figure above adapts: a language model as the agent's core controller, surrounded by memory, planning, and tool use.
- Xu et al., ReWOO: Decoupling Reasoning from Observations (2023). A deliberate contrast to ReAct: the model writes the whole plan up front and a separate executor runs it without re-querying the model between steps. Comparing the two reveals the trade-off between reactive replanning and pre-committed execution.
- Schick et al., Toolformer (2023). Covered again in 1c's references; revisit here for the underlying mechanism the ReAct loop exploits.
Comprehensive list at Going Further · References.