…
Module 1 · Part B of 3

The ReAct Loop

An agent built on a language model runs the same perceive-reason-act loop as anything else, with one detail that gives the pattern its name: the model reasons in words between the actions it takes. Yao et al. (2022) called this ReAct, for reasoning and acting. They found that interleaving short reasoning traces with tool use handles many tasks more reliably than planning everything first or acting with no reasoning step.

A tool is code your harness can run for the model: a named function with a description, an input schema, and a return value. A tool call is the model's structured request for that code. The harness parses the request, runs the function, appends the result, and calls the model again.

This loop is the foundation that most general-purpose agent frameworks are built on, LangChain's createReactAgent among them, and the rest of this module works up to it.

The loop is unchanged. With a language model in the reasoning step, the agent reads tool results and acts by calling tools.

The one field that tells the loop what to do next

Every chat-completions response carries a finish_reason field that tells you why the model stopped generating. Inside an agent loop, read it alongside the response content and tool requests to decide what to do next: run a tool, return the answer to the user, or handle an error. Before reading what its values mean, make two calls and compare their replies.

The first call offers no tools. In the second, the model is offered a tool and asked for the current time. It still chooses whether to request the clock tool. Compare the displayed questions, replies, and finish_reason. If both calls return "stop", inspect the second reply: did it admit it cannot check the time, or give an unsupported answer? Change the model or question and rerun. A tool request names work for your code; this preview does not execute it.

In modern ReAct, a short router reads that field. If the model asked for a tool, the router runs your function, appends the result, and calls the model again. If the model stopped, the router returns the answer.

The model can revise its plan after reading a tool result. Whenever a tool returns something it did not expect, that result joins the next round of reasoning, and the model can reconsider where it stands before deciding what to do next.

Map the model, memory, tools, and router?

The pieces an LLM agent is built from

Zoom in on the agent and you find the decomposition Lilian Weng drew in 2023. A language model sits at the center. Around it sit memory, planning, tools, and the actions that reach the world.

In the tool loop below, these responsibilities are divided as follows:

  • Memory is the messages[] array.
  • Planning happens when the model uses the supplied context to choose its next response or tool request.
  • Tools are the functions your harness exposes.
  • Action is the tool call the harness executes.

The harness maintains the messages and executes the tools. Its router reads finish_reason to decide whether to process tool requests or return an answer.

After Lilian Weng (2023): the model is the agent's core controller, with memory, planning, tools, and action wired around it.
The modules are independent. You can replace the model or the tools or the router on its own, and as long as each one still honors the contract the others expect of it, the rest of the loop keeps working.

When this loop fails, inspect the values passed between its components. Check for:

  • the model returning the wrong finish_reason,
  • a missing turn in the messages array,
  • the tool executor handing back an unexpected value,
  • or the router misreading the signal.

Building the loop from scratch with one tool

As an exercise, we can try to create a simple loop from scratch. All we need is a single tool named get_current_time plus one system prompt and one while loop. Run it with "What time is it in UTC?" to start, and see what else you can ask:

Think, act, observe, repeat. The turn exits the moment finish_reason is "stop".
What just happened inside the messages array.
When the model requests the clock, the loop adds these turns to state.messages: The next model call receives those messages, including the clock result. Inspect its response to see whether it answers or requests another tool. The growing array is the only "memory" this agent has. Saving and reloading that array preserves the supplied conversation history; external facts and model behavior may change.
Need context-rot evidence and mitigations?

Context rot: why a longer loop reads worse

Context rot is the tendency for contradictions, stale observations, and poor conventions to accumulate in the prompt history until the model's next decision gets worse. The effect also has a positional component: Liu et al. (2023) showed that models can underuse information buried in the middle of a long context. The further an important signal drifts from the current turn, the easier it is for the loop to ignore or contradict it.

Redrawn from Liu et al. (2023). As the loop appends turns, the answer you need drifts into the weak middle.

Common mitigations keep each decision close to the evidence it needs:

  • plan before the loop accumulates observations,
  • delegate sub-tasks so each sub-agent works on a short context,
  • or summarize earlier turns before they become a liability.

Applying the Same Loop to the Course

The loop above hard-wires one tiny tool. LangChain's createReactAgent generalizes the same machinery to support more tools, standard interfaces, and observability integrations. The agent below uses a single tool, read_course_page, which hands the agent any page of this course as markdown. The model decides which pages to read for your question, so give it a try.

Run the cell to open a chat panel, then ask “How does the clock loop decide when to stop?” You can leave the page selection empty; the agent will choose which course pages to read. To choose the context yourself, open Model and context options and select the page title under Pages before your first message. The title's tooltip shows its source ID, such as 01b-react. If you change the selected pages later, start a New chat. Expand a source chip to compare the answer with what the agent actually read.

Before you continue

Return to the clock-loop canvas and edit its inputs for these comparisons.

  1. Force the model to call the tool even when it does not need to. Try a question the tool genuinely cannot answer ("capital of France?") and watch finish_reason on step 1. Was it "stop" or "tool_calls"? Now edit the system prompt to require a tool call before answering ("You must always call get_current_time once, then answer."). Does the model comply? What does that tell you about how strongly the system prompt steers a model that has training-data confidence in another direction?
  2. Count the steps on a genuinely tool-needing question. "How many minutes until 17:00 UTC?" needs a clock reading and a calculation. Run it a few times. Did the model call get_current_time once, twice, or three times? Why would a model re-call a tool whose output it already has, and how would you discourage that in the system prompt?
What should I learn from these comparisons?

The first comparison separates prompt instructions from enforced control flow. Compare the requested tool and finish_reason before deciding whether the instruction was followed. A request alone is not proof that the clock ran.

The second comparison asks whether each extra tool call supplies useful evidence. Inspect the clock results and their order in messages. If the sequence is unclear, return to the clock-loop canvas and expand its trace before changing the prompt. Your result may differ from another run; explain it from your own trace.

Trace these mechanisms to primary sources?

References

Comprehensive list at Going Further · References.

← Module 1a: The Agent