…
Module 1 · Part A of 3

The Agent

An agent is an entity that interacts with an environment at any meaningful level. A person going about their day is an agent in this sense, and so is a thermostat holding a room at temperature or a self-driving car making its way through traffic.

All of them perceive their environment through sensors and act on it through actuators. Because each action feeds the next observation, perceiving and acting fold into one continuous cycle, a perceive-reason-act loop, that lasts as long as the agent continues to interact. The agents you will build all adopt the same cycle and are set apart only by how much they can perceive and how far their actions reach.

The three parts on the loop

A single trip around the loop runs through three parts, each with a plain name and a distinct job.

The agent perceives the environment through observations and acts on it through actions. The decision step between them can change.

Perception gathers what the agent can know, reasoning chooses what to do, and action changes the environment. Perception and action often come from an existing interaction model, whether it controls a vacuum, web automation, a game entity, or a humanoid. The reasoning step can change without changing that boundary, which separates capability from reach.

Explore the loop and classic architectures?

Open this section when you want concrete examples and the fuller architecture taxonomy.

  • Perception gives the agent everything it knows about its environment. A thermostat reads a single number: the current temperature. A self-driving car fuses cameras, radar, and a map into a model of the road. Whatever perception captures is all the agent has to work from, since nothing downstream ever touches the world directly. A gap in perception becomes a gap the rest of the agent cannot see around.
  • Reasoning turns a perception into a decision, which can vary widely between agents. A thermostat compares the temperature against a target and stops there, while a chess engine searches millions of moves and a self-driving car runs a learned policy. The agents you build here reason by calling a language model, and almost everything the course adds later (tools/memory/planning) hangs off this single step.
  • Action changes the environment for itself and other participants. An action can be as small as flipping a relay or as involved as steering a car. Later pages call code exposed to the model a tool, because from the model's side it is just another available action. Calling a tool means invoking that exposed code as the action. It also marks the outer edge of an agent's reach, because however sharp the reasoning, the agent affects the world only through its action space.
The reasoning step is replaceable. Perception and action do not care which decision function sits between them. A simple rule, a decision tree, a search routine, or a neural network can drive the same loop. Building the loop once as a fixed frame keeps two questions cleanly apart: how capable the reasoning is, and how far the actions reach.

Classic agent designs differ at the decision step. Russell and Norvig's Artificial Intelligence: A Modern Approach arranges them as a spectrum, from a bare reflex agent up to a learning agent that improves its own rules. This course starts with rigid logic, then adds a language model and engineered context to move toward a model-based, goal-oriented loop. Later modules keep that loop running over changing state, which is the pattern learning agents build on.

Four agent architectures, from direct input-to-action rules to an agent that improves its rules. Adapted from Russell and Norvig's Artificial Intelligence: A Modern Approach (Ch 1).
Run the smallest rule-based agent loop?

An agent with no language model

The smallest decision function is a single if statement. Here it drives a one-dimensional track bounded by walls at positions 0 and 9, which leaves the whole agent about as stripped down as one can get. Press Run and watch the position number bounce between those two walls as the loop runs.

Two things worth examining in the code above:

  • The environment (wall positions, current position) lives in state.env, while the agent's intrinsics (its direction) live in state.agent. Keeping them separate means a different agent with different intrinsic state can operate in the same environment cleanly.
  • The reasoning step is one if statement. Replacing it with something more powerful, whether a decision tree, a neural network, or a language model, only requires changing what sits between perception and action. The loop structure, the state layout, and the action step all remain identical.

What changes when reasoning calls an LLM

You call the chat-completions interface like any other function: hand it a list of messages and a model name, and it hands back one assistant message plus metadata.

About chat(). It is the course-provided wrapper you run live a little further down this page. The snippet below previews its call shape and the fields it returns, so the names are familiar by the time you reach the runnable cell.
const reply = await chat({
  model:    "nvidia/nemotron-3.5-lightning-30b-a3b",
  messages: [{ role: "user", content: "Hi" }],
});
// reply.choices[0].message  → { content, reasoning_content }
// reply.choices[0].finish_reason → "stop" | "tool_calls" | "length"

Three facts about that function shape how agents built on top of it have to work:

Compare where APIs keep conversation state?

This course uses the Chat Completions API throughout: you keep the conversation as a messages[] array in your own code and resend the full array on every turn. The newer Responses API can hold the thread server-side instead, so you send only the new input and a reference to the prior turn. Either way the model stays stateless and unchanged between calls. The conversation lives outside it, in state your code or the server keeps.

Both interfaces wrap the same stateless call. They differ only in who keeps the conversation between turns.

A real call, in your browser

Press Run to send a single request. It stands entirely on its own, with no loop wrapped around it and nothing from an earlier turn carried along.

Inspect the raw HTTP request and response?

What chat() actually sends

chat() is a convenience wrapper around one ordinary HTTP POST in the OpenAI-compatible wire format. That request carries three things:

  • the endpoint to send to,
  • request headers; a custom endpoint may require an Authorization header,
  • and a JSON body with model and messages.

The cell below makes that call by hand. Open request to inspect the URL, headers, and body, with any authorization value redacted. Open raw response to inspect the full object returned by chat().

The model endpoint uses this format. Module 3's OpenClaw gateway exposes a separate protocol for agent turns, sessions, and runtime operations.

Configure a compatible endpoint in the model panel before running the cell. NVIDIA-hosted endpoints require an NVIDIA API key. Check the selected service's access and quota. The panel shows the active endpoint and credential settings; inspect the redacted request below before sending.

The selected route determines which service receives the request and its credential. A browser call requires the endpoint to permit the course origin. The course can use its configured relay when the provider does not allow that direct browser request. Inspect the active route and request headers in the model panel.

Inspect the streaming helper. chatStream() parses answer text, any returned reasoning text, usage metadata, and the end-of-stream signal. It connects the response to the lesson panel and Stop button. Open the helper source from the cell menu to inspect the request and event handling.
Compare maze reasoning inside one fixed loop?

Swapping the reasoning step while the loop stays fixed

The next examples keep the perceive-reason-act loop and vary its decision function. Compare fixed rules and search algorithms before using chat() to choose a move.

Run the nodes in order. Each puts a different decision rule into the same maze-solving loop:

  • The constant "always east" rule reaches the goal on a straight corridor, then stalls at the first wall in the maze, since a fixed heading cannot turn.
  • DFS dives and backtracks out of dead ends.
  • BFS fans out in rings before walking the shortest path.
  • A* biases those same rings toward the goal.
The decision step is the only part that varies. DFS, BFS, and A* all read the current maze state and return a move. One line in the last node chooses the decision function; the loop below runs the same way for each choice. The next example gives the model a choice of branches while a controller records explored cells and the return path.

Using a language model to choose maze moves

The controller records visited cells and a return path, filters out visited branches, and walks corridors. At junctions, the model reads the map and chooses the branch order through choose_direction. The controller manages depth-first exploration; it does not supply the model with a solved route.

Run the cells in this order, then vary one control at a time:

  1. Run the engine node first. It exports the maze utilities and one shared controller, state.runMaze, that walks corridors automatically and calls the model only at junctions.
  2. Run the editable maze-agent node below. DIRECTIVE is the core system instruction, and the options object controls the model, maze size, text map, move history, coordinate trail, and repeated tool schema.
  3. Start with the working defaults. Then switch model to compare capability-checked routes, remove history, or set advanced to grow the maze after the simple case works.
Validate the requested action. The choose_direction schema declares the allowed direction values. The model requests a direction, and the controller checks the tool name and arguments before moving. Inspect the request and response to distinguish the declared schema from the validation your code performs.

The loop, drawn locally

In the maze example, the controller builds a request for chat() and validates the returned move before applying it. The stages below trace that exchange.

  • Locally-sample. Your code reads a slice of the environment. Here it reads the current maze cell and its valid moves. The request can also include the full map and prior turns, depending on the selected controls.
  • Locally-perceive. That slice is rendered into a full engineered context: the system prompt, the running messages[] history, prior tool results, and the slice you just sampled. This example represents the maze as text in the messages sent to the model.
  • Locally-act. The model maps that text to more text: an intent signal, ideally a typed tool call like { "move": "east" } rather than free prose. The intent only requests an action, and it never performs one. Your code reads it and applies one bounded update to the environment (doMove(...)), validating it first.
  • Locally-impact. The update changes the world. Nothing persisted inside the model. The only durable state is the environment and the context your code will rebuild. The next turn's locally-sample reads the changed world, and the cycle repeats.

In this example, chat() receives text describing the maze and returns a requested move. The controller supplies the observations, validates the request, and changes the maze state. The model call does not execute the move.

Inspect both boundaries: the context supplied to the model and the validation your code performs before applying a requested action.

sample → perceive → (text → text) → act → impact
The maze request can include the full map and prior turns. Open its request details to see which context is present. Module 1b builds the tool loop with an explicit messages array.

Before you continue

Each question is answerable by editing a knob in the LLM-maze cell above and rerunning it.

  1. Compare conversation history. Set includeHistory:false, then rerun with the same profile and maze. Compare branch order, junction decisions, and elapsed time with history enabled. The controller retains visited cells and the return path in both runs.
  2. Remove the strategy hint from the DIRECTIVE. Delete the line that says to prefer unvisited branches over visited cells, then rerun with the same profile and maze. Compare branch order, junction decisions, and elapsed time. The controller still filters visited branches; use repeated runs to check whether any difference persists.
  3. Compare model profiles. Start with model:"direct", then compare model:"reasoning" and model:"omni" on the same maze. Inspect each request's model, token budget, temperature, and reasoning settings. These profiles change several settings; compare branch order and latency without attributing every difference to model identity alone.

Try it · compare conversation history

This chat starts with conversation history enabled. Tell it your name, then ask what your name is. Turn the memory toggle off and ask again. Each chat() call receives the messages assembled by the interface; with memory off, earlier turns are omitted. Compare the replies and inspect how the code builds the request.

← Setup