… …
Module 4 · Part A of 3

The OpenShell Sandbox

In Module 3, files supplied persistent context and scheduled jobs started unattended work. Those operations depend on the authority available to the agent process. OpenShell restricts that authority through the active policy and its enforcement.

Run commands in the sandbox, inspect their output and exit status, and investigate the cause. Then survey read access through the operator control plane.

OpenShell separates control entry, policy enforcement, agent state, tool access, and model routing. Keep this map in view as the page moves from the sandbox argument to live policy checks.
Keep the three product roles separate as you trace enforcement.

OpenShell is agent-agnostic. The same boundary can enclose OpenClaw, Hermes, LangChain deepagents, or another process that needs operating-system-level containment.

Where each defense layer intervenes

Suppose untrusted content tells an agent to read a credential and send it away. Three different layers get a chance to stop that action, and each layer answers a different question.

Why enforcement belongs below the model

Assume a valid-looking request can still be harmful

Zero trust rejects implicit trust based only on where a request came from. OWASP's prompt-injection guidance applies the same discipline to agents: constrain tools, restrict privileges, and assume external content can steer the model.

Grant only the authority the task needs

Saltzer and Schroeder described this as the principle of least privilege: each component receives only the permissions needed for its work. Because the discipline predates agents, it transfers cleanly to this setting.

Measure boundaries instead of trying to prove intent

Rice's theorem rules out a general test for non-trivial behavior of arbitrary programs. An agent with code execution deserves the same caution: instrument the boundaries around what runs instead of assuming a static check can prove every future action safe.

How the sandbox mechanisms answer testable questions

This launchable combines four operating-system mechanisms. Each answers a question you can test, and together they apply policy at more than one boundary. Read the live policy later on this page before treating any one guarantee as universal.

The blocked arrow illustrates a possible denied action. A command result alone does not identify the enforcing mechanism.

1 · Where can this process connect?

A network namespace (netns) gives the agent an isolated network stack. Outbound access starts closed. An HTTP CONNECT proxy provides the only route out. Open Policy Agent (OPA) checks the requesting binary, host, port, and path against policy. Whoever edits that policy decides which connections are allowed.

2 · Which files can it reach?

Landlock is a Linux Security Module that restricts future filesystem access. Here it limits new file operations to paths granted by policy. That is strong enforcement, but compatibility mode, OverlayFS, and handles opened before restriction still matter. Read the live policy and test the path you intend to expose.

3 · Which kernel operations can it request?

A system call, or syscall, asks the Linux kernel to perform an operation. seccomp is Linux syscall filtering; its BPF filter checks each request against an allowlist. A denied call returns EPERM, meaning "operation not permitted," before the operation runs. This launchable excludes dangerous primitives such as ptrace, mount, and setuid.

4 · How much authority does it start with?

The agent runs as a non-root process under the unprivileged sandbox user. A successful exploit therefore starts with less authority. The lower starting authority reduces blast radius, though it does not prove escalation is impossible. The host, container configuration, and syscall policy remain part of the security boundary.

Vocabulary in these four mechanisms

Network policy

Network namespace / netns
Isolated Linux network stack: devices, routes, firewall rules, and socket ports. network_namespaces(7)
Ingress / egress
Ingress is traffic entering the sandbox; egress is traffic leaving it. Default-deny means neither direction is open unless policy grants it.
CONNECT proxy
HTTP tunnel proxy to a requested host and port; safe deployments restrict allowed targets. MDN CONNECT
OPA
Policy engine that evaluates structured input and returns decisions to enforcement code. OPA docs
Binary, host, port, path
Egress policy inputs: executable identity, target name, TCP port, and requested URL path.

Filesystem boundary

Landlock
Linux Security Module that lets an unprivileged process restrict its own future access. Kernel docs
Ambient rights
Access a process would otherwise inherit from the OS, parent process, mounts, or credentials.
Workspace tree
The directory subtree a policy grants. Verify the live policy and kernel compatibility before treating it as a complete filesystem boundary.
Bind mount
A filesystem subtree made visible at another path, including across chroot boundaries. mount(2)
Symlink race
A checked path changes through a symbolic link before use. symlink(7)

Syscall boundary

seccomp
Linux syscall filtering that reduces kernel surface reachable by a process. Kernel docs
Syscall allowlist
The syscall set allowed to run; everything else receives a policy action.
BPF layer
Filter program that checks syscall numbers and arguments before execution. Kernel docs
TOCTOU / inspection race
Time-of-check/time-of-use bug; seccomp avoids pointer dereference races inside the filter. Kernel docs
EPERM / EACCES
Linux errno names for operation not permitted and permission denied. errno(3)

Process identity

Agent process
The OS process running the agent and tools inside these boundaries.
Non-root / unprivileged user
A nonzero UID subject to normal kernel permission checks. capabilities(7)
Linux capability
A split-out root power such as CAP_SYS_ADMIN or CAP_SETUID. capabilities(7)
ptrace, mount, setuid
Dangerous primitives for process control, filesystem attachment, and identity change. ptrace(2) · mount(2) · setuid(2)
Privilege escalation
Gaining authority beyond assignment: root, extra capabilities, broader file access, or broader network access.

Shape the runtime first

Before you drive the sandbox yourself, make the agent healthy by recovering the runtime if it is degraded, so that every live cell below has a working agent to talk to.

The diagram sketches enforcement boundaries. The following commands return observations; identifying the responsible mechanism requires policy or audit evidence.

Same runtime as Kickstart. Every cell here drives the same OpenClaw agent you connected on the Kickstart page.

The blocked arrow illustrates a possible denied action. A command result alone does not identify the enforcing mechanism.

Inspect command outcomes

Run one command at a time and inspect its output, exit status, and completion. A failure can come from policy, missing software, an invalid request, or the service itself. Use the evidence and live policy to investigate the cause.

The starting catalog walks the questions you would ask of a box like this, ranging from reaching the open internet all the way to rewriting the agent's own persona, and you are free to edit any command or add your own.

Run a multi-command sandbox trajectory?

Going further: a whole trajectory

An agent can plan several commands for one goal. Inspect its plan, then compare each command with the output and exit status returned by your sandbox. Determine which results provide access-denial evidence and which need further investigation.

A secret printed by an allowed command has already reached the output channel. Inspect successful steps as well as failures; a later network denial cannot undo an earlier disclosure.

Read your launchable's live policy, then see it as a map

Read the runtime policy before interpreting the command results. The cell calls helpers.policyGet() through the operator terminal and displays its status and parsed YAML. Check whether the returned policy is active.

The map evaluates candidate destinations against the loaded policy. Change the calling binary to inspect how its grants differ, then inspect the matched rule. These are browser predictions; live requests provide a separate observation.

Clicking any target opens the rule that governs it, so that you can read which binaries that endpoint allows alongside the rule used by the browser evaluator. The View policy source control prints the YAML it drew from, which is the same text you fetched a moment ago.

Operator-token plane. The amber band sits above the checkpoint, the escalation axis the egress rules never see.

Compare a policy prediction with a live request

Choose an action, predict its result from the loaded policy, and run the same request through the sandbox. Compare the evidence without assuming that a command error proves enforcement or that a matching result proves every policy rule.

The prediction considers binary identity, destination, port, method, and path. Confirm executes the selected curl request. Other binary identities remain predictions until you test that program through its own interface.

Inspect both observations. The browser evaluates the policy while openshell sandbox exec returns command output and exit status. An HTTP error can come from a proxy or the target application; identify its source before assigning a cause.
Agreement supports the prediction for the tested action. A mismatch is a reason to check the request, active policy, response source, and evaluator assumptions.

Above the sandbox: the control plane

The cell requests operator.admin with the existing gateway token and calls read methods for persona, schedules, sessions, models, and configuration. Inspect which requests succeed. Administrative writes may be available under the server configuration, but this survey does not test them. It sends no edit requests; connections and reads can still produce logs.

What the sandbox cannot catch

The read survey does not establish write authority. Separately, harmful workflows can combine individually allowed reads and outbound channels.

Untrusted input, private data, and external communication together form the lethal trifecta.

Risk appears when these permissions compose. A policy can correctly allow one outbound endpoint, but that endpoint is still full-duplex: the request leaves and the response comes back as more input. Once hostile content can steer the model, it can ask for private data and send that data through a channel the sandbox regards as allowed. A network allowlist approves the destination, not the meaning of bytes crossing an allowed connection. Mediating any one leg above the sandbox removes the path the exploit needs.

The red center is a path across allowed capabilities. Hostile content can steer a runtime that reads protected state and uses an approved, two-way channel.

Consider these failure modes against the actual configuration:

Choose the control by failure mode

A sandbox limits what a compromised process can reach. The application and operating practice still own decisions whose meaning is invisible to the kernel.

Failure modeWhat the sandbox contributesControl still needed above it
Goal hijack or prompt injectionLimits reachable files, processes, and destinations after the model is steeredInput provenance, approval for consequential actions, and behavioural evaluation
Tool misuseRestricts the OS and network envelope in which the tool runsPer-tool authorization, argument validation, rate limits, and confirmation for irreversible calls
Identity or privilege abuseStarts the process as a non-root identity and blocks selected escalation pathsWorkload-specific identities, minimal token scopes, short lifetimes, rotation, and audit
Supply-chain or unexpected code executionReduces the files, syscalls, and egress available to installed codePinned and verified dependencies, install policy, and integrity monitoring
Memory or context poisoningCan protect selected files when policy makes them read-onlyWrite provenance, reviewable history, trust labels, and recovery from a known-good state

The categories align with the OWASP Top 10 for Agentic Applications. Use the table to pick an enforcement point instead of asking one layer to recognize every kind of harm.

References

Comprehensive list at Going Further · References.

← Module 3c: Always-On