The OpenShell Sandbox
In Module 3, files supplied persistent context and scheduled jobs started unattended work. Those operations depend on the authority available to the agent process. OpenShell restricts that authority through the active policy and its enforcement.
Run commands in the sandbox, inspect their output and exit status, and investigate the cause. Then survey read access through the operator control plane.
- OpenClaw is the agent framework. It decides which tool to use and what action to attempt.
- OpenShell is the security runtime. It limits the files, network destinations, and operating-system operations available to that action.
- NemoClaw is the reference stack that packages OpenClaw and OpenShell with lifecycle commands, policy presets, inference guidance, and deployment defaults.
OpenShell is agent-agnostic. The same boundary can enclose OpenClaw, Hermes, LangChain deepagents, or another process that needs operating-system-level containment.
Where each defense layer intervenes
Suppose untrusted content tells an agent to read a credential and send it away. Three different layers get a chance to stop that action, and each layer answers a different question.
- The prompt shapes intent. A rule such as "do not reveal credentials" tells the model how it should behave. Models can still be deceived by Adversarial retrieved content, malicious tool output, or another agent in the loop. This layer influences the decision. The agent retains whatever authority the runtime granted it.
- The agent harness validates the request. OpenClaw or Hermes can reject an unsupported shell form, a forbidden tool call, or a malformed argument before the request reaches the operating system. That deterministic check is useful, but it remains application logic that can be misconfigured, outdated, or compromised.
- The sandbox limits the result. OpenShell gives the process only the authority allowed by policy. If the prompt and harness both fail, the operating system can still deny a file, network, or process operation before it completes.
How the sandbox mechanisms answer testable questions
This launchable combines four operating-system mechanisms. Each answers a question you can test, and together they apply policy at more than one boundary. Read the live policy later on this page before treating any one guarantee as universal.
1 · Where can this process connect?
A network namespace (netns) gives the
agent an isolated network stack. Outbound access starts closed. An
HTTP CONNECT proxy provides the only route out.
Open Policy Agent (OPA)
checks the requesting binary, host, port, and path against policy.
Whoever edits that policy decides which connections are allowed.
2 · Which files can it reach?
Landlock is a Linux Security Module that restricts future filesystem access. Here it limits new file operations to paths granted by policy. That is strong enforcement, but compatibility mode, OverlayFS, and handles opened before restriction still matter. Read the live policy and test the path you intend to expose.
3 · Which kernel operations can it request?
A system call, or syscall, asks the Linux kernel to
perform an operation. seccomp
is Linux syscall filtering; its BPF filter checks each request against
an allowlist. A denied call returns EPERM, meaning
"operation not permitted," before the operation runs. This launchable
excludes dangerous primitives such as ptrace,
mount, and setuid.
4 · How much authority does it start with?
The agent runs as a non-root process under the
unprivileged sandbox user. A successful exploit therefore
starts with less authority. The lower starting authority reduces
blast radius, though it does not prove escalation is impossible. The
host, container configuration, and syscall policy remain part of the
security boundary.
Shape the runtime first
Before you drive the sandbox yourself, make the agent healthy by recovering the runtime if it is degraded, so that every live cell below has a working agent to talk to.
The diagram sketches enforcement boundaries. The following commands return observations; identifying the responsible mechanism requires policy or audit evidence.
Same runtime as Kickstart. Every cell here drives the same OpenClaw agent you connected on the Kickstart page.
Inspect command outcomes
Run one command at a time and inspect its output, exit status, and completion. A failure can come from policy, missing software, an invalid request, or the service itself. Use the evidence and live policy to investigate the cause.
The starting catalog walks the questions you would ask of a box like this, ranging from reaching the open internet all the way to rewriting the agent's own persona, and you are free to edit any command or add your own.
Run a multi-command sandbox trajectory?
Going further: a whole trajectory
An agent can plan several commands for one goal. Inspect its plan, then compare each command with the output and exit status returned by your sandbox. Determine which results provide access-denial evidence and which need further investigation.
A secret printed by an allowed command has already reached the output channel. Inspect successful steps as well as failures; a later network denial cannot undo an earlier disclosure.
Read your launchable's live policy, then see it as a map
Read the runtime policy before interpreting the command results. The cell calls
helpers.policyGet() through the operator terminal and displays its status
and parsed YAML. Check whether the returned policy is active.
The map evaluates candidate destinations against the loaded policy. Change the calling binary to inspect how its grants differ, then inspect the matched rule. These are browser predictions; live requests provide a separate observation.
Clicking any target opens the rule that governs it, so that you can read which binaries that endpoint allows alongside the rule used by the browser evaluator. The View policy source control prints the YAML it drew from, which is the same text you fetched a moment ago.
Compare a policy prediction with a live request
Choose an action, predict its result from the loaded policy, and run the same request through the sandbox. Compare the evidence without assuming that a command error proves enforcement or that a matching result proves every policy rule.
The prediction considers binary identity, destination, port, method, and path. Confirm executes the selected curl request. Other binary identities remain predictions until you test that program through its own interface.
openshell sandbox exec returns command output and exit status. An HTTP error
can come from a proxy or the target application; identify its source before assigning a cause.
Above the sandbox: the control plane
The cell requests operator.admin with the existing gateway token and calls read methods for persona, schedules, sessions, models, and configuration. Inspect which requests succeed. Administrative writes may be available under the server configuration, but this survey does not test them. It sends no edit requests; connections and reads can still produce logs.
What the sandbox cannot catch
The read survey does not establish write authority. Separately, harmful workflows can combine individually allowed reads and outbound channels.
Untrusted input, private data, and external communication together form the lethal trifecta.
- Private data is anything the agent may read but should not leak: workspace files, emails, tokens, memory, or internal records.
- Untrusted input is content from outside the operator's authorship that the model may still treat as instructions: a web page, issue, document, tool result, or peer-agent response.
- External communication is any path that can carry bytes out: HTTP, a pull request, chat, a log line, an image load, or even a link the user might click.
Risk appears when these permissions compose. A policy can correctly allow one outbound endpoint, but that endpoint is still full-duplex: the request leaves and the response comes back as more input. Once hostile content can steer the model, it can ask for private data and send that data through a channel the sandbox regards as allowed. A network allowlist approves the destination, not the meaning of bytes crossing an allowed connection. Mediating any one leg above the sandbox removes the path the exploit needs.
Consider these failure modes against the actual configuration:
- Persona tamper. If the agent can modify SOUL.md, inspect ownership and write permissions. Policy, host-managed file protection, and reviewable history can help protect persona changes.
- Control-plane authority. Token scope and server configuration determine which administrative methods are available. Successful reads do not prove permission to change the persona, schedule, or conversation.
- Inbound injection. An allowed channel can carry adversarial instructions. Network restrictions do not establish that incoming content is trustworthy; evaluate input handling and consequential actions separately.
- Exfiltration on an allowed channel. A permitted endpoint is a path an exfiltration can ride. The policy proves the channel is allowed and says nothing about whether the bytes on it are benign.
- Hidden intent. Individually permitted actions can combine into a harmful workflow. Review their combined effects when assessing a tool log.
- Peer preservation. An agent's verdict about a peer can change with social context. Compare its judgments across controlled variations of that context.
Choose the control by failure mode
A sandbox limits what a compromised process can reach. The application and operating practice still own decisions whose meaning is invisible to the kernel.
| Failure mode | What the sandbox contributes | Control still needed above it |
|---|---|---|
| Goal hijack or prompt injection | Limits reachable files, processes, and destinations after the model is steered | Input provenance, approval for consequential actions, and behavioural evaluation |
| Tool misuse | Restricts the OS and network envelope in which the tool runs | Per-tool authorization, argument validation, rate limits, and confirmation for irreversible calls |
| Identity or privilege abuse | Starts the process as a non-root identity and blocks selected escalation paths | Workload-specific identities, minimal token scopes, short lifetimes, rotation, and audit |
| Supply-chain or unexpected code execution | Reduces the files, syscalls, and egress available to installed code | Pinned and verified dependencies, install policy, and integrity monitoring |
| Memory or context poisoning | Can protect selected files when policy makes them read-only | Write provenance, reviewable history, trust labels, and recovery from a known-good state |
The categories align with the OWASP Top 10 for Agentic Applications. Use the table to pick an enforcement point instead of asking one layer to recognize every kind of harm.
References
- Simon Willison, The lethal trifecta (2025). The articulation of the private-data, untrusted-content, network-access exploit class the lethal-trifecta diagram in Step 3 draws.
- Greshake et al. (2023) Not what you've signed up for. The canonical indirect prompt-injection paper: attacks delivered through retrieved or rendered content the model trusts.
- Saltzer and Schroeder, The Protection of Information in Computer Systems (1975). Where least privilege was first named; the principle behind deny-by-default containment.
- seccomp filter and Landlock. The kernel syscall-allowlist and filesystem-capability primitives
openshelluses for mechanisms 2 and 3. - Open Policy Agent. The policy-as-data engine behind the OPA-evaluated CONNECT proxy that gates every outbound connection.
- NeMo Guardrails. The application-layer complement to kernel containment, catching policy violations the kernel cannot see.
- OWASP Top 10 for Agentic Applications. A current taxonomy for goal hijack, tool misuse, identity abuse, supply-chain compromise, unexpected code execution, and memory poisoning.
Comprehensive list at Going Further · References.