# Introduction (/docs)
## What is Agent Sandbox?
**Agent Sandbox** is an open-source sandbox engine for AI agents. It is purpose-built for three classes of workload:
- **Lightning Fast** — pre-warmed pools keep isolated environments on standby, eliminating cold-start latency for high-frequency agent loops, evaluations, and RL rollouts
- **Enterprise Grade** — deploy on any cloud using native Kubernetes CRDs, RBAC, and multi-cluster routing, without vendor lock-in
- **Agentic RL** — stateful environments with deterministic resets and any-image runtimes, built for complex multi-turn agent training
---
## Key Features
| | Feature | Description |
|---|---------|-------------|
| ⚡ | **Speed — Sub-60ms allocation** | Pre-warmed pools deliver idle sandboxes instantly, unblocking high-volume agent loops and multi-turn RL rollouts |
| ☸️ | **Infrastructure — Containers or microVMs** | Run on your existing estate using CRDs, namespaces, RBAC, and autoscaling to manage warm capacity efficiently |
| 🌐 | **Routing — Cross-region and cross-cloud** | Dispatch requests across clouds, clusters, and regions without forcing application teams to manage routing logic |
| 🧪 | **Runtime — Zero-rebuild runtimes** | Run any Docker image for SWE tasks, RL environments, and internal tools without building custom VM images |
| 🔌 | **Ecosystem — Drop-in agent SDKs** | Seamless compatibility with E2B clients, SWE-ReX workflows, and popular reinforcement learning frameworks |
| 📊 | **Observability — Console-grade visibility** | Complete view of pools, active sessions, logs, and metrics through a unified product console |
---
## Use Cases
### Reinforcement Learning at Scale
RL training requires thousands of environment resets per hour. Agent Sandbox pre-warms a pool of sandboxes so each rollout worker gets a fresh, isolated environment in milliseconds — removing the environment-reset bottleneck from your training loop. Supports SWE-bench Verified, SWE-Gym, Terminal-bench, and custom task distributions.
### AI Coding Agents & Evaluations
Give every agent turn or eval call its own isolated execution environment. The E2B-compatible API means existing SWE-agent, SWE-ReX, and similar frameworks work without modification.
### Enterprise Multi-Cluster Deployment
Deploy sandbox pools across multiple clouds or regions. The built-in ExtProc routing layer dispatches requests to the most available cluster transparently — no routing logic required in application code. Supported cloud providers: AWS, Google Cloud, Azure, Alibaba Cloud, Volcengine, Cloudflare.
**Isolation**
> Sandboxes are containers by default. Where the code is untrusted, run them in
> their own kernel instead — a microVM runtime class is one field on the template:
> [E2B Kata](/docs/examples/templates/e2b-kata).
## Start here
- [Installation](/docs/installation) — the platform on a cluster, then the CLI
and its skills
- [Examples](/docs/examples/templates) — templates and environments to apply
- [Concepts](/docs/concepts) — what a template, an env and a pool are
- [CLI guide](/docs/tutorials/cli) — driving the platform from a shell or an agent
---
# Installation (/docs/installation)
Agent Sandbox is two halves, and most people install both:
- the **platform** — an operator, an API and the data plane that runs the
sandbox Pods — which goes on a Kubernetes cluster, with Helm;
- the **client** — the `abx` CLI and the skills an agent reads — which goes on
your machine, or in the image your agent runs in.
The [CLI guide](/docs/tutorials/cli) and the [concepts](/docs/concepts) assume
both are done.
## What you need
| | |
|---|---|
| A Kubernetes cluster | 1.28 or newer; the charts install CRDs and cluster-scoped RBAC, so you need cluster-admin |
| `kubectl` | pointed at that cluster |
| Helm | 3.14 or newer — the charts are published as OCI artifacts |
Every image is public on GHCR, so there is no registry login to perform. The
charts default to `latest`; pin the chart version and the image tags for
anything you intend to keep.
## 1. The platform
```bash
helm upgrade --install agent-sandbox-worker \
oci://ghcr.io/scitix/agent-sandbox-worker \
--version 0.1.0 \
--namespace agentbox-system --create-namespace \
--set controller.localClusterId=YOUR_CLUSTER_ID
```
One release is the whole platform on one cluster:
| Installs | What it is |
|---|---|
| three CRDs | `SandboxTemplate`, `SandboxEnv`, `SandboxPool` — [the object model](/docs/concepts/index) |
| the controller | the operator that renders environments and pools, and the API `abx` talks to |
| ExtProc + Envoy | the data plane, the gateway a sandbox's ports are published through |
`controller.localClusterId` names this cluster and is **required**: it is how a
member Pool is placed on a cluster segment, and how the console labels the
cluster it is showing. Left empty, Pool writes answer `503 server-misconfigured`
and the environment reconciler has nowhere to put a member.
**Pin the version**
> `--version 0.1.0` is the chart; the images inside it default to `latest`. For
> anything you keep, pin both — `--set controller.image.tag=…`,
> `--set extproc.image.tag=…` — so a rollout is something you chose.
### Check it came up
```bash
kubectl -n agentbox-system get pods
kubectl get crd | grep agents.navix.sh
```
### Reach it
Everything a client needs is a Service in that namespace:
| Service | Port | What it is |
|---|---|---|
| `agent-sandbox-api` | 80 | the native API — what `abx` talks to |
| `agent-sandbox-e2b-api` | 80 | the E2B-compatible API |
| `agent-sandbox-data-plane` | 80 | the gateway that sandbox ports are reached through |
From outside the cluster, either publish the ingress the chart ships
(`--set extproc.ingress.host=agentbox.example.com`, with an ingress controller
installed) or, for a first look, forward one port:
```bash
kubectl -n agentbox-system port-forward svc/agent-sandbox-api 8080:80
```
## 2. The console (optional)
The console is a separate release: a dashboard that reads the same API. Install
it in the same cluster, or in another one that can reach the worker.
```bash
helm upgrade --install agent-sandbox-hub \
oci://ghcr.io/scitix/agent-sandbox-hub \
--version 0.1.0 \
--namespace agentbox-system \
--set env.secret="$(openssl rand -hex 32)" \
--set clusters[0].id=YOUR_CLUSTER_ID \
--set clusters[0].name="YOUR_CLUSTER" \
--set clusters[0].url=http://agent-sandbox-api.agentbox-system.svc.cluster.local \
--set ingress.host=agentbox.example.com
```
Two settings matter more than the rest:
- **`env.secret`** is the shared secret between the console and the worker
(`controller.secrets.secret`). Generate it once and give both releases the
same value; the console cannot authenticate against a worker that does not
share it.
- **`clusters[]`** is the list of workers the console knows about. A console
only knows the clusters configured here — that is the list its cluster picker
shows, and the list `abx clusters` will print if you point `abx` at the
console's address.
The console's built-in assistant is off by default (`assistant.enabled`), and
its images are not part of the public release. Everything else — environments,
pools, autoscaling, quotas, templates, the vault — is the same API the CLI uses.
## 3. The CLI, and the skills
`abx` is a single binary with no runtime dependencies, and it ships nine skills
— Markdown files an agent reads before it touches the platform:
```bash
curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh | sh
```
That writes `~/.local/bin/abx` and `~/.agents/skills/`. Then name the deployment
you are talking to:
```bash
abx context set YOUR_DEPLOYMENT \
--endpoint 'https://YOUR_CONSOLE/agentbox' \
--api-key agbx_...
abx clusters
```
The endpoint is the console's address, or a worker's API address if you are not
running the console; the key is issued in the console under **API Keys** (or, on
a worker with no console, by whoever created the release). The
[CLI guide](/docs/tutorials/cli) takes it from there, and
[Skills](/docs/skills) explains what was installed for your agent.
## First sandbox
An environment is created from a **template**, and a sandbox from an
environment's **pool** — nothing exists to create sandboxes from until you add
both. The repository carries three to start with, each with a matching
environment:
```bash
kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml # the template
kubectl apply -f config/samples/e2b_sandboxenv.yaml # a warm pool on it
```
[Examples](/docs/examples/templates) has the other two — Docker inside the
sandbox, and a microVM-isolated one — with the manifests rendered on the page.
From there the [CLI guide](/docs/tutorials/cli#5-run-code-in-a-sandbox) takes
over: read the environment's own documentation for the endpoints of your
cluster, then create a sandbox. The [concepts](/docs/concepts) explain what the
two objects are for, and why a template is not the thing you pass to
`Sandbox.create()`.
## Upgrading and removing
```bash
helm upgrade agent-sandbox-worker oci://ghcr.io/scitix/agent-sandbox-worker \
--namespace agentbox-system --reuse-values --version 0.1.1
helm uninstall agent-sandbox-worker --namespace agentbox-system
```
The CRDs carry `helm.sh/resource-policy: keep`, so `uninstall` leaves them — and
every environment, pool and template you created — in place. That is deliberate:
deleting a CRD deletes its objects, and with them the record of what your
sandboxes were. Remove them explicitly when you mean it:
**Uninstall does not uninstall everything**
> `helm uninstall` removes the controller and the data plane. The CRDs, and every
> `SandboxTemplate`, `SandboxEnv` and `SandboxPool` on the cluster, stay until you
> delete the CRDs themselves — which is the command below, and is irreversible.
```bash
kubectl delete crd sandboxtemplates.agents.navix.sh sandboxenvs.agents.navix.sh sandboxpools.agents.navix.sh
```
---
# Autoscaling (/docs/concepts/autoscaling)
Autoscaling answers two questions about a set of
[pools](/docs/concepts/pools): *how few Pods may I keep when nothing is
running*, and *how many may I add when everything is busy*.
## Groups, not pools
Pools of the same resource shape inside one env form a **scaling group**, and
the policy lives on the group. The group's name is derived from the shape
(`4c64gi`-style), which is why you never create one: a group appears when a
member pool declares its shape, and is collected when the last one goes away.
```bash
abx envs YOUR_ENV scaling-groups --cluster YOUR_CLUSTER
abx envs YOUR_ENV scaling-groups YOUR_GROUP --cluster YOUR_CLUSTER
```
## The policy
| Field | What it means |
|---|---|
| `enabled` | whether the autoscaler acts on this group at all; off means the pool's `replicas` is yours to set |
| `minReplicas`, `maxReplicas` | the floor and the ceiling on the group's aggregate replicas |
| `scaleUpPolicy.mode` | `Conservative`, `Default` or `Aggressive` — the step size in both directions |
| `scaleUpPolicy.cooldownSeconds` | minimum gap between two scale-ups |
| `scaleUpPolicy.idleThresholdSeconds` | how long aggregate idle must stay at zero before a proactive scale-up |
| `scaleUpPolicy.idleZeroQuietWindowSeconds` | suppresses that trigger when nothing has been claimed for this long |
| `scaleUpPolicy.saturationCooldownSeconds` | how long a member stays marked saturated after a failed probe |
| `scaleDownPolicy.idleTimeoutSeconds` | how long a Pod idles before it counts as removable |
| `scaleDownPolicy.stabilizationSeconds` | minimum gap between two scale-downs |
| `scaleDownPolicy.protectionWindowSeconds` | how long a marked Pod can still be claimed, cancelling its deletion |
Every member pool may also carry its own `minReplicas` and `maxReplicas`, which
is how one shape is held larger than its siblings inside the same group.
## Reading the policy against the two questions
**Cost.** `minReplicas = 0` plus a short `idleTimeoutSeconds` is "keep nothing
warm"; `minReplicas = N` is "always have N claims ready", and it is a floor on
spend as much as on capacity.
**Concurrency.** The ceiling is `maxReplicas` — and, separately, your quota.
Quota is a hard cap the autoscaler cannot argue with: a group may sit below its
ceiling indefinitely while the pool reports `ResourceQuotaExhausted`. Check
both before promising a number of parallel sandboxes.
**Speed.** Scaling up creates Pods, and a Pod that has just been created still
has to start. The autoscaler removes the wait for *capacity*; it does not make
a brand-new Pod faster than a pre-warmed one, which is why a floor above zero
is what keeps first-claim latency flat.
## Changing one
```bash
abx envs YOUR_ENV scaling-groups YOUR_GROUP --editable --cluster YOUR_CLUSTER > group.json
# edit group.json
abx update envs YOUR_ENV scaling-groups YOUR_GROUP -f group.json --cluster YOUR_CLUSTER
```
Like every write, this is a PUT: the file is the whole policy. A field left out
is a field you are asking to remove — dropping `maxReplicas` removes the
ceiling rather than keeping it.
## When it does not behave
| Symptom | Usual cause |
|---|---|
| never scales down | running sandboxes, `idleTimeoutSeconds`, or the protection window still counting |
| scales down under load | `idleTimeoutSeconds` shorter than the gap between real claims |
| scales up in bursts | `idleThresholdSeconds` / `idleZeroQuietWindowSeconds` too eager for the traffic pattern |
| sits under `minReplicas` | quota, or the instance type has no room on the cluster |
| a rolled-out pool refuses claims | the previous generation is still draining; `maxUnavailable` on the env governs how fast that is |
## See also
- [Pools](/docs/concepts/pools) — what is being scaled
- [Envs](/docs/concepts/envs) — the rollout policy that governs rolling replacements
- [Capacity and quotas](/docs/tutorials/cli) — reading quota from the CLI
---
# Cross-cluster (/docs/concepts/cross-cluster)
One Agent Sandbox address reaches a platform, and a platform can reach several
clusters. You do not open a second account to use the second cluster; you
address it.
```bash
abx clusters # every cluster behind this address
abx envs --cluster YOUR_CLUSTER # the envs on one of them
```
## Same name, several clusters
An env exists per cluster — pools, quota, instance types and images are local
to one — but two envs with the **same name** on different clusters are
associated, and the platform may serve a claim for that name from either.
To build one: create the env on the first cluster, then extend it to the next
one from the console (**Sandbox Envs → extend an environment to another
cluster**). The extension creates the env there; you then add member pools on
that cluster, because instance types and quota are properties of the cluster
it lives on, and registry credentials are entered per cluster.
Two conditions, and the second fails quietly:
1. The image must exist in each cluster's registry —
[images follow you](#the-same-image-in-every-region).
2. The team → namespace mapping must match across clusters. If it does not,
federation cannot pair `(namespace, env name)` and a bare env name will not
spread — name the cluster explicitly instead.
## Addressing
The E2B SDK's first argument is an address, and the address can carry a
cluster:
| You write | You get |
|---|---|
| `YOUR_ENV` | this cluster first; if it has no idle capacity and cannot scale, a cluster with the same-named env |
| `SOME_CLUSTER::YOUR_ENV` | that cluster's env, which then routes across its own member pools |
| `SOME_CLUSTER::YOUR_POOL` | that exact pool, with no routing at all |
| `YOUR_ENV//IMAGE` | as above, with the main container image replaced |
The `//IMAGE` suffix combines with any of the others. For a pipeline that
should not care where capacity is, `YOUR_ENV` is the right spelling; for a
result you may have to reproduce, `SOME_CLUSTER::YOUR_ENV` pins the cluster and
still survives a pool being replaced by one of the same shape.
In `abx`, a cluster is a flag rather than part of the address:
`abx envs YOUR_ENV --cluster SOME_CLUSTER`.
## The same image in every region
When you name an image that belongs to another cluster's private registry, the
platform rewrites the **registry host** to this cluster's equivalent — same
type of registry, same path, same tag — so a sandbox is never pulling across
regions:
```
you write: registry-REGION-A.example.com/agentbox/eval/thing:260328
pulled from: registry-REGION-B.example.com/agentbox/eval/thing:260328
```
The rules that follow from it:
- the same image has to exist at the **same path** in each region's registry;
- rewriting only happens between registries of the same configured type;
- public registries (`docker.io` and friends) are never rewritten;
- if this cluster has no registry of that type, the address is used as written
— which may work, and may be slow, and says nothing either way.
## What is shared and what is not
| Shared across the platform | Per cluster |
|---|---|
| the template catalogue (synced) | pools, their replicas and their autoscaling |
| an env's identity and name | the env's member pools on that cluster |
| API keys | quota, instance types, namespaces |
| | registry credentials, and the images behind them |
## Common misconceptions
- **"A pool spans clusters."** It does not. A pool belongs to one cluster; the
env is what spans them.
- **"Cross-cluster is a fallback for outages."** It is capacity routing, not
failover: a cluster that has no Pod to give you cannot be scaled into
existence instantly, and the request fails rather than waiting for one.
- **"The same env name is enough."** Without the env on the target cluster, the
request has nowhere to go; without the image, it lands and then fails to pull.
## See also
- [Envs](/docs/concepts/envs) — naming and per-cluster membership
- [Pools](/docs/concepts/pools) — where a cluster's capacity lives
- [E2B Python SDK](/docs/tutorials/e2b) — create parameters and routing
---
# Egress and secrets (/docs/concepts/egress-and-secrets)
Two decisions, and conflating them is the usual mistake:
```
SandboxEnv does this environment HAVE a gateway one switch
create call what THIS sandbox may reach, and what per sandbox
may be injected into its outbound requests
```
The env carries a switch; the rules belong to an individual sandbox and arrive
with the create call, in the E2B SDK's own vocabulary. There is no separate
Agent Sandbox dialect for them.
## Why the switch is on the env and the rules are not
Enabling the gateway injects a transparent proxy sidecar. That changes the Pod
spec, and a changed Pod spec rolls the env's pools — a real, environment-level
decision, taken once. Rules are per sandbox because an env is shared: an
env-wide allowlist would be a default every caller overrides anyway.
**It fails closed.** A create that carries filtering rules against an env with
no gateway is refused with `400`, not accepted and ignored. A Pod with no
sidecar has no redirection, so accepting the rules would mean they silently did
nothing — which for an evaluation is the worst outcome, because the run
finishes and the numbers are wrong.
```bash
abx update envs YOUR_ENV --help --cluster YOUR_CLUSTER # the env's one gateway field
```
## Cutting a sandbox off, and letting one thing through
The common case is an agent under test that must not fetch the answer and must
not install its way around a missing dependency. The shape is always the same:
**deny everything, then allow what the task genuinely needs.** A denylist is a
list of the routes somebody thought of.
```python
# nothing in, nothing out
sbx = Sandbox.create("YOUR_ENV", timeout=3000, secure=False, allow_internet_access=False)
# or: everything out except these
sbx = Sandbox.create(
"YOUR_ENV",
timeout=3000,
secure=False,
network={"allow_out": ["api.openai.com", "pypi.org", "*.pythonhosted.org"]},
)
```
Naming an allow list is what makes everything else a deny — the entries are the
whole of what is reachable. A CIDR or a bare IP works in the same list.
**An empty allow list is not a deny**
> `network={"allow_out": []}` declares no filtering at all, and egress stays
> unrestricted. To cut a sandbox off, say so: `allow_internet_access=False`, or
> `deny_out=["0.0.0.0/0"]`.
Three things worth checking before calling a run isolated:
1. **The env has a gateway.** Without it, a create carrying rules is refused —
so make sure you saw a sandbox, not a `400`.
2. **The package index is not on the allowlist** unless the task needs it. It is
the most common accidental hole: the agent cannot search, but it can
`pip install` something that can.
3. **Watch it fail from inside.**
```python
sbx.commands.run("curl -sS -m 5 https://example.com") # expected: it fails
```
An isolation you have not seen fail is an isolation you are assuming.
## Credentials a sandbox can use but cannot read
```
vault (write-only) → operator memory → sidecar tmpfs → outbound header
```
The value never enters the sandbox. The sandbox holds a **decoy** — a
placeholder that looks like a token — and the egress sidecar substitutes the
real one as the request leaves, for the hosts and headers the rule names. Code
inside runs unmodified: it reads `OPENAI_API_KEY`, sends it, and the sidecar
replaces it on the way out. Library code that has never heard of Agent Sandbox
works.
Secrets live in a **vault** that is write-only: you can list names, overwrite
them and delete them, never read them back. In the console it is the **Vault**
page; through the SDK it is the E2B `/secrets` surface, which the official
package speaks as `e2b.Secret` (that module arrived in **e2b 2.43**).
### Store it once
```bash
abx envs YOUR_ENV docs --cluster YOUR_CLUSTER # API URL and data-plane domain
```
```python
import os
os.environ["E2B_API_KEY"] = "agbx_..." # your platform key
os.environ["E2B_API_URL"] = "https://YOUR_GATEWAY/agent-sandbox/api/e2b"
os.environ["E2B_DOMAIN"] = "YOUR_GATEWAY/agent-sandbox/api/data" # no scheme
from agent_sandbox_e2b import patch_e2b # before the e2b import
patch_e2b()
from e2b import Secret
Secret.create("openai-api-key", os.environ["OPENAI_API_KEY"]) # write-only
# rotate it later with Secret.update("openai-api-key", new_value)
```
### Use it at create time
Reference it **by name**; the value never leaves the vault. `Secret.fill()` is a
local formatting helper — it returns the string the wire carries, and makes no
call of its own:
```python
from e2b import Sandbox, Secret
sbx = Sandbox.create(
"YOUR_ENV",
timeout=3000,
secure=False,
# What the code inside reads. It is a decoy: the real value is put in on the
# way out, so the sandbox cannot leak what it never had.
envs={"OPENAI_API_KEY": "decoy-not-a-real-key"},
network={
# A rule is a transform, not a permission — the host has to be allowed
# too, or the request never leaves.
"allow_out": ["api.openai.com"],
"rules": {
"api.openai.com": [
{
"transform": {
"headers": {
"Authorization": f"Bearer {Secret.fill('openai-api-key')}",
}
}
}
]
},
},
)
```
Ordinary code inside, no Agent Sandbox dialect:
```python
sbx.commands.run(
'curl -s https://api.openai.com/v1/models -H "Authorization: Bearer $OPENAI_API_KEY"'
) # succeeds
sbx.commands.run("echo $OPENAI_API_KEY") # prints the decoy, not the key
sbx.commands.run("curl -sS -m 5 https://example.com") # fails: not allowed
sbx.kill()
```
A plaintext credential in the rules is refused with `400` — deliberately,
because accepting it would put the value in the request body, the access log
and the caller's source, which is the exposure the feature exists to remove.
Wildcard hosts are refused for the same reason: whoever controls a matching
subdomain would receive the injected credential.
### The env has to have the gateway
All of the above is refused on an Env whose gateway is off, with a `400` that
says so:
```
a network policy was requested but this environment has no egress gateway;
enable it on the SandboxEnv (overrides.gateway.enabled) and let its pools roll
```
That is the switch from the top of this page, and turning it on is a change to
the Pod spec, so it rolls the Env's pools.
## When it silently does not work
The sandbox runs, the request goes out, and nothing is substituted. Check in
this order:
| Check | Why |
|---|---|
| the env has the gateway on | without it, a create with rules is refused — so if a sandbox exists, this one is satisfied |
| the host matches the rule exactly | a redirect to another host is not covered |
| the port is 80 or 443 | only those are parsed at layer 7; rules for other ports never fire and nothing says so |
| the secret name exists in **your** vault | a name that resolves to nothing leaves the decoy in place, and the upstream answers `401`, which reads like a bad key |
`abx whoami` tells you which identity the vault is read as; a secret stored by
one user is not visible to another.
## Agents, vaults and approval
Writing a vault secret **is** allowed to an agent credential, under the normal
approval gate: the credential lands in the acting person's own vault and widens
nobody's authority. Minting an Agent Sandbox API key is the act that is
refused. See [API key permissions](/docs/tutorials/cli#6-api-key-permissions).
## See also
- [Envs](/docs/concepts/envs) — the gateway switch and the rest of the settings
- [Pools](/docs/concepts/pools) — why enabling the gateway rolls them
- [In-place update](/docs/concepts/inplace-update) — what a claim does to a Pod
---
# Sandbox environments (/docs/concepts/envs)
A **SandboxEnv** is the addressable unit: the name you pass to
`Sandbox.create()`, the thing the console shows a page for, and the object
that binds exactly one [template](/docs/concepts/templates).
```python
sbx = Sandbox.create("YOUR_ENV", timeout=3000, secure=False)
```
An env is one class of runtime — "the E2B sandboxes for this team", "the
Docker-in-Docker ones for this evaluation" — and it fans out to member
[pools](/docs/concepts/pools) that hold the actual capacity. It holds no Pods
itself.
## What an env owns
These are the settings every sandbox claimed from the env inherits:
| Setting | What it does |
|---|---|
| `overrides.image` | replaces the template's main container image for every member pool |
| `overrides.podCreationImagePolicy` | whether a new Pod starts on the template's image or on the idle image |
| `overrides.defaultStartupTimeout` | how long a claim waits for the sandbox to become ready, when it does not say |
| `overrides.defaultIdleTimeout` | how long an untouched sandbox lives before it is reclaimed, when it does not say |
| `overrides.imagePullSecret` | credentials for a private registry, materialised into a Secret the member pools reference |
| `overrides.gateway` | the **switch** that gives the env's Pods an egress sidecar — see [egress and secrets](/docs/concepts/egress-and-secrets) |
| `overrides.volumes` | PersistentVolumeClaims mounted into every sandbox |
| `overrides.updateStrategy` | `autoUpdate` and `maxUnavailable` for rollouts |
| `labels`, `annotations` | metadata stamped on the env and its member pools |
Changing most of these changes the Pod spec, and a changed Pod spec means the
env's pools **roll**: existing idle Pods are replaced, one rollout at a time,
and running sandboxes are left alone until they are returned.
## Mode
| Mode | Behaviour |
|---|---|
| `WarmPool` | claims are served from the env's member pools. This is what the platform runs. |
| `OnDemandJob` | reserved: accepted by the API enum, not implemented by the controller. Do not build on it. |
## The name goes everywhere
An env name is an RFC 1123 DNS label, capped at 24 characters because pool and
Pod names are derived from it (`POOL = ENV + resourceKey (+ quotaShort)`,
`POD = POOL + uuid`) and those have to stay inside the 63-character limit.
The same name is the env in the E2B SDK, the console URL, and `abx`. Create the
same-named env in two clusters and they federate —
[cross-cluster](/docs/concepts/cross-cluster).
## Reading and writing one
```bash
abx envs --cluster YOUR_CLUSTER # every env on the cluster
abx envs YOUR_ENV --cluster YOUR_CLUSTER # one env: template, mode, pools, sizing rule
# the safe way to change one
abx envs YOUR_ENV --editable --cluster YOUR_CLUSTER > env.json
# edit env.json
abx update envs YOUR_ENV -f env.json --cluster YOUR_CLUSTER
```
`abx create envs --help` prints the whole body, field by field. Two things
about it are worth knowing before you edit a file:
- **A write is a PUT.** A field you leave out is a field you are asking to
remove, and `overrides` is replaced wholesale.
- **`imagePullSecret` cannot be read back.** `GET` reports
`imagePullSecretConfigured` instead of the credentials, and that flag is also
the only way to say "keep them": a PUT carrying neither the secret nor the
flag is asking for the stored credentials to be deleted. Editing an unrelated
setting with a file that dropped the flag revokes your registry access.
**The registry credentials are one flag away from being deleted**
> `--editable` prints `imagePullSecretConfigured: true`, not the secret. Send that
> file back unchanged and the credentials survive. Strip the flag while editing
> something else and the PUT means "delete them" — the next image pull fails, and
> nothing in the response says why.
## Common misconceptions
- **"The env is a namespace."** It is not; the namespace comes from your team.
An env is a runtime identity plus its settings.
- **"An env holds capacity."** Pools do. An env with no member pool has no
sandboxes to hand out, however it is configured.
- **"I can repoint an env at another template."** `templateRef` is fixed at
create. Moving runtimes means creating an env on the new template and moving
traffic.
## See also
- [Pools](/docs/concepts/pools) — where the capacity is
- [Autoscaling](/docs/concepts/autoscaling) — the size of that capacity over time
- [Egress and secrets](/docs/concepts/egress-and-secrets) — the env's gateway switch
---
# The object model (/docs/concepts)
Agent Sandbox is four objects and the chain between them. The console, `abx`
and the E2B SDK are three ways of driving that same chain.
```mermaid
flowchart TB
T["SandboxTemplate
the Pod shape and runtime
platform admin"]
E["SandboxEnv
the name you create sandboxes with
you"]
P["SandboxPool
pre-warmed Pods of one resource shape
you"]
S(["sandbox
your code, running"])
T -->|binds to one| E
E -->|owns members| P
P -->|a claim hands one over| S
```
| Object | What it decides | Who creates it | In the console |
|---|---|---|---|
| **SandboxTemplate** | the Pod shape, the idle image, default timeouts, an environment's own documentation | platform admin | Sandbox Templates |
| **SandboxEnv** | which runtime, the timeouts, egress, the rollout policy | you | Sandbox Envs |
| **SandboxPool** | how much warm capacity of one resource shape there is | you | the env's Pools tab |
| **Pod → sandbox** | one running sandbox, handed to one claim | the platform | Sandboxes |
## The same word means different things
Agent Sandbox runs on Kubernetes and serves the E2B API, so the two vocabularies
meet in one place — and one word is genuinely misleading:
| You say | E2B SDK | Agent Sandbox |
|---|---|---|
| template | a snapshot built from a Dockerfile | a **container image you bring**, plus the platform's `SandboxTemplate` that says how to run it — nothing is built |
| the first argument of `Sandbox.create()` | the template | the **env** |
| pool | — | a `SandboxPool`: warm Pods, which E2B has no equivalent of |
| sandbox | `Sandbox` | a claimed Pod |
Read the first two rows together: in the E2B SDK the argument is called
`template`, and on Agent Sandbox the value you put there is an env name. A
template is what an env is *made from*, not what you create sandboxes with —
[Sandbox templates](/docs/concepts/templates) is about that difference, and
about why there is no build step and no snapshot format here.
## Two facts the design follows from
**A sandbox is claimed, not created.** Pools hold Pods that already exist, so
serving a request is a swap rather than a build — which is where the
sub-second start comes from, and what
[in-place update](/docs/concepts/inplace-update) is about.
**Capacity is per cluster; identity is per name.** Quota, pools and images
belong to one cluster, and an env with the same name in two clusters is one
env the platform may dispatch to: [cross-cluster](/docs/concepts/cross-cluster).
## Which page answers which question
| Question | Page |
|---|---|
| What is a SandboxTemplate, and how does it differ from an E2B template? | [templates](/docs/concepts/templates) |
| What can I configure on an environment? | [envs](/docs/concepts/envs) |
| How many sandboxes can run at once, and what does one cost? | [pools](/docs/concepts/pools) |
| Why is my pool not scaling, or scaling too far? | [autoscaling](/docs/concepts/autoscaling) |
| Why is a claim fast, and what happens to the Pod afterwards? | [in-place update](/docs/concepts/inplace-update) |
| How do I run the same workload on another cluster? | [cross-cluster](/docs/concepts/cross-cluster) |
| How do I cut a sandbox off the network, or give it a credential it cannot read? | [egress and secrets](/docs/concepts/egress-and-secrets) |
| What vocabulary does my agent have for all of this? | [skills](/docs/skills) |
## Driving it
```bash
abx clusters # every cluster this address reaches
abx envs --cluster YOUR_CLUSTER # the envs on one of them
abx envs YOUR_ENV --cluster YOUR_CLUSTER # one env: its template, pools, sizing rule
```
Everything the console shows for these objects is readable from `abx`, and
almost all of it is writable from `abx` too. The
[CLI guide](/docs/tutorials/cli) covers the setup; the
[E2B SDK guide](/docs/tutorials/e2b) covers the sandbox side.
---
# In-place update (/docs/concepts/inplace-update)
An in-place update is how a claim is served. A
[pool](/docs/concepts/pools) holds Pods that already exist and already run a
runtime; a claim swaps the container image on one of them instead of
scheduling anything new. It is the mechanism behind the platform's start
latency, and behind "run any image without building a template".
## The life of a Pod
```mermaid
stateDiagram-v2
direction LR
[*] --> Idle : the pool creates it
Idle --> Running : a claim swaps the image in place
Running --> Idle : killed, or the idle timeout expires
Idle --> [*] : the pool is scaled down or rolled
```
The same three states, in the order they matter:
**Idle — the Pod runs the template's `idleImage`.** That is the sandbox runtime
sitting in front of no workload, which is what makes it cheap to keep Pods
around: twenty idle Pods are not twenty copies of your image.
**Claim — the image is swapped in place.** The platform picks an idle Pod and
replaces the main container's image with the one the create asked for.
Kubernetes restarts that container on the same Pod: no scheduling, no new Pod,
no volume re-attach.
**Running — your process starts**, and the sandbox is handed to the caller.
**Return — the Pod goes back to idle.** When the sandbox is killed or times
out, the container is restarted onto the idle image and the previous sandbox's
metadata is cleared, so the next claim does not inherit it.
## What that buys, and what it costs
| Buys | Costs |
|---|---|
| a claim does not wait for a scheduler, a node, or a volume | the container image still has to be pulled — first claim for a large image is slower |
| any image, no per-workload build | Pods cannot change *shape*: requests, limits, volumes and sidecars are fixed for the Pod's life |
| idle Pods are cheap, so pools can be deep | the image is chosen per claim, so two sandboxes from one pool may run different images |
The practical consequences:
- **Sizing is a pool decision, not a claim decision.** A different resource
shape is a different pool, never a resized Pod.
- **The image is resolved when you claim.** Edit the env's or template's image
and the next claim uses it. Nothing is rebuilt, and nothing that is already
running is disturbed — see [templates](/docs/concepts/templates).
- **Warm Pods make claims fast; images make them slower.** If first-claim
latency matters, that is an image-size question, and the pools holding that
image's shape are the ones worth keeping warm.
## Three timeouts, three different things
| Timeout | Counts | Where it is set |
|---|---|---|
| **Startup** | from the claim to the sandbox being ready — dominated by the image pull | the env's/template's default, or `agentbox.scitix.ai/startup-timeout` on the create |
| **Idle** | from the last activity to the sandbox being reclaimed | `timeout=` on the create, `sbx.set_timeout(…)` to extend, the env's default otherwise |
| **Pod lifetime** | not a user-facing timeout; a Pod is recycled by rollouts and by the pool | — |
## Rolling is not the same thing
An in-place update changes one Pod's image to serve one claim. A **roll** is
the platform replacing idle Pods because the Pod's identity changed — the idle
image, the Pod body, the gateway, or the template's metadata. Rolls are
governed by the env's `updateStrategy` (`autoUpdate`, `maxUnavailable`) and can
be watched on a pool as `updateRevision`/`updatedReplicas`, with Pods counting
through `startingReplicas` and `stoppingReplicas`.
The running image is deliberately excluded from that identity, which is why a
roll is not triggered by changing the workload image.
## See also
- [Pools](/docs/concepts/pools) — where the Pods are
- [Templates](/docs/concepts/templates) — idle image and defaults
- [E2B Python SDK](/docs/tutorials/e2b) — `timeout`, `set_timeout`, `kill`
---
# Sandbox pools (/docs/concepts/pools)
A **SandboxPool** is a set of Pods of one resource shape, kept warm for an
[env](/docs/concepts/envs). It is where capacity and cost live: an env with no
pool can be configured perfectly and still serve nothing.
## The numbers on a pool
| Field | Meaning |
|---|---|
| `replicas` | how many Pods this pool keeps — the pool's **theoretical maximum concurrency** |
| `idleReplicas` | Pods ready to claim right now |
| `runningReplicas` | Pods currently in use by a sandbox |
| `startingReplicas`, `stoppingReplicas` | Pods mid-transition, on their way in or out |
| `unavailableIdleReplicas` | idle Pods that are not Ready — counted as idle, but unable to take a claim |
| `pendingRequests` | claims queued against this pool |
`replicas` counts every state, so a claim that finds `idleReplicas = 0` waits,
and the wait ends when a Pod is returned or the pool grows
([autoscaling](/docs/concepts/autoscaling)).
## Names are derived, shapes are fixed
A pool's name is derived from the env name and the effective resources — as in
`YOUR_ENV-1c16gi-10-ondemand` — and the same derivation produces its scaling
group. Both follow from the shape, which is why the shape is **fixed at
create**: a pool does not resize, it gets replaced by one of a different shape.
That is also why the console asks you for the shape, not for a size: a pool's
identity *is* its shape.
## Which shape you may declare: `poolSizing`
Whether a pool names a quota and an instance type is a property of the env's
template, not a matter of taste. Read it off the env before writing a body:
```bash
abx envs YOUR_ENV --cluster YOUR_CLUSTER # prints poolSizing
```
| `poolSizing` | What a member pool must declare | What it refuses |
|---|---|---|
| `billed` | a quota label **and** `instanceType` (with optional `multiplier`) | a pool without them |
| `free-form` | `inlineResources` only | `instanceType`, `multiplier`, quota labels |
| `either` | the caller chooses | nothing — the deployment states no rule |
On a **billed** env the instance type is a billing *envelope*: `instanceType ×
multiplier` is reserved and charged, and `inlineResources` may then ask for less
than the envelope (rounded down is allowed, rounding up is refused). On a
**free-form** env, `inlineResources` is the whole size of the Pod.
The server enforces this rather than trusting the client, so the answer to
"which fields do I send" is always the env's, never a template you copied.
**Read the rule before writing a body**
> ```bash
abx envs YOUR_ENV --cluster YOUR_CLUSTER # poolSizing: billed | free-form | either
> ```
> `abx create envs pools -f` refuses the wrong shape with a `400` that names
> the field. The console's Pool form shows the same rule, so the two cannot
> disagree.
## Writing one
```bash
abx create envs YOUR_ENV pools --help # the body, field by field
abx envs YOUR_ENV pools --cluster YOUR_CLUSTER # what exists
abx create envs YOUR_ENV pools -f pool.json --cluster YOUR_CLUSTER
abx scale envs YOUR_ENV pools YOUR_POOL --replicas 4 --cluster YOUR_CLUSTER
```
`abx scale` changes size and nothing else: it re-sends the current bounds
unchanged. Everything else is a `create`/`update` with a file.
## Several pools, one env
One env may hold pools of different shapes and different quotas, and a claim
against the env is routed across them by the platform — by availability first,
with the pool's priority as a tie-break.
To make placement deterministic, name the shape you want instead of letting the
platform choose: pass `agentbox.scitix.ai/scaling-group` in the create metadata
(or address `cluster::pool` directly, [cross-cluster](/docs/concepts/cross-cluster)).
If that group has no member in the env, the create fails with `503` rather than
quietly landing on another shape — a deliberate choice, because a run that
silently half-succeeded on the wrong hardware is worse than a run that did not
start.
## When capacity is not there
| Symptom | Where to look |
|---|---|
| `idleReplicas = 0`, claims queue | add a pool, or let the group scale ([autoscaling](/docs/concepts/autoscaling)) |
| the pool will not grow to `replicas` | quota ([reading a quota](#reading-a-quota)): `abx quotas --cluster YOUR_CLUSTER`, and the pool's `ResourceQuotaExhausted` condition |
| idle Pods never become claimable | `unavailableIdleReplicas` — Pods that are not Ready; usually an image pull |
| an env has pools but no capacity | the pools may be on another cluster in the federation |
### Reading a quota
`abx quotas` answers one row per instance type, because that is the unit a pool
is sized in — a pool names a quota *and* a shape, and a quota's total across
shapes is not a number anyone can spend. Each row's `ceiling` is one of three
things, and which one decides whether a submission can succeed:
| `ceiling` | Meaning |
|---|---|
| a number | the enforced cap for that instance type |
| `0` | nothing allocated to you on this pool; a submission against it is refused |
| `unlimited` | this deployment skips the quota check for that pool — the ondemand and spot pools are built that way |
`unlimited` is not a promise of capacity. Nothing is capping you, so what decides
is the pool's own stock: it is worth trying, and worth retrying once other
tenants release Pods. The `name` column is the quota url a pool carries as
`labels["quota.scitix.ai/url"]`.
## See also
- [Autoscaling](/docs/concepts/autoscaling) — bounds, policies and cooldowns
- [In-place update](/docs/concepts/inplace-update) — what a claim does to a Pod
- [CLI guide](/docs/tutorials/cli) — endpoints, keys and the address grammar
---
# Sandbox templates (/docs/concepts/templates)
A **SandboxTemplate** is what a sandbox is made from: a Pod shape, an idle
image, a runtime, the default timeouts, and the documentation an environment
carries. It is a Kubernetes object, and it is the platform's, not yours.
## It is not an E2B template
This is the first thing E2B users ask, and the answer changes how you write
code:
| | E2B template | Agent Sandbox `SandboxTemplate` |
|---|---|---|
| What it is | a build artifact: a Dockerfile compiled into a snapshot | a Kubernetes object: Pod template, idle image, runtimes, defaults |
| How it comes to exist | you run `e2b template build` | the platform publishes it; nothing is built per workload |
| Where the workload image comes from | baked into the snapshot at build time | chosen when a sandbox is created, and swapped in on the claim |
| What you pass to `Sandbox.create()` | the template id | the **env** name |
| Versions | template builds | `spec.version` is a label a person maintains |
So there is no build step and no per-workload snapshot to maintain: an env can
serve any image you can pull, and changing that image does not rebuild
anything. The cost of that design is visible in
[in-place update](/docs/concepts/inplace-update).
## What a template decides
| Field | What it means |
|---|---|
| `template` | the Pod template: containers, sidecars, resources, volumes, node selectors |
| `idleImage` | what an **unclaimed** Pod runs. Defaults to the main container's image |
| `runtimes` | the sandbox runtime the Pod carries (the E2B runtime is the default) — that runtime is envd, [and these are the patches we carry](/docs/designs/envd) |
| default startup / idle timeouts | what a create means when it does not say |
| documentation | the text `abx envs docs` prints, rendered per cluster at read time |
| visibility | which teams and users see the template at all; empty means public |
| `version` | a human-maintained string, surfaced as the template's version |
## What a template does not decide
Replicas, quota and instance type belong to a pool, not to a template: one
template can back several envs, each with pools of different sizes. The env
may override a few template fields — the image, the image policy, the default
timeouts, private-registry credentials — uniformly for all of its pools; see
[envs](/docs/concepts/envs).
Resource *shape* is the exception worth knowing: Pods in one pool all have the
same requests and limits. Resizing a workload means another pool, not another
Pod — [pools](/docs/concepts/pools).
## Reading one
```bash
abx templates --cluster YOUR_CLUSTER # the catalogue
abx templates YOUR_TEMPLATE --cluster YOUR_CLUSTER # one, with its docs
```
`abx templates ` prints the template's documentation and its raw object,
which is what a diff against a rollout is read against. Writing a template is
`abx admin-templates` and needs an admin credential — the catalogue everyone
else reads is read-only.
## Versions, and what actually triggers a rollout
`spec.version` is a label a person keeps up to date; the platform does not
resolve it to a revision. The thing that decides whether idle Pods are stale
is a **revision hash** over the materialised idle Pod — the idle image, the Pod
body, the gateway and the template's metadata.
The running image is deliberately *not* part of that hash: an idle Pod is not
running your workload image, and the image it will run is resolved when a
sandbox is claimed. Edit a template's running image and the next claim picks it
up with no rollout; edit anything in the idle Pod's identity and the pools roll
onto the new revision, governed by each env's update strategy
([envs](/docs/concepts/envs), [in-place update](/docs/concepts/inplace-update)).
## Common misconceptions
- **"The template is my image."** The template carries a default image, but the
image a sandbox runs is chosen per create and can differ every time.
- **"A template bump upgrades running sandboxes."** It does not touch a running
sandbox. Rollouts replace **idle** Pods; claims that already happened keep
what they claimed.
- **"Templates are per env."** They are cluster-scoped and shared: several envs
may bind one template, and an admin editing it affects all of them.
## See also
- [Envs](/docs/concepts/envs) — what binds a template and why
- [In-place update](/docs/concepts/inplace-update) — how a claim uses the idle image
- [E2B Python SDK](/docs/tutorials/e2b) — creating sandboxes from an env
---
# Cross-cluster routing (/docs/designs/cross-cluster)
[Cross-cluster](/docs/concepts/cross-cluster) is what a user writes: a bare env
name, or `cluster::env`, or `cluster::pool`. This page is the machinery behind
those three, and the two planes it has to work on.
## Two planes
| Plane | What crosses | Decided by |
|---|---|---|
| **Control** — `Sandbox.create()` | one HTTP request, forwarded verbatim | the Env router on the cluster that received it |
| **Data** — exec, files, PTY, ports | every later connection to the sandbox | ExtProc and Envoy, using the cluster prefix inside the sandbox id |
Getting the first one right is the visible half. The second is the one that
makes a forwarded sandbox usable: a caller on cluster A talks to a sandbox the
gateway placed on cluster B for hours, without knowing it.
## Control: who decides where a create lands
A create arrives as a reference with three possible shapes. The first is
obeyed, the other two are resolved:
| The caller wrote | What happens |
|---|---|
| `SOME_CLUSTER::env` or `SOME_CLUSTER::pool` | forwarded to that cluster's E2B API, verbatim — no second-guessing |
| `env` (bare) | resolved locally, and forwarded only if the local answer is "not here" |
| `env//image` | as above; the image override travels with the forward |
The bare name is where the design lives. Resolution runs in one order, and each
step exists because the alternative is worse:
1. **local has an idle Pod** → serve it locally. A cross-cluster hop costs a
round trip; there is no reason to pay it when the answer is already here.
2. **a foreign member has an idle Pod** → forward, rewritten as
`cluster::pool` so the receiver claims that exact pool instead of routing
again.
3. **nobody has idle capacity, but the local cluster can still scale** → stay
local, and let it scale.
4. **only a foreign member can scale** → forward there.
5. **nobody can serve it** → park it locally. A request that waits is a request
the operator can see; one that bounced between clusters is not.
The distinction between 3 and 4 is the part that is easy to get wrong:
forwarding a request to a cluster that also has to scale buys a hop and a
second queue, without buying capacity.
The state this reads — local idle, foreign idle, whether the local cluster can
grow — is maintained per Env across clusters, keyed by `(namespace, env name)`.
That keying is why the same env in two clusters must live in the same namespace
name: change team or namespace and the two halves stop being one env, and a bare
name silently stops spreading. The explicit `cluster::env` form keeps working,
because it never consults the federation at all.
## Data: making a forwarded sandbox reachable
Once a sandbox exists on cluster B, every later call has to be able to reach
its ports from wherever the caller is. The sandbox id carries the cluster, so
nothing has to be looked up — but a plain `ORIGINAL_DST` load balancer on
cluster A would happily try to resolve a Pod that does not exist there.
ExtProc sits in that path and rewrites the request before Envoy routes it:
- `:authority` becomes the target gateway's host, for TLS SNI and HTTP Host
matching;
- `:path` gets the data-plane prefix;
- `x-agentbox-cross-cluster: true` makes Envoy match the header-routed
`ORIGINAL_DST` cluster instead of the local one.
The scheme cannot be assumed: Envoy applies a transport socket per cluster, not
per request, so a header carries `https` or `http` and selects between the TLS
and plain-text gateway clusters.
**A header name that is a loop guard**
> The resolved upstream host travels in `x-agentbox-upstream-host`, not in
> `x-envoy-original-dst-host`. The well-known name is read by *any* Envoy that
> receives it: if it leaked through an nginx-ingress to the remote cluster's own
> `original_dst_cluster`, that Envoy would dial the public IP it was handed —
> itself — and fail with `TLS_WRONG_VERSION_NUMBER` as a 503. A private name
> means the same value is meaningless on the far side.
## What is not on this path
- **Not failover.** A cluster with no Pod to give you is not rescued by another
one holding a spare — capacity routing is decided at create time, and a
cluster that cannot serve returns a failure rather than a long queue.
- **Not image distribution.** A sandbox runs where its image can be pulled;
the platform rewrites a private registry's *host* to this region's equivalent,
which is why the same image path has to exist in each.
- **Not the console's job.** The console lists the clusters its config names;
a console that does not know a cluster cannot offer it, and `abx clusters`
prints what the address behind it reaches.
## Where the code is
| Path | What it holds |
|---|---|
| [`pkg/e2bcompat/handlers/server.go`](https://github.com/scitix/Agent-Sandbox/blob/develop/pkg/e2bcompat/handlers/server.go) | the create path: parse the reference, forward when it names another cluster, resolve a bare name through the router |
| [`pkg/apiserver/service/envscheduler/`](https://github.com/scitix/Agent-Sandbox/tree/develop/pkg/apiserver/service/envscheduler) | the five-step order above, and the federation state it reads |
| [`pkg/apiserver/service/cross_cluster_forwarder.go`](https://github.com/scitix/Agent-Sandbox/blob/develop/pkg/apiserver/service/cross_cluster_forwarder.go) | the protocol-agnostic forwarder: method, path, query, headers and body verbatim, base URL swapped by kind |
| [`pkg/envoy/extproc/cross_cluster.go`](https://github.com/scitix/Agent-Sandbox/blob/develop/pkg/envoy/extproc/cross_cluster.go) | the data-plane rewrite: authority, path prefix, and the two private headers |
## See also
- [Cross-cluster](/docs/concepts/cross-cluster) — addressing, federation, images
- [Pools](/docs/concepts/pools) — what "idle" means, since it decides the route
- [The egress filter](/docs/designs/egress) — the other sidecar in the same Pod
---
# The egress filter (/docs/designs/egress)
[Egress and secrets](/docs/concepts/egress-and-secrets) is what a user asks
for: one switch on the Env, rules on the create call. This page is what
implements it, and where the sharp edges are.
The enforcement model follows
[E2B's tcpfirewall](https://github.com/e2b-dev/infra), adapted for a Kubernetes
Pod and hardened for evaluation: **the default action is deny**, not allow.
## Two containers, and why not zero
Turning the gateway on adds two things to every sandbox Pod:
| Container | What it does | Why separate |
|---|---|---|
| `egress-init` (init) | installs one `nat` OUTPUT chain in the Pod's network namespace | needs `CAP_NET_ADMIN`, and a network namespace is shared by every container in the Pod — one install covers the sandbox, whatever it runs |
| `egress-proxy` (native sidecar, uid 1337) | evaluates policy and injects headers | the redirect exempts uid 1337, which is what keeps the proxy's own upstream connections from being redirected back into itself |
Everything is expressed in the Pod's network namespace, so enforcement is
independent of two things that would otherwise matter: the sandbox's image
(the filter survives an in-place image swap, because it is not in that image)
and the cluster's CNI (the same rules work on Calico, ENI or anything else that
gives a Pod a netns).
The proxy listens on three ports — 15001 for `:80`, 15002 for `:443`, 15003 for
everything else — plus a health port, 15004, which is deliberately *not* one of
the three: a probe aimed at a data-plane port arrives indistinguishable from a
redirected sandbox connection, so it would be policy-evaluated and logged as a
denial on every interval.
**Only 80 and 443 are understood**
> `HTTP` rules match on the `Host` header, `TLS` rules on the SNI. Any other port
> is CIDR-only: a rule written for `:3000` never fires, and nothing says so — the
> traffic simply passes the filter unexamined. That is a property of layer-7
> filtering, not a bug to be waited out.
## The policy is a file, and absence means deny
The control plane writes a policy document into an `emptyDir` that only the
sidecar mounts; the proxy watches it with `fsnotify` and re-reads on change.
The contract is fail-closed at every step:
- **an absent, empty or unparseable file denies everything** except DNS, which
stays resolvable so lookups fail fast instead of hanging;
- `enforce: false` means the same as deny-all. The control plane flips it to
true only after it has resolved a concrete ruleset for a claimed sandbox —
so the window between a Pod starting and its policy arriving is closed by
default, not by timing.
## The SSRF baseline is two tiers
"Internal" covers two things that deserve different answers, and collapsing
them forces a bad trade — a sandbox that needs one internal service would have
to be handed the cloud metadata endpoint with it.
| Tier | What it is | Can a policy open it? |
|---|---|---|
| Always denied | instance-metadata and link-local (`169.254.0.0/16`, `100.100.100.200`, `fd00:ec2::254`, `fe80::/10`) and loopback | **no.** An unauthenticated `GET` to these hands out cloud credentials, and no sandbox workload has a legitimate reason to reach them |
| Default denied | RFC1918, CGNAT and ULA | yes — naming a host or CIDR in the allow list lifts the baseline for exactly that destination |
A wildcard does not lift the second tier: `allowOut: ["*"]` means the public
internet. `agentbox.scitix.ai/allow-private-networks` is the explicit way to say
"everything, including the cluster's own network".
## Credentials, and why the value never enters the sandbox
This is the part the design is shaped around: an agent with a shell must not be
able to read the credential it is allowed to use.
```
vault (write-only)
→ operator reads it during one reconcile
→ sidecar tmpfs (0600, never mounted into the sandbox)
→ header rewritten on the matching request
```
Three consequences fall out of that path:
- The operator resolves the CRD's `${e2b.secrets.NAME}` templates **before**
pushing, so credential *names* never reach the sidecar and it needs no
template engine. The wire between them carries values for one push and
nothing else.
- The sidecar file lives on the sidecar's own tmpfs and is **removed when the
sandbox is released**. It is never an annotation, an environment variable or
a log line — the CRD holds the reference, not the value.
- The sandbox sees a **decoy**: a placeholder in the environment it can read,
which the proxy replaces on the way out. Code inside runs unmodified; an
agent that prints its environment prints the decoy.
Interception needs TLS to stop being end-to-end, so the proxy mints a
**per-sandbox CA** and a leaf certificate per intercepted host (cached in
memory, 24-hour TTL, never written to a volume the sandbox can mount). The CA
certificate is installed into the sandbox's trust store through `/init` — which
is why a custom image needs `/etc/ssl/certs/ca-certificates.crt` to exist, and
why the container must not run as uid 1337: that uid is exempt from the
redirect, so a sandbox process running as it would bypass the filter entirely.
**What injection does not do**
> It rewrites **headers** on matching requests. It does not inspect bodies or
> query strings, it does not follow a redirect to another host, and a client with
> certificate pinning will fail rather than be helped. The sandbox can also use
> the injected credential as many times as it likes for the hosts the rule names —
> readable was the problem, not usable. Keep the rule's path and host as narrow as
> the workload allows.
## Where the code is
| Path | What it holds |
|---|---|
| [`pkg/egressproxy/`](https://github.com/scitix/Agent-Sandbox/tree/develop/pkg/egressproxy) | the proxy: policy, matching, the CA and MITM, the injection rules, the iptables redirect |
| [`pkg/framework/plugins/egress/`](https://github.com/scitix/Agent-Sandbox/tree/develop/pkg/framework/plugins/egress) | the `PreCreatePod` plugin that puts the two containers into the Pod |
| [`pkg/e2bcompat/handlers/egress.go`](https://github.com/scitix/Agent-Sandbox/blob/develop/pkg/e2bcompat/handlers/egress.go) | E2B's `network.allowOut` / `denyOut` / `rules` translated into that policy, including the refusals (literals, wildcards, identity tokens) |
| [`installer/dockerfile/Dockerfile.idleimage`](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/dockerfile/Dockerfile.idleimage) | the idle image, which carries `/egress-proxy` — the same image a Pod runs while idle is the one the sidecar executes |
## See also
- [Egress and secrets](/docs/concepts/egress-and-secrets) — the user-facing half
- [envd, and the patches we carry](/docs/designs/envd) — the filter's dependency on that runtime
- [Pools](/docs/concepts/pools) — why flipping the switch rolls them
---
# envd, and the patches we carry (/docs/designs/envd)
Every sandbox runs **envd**: the agent the E2B SDK talks to. Process exec,
filesystem, PTY — that is envd's surface, and it is
[upstream](https://github.com/e2b-dev/infra), not a fork of ours.
We do carry three patches, and this page is what they change and why, because
each one is a difference you can run into: a command that never starts, a
result that is wrong rather than an error, a log full of connection attempts to
an address that cannot answer.
| Patch | Fixes | Symptom without it |
|---|---|---|
| [0001](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/runtimes/envd/patches/0001-skip-oom-nice-wrapper-when-not-firecracker.patch) | skips the OOM/`nice` exec wrapper outside a microVM | every command exits with a permission error, or never runs at all |
| [0002](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/runtimes/envd/patches/0002-await-init-gate.patch) | refuses RPCs until the sandbox is armed | a command runs in a sandbox with no environment and no credentials — and succeeds |
| [0003](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/runtimes/envd/patches/0003-skip-mmds-poll-when-not-firecracker.patch) | stops the MMDS poll outside a microVM | 1200 requests per sandbox to an address that never answers |
The patches and the build script live in
[`installer/runtimes/envd/`](https://github.com/scitix/Agent-Sandbox/tree/develop/installer/runtimes/envd);
each patch file carries the longer version of what is below.
## 1. Not a Firecracker microVM, so do not pretend to be one
Before exec-ing a command, upstream envd wraps it as
`/bin/sh -c "echo 100 > /proc/$$/oom_score_adj && exec /usr/bin/nice -n N -- CMD"`.
That priming makes sense **inside** a Firecracker microVM: children would
otherwise inherit envd's protected `oom_score_adj=-1000` and `nice -20`.
Under Kubernetes it is wrong twice over:
- the kubelet manages `oom_score_adj` itself and pins a floor that forbids
lowering it, so the write fails with `EACCES` — and because the wrapper uses
`&&`, the user's command never runs;
- routing every exec through `/bin/sh` makes execution depend on the image
having a working one, which is how busybox multi-call images broke.
envd already takes a `-isnotfc` flag, which our entrypoint passes. The patch
threads it into `execcontext.Defaults.IsNotFC` and, in that mode, execs the
command directly. cgroups and uid setup are untouched: those are the real
resource controls here.
This replaced an earlier hack that rewrote `/bin/sh` inside every image, and it
is a candidate for upstreaming — gating a wrapper on "not a microVM" is what
upstream would do.
## 2. A sandbox is not ready until it is armed
A sandbox's environment variables, its injected trust-store certificate and its
egress credentials all arrive in a `POST /init` that happens **after** envd is
listening. A command accepted in that window runs in a sandbox that is not yet
the one the caller asked for.
It does not fail. It returns a wrong answer, which is the harder failure to
notice.
Two other layers already close this — the create call waits for the sandbox to
be armed, and the data plane refuses to route to one that is not — and the
patch is the third: the `-await-init` flag makes envd itself refuse process and
filesystem RPCs with `failed_precondition` until the first `/init` lands.
`/health` is deliberately not gated: the readiness probe is what drives the
phase transition that triggers that `/init`, so gating it would deadlock the
two against each other. It ships default-off for the same reason — the control
plane has to be sending an unconditional `/init` before a template turns it on.
## 3. Stop polling a metadata service that is not there
`169.254.169.254` is Firecracker's MMDS — the microVM's metadata service.
Outside a microVM nothing answers.
The boot-time poll is already gated on `-isnotfc`, but a second one started on
**every** `/init`, unconditionally. AgentBox sends an `/init` to every sandbox,
so every sandbox polled: one request every 50 ms for 60 seconds, 1200 of them,
each a real connection attempt that the pod's egress filter evaluated and
logged. It was found in a sidecar log that was almost entirely
```
level=INFO msg="egress denied" host=169.254.169.254 port=80 match=ssrf
```
at a steady 50 ms cadence for the first minute of every sandbox's life.
The patch gates that goroutine on the same flag. The two values it fetched
(`E2B_SANDBOX_ID`, `E2B_TEMPLATE_ID`) are known to the orchestrator, which puts
them in the `/init` body it was already sending.
## How these images are built
The build is deliberately boring, and the interesting part is what it refuses
to do:
- the upstream commit is **pinned** (`INFRA_REF` in
[`build-envd.sh`](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/runtimes/envd/build-envd.sh)),
so a patch applies deterministically and the version cannot drift under you;
- patches are applied in filename order onto one tree, so each is a diff against
the previous ones' result;
- **a patch that does not apply aborts the build.** Shipping unpatched envd is
not a fallback, it is the bug coming back;
- the image tag carries a rebuild suffix — `0.9.0-2` is the third image of
envd 0.9.0. That is not decoration: templates pull with
`imagePullPolicy: IfNotPresent`, so rebuilding at an unchanged tag would leave
every node that already has the image running the old code, with no error
anywhere.
The version the platform *reports* to the SDK is a separate constant,
`DefaultEnvdVersion` — and the [samples](/docs/examples/templates) are pinned to
the image that carries it, which `hack/sync-sample-images.py` keeps true.
## What this means when you run into it
- **A command fails with `EACCES` on `oom_score_adj`**, or an image with a
minimal `/bin/sh` misbehaves: the sandbox is running an envd without
patch 1. Check what the template pins.
- **A command succeeds with an empty environment**: the sandbox was used before
it was armed. The gate above is off by default; the create path should not
hand you a sandbox that early, and if it does, that is a bug worth reporting.
- **Egress logs full of `169.254.169.254`**: an envd without patch 3. It is
noise, not a leak — the SSRF baseline denies it every time.
---
# Design notes (/docs/designs)
These pages are the *why* behind behaviour described elsewhere on this site.
Each one starts from a problem, says what it cost to solve, and links the code
that carries it — including the parts that are not settled.
| Page | The question it answers |
|---|---|
| [envd, and the patches we carry](/docs/designs/envd) | which parts of the runtime inside every sandbox are upstream, which are ours, and what each change fixes |
| [The egress filter](/docs/designs/egress) | how a sandbox's traffic is filtered, and how a credential is injected without ever entering the sandbox |
| [Cross-cluster routing](/docs/designs/cross-cluster) | how a create lands on another cluster, and how every later connection follows it |
They are written for someone deciding whether to trust the platform with
something — a benchmark run, an untrusted model's code, a multi-cluster rollout
— rather than for someone about to change the code. Where a design has an open
question, it says so.
---
# What is here (/docs/examples)
Every page in this section is a file from
[`config/samples`](https://github.com/scitix/Agent-Sandbox/tree/develop/config/samples),
rendered as it ships: the manifest's own comments become the introduction, and
the page links to the same file on `develop`.
| Group | What it holds |
|---|---|
| [Templates](/docs/examples/templates) | three `SandboxTemplate`s to start from: [E2B Basic](/docs/examples/templates/e2b), [E2B Docker](/docs/examples/templates/e2b-docker), [E2B Kata](/docs/examples/templates/e2b-kata) |
| [Environments](/docs/examples/envs) | a `SandboxEnv` with a warm pool, and the one line that points it at another template |
Apply a template first, then an environment — an environment with no template
to bind has nothing to render. [Installation](/docs/installation#first-sandbox)
walks the two commands in order.
---
# abx-common (/docs/skills/abx-common)
**Generated from**
> [`plugin/skills/abx-common/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-common/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## abx: the shared half
`abx` addresses the **platform**: the environments, warm pools and autoscaling
groups that sandboxes are claimed from. The sandboxes themselves — creating
one, running commands in it, reading its filesystem — are the **E2B SDK's**
job, and that line is the product, not an omission. If the task is "run
something in a sandbox", reach for E2B. If it is "there is nowhere to run it
yet", or "it will not scale", you are in the right place.
## Do this first
```bash
abx agent-context # the whole tool as one JSON document
```
Every resource, every filter, every column, every write, and the grammar that
assembles them. It is generated from the same registry the CLI dispatches on,
so it cannot describe a command that does not exist. Reading it once costs less
than three `--help` calls and answers more.
## The grammar
```
abx list
abx get
abx sub-list
abx sub-get
abx create -f FILE make one; it must not exist yet
abx update - -f FILE change one; it must exist
abx delete
-
abx scale
- --replicas N pools only
```
**A verb is read in the first position and nowhere else.** Everything after it
is an address, which is why `abx envs apply` is the env *called* `apply` and
why a name can never collide with a write. Reads have no verb at all.
The address matches the console URL segment for segment, so
`/clusters/c/envs/e/pools/p` and `abx envs e pools p --cluster c` are the same
thing said twice.
## Driving an Env over E2B, and the one read that comes first
Sandboxes belong to the E2B SDK, and what that SDK needs and you cannot guess is
**where the endpoints are**: the E2B-compatible API URL, the data-plane domain,
and whether the data plane is http or https. Every Env carries its template's
documentation, already rendered for the cluster that Env lives on:
```bash
abx envs docs
```
**Read it before creating a sandbox against an Env you have not used.** It is
the only place that knows the endpoints of the cluster you are actually talking
to, and the failure it prevents is the quiet one: the SDK pointed at the wrong
host, or at e2b.dev, never connects and never says why.
When the document offers more than one way in — the same cluster, another
cluster, the public one — **take the public one unless you are told otherwise**
or you can see you are already inside that cluster's network. It is the only
path that does not depend on where the caller happens to be running.
The key in that document stays written as `${AGBX_API_KEY}`. That is not a
missing value: it is your own credential, the one `abx` is already
authenticating with, and the SDK reads the same key from `E2B_API_KEY`. No
rendered document ever carries a live token — for a person's key or an agent's —
which is what makes these pages safe to read, print and relay. The console fills
the placeholder in for a person reading it there; nothing else should.
## Finding a shape — never from memory, never from a document
Anything a command takes is answerable *by that command*, and asking is the
only way to get the answer for the build you are actually talking to. Prose —
this file included — is a copy, and a copy of a schema goes stale the first
time the schema moves:
```bash
abx create envs --help # the file create takes: example + every field
abx create envs --schema # …the same body as JSON, refs resolved, no key needed
abx update envs --help # the same file, plus how to obtain one
abx update envs --schema # …and the same body as JSON
abx envs pools --help # …and for a member pool, addressed the same way
abx --help # columns, filters, sub-resources, writes
abx envs --editable # the current values, in exactly that shape
abx agent-context # all of it as one JSON document
```
`abx create|update --help` is generated from the API schema at build
time, so its field list cannot drift from the server you are writing to. When
you are about to write a file, that page *is* the specification; when you are
about to change one, `--editable` is the file to start from.
The image carries the contracts themselves, read-only, for the shapes `abx`
does not own — a sandbox's own create call, for instance:
```
/opt/agentbox/source/pkg/openapi/native/openapi.yaml the platform API
/opt/agentbox/source/pkg/openapi/e2b/openapi.yaml the E2B surface a sandbox speaks
/opt/agentbox/source/sdk/ this platform's own SDKs
/opt/agentbox/source/cli/src/ this CLI's source
```
That is where to look for anything the CLI only *uses*: the exact fields a
sandbox create call accepts are in the E2B spec there, not in a skill. The
Python SDK is a third-party package installed in the sandbox, so ask it
directly — its signature is the version-matched answer:
```bash
python -c "import inspect, e2b; print(inspect.signature(e2b.Sandbox.create))"
```
## Authentication
Two settings are required, resolved as flag → environment →
`~/.config/abx/config.json`:
| Setting | Flag | Environment |
|---|---|---|
| Console address | `--endpoint` | `AGENTBOX_ENDPOINT` |
| Credential | `--api-key` | `AGENTBOX_API_KEY` |
That is the whole setup. The endpoint is the console's own address — the one a
person types in a browser — and **one address reaches every cluster** the
platform has.
Which header the key travels in follows from that, so there is no scheme to
set: the console takes `Authorization: Bearer`, a cluster API takes
`AGENTBOX-API-KEY`. `--auth-scheme` is an override for a deployment answering to
neither, and is otherwise unnecessary. Console links in output are derived from
the endpoint too.
**The exception**: a sandbox with no network route to the console is configured
with `--cluster-api` (`AGENTBOX_CLUSTER_API`) pointing at one cluster's own API
instead. Everything below about choosing a cluster then does not apply — that
address answers for one cluster and refuses any other. You will not choose this;
whoever deployed the platform did, and it shows up already set in the
environment.
Under the Claude Code plugin the key is in the OS keychain and a hook writes
the config file. **Do not read that file, echo the key, or pass it on a command
line** — it is deliberately kept out of the conversation, and putting it back
in defeats the arrangement.
`abx whoami` answers who the key acts as, and — the part worth checking before
a write — whether it is an `agent` key.
`abx` speaks as one tenant and has no flag to act as another. That is
deliberate: acting as somebody means holding their key, not asking yours to
pretend. Administrative work across tenants belongs in the console.
## Clusters
Management calls are per cluster, and the cluster is chosen **per command** with
`--cluster` (`AGENTBOX_CLUSTER`). It is not part of the context: a context names
a platform, and that platform may have several clusters, so a default would
quietly answer for whichever one happened to be set.
Through a console — the normal case — `--cluster` reaches any cluster the
platform has, and when there is exactly one it is filled in for you.
In the `--cluster-api` exception above, the address answers for its own cluster
and **refuses** a `--cluster` naming a different one. That refusal is
deliberate: returning the local cluster's rows under another cluster's name is
data that is confidently mislabelled, and a reader cannot tell. Treat the
refusal as the truth about that sandbox, not as something to work around — there
is no route to the other cluster from there.
```bash
abx clusters # what this deployment can reach
abx envs --cluster # choose one for this command
```
## Output
Three registers, and the middle one is the default:
- **table** — a header line, a count, a `view:` link to the same page in the
console where the console has a page for it, and `hint:` lines naming what to
do next. Truncated at 200 rows with the filters that would narrow it.
- `--json` — the raw API shape. No header, no hints. Use it when piping.
- `--csv` — flat, for a spreadsheet or `cut`.
`--wide` adds the columns held back by default. `--filter key=value` narrows a
list, where **key is a column heading** — the heading, the filter key and the
CSV column are one name on purpose.
You have a shell. Use it: `abx envs pools --json | jq`, loops, aggregation. That is
the point of a CLI over a tool-per-operation surface, and the sandbox you are
probably running in has no real credentials in it anyway — the egress sidecar
substitutes them on the way out.
## Writes, and approval
**Do not write the file from memory.** Ask the CLI for it:
```bash
abx create envs --help # create's file: address, example, fields
abx create envs --schema # the same body as JSON, every ref resolved
abx update envs --help # the same file, plus how to get one
abx create envs pools --help # …and the same for a member pool
```
That page is generated from the API schema, so it lists every field, which are
required, and which are **fixed at create** — it cannot be out of date in the
way a copy in a document can. `abx agent-context` carries the same data as
JSON, if you want to read it once rather than per command; `--schema` is the
same body for one address, and it works with no deployment or key, because the
shape comes from the build, not the server.
The two verbs are separate words now: `create` makes one (it must not exist
yet), `update` changes one (it must exist). They take the **same file** — the
file a create wrote is the file an update takes.
`update` is a **PUT of the desired state**: a field the file leaves out is a
field you are asking to **remove**. That is why the safe edit is to start from
what is there rather than from a blank page:
```bash
abx envs demo pools --json # find the pool's name
abx envs demo pools demo-1c2gi --editable > pool.json
# edit pool.json
abx update envs demo pools demo-1c2gi -f pool.json
```
The file is NOT the object `--json` prints: those keys are nested under
`spec.*`/`status.*`, the write reads top-level ones, and the ones it does not
find are the ones it clears. Piping `--json` straight into a write is the
fastest way to lose state.
`abx scale envs pools --replicas N` is the exception: one field, no
clearing, and it re-sends the current bounds unchanged. Use it when size is all
you are changing — and not at all when the pool's scaling group has autoscaling
enabled, where the autoscaler owns the number and the API says so.
**An `agent` key's writes are held for a person to release.** The command comes
back saying so, with a link. That is not an error to work around: re-run the
command after the person has acted.
**Three writes are refused outright, with no approval to wait for**: issuing an
API key, promoting one, and deciding an approval. An agent that could mint a
credential could mint one without the agent restriction and leave the gate
entirely, and no approval dialog conveys that. The answer is a console link for
the person to act on themselves. Do not queue, retry, or look for a flag.
For the same reason, an agent key gets key **metadata** without key
**material**: `abx api-keys` lists what exists and which keys are gated, and
the token field is absent. A rendered document — an Env's docs, a Template's
docs, the setup guide in the console — comes back with `${AGBX_API_KEY}` intact
rather than a live token, for every caller and not only for you. Relay the
document and say where the person's own key goes; there is nothing missing to
hunt for.
The refusal and the wait are different answers and the error codes say which:
`APPROVAL_REQUIRED` means a person is about to decide, so re-run it shortly.
`FORBIDDEN_FOR_AGENT` means no approval exists or ever will — stop, and hand
over the link.
## Errors carry the recovery
A rejection names the valid set and the next command. Read it before
reformulating — it usually contains the answer, and a guess costs a round trip.
## Where to go next
| You want to | Skill |
|---|---|
| Run RL rollouts at scale | `abx-reinforcement-learning` |
| Put sandboxes behind your own agent | `abx-managed-agent` |
| Size pools, fix autoscaling, read quota | `abx-resource-capacity` |
| Run a benchmark suite | `abx-harbor-framework` |
| Run Docker inside a sandbox | `abx-sandbox-docker` |
| Stop a sandbox reaching the internet | `abx-sandbox-network` |
| Give a sandbox a credential it cannot read | `abx-sandbox-secrets` |
| Work out why something is broken or slow | `abx-observe` |
## Read more
[The object model](https://scitix.github.io/Agent-Sandbox/docs/concepts/index.md)
---
# abx-harbor-framework (/docs/skills/abx-harbor-framework)
**Generated from**
> [`plugin/skills/abx-harbor-framework/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-harbor-framework/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## Running evaluations on AgentBox
[Harbor](https://github.com/harbor-framework/harbor) already knows how to drive
a benchmark. `agent-sandbox-harbor` is an environment plugin that makes it claim
sandboxes from a warm pool instead of building one per task — which is where the
time goes in a normal Harbor run.
**No Harbor fork, and no Template Build step.** The plugin attaches through
Harbor's official `--environment-import-path`, and because AgentBox pools swap
the image in place on an already-running Pod, a task starts with one API call
rather than an image build.
## The shape of a run
```bash
pip install 'harbor[e2b]' agent-sandbox-harbor
cat > agentbox.env <<'EOF'
E2B_API_KEY=agbx_…
E2B_API_URL=https:///agent-sandbox/api/e2b
E2B_DOMAIN=/agent-sandbox/api/data
AGBX_POOL_NAME=terminal-bench-pool
AGBX_CLUSTER_ID=cluster-a
AGBX_IMAGE_PREFIX=registry.internal/agent-sandbox
EOF
harbor run \
-d terminal-bench@2.0 -a oracle -n 16 -y \
--environment-import-path agent_sandbox_harbor:AgentSandboxEnvironment \
--env-file agentbox.env
```
The two endpoint lines are the env's, not a sketch: `abx envs docs` prints
them for the cluster you are actually reaching, including whether the data plane
is http or https, and `E2B_API_KEY` is the key you authenticate `abx` with (what
that document writes as `${AGBX_API_KEY}`). Where it offers several ways in, use
the public one unless you are already inside the cluster.
`-n 16` is concurrency, and it is the number that has to exist as **idle Pods**
before the run starts moving. Size the pool for it first:
```bash
abx envs pools # idleReplicas is the real answer
abx scale envs pools --replicas 16
```
See `abx-resource-capacity` if it will not grow.
## Images: the part that actually bites
Every task needs a **pre-built** image. This environment does not build from a
Dockerfile and does not mutate a running sandbox, so an image is chosen in
exactly this order:
1. **`AGBX_IMAGE_MAP`** — a ` ` file. Used verbatim, no
rewriting. This is how you run a dataset whose `task.toml` has no
`docker_image` at all, which is the case for **SWE-bench**, where the task
*is* a Dockerfile.
2. **`task.toml`'s `docker_image`** — Terminal-Bench's case. Rewritten by
`AGBX_IMAGE_PREFIX` (with `docker.io/` stripped first) and `AGBX_IMAGE_TAG`.
3. Neither → **the task is rejected**, deliberately and loudly.
So a SWE-bench run is really two jobs: mirror or build the images once and write
the map file; then run Harbor against it. Budget for the first.
## Settings that matter under load
| Variable | Why you would touch it |
|---|---|
| `AGBX_STARTUP_TIMEOUT` | default 300s; raise for heavy images |
| `AGBX_READY_TIMEOUT` | default 600s; a cold SWE-bench image can exceed it |
| `AGBX_IMAGE_PREFIX` | point every `docker.io/…` at an internal mirror |
| `AGBX_HTTPS` | `false` when the data plane is plain HTTP — a mismatch shows up as "never connects", never as a scheme error |
One version note worth checking before blaming the platform: e2b SDK ≥ 2.24
rejects non-`e2b_` keys client-side. `agent-sandbox-e2b >= 0.0.4` neutralises
that so `agbx_` keys work, and `harbor >= 0.13` pulls a new enough e2b to need it.
## When tasks fail rather than the run
```bash
abx sandboxes --filter status=Failed
abx sandboxes logs
abx envs events
```
A whole dataset failing the same way is almost always the image map or the
registry; individual tasks failing is usually the task.
## Related
- `abx-resource-capacity` — making the concurrency you asked for exist
- `abx-observe` — reading what failed
- Reference: `sdk/python/harbor/README.md` and `INTEGRATION.md` in the
agent-sandbox repository
## Read more
[Sandbox pools](https://scitix.github.io/Agent-Sandbox/docs/concepts/pools.md)
---
# abx-managed-agent (/docs/skills/abx-managed-agent)
**Generated from**
> [`plugin/skills/abx-managed-agent/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-managed-agent/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## Sandboxes as your agent's hands
Your agent keeps running where it runs. Its `bash`, `read`, `write`, `edit`,
`grep`, `glob` and `apply_patch` stop touching that machine and start acting on
a sandbox bound to the conversation — so the work survives a process restart,
can be browsed from a file UI, and is reclaimed on a timer instead of
accumulating in someone's home directory.
```
your agent process AgentBox
┌──────────────────────────┐ ┌────────────────────────┐
│ harness │ │ sandbox for this │
│ ↓ tool call │ │ session │
│ hands binding │ │ │
│ ↓ HTTP │ │ │
│ hands daemon ───────────┼── E2B API ─┼─→ bash / files │
└──────────────────────────┘ └────────────────────────┘
```
## Read this before describing it to anyone
**Confinement is not isolation.** The agent process still runs on your machine,
with your files, your environment and your credentials. What moves into the
sandbox is where the agent's *tools* act. That is genuinely worth having, but an
agent that can install a package or load a plugin can reach the host again. For
real isolation, run the harness itself in a container — orthogonal, and they
compose.
Saying otherwise to a user is the one mistake here that matters.
## Three bindings, one behaviour
| Harness | Binding |
|---|---|
| Claude Agent SDK | `sdk/hands/typescript/src/harness/claude-code/` |
| OpenCode | `sdk/hands/typescript/src/harness/opencode/` |
| anything else | generic MCP: `sdk/hands/typescript/src/harness/mcp/` |
`core/` decides what the tools do; a binding only says "replace these built-ins
with these tools" in one harness's vocabulary. **A binding that reimplements
behaviour from `core` is a bug** — two copies drift, and the drift is invisible
from the signatures.
The session-binding daemon and workspace file API are the Python half
(`sdk/hands/python/agentbox_hands/`).
## Ask two things before configuring
1. **Which harness** — Claude Agent SDK, OpenCode, or something else that needs
the generic MCP binding. The binding decides the whole integration and there
is no useful generic answer.
2. **How many concurrent conversations**, not how many users. Ten people with
one session each is ten; the pool is sized against that number.
Then say the confinement sentence below **before** they build anything on it.
## What has to exist on the platform first
1. **An env** whose template is what your agent's tools need — an E2B-compatible
image with a shell and the language runtimes the work requires.
2. **A pool with idle Pods**, sized to concurrent *conversations*, not total
users. Ten people chatting with one session each is ten.
3. **A key for the right identity.** Which identity a sandbox is created as
decides whose namespace and whose quota it lands in.
```bash
abx envs
abx envs pools
abx scale envs pools --replicas 10
abx whoami # which identity, and whether this key's writes need approval
abx envs docs # where the SDK points: E2B API URL, data domain, scheme
```
That last one is not optional reading before wiring a binding up. The agent's
tools reach the sandbox through the E2B API, and which URL that is depends on
the cluster the env lives on — `abx envs docs` is where the platform has
already worked it out. Prefer the public entry it prints unless the agent
process itself runs inside the cluster.
## Identity: the decision to make deliberately
Two coherent models, and mixing them is where the confusion comes from:
- **Per-person** — each session is created with that person's own credential.
Sandboxes land in their namespace and count against their quota, and their
sandbox list shows their own work.
- **Service-owned** — every session is created as the env's owner, and *who the
conversation is for* is recorded in sandbox `metadata`. One quota, one
namespace, and the list needs the metadata column to tell sessions apart.
Pick one. Under the second, all usage bills to one tenant — which is a choice,
not a bug, but only if it was chosen.
Whichever you pick, **a tenant key, never an admin key.** A service holding an
admin credential means every conversation runs as admin, and nothing in the
behaviour reveals it until something goes wrong.
## Credentials inside the sandbox
The sandbox is a remote execution environment that holds **no real
credentials** — decoy values sit in the environment and the egress sidecar
substitutes the real ones on the way out, per host and per header. So an agent
can run arbitrary shell there without that shell being a way to exfiltrate a
key. When you need a third-party token available to the agent's code, that is
what the vault and injection rules are for, not an environment variable.
## Related
- `abx-resource-capacity` — sizing the pool for concurrent sessions
- `abx-observe` — why a session's sandbox failed
- `abx-common` — endpoint, key, cluster, approval
- Reference: `sdk/hands/README.md` in the agent-sandbox repository
## Read more
[Sandbox environments](https://scitix.github.io/Agent-Sandbox/docs/concepts/envs.md)
---
# abx-observe (/docs/skills/abx-observe)
**Generated from**
> [`plugin/skills/abx-observe/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-observe/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## Observing: what broke, and what it is costing
## Start from the object, not the symptom
```bash
abx envs demo # is the env itself ready
abx envs demo pools # desired vs idle vs running, per shape
abx envs demo events # what Kubernetes said, most recent first
abx sandboxes --filter status=Failed
abx sandboxes logs
```
`events` is the highest-yield of these and the most often skipped. A pool that
will not grow almost always has a Warning event saying why, in Kubernetes'
words rather than the platform's.
## Metrics come from E2B, not from abx
The platform serves the **E2B-compatible** metrics endpoints, so the way to
read resource usage is the E2B SDK you already have in the sandbox — not a
separate `abx` command, and not a Prometheus query you have to construct:
Which API that is depends on the env's cluster: `abx envs docs` prints the
E2B API URL and data-plane domain the SDK should be pointed at.
```python
from e2b import Sandbox
sbx = Sandbox.connect(sandbox_id)
metrics = sbx.get_metrics() # cpu, memory, disk over the sandbox's life
```
Two endpoints, both E2B's own:
| | |
|---|---|
| `GET /sandboxes/{id}/metrics` | a time series for one sandbox |
| `GET /sandboxes/metrics?sandbox_ids=…` | the latest point for several |
**`/teams/{id}/metrics` is not implemented** — it answers `501`, deliberately:
team and cluster administration is not exposed through the E2B surface. For
usage across a team, read `abx statistics` or the console.
Deliberately no new convention for the two that do exist: a caller who already
speaks E2B needs nothing extra, and one who does not is better served learning
E2B than a bespoke metrics dialect.
**Units, because the two surfaces differ and it has caught people out:**
- `cpuUsedPct` from the API is a **percentage of the sandbox's own cores**
(`cores_used / cpuCount * 100`). A 16-core sandbox with 8 busy cores reads
`50.08`, not `8` and not `5008`.
- The console's charts plot **cores**, not a percentage. The same moment reads
`50%` in one place and `8` in the other, and both are right.
- `diskTotal` / `diskUsed` are `0` wherever the backend does not collect
`container_fs_*`. Zero here means "not collected", not "empty disk".
Time-series **charts** live in the console. `abx` prints a `view:` link on
every command, and for an env or pool that link lands on the page with the
charts. When asked for a trend rather than a number, hand over the link.
## The failures that look like something else
| Symptom | Usually |
|---|---|
| Sandbox created, commands fail to connect | the runtime never came up inside the Pod — check `abx sandboxes logs` before anything else |
| Pool stuck at 0 available | reservation or quota refused the Pods; see `abx-resource-capacity` — reduce the target, then grow |
| Env lists a running sandbox the sandbox list does not show | two different counts: the env's is cluster-wide, the list is filtered to your tenant |
| Everything empty but nothing errors | the credential's namespace does not exist on THIS cluster — `abx whoami` shows which namespace it resolved to |
| `--cluster` refused | that endpoint serves one cluster; use an endpoint whose path carries `{cluster}` |
## Reading a pool's status honestly
```bash
abx envs demo pools demo-1c2gi --json | jq '.status'
```
`idleReplicas` is the only number that answers "can I claim one right now".
`replicas` is intent, `updatedReplicas` is rollout progress, and a pool can
report `Ready` while having nothing claimable.
## Related
- `abx-resource-capacity` — the fix, once you know it is capacity
- `abx-common` — endpoint, key, cluster, approval
## Read more
[Sandbox pools](https://scitix.github.io/Agent-Sandbox/docs/concepts/pools.md)
---
# abx-reinforcement-learning (/docs/skills/abx-reinforcement-learning)
**Generated from**
> [`plugin/skills/abx-reinforcement-learning/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-reinforcement-learning/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## RL rollouts on AgentBox
## Before configuring anything, ask what concurrency they need
It is the only number that decides the whole shape of the answer, it is never
in the question, and guessing it wastes the conversation: a pool sized for 8
when they wanted 200 looks like it worked right up until the run stalls.
Ask for **peak concurrent sandboxes**, not total episodes. People usually know
the second and have to be walked to the first: 10,000 episodes at 64 in flight
needs 64.
Two more defaults worth stating rather than deciding silently:
- **Leave autoscaling on**, with `maxReplicas` at their peak. A fixed pool
holds capacity between runs and bills for it; a group with a ceiling drains
and comes back.
- **Set `minReplicas` to the steady-state floor** when the run ramps faster
than the scale-up cooldown, so the autoscaler only handles the tail.
## The division of labour
**`abx` provisions, the E2B SDK executes.** You use `abx` once to make sure
there is capacity, and then your trainer talks E2B for the rest of the run.
```
trainer process AgentBox
├─ abx scale envs … pools … provision the pool once, up front
└─ e2b.Sandbox(...) × N claim / run / discard, per episode
```
Sandboxes are **claimed from a warm pool**, not built. That is why a rollout
gets an environment in about a second instead of a minute, and it is also why
the pool has to be the right size before the trainer starts.
## Driving sandboxes
Standard E2B — the platform serves the E2B-compatible API, so the SDK you would
already reach for works unchanged. The call's fields belong to that SDK, not to
this CLI: read them from the SDK source the image ships
(`/opt/agentbox/source/sdk/`) or the E2B spec beside it, rather than from a
sketch in a document.
Before the first sandbox, read where to point it. `abx envs docs` prints
the env's documentation, rendered for its cluster: the E2B API URL, the
data-plane domain, and the scheme — the three things the SDK cannot guess and
will not complain about. Where more than one way in is offered, use the public
one unless the trainer runs inside the cluster. The key stays `${AGBX_API_KEY}`
in that document; the SDK reads the same one from `E2B_API_KEY`.
```python
from e2b import Sandbox
sbx = Sandbox.create("", …) # the env name; the rest is the SDK's
result = sbx.commands.run("python solve.py")
sbx.kill()
```
Two things to carry over whatever the signature says: the `template` is the
**env name** (`abx envs` lists them), and tagging each sandbox is worth the
trouble — it is what tells two of them apart later, and `abx sandboxes --wide`
shows it. Both are the SDK's fields, so take their names from the SDK and not
from this line.
> SWE ReX is **deprecated** — do not reach for it. E2B is the interface.
## Sizing the pool
```bash
abx envs # what exists
abx envs pools # idleReplicas is what you can claim now
abx instancetypes # shapes and relative cost
abx quotas # your ceiling
abx scale envs pools --replicas 64
```
Then leave autoscaling on with a ceiling at your peak, so the pool drains
between runs rather than holding capacity idle:
```bash
abx envs scaling-groups --editable > g.json
# edit it — the fields, and what each one does, are in:
# abx update envs scaling-groups --help
abx update envs scaling-groups -f g.json
```
`update` is a whole-object PUT, and `--editable` is the file it takes — the
object `--json` prints is a different document (nested under `spec.*`), and a
file built from it clears whatever it does not carry.
## The three stalls, in the order they happen
1. **Claims queue and the pool does not grow.** Demand-anchored scaling grows
toward what is actually being asked for; a wide `maxReplicas` is a ceiling,
not a request. If the pool is stuck, **reduce the replica target and grow
again** — a pool asking for more than the cluster can place stays stuck
asking. See `abx-resource-capacity`.
2. **Quota is the ceiling.** `abx quotas`. No pool setting fixes this.
3. **Sandboxes are created but commands fail.** The runtime inside the Pod did
not come up. `abx sandboxes logs` first; see `abx-observe`.
## Running a rollout loop against a pool that is also being scaled
Claims are served from idle Pods, so a trainer at steady state and an autoscaler
adjusting the pool do not fight — but a trainer that ramps faster than the
scale-up cooldown will see queueing. If the run's concurrency is known up front,
set the floor with `minReplicas` on the group and let the autoscaler only handle
the tail.
## Related
- `abx-resource-capacity` — the full autoscaler model and the recovery procedure
- `abx-harbor-framework` — running an actual benchmark rather than free-form rollouts
- `abx-observe` — logs, events and E2B-native metrics
- `abx-common` — endpoint, key, cluster, approval
## Read more
[In-place update](https://scitix.github.io/Agent-Sandbox/docs/concepts/inplace-update.md)
---
# abx-resource-capacity (/docs/skills/abx-resource-capacity)
**Generated from**
> [`plugin/skills/abx-resource-capacity/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-resource-capacity/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## Capacity: replicas, autoscaling, quota
## The objects, in the order they matter
```
SandboxEnv what a sandbox is made from (one template)
└─ SandboxPool a warm pool of one resource shape — this is what has a size
└─ group the autoscaling policy several pools share
```
A pool holds **idle** Pods. Claiming one is fast precisely because it already
exists; that is the whole design. `replicas` is how many Pods the pool keeps,
`idleReplicas` how many are claimable right now, `runningReplicas` how many are
in use.
```bash
abx envs
abx envs demo pools
abx envs demo scaling-groups
```
## Setting a size
If the pool's group has autoscaling **off**, the size is yours:
```bash
abx scale envs demo pools demo-1c2gi --replicas 40
```
If the group is **on**, `replicas` belongs to the autoscaler and setting it is
refused — the lever is the group's bounds instead:
```bash
abx envs demo scaling-groups 1c2gi --editable > g.json
# edit it — `abx update envs demo scaling-groups 1c2gi --help` lists the fields
abx update envs demo scaling-groups 1c2gi -f g.json
```
Remember `update` is the whole object: a bound you delete from the file is a
bound you are removing. That is how you take a ceiling **off**, which is worth
knowing because it is the one thing an "update just this field" API could never
express.
## How the autoscaler decides
Scale-up is **demand-anchored**, not step-anchored. It looks at what is
actually being asked for — running sandboxes plus queued requests — and grows
toward that, plus a buffer, capped by a per-mode ceiling relative to demand.
The mode sets the aggressiveness of both directions:
| mode | scale-up buffer | scale-down step |
|---|---|---|
| `Conservative` | smallest | 1 replica per window |
| `Default` | proportional | a quarter of the pool |
| `Aggressive` | largest | half the pool |
Two consequences worth carrying:
- **A wide `maxReplicas` is not an instruction.** Demand is the anchor; the
ceiling only stops growth. Setting 0–2560 does not ask for 2560.
- **Scale-down is proportional**, so a pool that over-grew drains in minutes
rather than one replica per stabilisation window.
## When a pool will not grow
Check in this order — the first two are most of the cases:
```bash
abx envs demo pools demo-1c2gi # phase, and desired vs actual
abx envs demo pools demo-1c2gi --json | jq '.status'
abx quotas # is the team's ceiling in the way
abx envs demo events # what Kubernetes said about it
```
**The recovery is counter-intuitive and worth stating plainly: reduce the
replica count, then grow again.** A pool asking for more than the cluster can
place stays stuck asking; nothing retries it into existence. Dropping the
target below what is available lets the reservation succeed, and you climb from
there. Doubling down on the number that already failed does nothing.
If quota is the limit, no amount of pool configuration helps — that is a
request to whoever owns the quota, not a setting.
Read the `ceiling` column the way `abx quotas` means it: a number is the cap,
`0` is a hard zero (nothing allocated — do not submit, ask for an allocation),
and `unlimited` means the deployment skips the quota check for that pool, so
nothing caps you and the pool's own stock decides. `unlimited` is worth
retrying; `0` is not.
## Defaults worth stating out loud
- **Autoscaling on, with a ceiling.** A fixed pool holds capacity nobody is
using between runs. A group with `maxReplicas` at the expected peak drains
and comes back, and the ceiling is what stops a runaway — not a substitute
for asking how much they need.
- **A wide ceiling is not a request.** Scale-up is anchored on demand; setting
0–2560 does not ask for 2560, and someone who read it as a target has the
wrong model of the autoscaler.
- **Never raise a target that just failed.** Reduce it, let the reservation
succeed, then climb. This is the one procedure people reliably get backwards.
## Planning for N concurrent sandboxes
You need `N` claimable Pods at peak, not `N` over the run:
1. `abx instancetypes` — what shapes exist and what each costs per unit.
2. `abx quotas` — what is committed and what the ceiling is.
3. Size the pool for peak concurrency, not total work. A rollout that runs
1000 episodes 50 at a time needs 50.
4. Leave the group enabled with a `maxReplicas` at your peak, so the pool
drains between runs instead of holding capacity nobody is using.
Once the Pods are there, the caller has to reach them: `abx envs docs`
prints the E2B endpoints of that env's cluster, which is what the SDK is
pointed at before the first create.
## Related
- `abx-observe` — the metrics and logs behind "it is slow" or "it failed"
- `abx-reinforcement-learning` — the rollout-shaped version of this
- `abx-common` — endpoint, key, cluster, approval
## Read more
[Autoscaling](https://scitix.github.io/Agent-Sandbox/docs/concepts/autoscaling.md)
---
# abx-sandbox-docker (/docs/skills/abx-sandbox-docker)
**Generated from**
> [`plugin/skills/abx-sandbox-docker/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-sandbox-docker/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## Docker inside a sandbox
A sandbox can run `docker` and `docker compose`. It is not on by default: the
template decides it, because a container runtime changes the Pod spec — a
dockerd, a data directory on node disk, and either a microVM or a privileged
container around it.
## The choice that matters, and it is not a preference
| Runtime | Isolation | Use it |
|---|---|---|
| **kata** (Firecracker microVM) | dockerd runs inside a guest VM; `privileged` applies to the guest | **This one.** Full functionality, compose included. |
| **runc + privileged** | dockerd shares the host kernel and holds every host capability | Only where the cluster has no kata runtime. |
Say the second one's consequence plainly whenever you recommend it: **escape
means the host**. It is not "slightly less isolated", and a user choosing it
should be choosing it knowingly.
Rootless dind is not the middle ground people expect — it has been measured on
these clusters and does not work under runc. The way to avoid dangerous
privilege is kata, where the privilege is confined to a microVM, not a rootless
daemon.
## Finding out what this deployment has
```bash
abx templates # which templates exist here
abx templates # what it carries
```
Look for a template whose description names dind or kata. If there is none, the
answer is that this deployment has not published one — not that you should hand
the user a template to install, which is an admin action.
## What it looks like from the SDK
Nothing special. It is the same E2B create; the template is what differs. The
call's fields are the SDK's, not this CLI's — read them off the E2B spec in the
image or the installed SDK, the way `abx-common` describes. What matters here
is one number:
```python
sbx = Sandbox.create("", …) # with a longer timeout than a plain sandbox
print(sbx.commands.run("docker version").stdout)
sbx.commands.run("docker compose up -d", cwd="/home/user/project")
```
The endpoint this factory talks to is the env's: `abx envs docs` prints it
for the cluster the env lives on — the E2B API URL, the data-plane domain and
the scheme. Take the public entry from that document unless you are already
inside the cluster.
Give it a longer timeout than you would a plain sandbox: dockerd starts in the
background while envd comes up in front, and the first `docker` call can arrive
before the daemon is listening. A short retry around the first command is
ordinary, not a symptom.
## Registries
Public images may not be reachable — many deployments are on internal networks.
Two things to check before concluding the image is broken:
- an internal mirror, with a prefix to rewrite `docker.io/...` onto;
- registry credentials on the env, which the platform materialises as an
image-pull secret. Those cover the **sandbox's own** image, not what `docker`
pulls from inside it — inside, `docker login` is the user's to run.
## When `docker` is not found
The env is on a template without it. Check which:
```bash
abx envs --json | jq '.spec.templateRef'
abx templates
```
Moving an env to another template is a template change, not a sandbox one, and
it rolls the pool.
## Related
- `abx-sandbox-network` — letting the sandbox reach a registry while the agent
inside it cannot reach the internet
- `abx-resource-capacity` — a dind sandbox is a bigger sandbox; size for it
- `abx-common` — endpoint, key, cluster, approval
## Read more
[Sandbox templates](https://scitix.github.io/Agent-Sandbox/docs/concepts/templates.md)
---
# abx-sandbox-network (/docs/skills/abx-sandbox-network)
**Generated from**
> [`plugin/skills/abx-sandbox-network/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-sandbox-network/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## What a sandbox can reach
Two separate decisions, and conflating them is the usual mistake:
```
SandboxEnv the env carries it does this environment HAVE a gateway
create call the sandbox carries it what THIS sandbox may reach
```
The env carries a switch and no rules. Rules belong to the individual sandbox
and arrive with the create call, in the E2B SDK's own vocabulary — there is no
AgentBox dialect to learn. The env's side is `abx update envs --help`
(one field, and that page names it); the sandbox's side is the SDK, whose shape
is in the image rather than in this file.
## Why the switch is on the env and the rules are not
Installing the proxy sidecar changes the Pod spec, so it rolls the pool. That
genuinely is an environment-level decision. Rules are per sandbox because an
environment is shared: an env-level allowlist would be a default that every
create overrides anyway, so it would buy nothing and cost a second configuration
surface.
**Fail-closed, deliberately:** a create that carries filtering rules against an
env with no gateway is **refused with 400**, not accepted-and-ignored. A Pod
without the sidecar has no redirection either, so accepting it would mean the
rules silently did nothing — which for an evaluation is the worst possible
outcome, because the run completes and the numbers are wrong.
## Cutting an evaluation off from the internet
This is the common case: the agent under test must not fetch the answer, and
must not install its way around a missing dependency.
The create call takes a network policy; **ask for its shape rather than
recalling it** — the fields are in the E2B spec the image ships
(`/opt/agentbox/source/pkg/openapi/e2b/openapi.yaml`), and the installed SDK
will tell you its own signature. `abx-common` has the whole recipe; the short
version is `abx update envs --help` for anything the env owns and the spec
above for anything the sandbox owns.
The call itself needs somewhere to go first: `abx envs docs` prints the
E2B API URL and data-plane domain for the env's cluster, which is what the
create is sent to. Take the public entry it offers unless the caller runs inside
the cluster.
What matters, and does not change with the field names: **deny-all-then-allow**,
never allow-all-then-deny. A denylist is a list of the routes you thought of.
Three things worth checking before declaring a run isolated:
1. **The env has a gateway.** Without it the create is refused — verify you saw
a sandbox, not a 400.
2. **The package index is not on the allowlist** unless the task needs it. It is
the most common accidental hole: the agent cannot search, but it can `pip
install` something that can.
3. **Test it from inside.** `sbx.commands.run("curl -sS -m 5 https://example.com")`
should fail. An isolation you did not observe failing is an isolation you are
assuming.
## Allowing one thing and nothing else
An evaluation that needs a model API but nothing else is the same shape:
deny everything, then allow that one host. If the credential for it must not be
readable inside the sandbox, that is `abx-sandbox-secrets` — the sidecar can
inject it on the way out, so the sandbox reaches the API while holding no key.
## When traffic gets through anyway
- Only `:80` and `:443` are parsed at layer 7. Traffic on another port passes
through without inspection, so a rule written for `:3000` does nothing.
- Check the env actually has the gateway on, and that the sandbox you are
testing came from that env.
## Related
- `abx-sandbox-secrets` — reaching a service without holding its credential
- `abx-harbor-framework` — running an evaluation suite on top of this
- `abx-common` — endpoint, key, cluster, approval
## Read more
[Egress and secrets](https://scitix.github.io/Agent-Sandbox/docs/concepts/egress-and-secrets.md)
---
# abx-sandbox-secrets (/docs/skills/abx-sandbox-secrets)
**Generated from**
> [`plugin/skills/abx-sandbox-secrets/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-sandbox-secrets/SKILL.md) — the same file the installer
> writes to `~/.agents/skills`. Edit it there; this page follows.
## Secrets a sandbox can use but cannot read
The mechanism, and the reason it is shaped this way:
```
vault (write-only) → operator memory → sidecar tmpfs → outbound header
```
The value never enters the sandbox. The sandbox holds a **decoy** — a
placeholder that looks like a token — and the egress sidecar substitutes the
real one as the request leaves, for the hosts and headers the rule names.
So an agent that can run arbitrary shell in there still cannot print the
credential, because it is not in there. That is the property worth protecting,
and it is why "just put it in envVars" is the wrong answer even when it works.
## Storing one
The vault is the E2B `/secrets` surface, so the SDK you already have speaks it.
Values are write-only: you can list names and overwrite, never read back.
Both surfaces are reached through the same E2B API, whose address and data-plane
domain are the env's: `abx envs docs` prints them for the cluster the env
lives on. Use the public entry it gives unless the caller is inside the cluster.
```python
sbx_secrets.set("OPENAI_KEY", "sk-…") # stored; not readable afterwards
```
In the console it is the **Vault** page.
## Using one
Reference it by name on the create call. **A plaintext value in `network.rules`
is refused with 400** — deliberately, because accepting it would put the
credential in the request body, the access log, and the caller's source, which
is the exposure the whole feature exists to remove.
The create call carries the sandbox's environment and the injection rules.
**Look its shape up rather than recalling it**: it belongs to the E2B surface,
so it is in `/opt/agentbox/source/pkg/openapi/e2b/openapi.yaml`, and the
installed SDK will print its own signature. `abx-common` has the general recipe
for finding any shape this way.
Two things about it are not about the field names, and are the part to get
right: the value the sandbox's code reads is a **decoy**, and the real one is
referenced by *name* — the sidecar substitutes it on the way out. A literal
credential in the request is a 400, because it would put the secret in the
request body, the access log and the caller's source.
The code inside runs unmodified: it reads `OPENAI_API_KEY`, sends it, and the
sidecar replaces it. Library code that has never heard of AgentBox works.
## Prerequisites, each of which fails silently
The sandbox runs, the request goes out, and the credential is simply not
substituted. Check in this order:
1. **The env has the gateway on** (`overrides.gateway.enabled`). Without it, a
create carrying rules is refused — so if you have a running sandbox, this
one is satisfied.
2. **The host matches the rule exactly.** The rule keys on the host it was
written for; a redirect elsewhere is not covered.
3. **The port is 80 or 443.** Only those are parsed at layer 7. A rule for a
service on another port never fires, and nothing says so.
4. **The secret name exists in the vault of the acting user.** A name that
resolves to nothing leaves the decoy in place, and the upstream returns 401 —
which reads as a bad key rather than a missing one.
`abx whoami` tells you which identity the vault is being read as; a secret
stored by one user is not visible to another.
## For an agent doing an integration
Writing vault secrets **is** allowed for an agent credential, under the normal
approval gate. This is deliberate and worth knowing: the credential lands in
the acting person's own vault and widens nobody's authority, and finishing an
integration end to end is exactly what people want an agent for. Minting an
AgentBox API key is the thing that is refused outright — different act,
different answer. See `abx-common`.
## Related
- `abx-sandbox-network` — deny everything, then allow the one host this rule
targets
- `abx-managed-agent` — the same mechanism is how the platform's own assistant
holds no real credential
- `abx-common` — endpoint, key, cluster, approval
## Read more
[Egress and secrets](https://scitix.github.io/Agent-Sandbox/docs/concepts/egress-and-secrets.md)
---
# Using the skills (/docs/skills)
A **skill** is a Markdown file an agent reads before it touches the platform. It
carries the part a `--help` page cannot: when to reach for a command at all, and
what the objects mean once you have.
`abx` ships nine of them. They are not documentation *about* the product — they
are the vocabulary the platform hands an agent, which is why this site renders
them here: what you read on these pages is the file the installer writes to
`~/.agents/skills`, generated from that same file at build time.
## Install
The skills arrive with the CLI. There is nothing to install separately.
**Installer**
`abx`'s installer writes the binary to `~/.local/bin` and the nine skills to
`~/.agents/skills`:
```bash
curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh | sh
```
Both paths are overridable, which is what an image build wants:
```bash
AGBX_BIN_DIR=/usr/local/bin AGBX_SKILL_DIR=/opt/skills sh install.sh
```
**Claude Code**
The plugin installs the skills and keeps the API key in the OS keychain, where
the agent never sees it:
```bash
/plugin marketplace add scitix/agent-sandbox
/plugin install agentbox
```
**Sandbox image**
A sandbox that is meant to drive the platform itself has the same two things to
do — the binary on `PATH`, the skills somewhere readable. `/opt/skills` is the
predictable place, and any harness told to read that directory will find them:
```dockerfile
RUN curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh \
| AGBX_BIN_DIR=/usr/local/bin AGBX_SKILL_DIR=/opt/skills sh
```
## The nine
| Skill | Read it when |
|---|---|
| [`abx-common`](/docs/skills/abx-common) | reaching a platform at all: the endpoint, the key, which cluster, and what happens to a write that needs a person's approval. Every other skill defers here. |
| [`abx-resource-capacity`](/docs/skills/abx-resource-capacity) | sizing: how many sandboxes fit, why a pool will not grow, how autoscaling and quota interact. |
| [`abx-observe`](/docs/skills/abx-observe) | something is failing or slow — events, pool status, sandbox logs, metrics. |
| [`abx-harbor-framework`](/docs/skills/abx-harbor-framework) | running a benchmark suite (Terminal-Bench, SWE-bench, a custom dataset) against pre-warmed pools. |
| [`abx-reinforcement-learning`](/docs/skills/abx-reinforcement-learning) | rollouts for training: concurrency, environment resets, collecting trajectories. |
| [`abx-managed-agent`](/docs/skills/abx-managed-agent) | giving your own agent a sandbox for its file and shell tools, durably across restarts. |
| [`abx-sandbox-docker`](/docs/skills/abx-sandbox-docker) | the workload needs Docker inside the sandbox. |
| [`abx-sandbox-network`](/docs/skills/abx-sandbox-network) | cutting a sandbox off the network, or allowing exactly one host. |
| [`abx-sandbox-secrets`](/docs/skills/abx-sandbox-secrets) | a sandbox needs a credential it must not be able to read. |
## How an agent uses them
A skill describes a task; the tool it reaches for is the CLI, whose shape is one
document:
```bash
abx agent-context # every resource, filter, column and write body, as JSON
```
Two habits are worth engineering into whatever agent you run, and both are in
`abx-common`: have it read that document once rather than guessing at commands,
and have it read the environment's own documentation
(`abx envs YOUR_ENV docs --cluster YOUR_CLUSTER`) rather than carrying endpoints
in its prompt.
Give an agent an **agent**-mode key
([section 6 of the CLI guide](/docs/tutorials/cli#6-api-key-permissions)), never
an unrestricted one: the platform's writes then wait for your approval, and
sandbox work is unaffected.
## See also
- [CLI guide](/docs/tutorials/cli) — installing `abx` itself, and using it from
an agent
- [E2B SDK guide](/docs/tutorials/e2b) — the sandbox half these skills assume
---
# Agent Sandbox CLI Guide (/docs/tutorials/cli)
Agent Sandbox runs on the command line, which is what makes it usable by an
agent as well as by a person: everything the console does, `abx` does, and
everything a sandbox does, the E2B SDK does. Hand this page to an agent and it
can create, use and reclaim its own sandboxes without a browser.
Two tools, and the split between them is the product:
| Tool | Covers |
|---|---|
| `abx` | the platform — environments, warm pools, autoscaling, quotas, templates |
| E2B SDK | the sandboxes themselves — create, exec, files, network |
The objects those commands address — templates, envs, pools — are described in
[Concepts](/docs/concepts). This page is how to drive them.
Values written in `UPPER_CASE` are placeholders. The console's own copy of this
guide — the floating **CLI** button on any page — arrives with the address, the
cluster and the E2B endpoints already filled in for your deployment.
## 1. Get an API key
Create one in the console, under **API Keys**. The plaintext is shown once at
creation, and the platform keeps a copy you can retrieve from that page at any
time.
If the key is going to an unattended agent, issue it in **agent** mode — see
[section 6](#6-api-key-permissions).
## 2. Install `abx`
`abx` is a single binary with nothing behind it: no Node.js, no virtualenv. One
script installs it on either platform and writes nine skills to
`~/.agents/skills` — see [section 4](#4-using-it-from-an-agent). It needs a
platform to talk to; if you have not installed one yet, start at
[Installation](/docs/installation).
**macOS**
```bash
curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh | sh
```
The binary lands in `~/.local/bin`. New shells find it through `~/.zprofile`;
for the one you already have open:
```bash
export PATH="$HOME/.local/bin:$PATH"
```
**Linux**
```bash
curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh | sh
```
The binary lands in `~/.local/bin`. Not every distribution puts that on `PATH`;
the script tells you when yours does not, and this fixes the shell you are in:
```bash
export PATH="$HOME/.local/bin:$PATH"
```
## 3. Initialise `abx`
```bash
abx context set YOUR_DEPLOYMENT \
--endpoint 'https://YOUR_CONSOLE/agentbox' \
--api-key agbx_...
abx clusters
```
That is the entire configuration: the console address and the key.
- The **address** reaches the platform, not one cluster: `abx clusters` lists
every cluster it has, and any command takes `--cluster `; a single-cluster
platform fills it in for you.
- The **first command** is therefore `abx clusters`, not `abx envs`: listing
clusters needs no `--cluster`, so it is the one command that cannot fail for
want of an id — and it prints the ids every other command takes.
- The **auth header** follows from the address, so there is nothing else to
configure.
- To work with several Agent Sandbox platforms, save each as its own context and
switch with `abx context use `. The context is which platform;
`--cluster` is which of its clusters, per command.
```bash
abx agent-context # the whole CLI as JSON — hand this to an agent
abx envs YOUR_ENV pools --cluster YOUR_CLUSTER
```
In CI or other unattended settings, `AGENTBOX_ENDPOINT` and `AGENTBOX_API_KEY`
stand in for a context.
## 4. Using it from an agent
The install above is the whole prerequisite: the binary, and nine skills that
give an agent the platform's vocabulary — rollouts, capacity, evaluations,
Docker-in-sandbox, egress isolation, vault secrets, approvals. What differs is
where your agent looks for them.
**Claude Code**
The plugin installs the skills and keeps your key in the OS keychain, where the
agent never sees it:
```bash
/plugin marketplace add scitix/agent-sandbox
/plugin install agentbox
```
Then set the endpoint and key in the plugin's settings; a hook writes them to
`~/.config/abx/config.json` for the CLI to read.
**Codex**
Codex reads `~/.agents/skills` directly, so the install already did it:
```bash
ls ~/.agents/skills # abx-common, abx-resource-capacity, abx-observe, …
```
Give it a context ([section 3](#3-initialise-abx)) and the skills are live; the
only other thing worth telling it is to read `abx agent-context` rather than
guess at commands.
**OpenCode**
The same directory, named from OpenCode's own instructions (an `AGENTS.md`, or
whatever file your project uses for agent guidance):
```markdown
Platform access is documented in ~/.agents/skills/abx-common/SKILL.md.
Run `abx agent-context` before composing an abx command.
```
Nothing about the platform changes with the tool: same binary, same key, same
commands.
**Anything else**
Any agent that can run a shell needs two things: the binary on its `PATH`, and
one of these two commands, whose output is written to be read by a model rather
than parsed by a person.
```bash
abx agent-context # the whole CLI as JSON — the entry point
abx envs YOUR_ENV docs --cluster YOUR_CLUSTER # this cluster's endpoints
```
Point it at `~/.agents/skills` if it reads a skills directory, and at
`abx agent-context` if it does not.
Whichever it is, give the agent an **agent**-mode key
([section 6](#6-api-key-permissions)), never an unrestricted one — the platform's
writes then wait for your approval, while sandbox work is unaffected. And have it
read the environment's own documentation ([section 5](#5-run-code-in-a-sandbox))
rather than carrying endpoints in its prompt: that document is rendered for the
cluster it belongs to, and the prompt is not.
All nine are also rendered on this site, one page each, for reading rather than
installing: [Skills](/docs/skills).
## 5. Run code in a sandbox
Every environment carries the documentation its template was written with,
rendered for the cluster it belongs to — the E2B API URL, the data-plane domain,
and the scheme to use:
```bash
abx envs YOUR_ENV docs --cluster YOUR_CLUSTER
```
Read it before driving an environment through E2B: it is the only place that
knows which endpoints your cluster answers on, and it covers what this page does
not — the two access paths, pool and scaling-group selection, cross-cluster
routing, and how images are rewritten between regions. The
[E2B SDK guide](/docs/tutorials/e2b) goes through those in full.
Sandboxes are E2B's surface, so the official SDK works unchanged — one call
before the import points it at Agent Sandbox instead of `e2b.dev`.
```bash
uv venv && source .venv/bin/activate
uv pip install e2b agent-sandbox-e2b
```
```python
from agent_sandbox_e2b import patch_e2b
patch_e2b(
api_url="https://YOUR_GATEWAY/agent-sandbox/api/e2b",
domain="YOUR_GATEWAY/agent-sandbox/api/data",
)
# patch_e2b() must run BEFORE this import, or the SDK talks to e2b.dev.
from e2b import Sandbox
sbx = Sandbox.create("YOUR_ENV", timeout=3600, secure=False)
print(sbx.commands.run("python -V").stdout)
sbx.kill()
```
## 6. API key permissions
An API key acts as **you** — your team, your namespace, your quota. It cannot
list or reach another user's sandboxes. Delete it in the console and it stops
working everywhere, including in any agent you handed it to.
A key is issued in one of two modes:
| Mode | What it may do |
|---|---|
| **Unrestricted** | everything you can do, with no further confirmation. The right mode on your own machine. |
| **Agent** | the same, except that the platform's writes — creating an environment, scaling or deleting anything — wait for your approval in the console. |
Sandbox work is identical in both modes: an agent key starts sandboxes, runs
commands and moves files exactly as an unrestricted one does. The gate is on the
platform's write surface, which is the part that outlives the sandbox.
An agent key also cannot issue itself another key, and cannot read key material —
not from `abx api-keys`, and not from a rendered document such as this one.
---
# E2B Python SDK (/docs/tutorials/e2b)
Agent Sandbox is a sandbox service for agentic workloads — reasoning
evaluation, training rollouts, managed agents — with secure isolation and
flexible deployment. [E2B](https://e2b.dev/docs) is the sandbox provider behind
Manus, and `envd` is the runtime E2B provides; Agent Sandbox serves the
E2B-compatible API, so the official SDK reaches it unchanged.
This page is the SDK side of the platform: how to point the SDK at a cluster,
how to get a sandbox out of a warm pool, and what the create call takes. The
platform side — environments, pools, autoscaling, quotas — is `abx`; see the
[CLI guide](/docs/tutorials/cli), and [Concepts](/docs/concepts) is the object
model behind both.
Every deployment renders its own copy of this material with the real addresses
filled in, and that copy is authoritative for the cluster you are using:
```bash
abx envs YOUR_ENV docs --cluster YOUR_CLUSTER
```
On the console it is the **Env Docs** panel on the environment's page, key
included. This page keeps placeholders where that one has values.
## 1. Install
```bash
uv pip install 'agent-sandbox-e2b[e2b]'
```
The `[e2b]` extra pulls in the official `e2b` package. Two version notes worth
knowing: **`agent-sandbox-e2b >= 0.0.6` is required for public access** (older
builds do not assemble the gateway path, so sandbox ports do not connect), and
the server rejects clients below its minimum supported version with `426
Upgrade Required`.
## 2. Quick start
`patch_e2b()` must run **before** `from e2b import Sandbox`; otherwise the SDK
connects to the official E2B service instead of Agent Sandbox.
```python
import os
os.environ["E2B_API_KEY"] = "agbx_..." # your API key
os.environ["E2B_API_URL"] = "https://YOUR_GATEWAY/agent-sandbox/api/e2b"
os.environ["E2B_DOMAIN"] = "YOUR_GATEWAY/agent-sandbox/api/data" # no scheme
os.environ["E2B_HTTPS"] = "true" # "false" for a plain-http data plane
from agent_sandbox_e2b import patch_e2b
patch_e2b()
from e2b import Sandbox
sbx = Sandbox.create("YOUR_ENV", timeout=3000, secure=False)
print(sbx.is_running())
sbx.kill()
```
The same three values can be passed as arguments instead, which take precedence
over the environment:
```python
patch_e2b(
api_url="https://YOUR_GATEWAY/agent-sandbox/api/e2b",
domain="YOUR_GATEWAY/agent-sandbox/api/data",
https=True,
)
```
`secure=False` skips E2B's signed handshake: Agent Sandbox authenticates with an
API key, so it stays `False`.
## 3. Access paths
There are three ways in, and they differ only in how much of the network the
request crosses. The environment's own documentation names the values for the
cluster you are on.
| Path | Who it is for | What changes |
|---|---|---|
| **Public** | your laptop, CI, anything outside the cluster | the full set of values above: `E2B_API_URL`, `E2B_DOMAIN`, `E2B_HTTPS` |
| **Internal network** | a machine on the same private network as the cluster | point the gateway hostname at the cluster's internal IP — no code change |
| **In-cluster** | code running inside the cluster | nothing: `patch_e2b()` falls back to the cluster's own services |
**Public** is the default and the one to reach for unless you know you are
inside. It is the longest path, but it depends on nothing about where the caller
runs.
**Internal** skips the public hop by resolving the gateway host to the cluster's
internal IP. For a single machine, add the alias to `/etc/hosts`:
```bash
echo "YOUR_INNER_IP YOUR_GATEWAY_HOST" | sudo tee -a /etc/hosts
```
For a workload that is itself a Pod, `hostAliases` does the same without
touching a shared file:
```yaml
spec:
hostAliases:
- ip: "YOUR_INNER_IP"
hostnames:
- "YOUR_GATEWAY_HOST"
```
**In-cluster** is the shortest path: leave `E2B_API_URL`, `E2B_DOMAIN` and
`E2B_HTTPS` unset — `patch_e2b()` falls back to
`agentbox-e2b-api.agentbox-system.svc.cluster.local` and
`agentbox-data-plane.agentbox-system.svc.cluster.local` — and keep only
`E2B_API_KEY`.
## 4. What a sandbox comes from: template, env, pool
A sandbox is not built from scratch at create time. The platform keeps a set of
**pre-warmed Pods**, and `Sandbox.create()` claims one and swaps in your image,
which is why it takes seconds rather than minutes.
| Object | What it is | Who makes it |
|---|---|---|
| `SandboxTemplate` | the Pod shape, the runtime, this documentation | platform administrators |
| `SandboxEnv` | your entry name (`YOUR_ENV`), bound to one template | you |
| `SandboxPool` | a group of pre-warmed Pods under an env, one per resource shape | you |
Create them in the console: **Sandbox Envs → new environment** (name, template,
optional overrides, autoscaling, image pull secrets, network policy), then
**add pool** on the environment (quota, resource mode — instance type ×
multiplier, or explicit CPU and memory — and replica count). Replicas are the
concurrency ceiling: tasks beyond the number of idle Pods queue.
The environment's page shows `idle` / `running` / `desired` per member pool.
`idle > 0` is the point at which `Sandbox.create()` returns immediately.
## 5. `Sandbox.create()` parameters
The first argument (E2B calls it `template`) has four spellings here:
| Form | Meaning |
|---|---|
| `"YOUR_ENV"` | **recommended** — the environment in the current cluster; the platform schedules across its member pools |
| `"YOUR_ENV//IMAGE"` | the same, with the main container image replaced |
| `"CLUSTER_ID::YOUR_ENV"` | a named cluster plus environment (cross-cluster — see section 6) |
| `"CLUSTER_ID::POOL_NAME"` | a named cluster plus one specific pool, skipping env scheduling |
The `//IMAGE` suffix composes with any of them, e.g.
`"CLUSTER_ID::YOUR_ENV//docker.io/library/ubuntu:24.04"`.
Other parameters:
| Parameter | Meaning |
|---|---|
| `timeout=3000` | **idle** timeout in seconds. A sandbox with no activity for that long is reclaimed; extend a live one with `sandbox.set_timeout(1800)` |
| `secure=False` | keep it `False` — authentication is the API key |
| `metadata={...}` | your own labels, readable from the sandbox object. Three keys are reserved by the platform (below) |
Reserved metadata keys:
| Key | Effect | Example |
|---|---|---|
| `agentbox.scitix.ai/image` | overrides the main container image (same as `//IMAGE`, and wins over it) | `"registry.example.com/my/img:v1"` |
| `agentbox.scitix.ai/startup-timeout` | **startup** timeout in seconds — how long create waits for Ready; raise it for large images | `"900"` |
| `agentbox.scitix.ai/scaling-group` | **routes by scaling group**: only member pools in that group are eligible | `"1c16gi"` |
```python
sbx = Sandbox.create(
"YOUR_ENV",
timeout=3000,
secure=False,
metadata={
"agentbox.scitix.ai/scaling-group": "1c16gi", # only 1c16gi pools
"agentbox.scitix.ai/startup-timeout": "900", # large image, give it time
"run_id": "eval-2026-07-31", # ordinary label
},
)
```
A scaling group the environment has no pool for **fails the create (503)**
rather than falling back to another shape. That is deliberate: during an
evaluation, a size that silently drifts is harder to find than a request that
refuses.
## 6. Cross-cluster
An environment belongs to one cluster; same-named environments in different
clusters federate, so the platform can route a create to whichever side has
idle Pods.
| Target | Spelling | Behaviour |
|---|---|---|
| a specific pool in a specific cluster | `"other-cluster::POOL_NAME"` | straight to that pool; cluster and size both pinned |
| a specific cluster's environment | `"other-cluster::YOUR_ENV"` | forwarded to that cluster, which schedules across the environment's member pools (**the recommended cross-cluster form** — pool names can change) |
| either side, whichever is free | `"YOUR_ENV"` | local first; when the local side has no idle Pods and cannot scale, the request is forwarded to a cluster that has them |
The third form needs the environment to exist in every participating cluster:
in the console, **new environment → extend another cluster's environment**,
pick the existing one, and add member pools on the new side too.
Two prerequisites for cross-cluster scheduling:
1. **The image must exist in every region.** A sandbox lands where there is
capacity, and it has to be able to pull from there — see section 7.
2. **Team-to-namespace mapping must match** across clusters, or federation
cannot pair `(namespace, env)` and a bare name will not spread. Spell the
target out (`"other-cluster::YOUR_ENV"`) when that is the case.
## 7. Images and regions
Each cluster declares its own region's image registry. When the image you name
belongs to **another cluster's** private registry, the platform rewrites the
registry host to this cluster's equivalent, so a sandbox pulls from its own
region instead of across the world:
```text
you write: registry-region-a.example.com/team/swebench:260328
actually pulled: YOUR_REGISTRY_HOST/team/swebench:260328
```
The rules, in short:
- only the **host** is rewritten — the path and tag survive, so the same image
has to exist under the same path in each region's registry;
- the rewrite happens only between registries of the **same type** (the platform
configuration's `type`); public registries such as `docker.io` are never
rewritten;
- when no registry of that type exists in this cluster, the image is pulled from
the address you wrote — possibly across regions, and possibly slowly.
So before a cross-cluster evaluation, confirm the image has been synced to every
region you expect to run in.
## 8. Common operations
```python
sbx.is_running()
# run commands — as the standard user, or as root
sbx.commands.run("echo hello && python --version")
sbx.commands.run("id", user="root")
# files
sbx.files.write("/tmp/script.py", b"print('hello from AgentBox')\n")
content = sbx.files.read("/tmp/script.py")
# extend the idle timeout
sbx.set_timeout(1800)
# always destroy it: a running sandbox holds a warm Pod
sbx.kill()
```
## 9. Troubleshooting
| Symptom | Cause and fix |
|---|---|
| `400 sandbox env or pool "x" not found` | the name is wrong, or the environment is not in this cluster — for another cluster write `CLUSTER_ID::x` |
| `503 ... has no eligible members` | the environment has no live member pool here: add one, and if you set `scaling-group`, check that the group exists |
| a placeholder appears instead of an API key | the server never renders a key. The console fills it in from the key you select; a CLI or API caller replaces it with its own |
| `426 Upgrade Required` | the client is below the server's minimum version: upgrade `agent-sandbox-e2b` |
| the sandbox is created but its ports do not connect | usually `agent-sandbox-e2b < 0.0.6`, or `E2B_DOMAIN` missing the gateway path (it should be `YOUR_GATEWAY/agent-sandbox/api/data`) |
| creates keep timing out | the image is large or being pulled for the first time: raise `agentbox.scitix.ai/startup-timeout`, and check the image is present in this region's registry |
| `Sandbox.create()` hangs | it is waiting for a Pod to become Ready, from the warm pool or a cold start. Set the startup timeout so it fails with a reason instead of waiting |
| a sandbox disappeared | the idle `timeout` elapsed during a quiet stretch; call `set_timeout()` before that stretch begins |
---
# A warm pool (/docs/examples/envs/e2b)
**Source**
> [`config/samples/e2b_sandboxenv.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/config/samples/e2b_sandboxenv.yaml) on the `develop`
> branch — the file this page is rendered from, and the one to apply.
An Env binds one template and fans out to member pools. This one keeps two
Pods warm on the cluster you name, so `Sandbox.create("e2b")` is served by a
Pod that already exists rather than by a cold start.
`kubectl apply -f config/samples/e2b_sandboxenv.yaml`
Two things have to be true first:
- the template exists — `kubectl apply -f
config/samples/e2b-basic_sandboxtemplate.yaml`
- clusterID below is this cluster's id (the worker's `localClusterId`), and
the namespace is one your team owns
The other two templates take this file unchanged except for one line: point
templateRef at e2b-envd-docker or e2b-envd-kata, and give the member a name
and a size that match what it runs — a container engine wants more memory than
a bare sandbox, and a microVM carries its own kernel inside the same numbers.
A member's `spec` is the pool's frozen snapshot; leaving it out of the file is
deliberate — the Env renders it from the template on the next reconcile, so
what is written here stays the part a person decides: shape, size, replicas.
On a deployment that bills by quota (`abx envs ` prints poolSizing), a
pool must name a quota and an instance type instead of inlineResources.
```yaml
apiVersion: agents.navix.sh/v1alpha1
kind: SandboxEnv
metadata:
name: e2b
spec:
templateRef:
name: e2b-envd
mode: WarmPool
autoscaling:
groups:
- name: 1c2gi
enabled: true
minReplicas: 0
maxReplicas: 4
clusters:
- clusterID: YOUR_CLUSTER_ID
members:
- name: e2b-1c2gi
config:
scalingGroup: 1c2gi
# The Pod's size. A billed deployment names instanceType and
# multiplier here instead.
inlineResources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "1"
memory: 2Gi
spec:
# Warm capacity: how many sandboxes this pool can hand out at once
# before anyone has to wait for one to come back.
replicas: 2
```
```bash
kubectl apply -f config/samples/e2b_sandboxenv.yaml
```
---
# Using the environments (/docs/examples/envs)
A [SandboxEnv](/docs/concepts/envs) binds one template and fans out to member
pools; the pool is what holds the Pods a sandbox is claimed from. One
environment is enough to show how they are put together — the page is the
manifest from
[`config/samples`](https://github.com/scitix/Agent-Sandbox/tree/develop/config/samples),
and the other two templates take it with a one-line change.
| Environment | On which template |
|---|---|
| [A warm pool](/docs/examples/envs/e2b) | [E2B Basic](/docs/examples/templates/e2b) — two Pods kept warm |
Switching it to [E2B Docker](/docs/examples/templates/e2b-docker) or
[E2B Kata](/docs/examples/templates/e2b-kata) is `templateRef.name`, plus a
member name and a size that match what the sandbox actually runs: a container
engine needs more memory than a bare sandbox, and a microVM carries its own
kernel inside the same numbers.
## Order of application
The template has to exist first, and the cluster id has to be the one the
worker was installed with (`controller.localClusterId` and
`spec.clusters[].clusterID` are the same value):
```bash
kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml
kubectl apply -f config/samples/e2b_sandboxenv.yaml
```
## What the sample decides, and what the platform does
The member declares the shape (`inlineResources`), how it scales
(`config.scalingGroup`, matched by an entry in `spec.autoscaling.groups`), and
how many Pods to keep warm (`spec.replicas`). It does **not** repeat the Pod
spec: the environment renders each member from the template on the next
reconcile, so the file stays the part a person decides.
On a deployment that bills by quota (`abx envs ` prints `poolSizing`), a
member must name an instance type and a quota label instead of
`inlineResources` — the CLI guide's [pools section](/docs/concepts/pools) has
the three shapes.
## Check it
```bash
kubectl get sandboxenv e2b
abx envs --cluster YOUR_CLUSTER # the env, and its pools
abx envs e2b --cluster YOUR_CLUSTER # one env, member by member
```
`idleReplicas` above zero means a claim will be served instantly; until then
the Pods are still starting.
---
# E2B Docker (/docs/examples/templates/e2b-docker)
**Source**
> [`config/samples/e2b-docker_sandboxtemplate.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/config/samples/e2b-docker_sandboxtemplate.yaml) on the `develop`
> branch — the file this page is rendered from, and the one to apply.
Same envd runtime as e2b-basic_sandboxtemplate.yaml, but the Pod runs a container
engine of its own: dockerd starts in the background, and code inside the
sandbox can `docker build`, `docker run` and `docker compose up`. That is what
a workload needs when the task is "build this image", "bring up this compose
file" or "run a database".
`kubectl apply -f config/samples/e2b-docker_sandboxtemplate.yaml`
WARNING: privileged. A container engine needs more of the kernel than a
container is usually given, so this template runs the sandbox Pod privileged
(and starts dockerd inside it). A container escape from here reaches the
node. It is the right trade for a trusted workload, and the wrong one for
untrusted code. On a cluster with a microVM runtime class — kata, or anything
that can run a Pod in its own kernel — add `runtimeClassName` to the Pod spec
below and the same privileged dockerd is confined to a guest VM instead of
sharing the node's kernel. e2b-kata_sandboxtemplate.yaml is that template with
the field already set.
Every image is public: the container engine comes from Docker Hub's own dind
image, and the runtime pieces from this project's GHCR packages.
```yaml
apiVersion: agents.navix.sh/v1alpha1
kind: SandboxTemplate
metadata:
name: e2b-envd-docker
spec:
version: 0.0.1
description: E2B-compatible sandbox with Docker inside — dockerd and the compose plugin in the sandbox.
idleImage: ghcr.io/scitix/agent-sandbox-idle:0.0.10
runtimes:
- name: envd
port: 49983
protocol: TCP
description: E2B envd runtime
readinessProbe:
httpGet:
port: 49983
path: /health
template:
spec:
automountServiceAccountToken: false
enableServiceLinks: false
shareProcessNamespace: false
containers:
- name: sandbox
# A container engine, not a workload image: dockerd, the docker CLI
# and the compose plugin are all in this one.
image: docker:29-dind
imagePullPolicy: IfNotPresent
command:
- /mnt/agentbox/tini
- -g
- --
- /bin/sh
- -c
- |
# An idle Pod wears the idle image, which has no dockerd to start.
if [ "$AGENTBOX_IS_IDLE_IMAGE" = "true" ] || [ -f /etc/agentbox_is_idle_image ]; then
echo "[dind] idle Pod; waiting for a claim."
exec sleep infinity
fi
# Plain unix socket: the socket is inside this Pod, and nothing
# outside it needs to reach the daemon.
export DOCKER_TLS_CERTDIR=""
echo "[dind] starting dockerd"
dockerd --host=unix:///var/run/docker.sock >/var/log/dockerd.log 2>&1 &
# Best effort: the first `docker` call should not race the daemon.
# Not fatal — a failure here is visible in the log above.
for _ in $(seq 1 30); do
if docker version >/dev/null 2>&1; then
echo "[dind] dockerd is ready"
break
fi
sleep 1
done
# Hand the foreground to envd, which is what the E2B SDK talks to.
exec /bin/sh /mnt/agentbox/agentbox-entrypoint.sh
env:
- name: AGENTBOX_DIR
value: /mnt/agentbox
- name: PORT
value: "49999"
# Scrub the service environment Kubernetes injects, so a command
# inside the sandbox cannot see the cluster's addresses.
- name: KUBERNETES_SERVICE_HOST
- name: KUBERNETES_SERVICE_PORT
- name: KUBERNETES_SERVICE_PORT_HTTPS
- name: KUBERNETES_PORT
- name: KUBERNETES_PORT_443_TCP
- name: KUBERNETES_PORT_443_TCP_ADDR
- name: KUBERNETES_PORT_443_TCP_PORT
- name: KUBERNETES_PORT_443_TCP_PROTO
ports:
- containerPort: 49983
name: envd
protocol: TCP
- containerPort: 49999
name: app
protocol: TCP
# A container engine plus whatever it builds: size for both. As with
# the E2B template, a Pool's own sizing overrides this.
resources:
limits:
cpu: "2"
memory: 4Gi
requests:
cpu: "2"
memory: 4Gi
securityContext:
privileged: true
runAsUser: 0
runAsGroup: 0
volumeMounts:
- name: shared-bin
mountPath: /mnt/agentbox
readOnly: true
# Node disk, not the container's own overlay root: overlay2 on top
# of overlayfs is something dockerd refuses to start with.
- name: docker-storage
mountPath: /var/lib/docker
initContainers:
- name: tini-injector
image: ghcr.io/scitix/agent-sandbox-tini:v0.19.0-static
imagePullPolicy: IfNotPresent
command: [sh, -c]
args:
- |
TARGET_DIR="/mnt/agentbox"
mkdir -p "$TARGET_DIR"
if [ -f "$TARGET_DIR/tini" ]; then
echo "[Init] tini already exists at $TARGET_DIR. Skipping copy."
else
cp /workspace/tini "$TARGET_DIR/tini"
chmod +x "$TARGET_DIR/tini"
echo "[Init] tini static injected successfully."
fi
volumeMounts:
- name: shared-bin
mountPath: /mnt/agentbox
- name: envd-injector
image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2
imagePullPolicy: IfNotPresent
command: [sh, -c]
args:
- |
TARGET_DIR="/mnt/agentbox"
mkdir -p "$TARGET_DIR"
if [ -f "$TARGET_DIR/envd" ] && [ -f "$TARGET_DIR/agentbox-entrypoint.sh" ]; then
echo "[Init] envd and scripts already exist in $TARGET_DIR. Skipping copy."
else
cp /workspace/envd "$TARGET_DIR/"
cp /workspace/agentbox-entrypoint.sh "$TARGET_DIR/"
chmod +x "$TARGET_DIR/envd" "$TARGET_DIR/agentbox-entrypoint.sh"
echo "[Init] envd & scripts successfully injected."
fi
volumeMounts:
- name: shared-bin
mountPath: /mnt/agentbox
volumes:
- name: shared-bin
emptyDir: {}
- name: docker-storage
emptyDir: {}
```
```bash
kubectl apply -f config/samples/e2b-docker_sandboxtemplate.yaml
```
---
# E2B Kata (/docs/examples/templates/e2b-kata)
**Source**
> [`config/samples/e2b-kata_sandboxtemplate.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/config/samples/e2b-kata_sandboxtemplate.yaml) on the `develop`
> branch — the file this page is rendered from, and the one to apply.
The same sandbox as e2b-basic_sandboxtemplate.yaml, with one line that matters:
`runtimeClassName`, which asks the cluster to run the Pod in its own kernel
(Kata Containers, Firecracker, or whatever your cluster publishes) instead of
sharing the node's. A container escape from a shared-kernel Pod lands on the
node; from here it lands in a virtual machine.
`kubectl apply -f config/samples/e2b-kata_sandboxtemplate.yaml`
Two things to check before you use it:
- the class exists: `kubectl get runtimeclass`. The name below is the
conventional one; use whatever yours is called, or the Pod never schedules.
- the runtime has the RAM: a microVM carries its own kernel, so a sandbox
that fits in 2Gi on runc may want more here.
Cost is the other half of the trade: start-up is slower and each sandbox
occupies more than a container would. Pay it where the code is untrusted —
evaluations of a model you do not control, agents given a shell and a network.
Every image below is public, like the other examples.
```yaml
apiVersion: agents.navix.sh/v1alpha1
kind: SandboxTemplate
metadata:
name: e2b-envd-kata
spec:
version: 0.0.1
description: E2B-compatible sandbox in a microVM — envd on a runtime class with its own kernel.
idleImage: ghcr.io/scitix/agent-sandbox-idle:0.0.10
runtimes:
- name: envd
port: 49983
protocol: TCP
description: E2B envd runtime
readinessProbe:
httpGet:
port: 49983
path: /health
template:
spec:
# The whole difference from the basic template. Set this to the class
# `kubectl get runtimeclass` shows on your cluster.
runtimeClassName: kata-fc
automountServiceAccountToken: false
enableServiceLinks: false
shareProcessNamespace: false
containers:
- name: sandbox
image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2
imagePullPolicy: IfNotPresent
command:
- /mnt/agentbox/tini
- -g
- --
- /bin/sh
- /mnt/agentbox/agentbox-entrypoint.sh
env:
- name: AGENTBOX_DIR
value: /mnt/agentbox
- name: PORT
value: "49999"
ports:
- containerPort: 49983
name: envd
protocol: TCP
- containerPort: 49999
name: app
protocol: TCP
# More headroom than the basic template on purpose: the guest kernel
# and its own page cache live inside these numbers.
resources:
limits:
cpu: "1"
memory: 4Gi
requests:
cpu: "1"
memory: 4Gi
securityContext:
runAsGroup: 0
runAsUser: 0
volumeMounts:
- mountPath: /mnt/agentbox
name: shared-bin
readOnly: true
initContainers:
- name: tini-injector
image: ghcr.io/scitix/agent-sandbox-tini:v0.19.0-static
imagePullPolicy: IfNotPresent
command: [sh, -c]
args:
- |
TARGET_DIR="/mnt/agentbox"
mkdir -p "$TARGET_DIR"
if [ -f "$TARGET_DIR/tini" ]; then
echo "[Init] tini already exists at $TARGET_DIR. Skipping copy."
else
cp /workspace/tini "$TARGET_DIR/tini"
chmod +x "$TARGET_DIR/tini"
echo "[Init] tini static injected successfully."
fi
volumeMounts:
- name: shared-bin
mountPath: /mnt/agentbox
- name: envd-injector
image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2
imagePullPolicy: IfNotPresent
command: [sh, -c]
args:
- |
TARGET_DIR="/mnt/agentbox"
mkdir -p "$TARGET_DIR"
if [ -f "$TARGET_DIR/envd" ] && [ -f "$TARGET_DIR/agentbox-entrypoint.sh" ]; then
echo "[Init] envd and scripts already exist in $TARGET_DIR. Skipping copy."
else
cp /workspace/envd "$TARGET_DIR/"
cp /workspace/agentbox-entrypoint.sh "$TARGET_DIR/"
chmod +x "$TARGET_DIR/envd" "$TARGET_DIR/agentbox-entrypoint.sh"
echo "[Init] envd & scripts successfully injected."
fi
volumeMounts:
- name: shared-bin
mountPath: /mnt/agentbox
volumes:
- name: shared-bin
emptyDir: {}
```
```bash
kubectl apply -f config/samples/e2b-kata_sandboxtemplate.yaml
```
---
# E2B Basic (/docs/examples/templates/e2b)
**Source**
> [`config/samples/e2b-basic_sandboxtemplate.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/config/samples/e2b-basic_sandboxtemplate.yaml) on the `develop`
> branch — the file this page is rendered from, and the one to apply.
The base template: a Pod that runs envd, the agent the E2B SDK talks to. A
sandbox created from an Env on this template can run any image the cluster can
pull — the image is swapped in on the claim, so there is nothing to rebuild
when the workload changes.
`kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml`
Every image below is public, so this works on a cluster that has never seen
this project before. Pinning: the tags are the ones the platform runs; move
them together with your platform release.
```yaml
apiVersion: agents.navix.sh/v1alpha1
kind: SandboxTemplate
metadata:
name: e2b-envd
spec:
version: 0.0.1
description: E2B-compatible sandbox — envd inside a Pod, any image on claim.
idleImage: ghcr.io/scitix/agent-sandbox-idle:0.0.10
runtimes:
- name: envd
port: 49983
protocol: TCP
description: E2B envd runtime
readinessProbe:
httpGet:
port: 49983
path: /health
template:
spec:
automountServiceAccountToken: false
containers:
- command:
- /mnt/agentbox/tini
- -g
- --
- /bin/sh
- /mnt/agentbox/agentbox-entrypoint.sh
env:
- name: AGENTBOX_DIR
value: /mnt/agentbox
- name: PORT
value: "49999"
image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2
imagePullPolicy: IfNotPresent
name: sandbox
ports:
- containerPort: 49983
name: envd
protocol: TCP
- containerPort: 49999
name: app
protocol: TCP
# The Pod's size when nothing else decides it. A Pool on this Env
# usually does: on a billed deployment its instance type (times the
# multiplier) is the envelope, and inlineResources is what the Pod
# actually asks for. This is the default for a free-form Pool.
resources:
limits:
cpu: "1"
memory: 2Gi
requests:
cpu: "1"
memory: 2Gi
securityContext:
runAsGroup: 0
runAsUser: 0
volumeMounts:
- mountPath: /mnt/agentbox
name: shared-bin
readOnly: true
enableServiceLinks: false
initContainers:
- args:
- >
TARGET_DIR="/mnt/agentbox"
mkdir -p "$TARGET_DIR"
if [ -f "$TARGET_DIR/tini" ]; then
echo "[Init] tini already exists at $TARGET_DIR. Skipping copy."
else
cp /workspace/tini "$TARGET_DIR/tini"
chmod +x "$TARGET_DIR/tini"
echo "[Init] tini static injected successfully."
fi
command:
- sh
- -c
image: ghcr.io/scitix/agent-sandbox-tini:v0.19.0-static
imagePullPolicy: IfNotPresent
name: tini-injector
resources: {}
volumeMounts:
- mountPath: /mnt/agentbox
name: shared-bin
- args:
- >
TARGET_DIR="/mnt/agentbox"
mkdir -p "$TARGET_DIR"
if [ -f "$TARGET_DIR/envd" ] && [ -f
"$TARGET_DIR/agentbox-entrypoint.sh" ]; then
echo "[Init] envd and scripts already exist in $TARGET_DIR. Skipping copy."
else
cp /workspace/envd "$TARGET_DIR/"
cp /workspace/agentbox-entrypoint.sh "$TARGET_DIR/"
chmod +x "$TARGET_DIR/envd" "$TARGET_DIR/agentbox-entrypoint.sh"
echo "[Init] envd & scripts successfully injected."
fi
command:
- sh
- -c
image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2
imagePullPolicy: IfNotPresent
name: envd-injector
resources: {}
volumeMounts:
- mountPath: /mnt/agentbox
name: shared-bin
shareProcessNamespace: false
volumes:
- emptyDir: {}
name: shared-bin
```
```bash
kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml
```
---
# Using the templates (/docs/examples/templates)
A [SandboxTemplate](/docs/concepts/templates) is what a sandbox is made from:
the Pod shape, the idle image, the runtime, and the defaults a claim inherits.
These pages are the real files — each one is the manifest in
[`config/samples`](https://github.com/scitix/Agent-Sandbox/tree/develop/config/samples)
on the `develop` branch, rendered here so it can be read before it is applied.
All of the images are public (this project's GHCR packages, and Docker Hub's own
images), so each works on a cluster that has never seen Agent Sandbox before.
| Template | What it is for |
|---|---|
| [E2B Basic](/docs/examples/templates/e2b) | the sandbox itself: envd in a Pod, any image swapped in on claim. Start here. |
| [E2B Docker](/docs/examples/templates/e2b-docker) | the same, plus a container engine inside the sandbox — `docker build`, `docker run`, `docker compose up`. |
| [E2B Kata](/docs/examples/templates/e2b-kata) | the same sandbox in its own kernel, on a microVM runtime class. For code you do not trust. |
## Apply one
```bash
kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml
```
Then give it an [environment](/docs/examples/envs) to hold capacity — a
template on its own creates nothing.
## Reading a template
Three fields decide almost everything, and each example page comments the rest:
| Field | What it decides |
|---|---|
| `idleImage` | what a Pod runs while no sandbox holds it. Keeping this small is what makes a deep warm pool affordable. |
| `runtimes` | the runtime inside the Pod, and the ports it answers on. `envd` is the one the E2B SDK speaks to. |
| `template.spec` | the Pod: images, resources, security context, volumes. Anything you would write in a Pod spec, plus the `runtimeClassName` that picks the isolation level. |
The one field worth understanding before you edit an example: the running image
is chosen when a sandbox is *claimed*, not when the template is applied, so a
template does not pin your workload image — see
[in-place update](/docs/concepts/inplace-update).
---
# Brain and Hands (/docs/tutorials/managed-agents/brain-hands)
An agent harness has two jobs: decide what to do, and do it. The first is a
model loop with your prompt and your tools. The second is `bash`, `read`,
`write`, `edit`, `grep`, `glob`, `apply_patch` — and today it happens on
whatever machine the harness is running on.
**Brain and Hands** separates those two. The brain keeps running wherever you
run it; the hands become a sandbox bound to the conversation.
```mermaid
flowchart LR
subgraph host["wherever your agent runs"]
H["harness
Claude Code · OpenCode · your loop"]
B["hands binding
the seven tools"]
D["hands daemon
session → sandbox"]
H -->|tool call| B
B -->|HTTP| D
end
S["sandbox for this session
/home/agents/…"]
D -->|E2B API| S
```
## Why the split is worth it
| Without it | With it |
|---|---|
| the work lands in someone's home directory and stays there | it lands in a sandbox that is reclaimed on a timer |
| a restarted process loses its scratch space | the sandbox is bound to the conversation, not the process |
| the agent's `rm -rf` is aimed at the machine you develop on | it is aimed at a machine that exists for this task |
| the filesystem is invisible unless you go and look | it is inspectable: the same sandbox is a page in the console |
**This is confinement, not a security boundary**
> The agent process still runs on your machine, with your files, your environment
> and your credentials. What moves into the sandbox is where the agent's *tools*
> act. An agent that can install a package or load a plugin can reach the host
> again — cooperative confinement, not isolation. If you need the stronger thing,
> run the harness itself in a container or a sandbox; the two compose, and
> [E2B Kata](/docs/examples/templates/e2b-kata) is how you get the sandbox half
> to carry its own kernel.
## Three layers, and the rule that holds them together
`sdk/hands/` is one package with a deliberate seam:
| Layer | Path | What it does |
|---|---|---|
| **core** | `typescript/src/core/` | the seven tools and their behaviour contract — harness-neutral, knows nothing about harnesses |
| **binding** | `typescript/src/harness/{claude-code,opencode,mcp}/` | expresses "replace these built-ins with these tools" in one harness's vocabulary |
| **daemon** | `python/agentbox_hands/` | binds a stable session id to a sandbox, and serves the workspace file API |
The rule: **a binding must not reimplement behaviour from core.** A binding that
does drifts from the contract, and the drift is invisible from the signatures.
How completely the built-ins can be *taken away* differs by harness, and it is
worth knowing which one you are on before trusting the confinement:
| Harness | Mechanism | How complete |
|---|---|---|
| Claude Code | `tools: []` + `disallowedTools` + `mcpServers` + `toolAliases`, together | complete; subagents inherit it, and a test drives a real session to check |
| OpenCode | a tool whose name matches a built-in replaces it | **unverified** — the override appears to register alongside the built-in |
| Anything over MCP | the generic binding | MCP can *add* tools; it cannot remove a harness's built-ins |
That second row is the reason the console's own assistant runs on Claude Code:
an override that registers alongside the built-in leaves the built-in in play,
which is the one outcome this package exists to prevent.
## Plugging it in
**Claude Agent SDK**
```ts
import { sandboxToolOptions } from '@scitix/agentbox-hands/claude-code'
for await (const msg of query({
prompt,
options: { ...sandboxToolOptions({ sessionKey: threadId }) },
})) { … }
```
`sandboxToolOptions` returns all five tool-related options as one object on
purpose: applying four of five loses the guarantee and nothing complains.
**OpenCode**
Point the config directory's `tools/.ts` at the binding — one line each,
and the filename is what does the overriding:
```ts
// ~/.config/opencode/tools/bash.ts
export { default } from '@scitix/agentbox-hands/opencode/tools/bash'
```
Raise `tool_output` in `opencode.json` at the same time, or the confinement
leaks: OpenCode truncates an oversized tool result by writing the full text to a
file on the machine running the harness and handing the agent that path. The
binding's own offload writes into the sandbox instead, and needs the limit out
of the way.
**Any agent, over MCP**
```ts
import { createServer } from 'node:http'
import { handsMcpHttpHandler } from '@scitix/agentbox-hands/mcp/http'
createServer(handsMcpHttpHandler()).listen(8766)
```
Every request must carry the conversation's own stable id in `X-Hands-Session`.
The MCP transport's session id looks right and is not: it changes on reconnect
(the conversation quietly moves to a new sandbox) and is shared when one
connection serves several conversations (they end up on one filesystem).
## Where the code is
The package is in [`sdk/hands/`](https://github.com/scitix/Agent-Sandbox/tree/develop/sdk/hands),
with Claude Agent SDK, OpenCode and MCP bindings and its own README — which is
more precise about the open questions than this page can be.
**The platform's own assistant is the reference implementation**, and it runs
this architecture end to end: the console's floating assistant is a brain in
the browser, a gateway, and a sandbox per conversation.
| Piece | Where |
|---|---|
| The brain's instructions | [`brain/AGENTS.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/brain/AGENTS.md) — the prompt a sandbox-backed agent runs with |
| The gateway that drives the harness | [`brain/gateway/`](https://github.com/scitix/Agent-Sandbox/tree/develop/brain/gateway) |
| The conversation UI and its tool cards | [`dashboard/components/assistant-ui/`](https://github.com/scitix/Agent-Sandbox/tree/develop/dashboard/components/assistant-ui) |
| The session → sandbox binding | [`sdk/hands/python/agentbox_hands/`](https://github.com/scitix/Agent-Sandbox/tree/develop/sdk/hands/python/agentbox_hands) |
| The pool its sandboxes come from | [`installer/helm/agent-sandbox-hub/values.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/helm/agent-sandbox-hub/values.yaml) — `assistant.sandbox` names the Env and the image |
Reading those five together is the shortest path from "I have a harness" to "my
harness's tools act on a sandbox": the prompt shows what the brain is told, the
gateway shows the loop, and the hands package shows the tool layer underneath.
## The patterns behind it
The split is not ours to invent — it is where the ecosystem has been heading,
and these are the pieces worth reading alongside it:
- **OpenAI, [Agents SDK](https://openai.github.io/openai-agents-python/)** —
agents as a loop with tools, handoffs and guardrails.
- **Anthropic, [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents)** —
when a workflow beats an agent, and why tools are the interface.
- **Anthropic, [Writing effective tools for agents](https://www.anthropic.com/engineering/writing-tools-for-agents)** —
tool shape is prompt engineering; the seven tools here are the result of that
thinking applied to a filesystem and a shell.
- **Anthropic, [Code execution with MCP](https://www.anthropic.com/engineering/code-execution-with-mcp)** —
the argument for putting tool execution *somewhere else* rather than in the
agent's own process, which is the premise of this page.
- **[Claude Agent SDK](https://docs.claude.com/en/api/agent-sdk/overview)** and
**[Model Context Protocol](https://modelcontextprotocol.io/)** — the two
integration surfaces the bindings target.
## See also
- [`abx-managed-agent`](/docs/skills/abx-managed-agent) — the same material written for an agent
- [E2B Python SDK](/docs/tutorials/e2b) — how a sandbox is created, which is what the daemon does
- [In-place update](/docs/concepts/inplace-update) — why handing a session its own sandbox is cheap
---
# Harbor benchmarks (/docs/tutorials/evals/harbor)
[Harbor](https://github.com/harbor-framework/harbor) already knows how to drive
a benchmark: datasets, agents, verifiers, results. `agent-sandbox-harbor` is an
environment plugin that changes one thing — where the sandbox comes from. Each
task claims one from a warm pool instead of building and starting one of its
own, and that difference is most of the wall-clock time of a normal run.
There is no Harbor fork to maintain and no template build step: Agent Sandbox
swaps the workload image into a Pod that already exists, so a task starts with
one API call.
## Before you start
You need a platform with somewhere to run, which is two objects: an environment
and a member pool with idle capacity. The [installation
guide](/docs/installation) covers the platform, and
[Examples](/docs/examples/templates) has templates and environments to apply if
you have none yet.
Then read the environment's own documentation — it is the only place that knows
which endpoints *your* cluster answers on:
```bash
abx envs YOUR_ENV docs --cluster YOUR_CLUSTER
```
Use the public entry it offers; the in-cluster one is for a client running
inside the cluster.
## Install
```bash
uv pip install 'harbor[e2b]' agent-sandbox-harbor
```
The plugin attaches through Harbor's own `--environment-import-path`, so the
Harbor you already have stays the Harbor you have.
## Size the pool for the concurrency
`-n` is how many tasks run at once, and every one of them needs a Pod that is
**already idle** when the run starts. The number that matters is not the pool's
`replicas` but its `idleReplicas`:
```bash
abx envs YOUR_ENV pools --cluster YOUR_CLUSTER # idleReplicas is the real answer
abx scale envs YOUR_ENV pools YOUR_POOL --replicas 16 --cluster YOUR_CLUSTER
```
A pool that is short does not fail — the run just sits waiting for sandboxes to
come back, which looks like a slow model and is not. If the pool will not grow,
that is quota or the autoscaler: see [Autoscaling](/docs/concepts/autoscaling).
## Run it
The environment file is where the credentials and endpoints go:
```bash
cat > agentbox.env <<'EOF'
E2B_API_KEY=agbx_…
E2B_API_URL=https://YOUR_GATEWAY/agent-sandbox/api/e2b
E2B_DOMAIN=YOUR_GATEWAY/agent-sandbox/api/data
AGBX_POOL_NAME=YOUR_POOL
AGBX_CLUSTER_ID=YOUR_CLUSTER
AGBX_IMAGE_PREFIX=registry.example.com/agent-sandbox
EOF
harbor run \
-d terminal-bench@2.0 -a oracle -n 16 -y \
--environment-import-path agent_sandbox_harbor:AgentSandboxEnvironment \
--env-file agentbox.env
```
`E2B_DOMAIN` carries no scheme — the plugin adds it, and `AGBX_HTTPS=false` is
how you say the data plane is plain HTTP. A mismatch there is worth watching
for: it shows up as "the sandbox never connects", never as a scheme error.
## Images are the part that bites
Every task needs a **pre-built image**. This environment does not build from a
Dockerfile and does not mutate a running sandbox, so an image is chosen in
exactly this order:
1. **`AGBX_IMAGE_MAP`** — a file of ` ` lines, used verbatim.
This is how a dataset whose tasks have no `docker_image` at all gets run,
which is the case for **SWE-bench**, where the task *is* a Dockerfile.
2. **the task's own `docker_image`** — Terminal-Bench's case, rewritten by
`AGBX_IMAGE_PREFIX` (with `docker.io/` stripped first) and `AGBX_IMAGE_TAG`.
3. neither → **the task is rejected**, loudly and on purpose.
So a SWE-bench run is really two jobs: mirror or build the images once and write
the map file; then run Harbor against that map. Budget for the first — it is
minutes to hours, and it is done once per dataset version.
**A dataset with no images is not a slow run, it is a rejected one**
> Every task is rejected before Harbor starts it if neither the map nor the task
> names an image. `AGBX_IMAGE_MAP` is what makes SWE-bench runnable at all, and
> writing it is the step people skip because the run otherwise looks like it is
> starting fine.
## Settings worth knowing
| Variable | Why you would touch it |
|---|---|
| `AGBX_STARTUP_TIMEOUT` | default 300s. Raise it for heavy images: this is the wait for a claimed Pod to become ready. |
| `AGBX_READY_TIMEOUT` | default 600s. A cold SWE-bench image can exceed it on a first pull. |
| `AGBX_IMAGE_PREFIX` | points every `docker.io/…` image at a mirror, which is what keeps a 500-task run from being rate-limited. |
| `AGBX_HTTPS` | `false` when the data plane is plain HTTP. |
One version note, so you do not blame the platform: E2B's SDK ≥ 2.24 rejects
keys that do not start with `e2b_` on the client side. `agent-sandbox-e2b >=
0.0.4` neutralises that so an `agbx_` key works, and `harbor >= 0.13` pulls a
new enough E2B SDK to need it.
## When tasks fail, and not the run
```bash
abx sandboxes --filter status=Failed --cluster YOUR_CLUSTER
abx sandboxes YOUR_SANDBOX_ID logs --cluster YOUR_CLUSTER
abx envs YOUR_ENV events --cluster YOUR_CLUSTER
```
A whole dataset failing the same way is almost always the image map or the
registry; individual tasks failing is usually the task itself.
## See also
- [E2B Python SDK](/docs/tutorials/e2b) — the client the plugin uses underneath
- [Autoscaling](/docs/concepts/autoscaling) — making the concurrency exist
- [In-place update](/docs/concepts/inplace-update) — why a claim is fast, and
what it costs
---
# mini SWE Agent (/docs/tutorials/evals/mini-swe-agent)
Scitix AgentBox is a sandbox service designed for Agentic scenarios. It provides secure isolation and flexible deployment capabilities, catering to requirements such as inference evaluation and training rollouts.
For SWE-Agent inference evaluation/streaming scenarios, we provide a pre-packaged **mini-SWE-Agent** that connects directly to the AgentBox sandbox resource pool, eliminating the need for manual sandbox lifecycle management.
---
## Installation
Install the mini-SWE-Agent compatible with Scitix AgentBox.
---
## Quick Start
Refer to the aforementioned documentation to apply for an API Key and create a sandbox warm-up pool based on the SWE template. Then, set the following environment variables:
```bash
export SCITIX_API_KEY="${AGBX_API_KEY}" # Please apply for the API Key via the platform
export SCITIX_POOL_NAME="${AGBX_ENV_NAME}" # The sandbox env to allocate from (a pool name also works)
```
Select the corresponding SWE-Bench image registry based on your current cluster:
```bash
export SWEBENCH_REGISTRY="docker.io/swebench"
export SWEBENCH_IMAGE_TAG="latest"
```
Then, run `mini-extra swebench`:
```bash
mini-extra swebench \
--subset verified \
--split test \
-m openai/zai-org/GLM-4.7 \
-c swebench \
-c swebench_scitix \
-c "environment.idle_timeout=30m" \
-w 10 \
-c "model.model_kwargs.api_base=XXXXXXXX" \
-c "model.model_kwargs.api_key=XXXXXXXX" \
-c "model.cost_tracking=ignore_errors" \
-c "agent.step_limit=1000000" \
-c "agent.cost_limit=1000000"
```
Once running, logs related to Scitix Sandbox creation should appear, indicating successful execution.
---
## Parameter Description
### `-c` Configuration Options
| Parameter | Meaning |
|---|---|
| `-c swebench` | Uses official SWE-Agent default configuration; **must be included** |
| `-c swebench_scitix` | Uses Scitix custom initialization configuration (sets Idle Timeout, Startup Timeout, etc.) |
| `-c "environment.idle_timeout=30m"` | Overrides the default Idle Timeout (default is 5m); set to an appropriate value |
**Order matters**
> `-c swebench_scitix` must be added *after* `-c swebench`. Do not replace the
> original `-c swebench`.
### Worker Count
The `-w` parameter specifies the number of concurrent workers. It is recommended to keep this consistent with the size of the warm-up pool to avoid frequent cold starts.
---
## Cross-Cluster Usage
If the sandbox env is in a different cluster, prefix `SCITIX_POOL_NAME` with the cluster ID in the format `clusterId::name`, where `name` is an env name (preferred — the receiving cluster then picks a member pool for you) or a concrete pool name:
```bash
export SCITIX_POOL_NAME="${AGBX_CLUSTER_ID}::${AGBX_ENV_NAME}"
```
**Cross-cluster**
> Cross-cluster requests are forwarded by the AgentBox control plane through the
> gateway. Authentication is the same as within the local cluster, with no extra
> configuration; if you get an authentication error, re-issue the API key on the
> platform.
---
## FAQ
**Q: No Scitix Sandbox creation logs appear after running?**
Check if `-c swebench_scitix` has been added and ensure that the three environment variables `SCITIX_ENDPOINT`, `SCITIX_API_KEY`, and `SCITIX_POOL_NAME` are set correctly.
**Q: Sandboxes are frequently timing out or being reclaimed?**
The default `idle_timeout` is 5 minutes. If a single episode in your streaming task exceeds this duration, the sandbox will be reclaimed prematurely. It is recommended to adjust this using `-c "environment.idle_timeout=30m"`.
**Q: Receiving "no idle sandbox" errors during concurrent tasks?**
There are insufficient available sandboxes in the warm-up pool. Check if the number of replicas configured for the warm-up pool is equal to or greater than the number of workers specified by `-w`. Consider expanding the warm-up pool or reducing the concurrency.