# Introduction (/docs) ## What is Agent Sandbox? **Agent Sandbox** is an open-source sandbox engine for AI agents. It is purpose-built for three classes of workload: - **Lightning Fast** — pre-warmed pools keep isolated environments on standby, eliminating cold-start latency for high-frequency agent loops, evaluations, and RL rollouts - **Enterprise Grade** — deploy on any cloud using native Kubernetes CRDs, RBAC, and multi-cluster routing, without vendor lock-in - **Agentic RL** — stateful environments with deterministic resets and any-image runtimes, built for complex multi-turn agent training --- ## Key Features | | Feature | Description | |---|---------|-------------| | ⚡ | **Speed — Sub-60ms allocation** | Pre-warmed pools deliver idle sandboxes instantly, unblocking high-volume agent loops and multi-turn RL rollouts | | ☸️ | **Infrastructure — Containers or microVMs** | Run on your existing estate using CRDs, namespaces, RBAC, and autoscaling to manage warm capacity efficiently | | 🌐 | **Routing — Cross-region and cross-cloud** | Dispatch requests across clouds, clusters, and regions without forcing application teams to manage routing logic | | 🧪 | **Runtime — Zero-rebuild runtimes** | Run any Docker image for SWE tasks, RL environments, and internal tools without building custom VM images | | 🔌 | **Ecosystem — Drop-in agent SDKs** | Seamless compatibility with E2B clients, SWE-ReX workflows, and popular reinforcement learning frameworks | | 📊 | **Observability — Console-grade visibility** | Complete view of pools, active sessions, logs, and metrics through a unified product console | --- ## Use Cases ### Reinforcement Learning at Scale RL training requires thousands of environment resets per hour. Agent Sandbox pre-warms a pool of sandboxes so each rollout worker gets a fresh, isolated environment in milliseconds — removing the environment-reset bottleneck from your training loop. Supports SWE-bench Verified, SWE-Gym, Terminal-bench, and custom task distributions. ### AI Coding Agents & Evaluations Give every agent turn or eval call its own isolated execution environment. The E2B-compatible API means existing SWE-agent, SWE-ReX, and similar frameworks work without modification. ### Enterprise Multi-Cluster Deployment Deploy sandbox pools across multiple clouds or regions. The built-in ExtProc routing layer dispatches requests to the most available cluster transparently — no routing logic required in application code. Supported cloud providers: AWS, Google Cloud, Azure, Alibaba Cloud, Volcengine, Cloudflare. **Isolation** > Sandboxes are containers by default. Where the code is untrusted, run them in > their own kernel instead — a microVM runtime class is one field on the template: > [E2B Kata](/docs/examples/templates/e2b-kata). ## Start here - [Installation](/docs/installation) — the platform on a cluster, then the CLI and its skills - [Examples](/docs/examples/templates) — templates and environments to apply - [Concepts](/docs/concepts) — what a template, an env and a pool are - [CLI guide](/docs/tutorials/cli) — driving the platform from a shell or an agent --- # Installation (/docs/installation) Agent Sandbox is two halves, and most people install both: - the **platform** — an operator, an API and the data plane that runs the sandbox Pods — which goes on a Kubernetes cluster, with Helm; - the **client** — the `abx` CLI and the skills an agent reads — which goes on your machine, or in the image your agent runs in. The [CLI guide](/docs/tutorials/cli) and the [concepts](/docs/concepts) assume both are done. ## What you need | | | |---|---| | A Kubernetes cluster | 1.28 or newer; the charts install CRDs and cluster-scoped RBAC, so you need cluster-admin | | `kubectl` | pointed at that cluster | | Helm | 3.14 or newer — the charts are published as OCI artifacts | Every image is public on GHCR, so there is no registry login to perform. The charts default to `latest`; pin the chart version and the image tags for anything you intend to keep. ## 1. The platform ```bash helm upgrade --install agent-sandbox-worker \ oci://ghcr.io/scitix/agent-sandbox-worker \ --version 0.1.0 \ --namespace agentbox-system --create-namespace \ --set controller.localClusterId=YOUR_CLUSTER_ID ``` One release is the whole platform on one cluster: | Installs | What it is | |---|---| | three CRDs | `SandboxTemplate`, `SandboxEnv`, `SandboxPool` — [the object model](/docs/concepts/index) | | the controller | the operator that renders environments and pools, and the API `abx` talks to | | ExtProc + Envoy | the data plane, the gateway a sandbox's ports are published through | `controller.localClusterId` names this cluster and is **required**: it is how a member Pool is placed on a cluster segment, and how the console labels the cluster it is showing. Left empty, Pool writes answer `503 server-misconfigured` and the environment reconciler has nowhere to put a member. **Pin the version** > `--version 0.1.0` is the chart; the images inside it default to `latest`. For > anything you keep, pin both — `--set controller.image.tag=…`, > `--set extproc.image.tag=…` — so a rollout is something you chose. ### Check it came up ```bash kubectl -n agentbox-system get pods kubectl get crd | grep agents.navix.sh ``` ### Reach it Everything a client needs is a Service in that namespace: | Service | Port | What it is | |---|---|---| | `agent-sandbox-api` | 80 | the native API — what `abx` talks to | | `agent-sandbox-e2b-api` | 80 | the E2B-compatible API | | `agent-sandbox-data-plane` | 80 | the gateway that sandbox ports are reached through | From outside the cluster, either publish the ingress the chart ships (`--set extproc.ingress.host=agentbox.example.com`, with an ingress controller installed) or, for a first look, forward one port: ```bash kubectl -n agentbox-system port-forward svc/agent-sandbox-api 8080:80 ``` ## 2. The console (optional) The console is a separate release: a dashboard that reads the same API. Install it in the same cluster, or in another one that can reach the worker. ```bash helm upgrade --install agent-sandbox-hub \ oci://ghcr.io/scitix/agent-sandbox-hub \ --version 0.1.0 \ --namespace agentbox-system \ --set env.secret="$(openssl rand -hex 32)" \ --set clusters[0].id=YOUR_CLUSTER_ID \ --set clusters[0].name="YOUR_CLUSTER" \ --set clusters[0].url=http://agent-sandbox-api.agentbox-system.svc.cluster.local \ --set ingress.host=agentbox.example.com ``` Two settings matter more than the rest: - **`env.secret`** is the shared secret between the console and the worker (`controller.secrets.secret`). Generate it once and give both releases the same value; the console cannot authenticate against a worker that does not share it. - **`clusters[]`** is the list of workers the console knows about. A console only knows the clusters configured here — that is the list its cluster picker shows, and the list `abx clusters` will print if you point `abx` at the console's address. The console's built-in assistant is off by default (`assistant.enabled`), and its images are not part of the public release. Everything else — environments, pools, autoscaling, quotas, templates, the vault — is the same API the CLI uses. ## 3. The CLI, and the skills `abx` is a single binary with no runtime dependencies, and it ships nine skills — Markdown files an agent reads before it touches the platform: ```bash curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh | sh ``` That writes `~/.local/bin/abx` and `~/.agents/skills/`. Then name the deployment you are talking to: ```bash abx context set YOUR_DEPLOYMENT \ --endpoint 'https://YOUR_CONSOLE/agentbox' \ --api-key agbx_... abx clusters ``` The endpoint is the console's address, or a worker's API address if you are not running the console; the key is issued in the console under **API Keys** (or, on a worker with no console, by whoever created the release). The [CLI guide](/docs/tutorials/cli) takes it from there, and [Skills](/docs/skills) explains what was installed for your agent. ## First sandbox An environment is created from a **template**, and a sandbox from an environment's **pool** — nothing exists to create sandboxes from until you add both. The repository carries three to start with, each with a matching environment: ```bash kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml # the template kubectl apply -f config/samples/e2b_sandboxenv.yaml # a warm pool on it ``` [Examples](/docs/examples/templates) has the other two — Docker inside the sandbox, and a microVM-isolated one — with the manifests rendered on the page. From there the [CLI guide](/docs/tutorials/cli#5-run-code-in-a-sandbox) takes over: read the environment's own documentation for the endpoints of your cluster, then create a sandbox. The [concepts](/docs/concepts) explain what the two objects are for, and why a template is not the thing you pass to `Sandbox.create()`. ## Upgrading and removing ```bash helm upgrade agent-sandbox-worker oci://ghcr.io/scitix/agent-sandbox-worker \ --namespace agentbox-system --reuse-values --version 0.1.1 helm uninstall agent-sandbox-worker --namespace agentbox-system ``` The CRDs carry `helm.sh/resource-policy: keep`, so `uninstall` leaves them — and every environment, pool and template you created — in place. That is deliberate: deleting a CRD deletes its objects, and with them the record of what your sandboxes were. Remove them explicitly when you mean it: **Uninstall does not uninstall everything** > `helm uninstall` removes the controller and the data plane. The CRDs, and every > `SandboxTemplate`, `SandboxEnv` and `SandboxPool` on the cluster, stay until you > delete the CRDs themselves — which is the command below, and is irreversible. ```bash kubectl delete crd sandboxtemplates.agents.navix.sh sandboxenvs.agents.navix.sh sandboxpools.agents.navix.sh ``` --- # Autoscaling (/docs/concepts/autoscaling) Autoscaling answers two questions about a set of [pools](/docs/concepts/pools): *how few Pods may I keep when nothing is running*, and *how many may I add when everything is busy*. ## Groups, not pools Pools of the same resource shape inside one env form a **scaling group**, and the policy lives on the group. The group's name is derived from the shape (`4c64gi`-style), which is why you never create one: a group appears when a member pool declares its shape, and is collected when the last one goes away. ```bash abx envs YOUR_ENV scaling-groups --cluster YOUR_CLUSTER abx envs YOUR_ENV scaling-groups YOUR_GROUP --cluster YOUR_CLUSTER ``` ## The policy | Field | What it means | |---|---| | `enabled` | whether the autoscaler acts on this group at all; off means the pool's `replicas` is yours to set | | `minReplicas`, `maxReplicas` | the floor and the ceiling on the group's aggregate replicas | | `scaleUpPolicy.mode` | `Conservative`, `Default` or `Aggressive` — the step size in both directions | | `scaleUpPolicy.cooldownSeconds` | minimum gap between two scale-ups | | `scaleUpPolicy.idleThresholdSeconds` | how long aggregate idle must stay at zero before a proactive scale-up | | `scaleUpPolicy.idleZeroQuietWindowSeconds` | suppresses that trigger when nothing has been claimed for this long | | `scaleUpPolicy.saturationCooldownSeconds` | how long a member stays marked saturated after a failed probe | | `scaleDownPolicy.idleTimeoutSeconds` | how long a Pod idles before it counts as removable | | `scaleDownPolicy.stabilizationSeconds` | minimum gap between two scale-downs | | `scaleDownPolicy.protectionWindowSeconds` | how long a marked Pod can still be claimed, cancelling its deletion | Every member pool may also carry its own `minReplicas` and `maxReplicas`, which is how one shape is held larger than its siblings inside the same group. ## Reading the policy against the two questions **Cost.** `minReplicas = 0` plus a short `idleTimeoutSeconds` is "keep nothing warm"; `minReplicas = N` is "always have N claims ready", and it is a floor on spend as much as on capacity. **Concurrency.** The ceiling is `maxReplicas` — and, separately, your quota. Quota is a hard cap the autoscaler cannot argue with: a group may sit below its ceiling indefinitely while the pool reports `ResourceQuotaExhausted`. Check both before promising a number of parallel sandboxes. **Speed.** Scaling up creates Pods, and a Pod that has just been created still has to start. The autoscaler removes the wait for *capacity*; it does not make a brand-new Pod faster than a pre-warmed one, which is why a floor above zero is what keeps first-claim latency flat. ## Changing one ```bash abx envs YOUR_ENV scaling-groups YOUR_GROUP --editable --cluster YOUR_CLUSTER > group.json # edit group.json abx update envs YOUR_ENV scaling-groups YOUR_GROUP -f group.json --cluster YOUR_CLUSTER ``` Like every write, this is a PUT: the file is the whole policy. A field left out is a field you are asking to remove — dropping `maxReplicas` removes the ceiling rather than keeping it. ## When it does not behave | Symptom | Usual cause | |---|---| | never scales down | running sandboxes, `idleTimeoutSeconds`, or the protection window still counting | | scales down under load | `idleTimeoutSeconds` shorter than the gap between real claims | | scales up in bursts | `idleThresholdSeconds` / `idleZeroQuietWindowSeconds` too eager for the traffic pattern | | sits under `minReplicas` | quota, or the instance type has no room on the cluster | | a rolled-out pool refuses claims | the previous generation is still draining; `maxUnavailable` on the env governs how fast that is | ## See also - [Pools](/docs/concepts/pools) — what is being scaled - [Envs](/docs/concepts/envs) — the rollout policy that governs rolling replacements - [Capacity and quotas](/docs/tutorials/cli) — reading quota from the CLI --- # Cross-cluster (/docs/concepts/cross-cluster) One Agent Sandbox address reaches a platform, and a platform can reach several clusters. You do not open a second account to use the second cluster; you address it. ```bash abx clusters # every cluster behind this address abx envs --cluster YOUR_CLUSTER # the envs on one of them ``` ## Same name, several clusters An env exists per cluster — pools, quota, instance types and images are local to one — but two envs with the **same name** on different clusters are associated, and the platform may serve a claim for that name from either. To build one: create the env on the first cluster, then extend it to the next one from the console (**Sandbox Envs → extend an environment to another cluster**). The extension creates the env there; you then add member pools on that cluster, because instance types and quota are properties of the cluster it lives on, and registry credentials are entered per cluster. Two conditions, and the second fails quietly: 1. The image must exist in each cluster's registry — [images follow you](#the-same-image-in-every-region). 2. The team → namespace mapping must match across clusters. If it does not, federation cannot pair `(namespace, env name)` and a bare env name will not spread — name the cluster explicitly instead. ## Addressing The E2B SDK's first argument is an address, and the address can carry a cluster: | You write | You get | |---|---| | `YOUR_ENV` | this cluster first; if it has no idle capacity and cannot scale, a cluster with the same-named env | | `SOME_CLUSTER::YOUR_ENV` | that cluster's env, which then routes across its own member pools | | `SOME_CLUSTER::YOUR_POOL` | that exact pool, with no routing at all | | `YOUR_ENV//IMAGE` | as above, with the main container image replaced | The `//IMAGE` suffix combines with any of the others. For a pipeline that should not care where capacity is, `YOUR_ENV` is the right spelling; for a result you may have to reproduce, `SOME_CLUSTER::YOUR_ENV` pins the cluster and still survives a pool being replaced by one of the same shape. In `abx`, a cluster is a flag rather than part of the address: `abx envs YOUR_ENV --cluster SOME_CLUSTER`. ## The same image in every region When you name an image that belongs to another cluster's private registry, the platform rewrites the **registry host** to this cluster's equivalent — same type of registry, same path, same tag — so a sandbox is never pulling across regions: ``` you write: registry-REGION-A.example.com/agentbox/eval/thing:260328 pulled from: registry-REGION-B.example.com/agentbox/eval/thing:260328 ``` The rules that follow from it: - the same image has to exist at the **same path** in each region's registry; - rewriting only happens between registries of the same configured type; - public registries (`docker.io` and friends) are never rewritten; - if this cluster has no registry of that type, the address is used as written — which may work, and may be slow, and says nothing either way. ## What is shared and what is not | Shared across the platform | Per cluster | |---|---| | the template catalogue (synced) | pools, their replicas and their autoscaling | | an env's identity and name | the env's member pools on that cluster | | API keys | quota, instance types, namespaces | | | registry credentials, and the images behind them | ## Common misconceptions - **"A pool spans clusters."** It does not. A pool belongs to one cluster; the env is what spans them. - **"Cross-cluster is a fallback for outages."** It is capacity routing, not failover: a cluster that has no Pod to give you cannot be scaled into existence instantly, and the request fails rather than waiting for one. - **"The same env name is enough."** Without the env on the target cluster, the request has nowhere to go; without the image, it lands and then fails to pull. ## See also - [Envs](/docs/concepts/envs) — naming and per-cluster membership - [Pools](/docs/concepts/pools) — where a cluster's capacity lives - [E2B Python SDK](/docs/tutorials/e2b) — create parameters and routing --- # Egress and secrets (/docs/concepts/egress-and-secrets) Two decisions, and conflating them is the usual mistake: ``` SandboxEnv does this environment HAVE a gateway one switch create call what THIS sandbox may reach, and what per sandbox may be injected into its outbound requests ``` The env carries a switch; the rules belong to an individual sandbox and arrive with the create call, in the E2B SDK's own vocabulary. There is no separate Agent Sandbox dialect for them. ## Why the switch is on the env and the rules are not Enabling the gateway injects a transparent proxy sidecar. That changes the Pod spec, and a changed Pod spec rolls the env's pools — a real, environment-level decision, taken once. Rules are per sandbox because an env is shared: an env-wide allowlist would be a default every caller overrides anyway. **It fails closed.** A create that carries filtering rules against an env with no gateway is refused with `400`, not accepted and ignored. A Pod with no sidecar has no redirection, so accepting the rules would mean they silently did nothing — which for an evaluation is the worst outcome, because the run finishes and the numbers are wrong. ```bash abx update envs YOUR_ENV --help --cluster YOUR_CLUSTER # the env's one gateway field ``` ## Cutting a sandbox off, and letting one thing through The common case is an agent under test that must not fetch the answer and must not install its way around a missing dependency. The shape is always the same: **deny everything, then allow what the task genuinely needs.** A denylist is a list of the routes somebody thought of. ```python # nothing in, nothing out sbx = Sandbox.create("YOUR_ENV", timeout=3000, secure=False, allow_internet_access=False) # or: everything out except these sbx = Sandbox.create( "YOUR_ENV", timeout=3000, secure=False, network={"allow_out": ["api.openai.com", "pypi.org", "*.pythonhosted.org"]}, ) ``` Naming an allow list is what makes everything else a deny — the entries are the whole of what is reachable. A CIDR or a bare IP works in the same list. **An empty allow list is not a deny** > `network={"allow_out": []}` declares no filtering at all, and egress stays > unrestricted. To cut a sandbox off, say so: `allow_internet_access=False`, or > `deny_out=["0.0.0.0/0"]`. Three things worth checking before calling a run isolated: 1. **The env has a gateway.** Without it, a create carrying rules is refused — so make sure you saw a sandbox, not a `400`. 2. **The package index is not on the allowlist** unless the task needs it. It is the most common accidental hole: the agent cannot search, but it can `pip install` something that can. 3. **Watch it fail from inside.** ```python sbx.commands.run("curl -sS -m 5 https://example.com") # expected: it fails ``` An isolation you have not seen fail is an isolation you are assuming. ## Credentials a sandbox can use but cannot read ``` vault (write-only) → operator memory → sidecar tmpfs → outbound header ``` The value never enters the sandbox. The sandbox holds a **decoy** — a placeholder that looks like a token — and the egress sidecar substitutes the real one as the request leaves, for the hosts and headers the rule names. Code inside runs unmodified: it reads `OPENAI_API_KEY`, sends it, and the sidecar replaces it on the way out. Library code that has never heard of Agent Sandbox works. Secrets live in a **vault** that is write-only: you can list names, overwrite them and delete them, never read them back. In the console it is the **Vault** page; through the SDK it is the E2B `/secrets` surface, which the official package speaks as `e2b.Secret` (that module arrived in **e2b 2.43**). ### Store it once ```bash abx envs YOUR_ENV docs --cluster YOUR_CLUSTER # API URL and data-plane domain ``` ```python import os os.environ["E2B_API_KEY"] = "agbx_..." # your platform key os.environ["E2B_API_URL"] = "https://YOUR_GATEWAY/agent-sandbox/api/e2b" os.environ["E2B_DOMAIN"] = "YOUR_GATEWAY/agent-sandbox/api/data" # no scheme from agent_sandbox_e2b import patch_e2b # before the e2b import patch_e2b() from e2b import Secret Secret.create("openai-api-key", os.environ["OPENAI_API_KEY"]) # write-only # rotate it later with Secret.update("openai-api-key", new_value) ``` ### Use it at create time Reference it **by name**; the value never leaves the vault. `Secret.fill()` is a local formatting helper — it returns the string the wire carries, and makes no call of its own: ```python from e2b import Sandbox, Secret sbx = Sandbox.create( "YOUR_ENV", timeout=3000, secure=False, # What the code inside reads. It is a decoy: the real value is put in on the # way out, so the sandbox cannot leak what it never had. envs={"OPENAI_API_KEY": "decoy-not-a-real-key"}, network={ # A rule is a transform, not a permission — the host has to be allowed # too, or the request never leaves. "allow_out": ["api.openai.com"], "rules": { "api.openai.com": [ { "transform": { "headers": { "Authorization": f"Bearer {Secret.fill('openai-api-key')}", } } } ] }, }, ) ``` Ordinary code inside, no Agent Sandbox dialect: ```python sbx.commands.run( 'curl -s https://api.openai.com/v1/models -H "Authorization: Bearer $OPENAI_API_KEY"' ) # succeeds sbx.commands.run("echo $OPENAI_API_KEY") # prints the decoy, not the key sbx.commands.run("curl -sS -m 5 https://example.com") # fails: not allowed sbx.kill() ``` A plaintext credential in the rules is refused with `400` — deliberately, because accepting it would put the value in the request body, the access log and the caller's source, which is the exposure the feature exists to remove. Wildcard hosts are refused for the same reason: whoever controls a matching subdomain would receive the injected credential. ### The env has to have the gateway All of the above is refused on an Env whose gateway is off, with a `400` that says so: ``` a network policy was requested but this environment has no egress gateway; enable it on the SandboxEnv (overrides.gateway.enabled) and let its pools roll ``` That is the switch from the top of this page, and turning it on is a change to the Pod spec, so it rolls the Env's pools. ## When it silently does not work The sandbox runs, the request goes out, and nothing is substituted. Check in this order: | Check | Why | |---|---| | the env has the gateway on | without it, a create with rules is refused — so if a sandbox exists, this one is satisfied | | the host matches the rule exactly | a redirect to another host is not covered | | the port is 80 or 443 | only those are parsed at layer 7; rules for other ports never fire and nothing says so | | the secret name exists in **your** vault | a name that resolves to nothing leaves the decoy in place, and the upstream answers `401`, which reads like a bad key | `abx whoami` tells you which identity the vault is read as; a secret stored by one user is not visible to another. ## Agents, vaults and approval Writing a vault secret **is** allowed to an agent credential, under the normal approval gate: the credential lands in the acting person's own vault and widens nobody's authority. Minting an Agent Sandbox API key is the act that is refused. See [API key permissions](/docs/tutorials/cli#6-api-key-permissions). ## See also - [Envs](/docs/concepts/envs) — the gateway switch and the rest of the settings - [Pools](/docs/concepts/pools) — why enabling the gateway rolls them - [In-place update](/docs/concepts/inplace-update) — what a claim does to a Pod --- # Sandbox environments (/docs/concepts/envs) A **SandboxEnv** is the addressable unit: the name you pass to `Sandbox.create()`, the thing the console shows a page for, and the object that binds exactly one [template](/docs/concepts/templates). ```python sbx = Sandbox.create("YOUR_ENV", timeout=3000, secure=False) ``` An env is one class of runtime — "the E2B sandboxes for this team", "the Docker-in-Docker ones for this evaluation" — and it fans out to member [pools](/docs/concepts/pools) that hold the actual capacity. It holds no Pods itself. ## What an env owns These are the settings every sandbox claimed from the env inherits: | Setting | What it does | |---|---| | `overrides.image` | replaces the template's main container image for every member pool | | `overrides.podCreationImagePolicy` | whether a new Pod starts on the template's image or on the idle image | | `overrides.defaultStartupTimeout` | how long a claim waits for the sandbox to become ready, when it does not say | | `overrides.defaultIdleTimeout` | how long an untouched sandbox lives before it is reclaimed, when it does not say | | `overrides.imagePullSecret` | credentials for a private registry, materialised into a Secret the member pools reference | | `overrides.gateway` | the **switch** that gives the env's Pods an egress sidecar — see [egress and secrets](/docs/concepts/egress-and-secrets) | | `overrides.volumes` | PersistentVolumeClaims mounted into every sandbox | | `overrides.updateStrategy` | `autoUpdate` and `maxUnavailable` for rollouts | | `labels`, `annotations` | metadata stamped on the env and its member pools | Changing most of these changes the Pod spec, and a changed Pod spec means the env's pools **roll**: existing idle Pods are replaced, one rollout at a time, and running sandboxes are left alone until they are returned. ## Mode | Mode | Behaviour | |---|---| | `WarmPool` | claims are served from the env's member pools. This is what the platform runs. | | `OnDemandJob` | reserved: accepted by the API enum, not implemented by the controller. Do not build on it. | ## The name goes everywhere An env name is an RFC 1123 DNS label, capped at 24 characters because pool and Pod names are derived from it (`POOL = ENV + resourceKey (+ quotaShort)`, `POD = POOL + uuid`) and those have to stay inside the 63-character limit. The same name is the env in the E2B SDK, the console URL, and `abx`. Create the same-named env in two clusters and they federate — [cross-cluster](/docs/concepts/cross-cluster). ## Reading and writing one ```bash abx envs --cluster YOUR_CLUSTER # every env on the cluster abx envs YOUR_ENV --cluster YOUR_CLUSTER # one env: template, mode, pools, sizing rule # the safe way to change one abx envs YOUR_ENV --editable --cluster YOUR_CLUSTER > env.json # edit env.json abx update envs YOUR_ENV -f env.json --cluster YOUR_CLUSTER ``` `abx create envs --help` prints the whole body, field by field. Two things about it are worth knowing before you edit a file: - **A write is a PUT.** A field you leave out is a field you are asking to remove, and `overrides` is replaced wholesale. - **`imagePullSecret` cannot be read back.** `GET` reports `imagePullSecretConfigured` instead of the credentials, and that flag is also the only way to say "keep them": a PUT carrying neither the secret nor the flag is asking for the stored credentials to be deleted. Editing an unrelated setting with a file that dropped the flag revokes your registry access. **The registry credentials are one flag away from being deleted** > `--editable` prints `imagePullSecretConfigured: true`, not the secret. Send that > file back unchanged and the credentials survive. Strip the flag while editing > something else and the PUT means "delete them" — the next image pull fails, and > nothing in the response says why. ## Common misconceptions - **"The env is a namespace."** It is not; the namespace comes from your team. An env is a runtime identity plus its settings. - **"An env holds capacity."** Pools do. An env with no member pool has no sandboxes to hand out, however it is configured. - **"I can repoint an env at another template."** `templateRef` is fixed at create. Moving runtimes means creating an env on the new template and moving traffic. ## See also - [Pools](/docs/concepts/pools) — where the capacity is - [Autoscaling](/docs/concepts/autoscaling) — the size of that capacity over time - [Egress and secrets](/docs/concepts/egress-and-secrets) — the env's gateway switch --- # The object model (/docs/concepts) Agent Sandbox is four objects and the chain between them. The console, `abx` and the E2B SDK are three ways of driving that same chain. ```mermaid flowchart TB T["SandboxTemplate
the Pod shape and runtime
platform admin"] E["SandboxEnv
the name you create sandboxes with
you"] P["SandboxPool
pre-warmed Pods of one resource shape
you"] S(["sandbox
your code, running"]) T -->|binds to one| E E -->|owns members| P P -->|a claim hands one over| S ``` | Object | What it decides | Who creates it | In the console | |---|---|---|---| | **SandboxTemplate** | the Pod shape, the idle image, default timeouts, an environment's own documentation | platform admin | Sandbox Templates | | **SandboxEnv** | which runtime, the timeouts, egress, the rollout policy | you | Sandbox Envs | | **SandboxPool** | how much warm capacity of one resource shape there is | you | the env's Pools tab | | **Pod → sandbox** | one running sandbox, handed to one claim | the platform | Sandboxes | ## The same word means different things Agent Sandbox runs on Kubernetes and serves the E2B API, so the two vocabularies meet in one place — and one word is genuinely misleading: | You say | E2B SDK | Agent Sandbox | |---|---|---| | template | a snapshot built from a Dockerfile | a **container image you bring**, plus the platform's `SandboxTemplate` that says how to run it — nothing is built | | the first argument of `Sandbox.create()` | the template | the **env** | | pool | — | a `SandboxPool`: warm Pods, which E2B has no equivalent of | | sandbox | `Sandbox` | a claimed Pod | Read the first two rows together: in the E2B SDK the argument is called `template`, and on Agent Sandbox the value you put there is an env name. A template is what an env is *made from*, not what you create sandboxes with — [Sandbox templates](/docs/concepts/templates) is about that difference, and about why there is no build step and no snapshot format here. ## Two facts the design follows from **A sandbox is claimed, not created.** Pools hold Pods that already exist, so serving a request is a swap rather than a build — which is where the sub-second start comes from, and what [in-place update](/docs/concepts/inplace-update) is about. **Capacity is per cluster; identity is per name.** Quota, pools and images belong to one cluster, and an env with the same name in two clusters is one env the platform may dispatch to: [cross-cluster](/docs/concepts/cross-cluster). ## Which page answers which question | Question | Page | |---|---| | What is a SandboxTemplate, and how does it differ from an E2B template? | [templates](/docs/concepts/templates) | | What can I configure on an environment? | [envs](/docs/concepts/envs) | | How many sandboxes can run at once, and what does one cost? | [pools](/docs/concepts/pools) | | Why is my pool not scaling, or scaling too far? | [autoscaling](/docs/concepts/autoscaling) | | Why is a claim fast, and what happens to the Pod afterwards? | [in-place update](/docs/concepts/inplace-update) | | How do I run the same workload on another cluster? | [cross-cluster](/docs/concepts/cross-cluster) | | How do I cut a sandbox off the network, or give it a credential it cannot read? | [egress and secrets](/docs/concepts/egress-and-secrets) | | What vocabulary does my agent have for all of this? | [skills](/docs/skills) | ## Driving it ```bash abx clusters # every cluster this address reaches abx envs --cluster YOUR_CLUSTER # the envs on one of them abx envs YOUR_ENV --cluster YOUR_CLUSTER # one env: its template, pools, sizing rule ``` Everything the console shows for these objects is readable from `abx`, and almost all of it is writable from `abx` too. The [CLI guide](/docs/tutorials/cli) covers the setup; the [E2B SDK guide](/docs/tutorials/e2b) covers the sandbox side. --- # In-place update (/docs/concepts/inplace-update) An in-place update is how a claim is served. A [pool](/docs/concepts/pools) holds Pods that already exist and already run a runtime; a claim swaps the container image on one of them instead of scheduling anything new. It is the mechanism behind the platform's start latency, and behind "run any image without building a template". ## The life of a Pod ```mermaid stateDiagram-v2 direction LR [*] --> Idle : the pool creates it Idle --> Running : a claim swaps the image in place Running --> Idle : killed, or the idle timeout expires Idle --> [*] : the pool is scaled down or rolled ``` The same three states, in the order they matter: **Idle — the Pod runs the template's `idleImage`.** That is the sandbox runtime sitting in front of no workload, which is what makes it cheap to keep Pods around: twenty idle Pods are not twenty copies of your image. **Claim — the image is swapped in place.** The platform picks an idle Pod and replaces the main container's image with the one the create asked for. Kubernetes restarts that container on the same Pod: no scheduling, no new Pod, no volume re-attach. **Running — your process starts**, and the sandbox is handed to the caller. **Return — the Pod goes back to idle.** When the sandbox is killed or times out, the container is restarted onto the idle image and the previous sandbox's metadata is cleared, so the next claim does not inherit it. ## What that buys, and what it costs | Buys | Costs | |---|---| | a claim does not wait for a scheduler, a node, or a volume | the container image still has to be pulled — first claim for a large image is slower | | any image, no per-workload build | Pods cannot change *shape*: requests, limits, volumes and sidecars are fixed for the Pod's life | | idle Pods are cheap, so pools can be deep | the image is chosen per claim, so two sandboxes from one pool may run different images | The practical consequences: - **Sizing is a pool decision, not a claim decision.** A different resource shape is a different pool, never a resized Pod. - **The image is resolved when you claim.** Edit the env's or template's image and the next claim uses it. Nothing is rebuilt, and nothing that is already running is disturbed — see [templates](/docs/concepts/templates). - **Warm Pods make claims fast; images make them slower.** If first-claim latency matters, that is an image-size question, and the pools holding that image's shape are the ones worth keeping warm. ## Three timeouts, three different things | Timeout | Counts | Where it is set | |---|---|---| | **Startup** | from the claim to the sandbox being ready — dominated by the image pull | the env's/template's default, or `agentbox.scitix.ai/startup-timeout` on the create | | **Idle** | from the last activity to the sandbox being reclaimed | `timeout=` on the create, `sbx.set_timeout(…)` to extend, the env's default otherwise | | **Pod lifetime** | not a user-facing timeout; a Pod is recycled by rollouts and by the pool | — | ## Rolling is not the same thing An in-place update changes one Pod's image to serve one claim. A **roll** is the platform replacing idle Pods because the Pod's identity changed — the idle image, the Pod body, the gateway, or the template's metadata. Rolls are governed by the env's `updateStrategy` (`autoUpdate`, `maxUnavailable`) and can be watched on a pool as `updateRevision`/`updatedReplicas`, with Pods counting through `startingReplicas` and `stoppingReplicas`. The running image is deliberately excluded from that identity, which is why a roll is not triggered by changing the workload image. ## See also - [Pools](/docs/concepts/pools) — where the Pods are - [Templates](/docs/concepts/templates) — idle image and defaults - [E2B Python SDK](/docs/tutorials/e2b) — `timeout`, `set_timeout`, `kill` --- # Sandbox pools (/docs/concepts/pools) A **SandboxPool** is a set of Pods of one resource shape, kept warm for an [env](/docs/concepts/envs). It is where capacity and cost live: an env with no pool can be configured perfectly and still serve nothing. ## The numbers on a pool | Field | Meaning | |---|---| | `replicas` | how many Pods this pool keeps — the pool's **theoretical maximum concurrency** | | `idleReplicas` | Pods ready to claim right now | | `runningReplicas` | Pods currently in use by a sandbox | | `startingReplicas`, `stoppingReplicas` | Pods mid-transition, on their way in or out | | `unavailableIdleReplicas` | idle Pods that are not Ready — counted as idle, but unable to take a claim | | `pendingRequests` | claims queued against this pool | `replicas` counts every state, so a claim that finds `idleReplicas = 0` waits, and the wait ends when a Pod is returned or the pool grows ([autoscaling](/docs/concepts/autoscaling)). ## Names are derived, shapes are fixed A pool's name is derived from the env name and the effective resources — as in `YOUR_ENV-1c16gi-10-ondemand` — and the same derivation produces its scaling group. Both follow from the shape, which is why the shape is **fixed at create**: a pool does not resize, it gets replaced by one of a different shape. That is also why the console asks you for the shape, not for a size: a pool's identity *is* its shape. ## Which shape you may declare: `poolSizing` Whether a pool names a quota and an instance type is a property of the env's template, not a matter of taste. Read it off the env before writing a body: ```bash abx envs YOUR_ENV --cluster YOUR_CLUSTER # prints poolSizing ``` | `poolSizing` | What a member pool must declare | What it refuses | |---|---|---| | `billed` | a quota label **and** `instanceType` (with optional `multiplier`) | a pool without them | | `free-form` | `inlineResources` only | `instanceType`, `multiplier`, quota labels | | `either` | the caller chooses | nothing — the deployment states no rule | On a **billed** env the instance type is a billing *envelope*: `instanceType × multiplier` is reserved and charged, and `inlineResources` may then ask for less than the envelope (rounded down is allowed, rounding up is refused). On a **free-form** env, `inlineResources` is the whole size of the Pod. The server enforces this rather than trusting the client, so the answer to "which fields do I send" is always the env's, never a template you copied. **Read the rule before writing a body** > ```bash abx envs YOUR_ENV --cluster YOUR_CLUSTER # poolSizing: billed | free-form | either > ``` > `abx create envs pools -f` refuses the wrong shape with a `400` that names > the field. The console's Pool form shows the same rule, so the two cannot > disagree. ## Writing one ```bash abx create envs YOUR_ENV pools --help # the body, field by field abx envs YOUR_ENV pools --cluster YOUR_CLUSTER # what exists abx create envs YOUR_ENV pools -f pool.json --cluster YOUR_CLUSTER abx scale envs YOUR_ENV pools YOUR_POOL --replicas 4 --cluster YOUR_CLUSTER ``` `abx scale` changes size and nothing else: it re-sends the current bounds unchanged. Everything else is a `create`/`update` with a file. ## Several pools, one env One env may hold pools of different shapes and different quotas, and a claim against the env is routed across them by the platform — by availability first, with the pool's priority as a tie-break. To make placement deterministic, name the shape you want instead of letting the platform choose: pass `agentbox.scitix.ai/scaling-group` in the create metadata (or address `cluster::pool` directly, [cross-cluster](/docs/concepts/cross-cluster)). If that group has no member in the env, the create fails with `503` rather than quietly landing on another shape — a deliberate choice, because a run that silently half-succeeded on the wrong hardware is worse than a run that did not start. ## When capacity is not there | Symptom | Where to look | |---|---| | `idleReplicas = 0`, claims queue | add a pool, or let the group scale ([autoscaling](/docs/concepts/autoscaling)) | | the pool will not grow to `replicas` | quota ([reading a quota](#reading-a-quota)): `abx quotas --cluster YOUR_CLUSTER`, and the pool's `ResourceQuotaExhausted` condition | | idle Pods never become claimable | `unavailableIdleReplicas` — Pods that are not Ready; usually an image pull | | an env has pools but no capacity | the pools may be on another cluster in the federation | ### Reading a quota `abx quotas` answers one row per instance type, because that is the unit a pool is sized in — a pool names a quota *and* a shape, and a quota's total across shapes is not a number anyone can spend. Each row's `ceiling` is one of three things, and which one decides whether a submission can succeed: | `ceiling` | Meaning | |---|---| | a number | the enforced cap for that instance type | | `0` | nothing allocated to you on this pool; a submission against it is refused | | `unlimited` | this deployment skips the quota check for that pool — the ondemand and spot pools are built that way | `unlimited` is not a promise of capacity. Nothing is capping you, so what decides is the pool's own stock: it is worth trying, and worth retrying once other tenants release Pods. The `name` column is the quota url a pool carries as `labels["quota.scitix.ai/url"]`. ## See also - [Autoscaling](/docs/concepts/autoscaling) — bounds, policies and cooldowns - [In-place update](/docs/concepts/inplace-update) — what a claim does to a Pod - [CLI guide](/docs/tutorials/cli) — endpoints, keys and the address grammar --- # Sandbox templates (/docs/concepts/templates) A **SandboxTemplate** is what a sandbox is made from: a Pod shape, an idle image, a runtime, the default timeouts, and the documentation an environment carries. It is a Kubernetes object, and it is the platform's, not yours. ## It is not an E2B template This is the first thing E2B users ask, and the answer changes how you write code: | | E2B template | Agent Sandbox `SandboxTemplate` | |---|---|---| | What it is | a build artifact: a Dockerfile compiled into a snapshot | a Kubernetes object: Pod template, idle image, runtimes, defaults | | How it comes to exist | you run `e2b template build` | the platform publishes it; nothing is built per workload | | Where the workload image comes from | baked into the snapshot at build time | chosen when a sandbox is created, and swapped in on the claim | | What you pass to `Sandbox.create()` | the template id | the **env** name | | Versions | template builds | `spec.version` is a label a person maintains | So there is no build step and no per-workload snapshot to maintain: an env can serve any image you can pull, and changing that image does not rebuild anything. The cost of that design is visible in [in-place update](/docs/concepts/inplace-update). ## What a template decides | Field | What it means | |---|---| | `template` | the Pod template: containers, sidecars, resources, volumes, node selectors | | `idleImage` | what an **unclaimed** Pod runs. Defaults to the main container's image | | `runtimes` | the sandbox runtime the Pod carries (the E2B runtime is the default) — that runtime is envd, [and these are the patches we carry](/docs/designs/envd) | | default startup / idle timeouts | what a create means when it does not say | | documentation | the text `abx envs docs` prints, rendered per cluster at read time | | visibility | which teams and users see the template at all; empty means public | | `version` | a human-maintained string, surfaced as the template's version | ## What a template does not decide Replicas, quota and instance type belong to a pool, not to a template: one template can back several envs, each with pools of different sizes. The env may override a few template fields — the image, the image policy, the default timeouts, private-registry credentials — uniformly for all of its pools; see [envs](/docs/concepts/envs). Resource *shape* is the exception worth knowing: Pods in one pool all have the same requests and limits. Resizing a workload means another pool, not another Pod — [pools](/docs/concepts/pools). ## Reading one ```bash abx templates --cluster YOUR_CLUSTER # the catalogue abx templates YOUR_TEMPLATE --cluster YOUR_CLUSTER # one, with its docs ``` `abx templates ` prints the template's documentation and its raw object, which is what a diff against a rollout is read against. Writing a template is `abx admin-templates` and needs an admin credential — the catalogue everyone else reads is read-only. ## Versions, and what actually triggers a rollout `spec.version` is a label a person keeps up to date; the platform does not resolve it to a revision. The thing that decides whether idle Pods are stale is a **revision hash** over the materialised idle Pod — the idle image, the Pod body, the gateway and the template's metadata. The running image is deliberately *not* part of that hash: an idle Pod is not running your workload image, and the image it will run is resolved when a sandbox is claimed. Edit a template's running image and the next claim picks it up with no rollout; edit anything in the idle Pod's identity and the pools roll onto the new revision, governed by each env's update strategy ([envs](/docs/concepts/envs), [in-place update](/docs/concepts/inplace-update)). ## Common misconceptions - **"The template is my image."** The template carries a default image, but the image a sandbox runs is chosen per create and can differ every time. - **"A template bump upgrades running sandboxes."** It does not touch a running sandbox. Rollouts replace **idle** Pods; claims that already happened keep what they claimed. - **"Templates are per env."** They are cluster-scoped and shared: several envs may bind one template, and an admin editing it affects all of them. ## See also - [Envs](/docs/concepts/envs) — what binds a template and why - [In-place update](/docs/concepts/inplace-update) — how a claim uses the idle image - [E2B Python SDK](/docs/tutorials/e2b) — creating sandboxes from an env --- # Cross-cluster routing (/docs/designs/cross-cluster) [Cross-cluster](/docs/concepts/cross-cluster) is what a user writes: a bare env name, or `cluster::env`, or `cluster::pool`. This page is the machinery behind those three, and the two planes it has to work on. ## Two planes | Plane | What crosses | Decided by | |---|---|---| | **Control** — `Sandbox.create()` | one HTTP request, forwarded verbatim | the Env router on the cluster that received it | | **Data** — exec, files, PTY, ports | every later connection to the sandbox | ExtProc and Envoy, using the cluster prefix inside the sandbox id | Getting the first one right is the visible half. The second is the one that makes a forwarded sandbox usable: a caller on cluster A talks to a sandbox the gateway placed on cluster B for hours, without knowing it. ## Control: who decides where a create lands A create arrives as a reference with three possible shapes. The first is obeyed, the other two are resolved: | The caller wrote | What happens | |---|---| | `SOME_CLUSTER::env` or `SOME_CLUSTER::pool` | forwarded to that cluster's E2B API, verbatim — no second-guessing | | `env` (bare) | resolved locally, and forwarded only if the local answer is "not here" | | `env//image` | as above; the image override travels with the forward | The bare name is where the design lives. Resolution runs in one order, and each step exists because the alternative is worse: 1. **local has an idle Pod** → serve it locally. A cross-cluster hop costs a round trip; there is no reason to pay it when the answer is already here. 2. **a foreign member has an idle Pod** → forward, rewritten as `cluster::pool` so the receiver claims that exact pool instead of routing again. 3. **nobody has idle capacity, but the local cluster can still scale** → stay local, and let it scale. 4. **only a foreign member can scale** → forward there. 5. **nobody can serve it** → park it locally. A request that waits is a request the operator can see; one that bounced between clusters is not. The distinction between 3 and 4 is the part that is easy to get wrong: forwarding a request to a cluster that also has to scale buys a hop and a second queue, without buying capacity. The state this reads — local idle, foreign idle, whether the local cluster can grow — is maintained per Env across clusters, keyed by `(namespace, env name)`. That keying is why the same env in two clusters must live in the same namespace name: change team or namespace and the two halves stop being one env, and a bare name silently stops spreading. The explicit `cluster::env` form keeps working, because it never consults the federation at all. ## Data: making a forwarded sandbox reachable Once a sandbox exists on cluster B, every later call has to be able to reach its ports from wherever the caller is. The sandbox id carries the cluster, so nothing has to be looked up — but a plain `ORIGINAL_DST` load balancer on cluster A would happily try to resolve a Pod that does not exist there. ExtProc sits in that path and rewrites the request before Envoy routes it: - `:authority` becomes the target gateway's host, for TLS SNI and HTTP Host matching; - `:path` gets the data-plane prefix; - `x-agentbox-cross-cluster: true` makes Envoy match the header-routed `ORIGINAL_DST` cluster instead of the local one. The scheme cannot be assumed: Envoy applies a transport socket per cluster, not per request, so a header carries `https` or `http` and selects between the TLS and plain-text gateway clusters. **A header name that is a loop guard** > The resolved upstream host travels in `x-agentbox-upstream-host`, not in > `x-envoy-original-dst-host`. The well-known name is read by *any* Envoy that > receives it: if it leaked through an nginx-ingress to the remote cluster's own > `original_dst_cluster`, that Envoy would dial the public IP it was handed — > itself — and fail with `TLS_WRONG_VERSION_NUMBER` as a 503. A private name > means the same value is meaningless on the far side. ## What is not on this path - **Not failover.** A cluster with no Pod to give you is not rescued by another one holding a spare — capacity routing is decided at create time, and a cluster that cannot serve returns a failure rather than a long queue. - **Not image distribution.** A sandbox runs where its image can be pulled; the platform rewrites a private registry's *host* to this region's equivalent, which is why the same image path has to exist in each. - **Not the console's job.** The console lists the clusters its config names; a console that does not know a cluster cannot offer it, and `abx clusters` prints what the address behind it reaches. ## Where the code is | Path | What it holds | |---|---| | [`pkg/e2bcompat/handlers/server.go`](https://github.com/scitix/Agent-Sandbox/blob/develop/pkg/e2bcompat/handlers/server.go) | the create path: parse the reference, forward when it names another cluster, resolve a bare name through the router | | [`pkg/apiserver/service/envscheduler/`](https://github.com/scitix/Agent-Sandbox/tree/develop/pkg/apiserver/service/envscheduler) | the five-step order above, and the federation state it reads | | [`pkg/apiserver/service/cross_cluster_forwarder.go`](https://github.com/scitix/Agent-Sandbox/blob/develop/pkg/apiserver/service/cross_cluster_forwarder.go) | the protocol-agnostic forwarder: method, path, query, headers and body verbatim, base URL swapped by kind | | [`pkg/envoy/extproc/cross_cluster.go`](https://github.com/scitix/Agent-Sandbox/blob/develop/pkg/envoy/extproc/cross_cluster.go) | the data-plane rewrite: authority, path prefix, and the two private headers | ## See also - [Cross-cluster](/docs/concepts/cross-cluster) — addressing, federation, images - [Pools](/docs/concepts/pools) — what "idle" means, since it decides the route - [The egress filter](/docs/designs/egress) — the other sidecar in the same Pod --- # The egress filter (/docs/designs/egress) [Egress and secrets](/docs/concepts/egress-and-secrets) is what a user asks for: one switch on the Env, rules on the create call. This page is what implements it, and where the sharp edges are. The enforcement model follows [E2B's tcpfirewall](https://github.com/e2b-dev/infra), adapted for a Kubernetes Pod and hardened for evaluation: **the default action is deny**, not allow. ## Two containers, and why not zero Turning the gateway on adds two things to every sandbox Pod: | Container | What it does | Why separate | |---|---|---| | `egress-init` (init) | installs one `nat` OUTPUT chain in the Pod's network namespace | needs `CAP_NET_ADMIN`, and a network namespace is shared by every container in the Pod — one install covers the sandbox, whatever it runs | | `egress-proxy` (native sidecar, uid 1337) | evaluates policy and injects headers | the redirect exempts uid 1337, which is what keeps the proxy's own upstream connections from being redirected back into itself | Everything is expressed in the Pod's network namespace, so enforcement is independent of two things that would otherwise matter: the sandbox's image (the filter survives an in-place image swap, because it is not in that image) and the cluster's CNI (the same rules work on Calico, ENI or anything else that gives a Pod a netns). The proxy listens on three ports — 15001 for `:80`, 15002 for `:443`, 15003 for everything else — plus a health port, 15004, which is deliberately *not* one of the three: a probe aimed at a data-plane port arrives indistinguishable from a redirected sandbox connection, so it would be policy-evaluated and logged as a denial on every interval. **Only 80 and 443 are understood** > `HTTP` rules match on the `Host` header, `TLS` rules on the SNI. Any other port > is CIDR-only: a rule written for `:3000` never fires, and nothing says so — the > traffic simply passes the filter unexamined. That is a property of layer-7 > filtering, not a bug to be waited out. ## The policy is a file, and absence means deny The control plane writes a policy document into an `emptyDir` that only the sidecar mounts; the proxy watches it with `fsnotify` and re-reads on change. The contract is fail-closed at every step: - **an absent, empty or unparseable file denies everything** except DNS, which stays resolvable so lookups fail fast instead of hanging; - `enforce: false` means the same as deny-all. The control plane flips it to true only after it has resolved a concrete ruleset for a claimed sandbox — so the window between a Pod starting and its policy arriving is closed by default, not by timing. ## The SSRF baseline is two tiers "Internal" covers two things that deserve different answers, and collapsing them forces a bad trade — a sandbox that needs one internal service would have to be handed the cloud metadata endpoint with it. | Tier | What it is | Can a policy open it? | |---|---|---| | Always denied | instance-metadata and link-local (`169.254.0.0/16`, `100.100.100.200`, `fd00:ec2::254`, `fe80::/10`) and loopback | **no.** An unauthenticated `GET` to these hands out cloud credentials, and no sandbox workload has a legitimate reason to reach them | | Default denied | RFC1918, CGNAT and ULA | yes — naming a host or CIDR in the allow list lifts the baseline for exactly that destination | A wildcard does not lift the second tier: `allowOut: ["*"]` means the public internet. `agentbox.scitix.ai/allow-private-networks` is the explicit way to say "everything, including the cluster's own network". ## Credentials, and why the value never enters the sandbox This is the part the design is shaped around: an agent with a shell must not be able to read the credential it is allowed to use. ``` vault (write-only) → operator reads it during one reconcile → sidecar tmpfs (0600, never mounted into the sandbox) → header rewritten on the matching request ``` Three consequences fall out of that path: - The operator resolves the CRD's `${e2b.secrets.NAME}` templates **before** pushing, so credential *names* never reach the sidecar and it needs no template engine. The wire between them carries values for one push and nothing else. - The sidecar file lives on the sidecar's own tmpfs and is **removed when the sandbox is released**. It is never an annotation, an environment variable or a log line — the CRD holds the reference, not the value. - The sandbox sees a **decoy**: a placeholder in the environment it can read, which the proxy replaces on the way out. Code inside runs unmodified; an agent that prints its environment prints the decoy. Interception needs TLS to stop being end-to-end, so the proxy mints a **per-sandbox CA** and a leaf certificate per intercepted host (cached in memory, 24-hour TTL, never written to a volume the sandbox can mount). The CA certificate is installed into the sandbox's trust store through `/init` — which is why a custom image needs `/etc/ssl/certs/ca-certificates.crt` to exist, and why the container must not run as uid 1337: that uid is exempt from the redirect, so a sandbox process running as it would bypass the filter entirely. **What injection does not do** > It rewrites **headers** on matching requests. It does not inspect bodies or > query strings, it does not follow a redirect to another host, and a client with > certificate pinning will fail rather than be helped. The sandbox can also use > the injected credential as many times as it likes for the hosts the rule names — > readable was the problem, not usable. Keep the rule's path and host as narrow as > the workload allows. ## Where the code is | Path | What it holds | |---|---| | [`pkg/egressproxy/`](https://github.com/scitix/Agent-Sandbox/tree/develop/pkg/egressproxy) | the proxy: policy, matching, the CA and MITM, the injection rules, the iptables redirect | | [`pkg/framework/plugins/egress/`](https://github.com/scitix/Agent-Sandbox/tree/develop/pkg/framework/plugins/egress) | the `PreCreatePod` plugin that puts the two containers into the Pod | | [`pkg/e2bcompat/handlers/egress.go`](https://github.com/scitix/Agent-Sandbox/blob/develop/pkg/e2bcompat/handlers/egress.go) | E2B's `network.allowOut` / `denyOut` / `rules` translated into that policy, including the refusals (literals, wildcards, identity tokens) | | [`installer/dockerfile/Dockerfile.idleimage`](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/dockerfile/Dockerfile.idleimage) | the idle image, which carries `/egress-proxy` — the same image a Pod runs while idle is the one the sidecar executes | ## See also - [Egress and secrets](/docs/concepts/egress-and-secrets) — the user-facing half - [envd, and the patches we carry](/docs/designs/envd) — the filter's dependency on that runtime - [Pools](/docs/concepts/pools) — why flipping the switch rolls them --- # envd, and the patches we carry (/docs/designs/envd) Every sandbox runs **envd**: the agent the E2B SDK talks to. Process exec, filesystem, PTY — that is envd's surface, and it is [upstream](https://github.com/e2b-dev/infra), not a fork of ours. We do carry three patches, and this page is what they change and why, because each one is a difference you can run into: a command that never starts, a result that is wrong rather than an error, a log full of connection attempts to an address that cannot answer. | Patch | Fixes | Symptom without it | |---|---|---| | [0001](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/runtimes/envd/patches/0001-skip-oom-nice-wrapper-when-not-firecracker.patch) | skips the OOM/`nice` exec wrapper outside a microVM | every command exits with a permission error, or never runs at all | | [0002](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/runtimes/envd/patches/0002-await-init-gate.patch) | refuses RPCs until the sandbox is armed | a command runs in a sandbox with no environment and no credentials — and succeeds | | [0003](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/runtimes/envd/patches/0003-skip-mmds-poll-when-not-firecracker.patch) | stops the MMDS poll outside a microVM | 1200 requests per sandbox to an address that never answers | The patches and the build script live in [`installer/runtimes/envd/`](https://github.com/scitix/Agent-Sandbox/tree/develop/installer/runtimes/envd); each patch file carries the longer version of what is below. ## 1. Not a Firecracker microVM, so do not pretend to be one Before exec-ing a command, upstream envd wraps it as `/bin/sh -c "echo 100 > /proc/$$/oom_score_adj && exec /usr/bin/nice -n N -- CMD"`. That priming makes sense **inside** a Firecracker microVM: children would otherwise inherit envd's protected `oom_score_adj=-1000` and `nice -20`. Under Kubernetes it is wrong twice over: - the kubelet manages `oom_score_adj` itself and pins a floor that forbids lowering it, so the write fails with `EACCES` — and because the wrapper uses `&&`, the user's command never runs; - routing every exec through `/bin/sh` makes execution depend on the image having a working one, which is how busybox multi-call images broke. envd already takes a `-isnotfc` flag, which our entrypoint passes. The patch threads it into `execcontext.Defaults.IsNotFC` and, in that mode, execs the command directly. cgroups and uid setup are untouched: those are the real resource controls here. This replaced an earlier hack that rewrote `/bin/sh` inside every image, and it is a candidate for upstreaming — gating a wrapper on "not a microVM" is what upstream would do. ## 2. A sandbox is not ready until it is armed A sandbox's environment variables, its injected trust-store certificate and its egress credentials all arrive in a `POST /init` that happens **after** envd is listening. A command accepted in that window runs in a sandbox that is not yet the one the caller asked for. It does not fail. It returns a wrong answer, which is the harder failure to notice. Two other layers already close this — the create call waits for the sandbox to be armed, and the data plane refuses to route to one that is not — and the patch is the third: the `-await-init` flag makes envd itself refuse process and filesystem RPCs with `failed_precondition` until the first `/init` lands. `/health` is deliberately not gated: the readiness probe is what drives the phase transition that triggers that `/init`, so gating it would deadlock the two against each other. It ships default-off for the same reason — the control plane has to be sending an unconditional `/init` before a template turns it on. ## 3. Stop polling a metadata service that is not there `169.254.169.254` is Firecracker's MMDS — the microVM's metadata service. Outside a microVM nothing answers. The boot-time poll is already gated on `-isnotfc`, but a second one started on **every** `/init`, unconditionally. AgentBox sends an `/init` to every sandbox, so every sandbox polled: one request every 50 ms for 60 seconds, 1200 of them, each a real connection attempt that the pod's egress filter evaluated and logged. It was found in a sidecar log that was almost entirely ``` level=INFO msg="egress denied" host=169.254.169.254 port=80 match=ssrf ``` at a steady 50 ms cadence for the first minute of every sandbox's life. The patch gates that goroutine on the same flag. The two values it fetched (`E2B_SANDBOX_ID`, `E2B_TEMPLATE_ID`) are known to the orchestrator, which puts them in the `/init` body it was already sending. ## How these images are built The build is deliberately boring, and the interesting part is what it refuses to do: - the upstream commit is **pinned** (`INFRA_REF` in [`build-envd.sh`](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/runtimes/envd/build-envd.sh)), so a patch applies deterministically and the version cannot drift under you; - patches are applied in filename order onto one tree, so each is a diff against the previous ones' result; - **a patch that does not apply aborts the build.** Shipping unpatched envd is not a fallback, it is the bug coming back; - the image tag carries a rebuild suffix — `0.9.0-2` is the third image of envd 0.9.0. That is not decoration: templates pull with `imagePullPolicy: IfNotPresent`, so rebuilding at an unchanged tag would leave every node that already has the image running the old code, with no error anywhere. The version the platform *reports* to the SDK is a separate constant, `DefaultEnvdVersion` — and the [samples](/docs/examples/templates) are pinned to the image that carries it, which `hack/sync-sample-images.py` keeps true. ## What this means when you run into it - **A command fails with `EACCES` on `oom_score_adj`**, or an image with a minimal `/bin/sh` misbehaves: the sandbox is running an envd without patch 1. Check what the template pins. - **A command succeeds with an empty environment**: the sandbox was used before it was armed. The gate above is off by default; the create path should not hand you a sandbox that early, and if it does, that is a bug worth reporting. - **Egress logs full of `169.254.169.254`**: an envd without patch 3. It is noise, not a leak — the SSRF baseline denies it every time. --- # Design notes (/docs/designs) These pages are the *why* behind behaviour described elsewhere on this site. Each one starts from a problem, says what it cost to solve, and links the code that carries it — including the parts that are not settled. | Page | The question it answers | |---|---| | [envd, and the patches we carry](/docs/designs/envd) | which parts of the runtime inside every sandbox are upstream, which are ours, and what each change fixes | | [The egress filter](/docs/designs/egress) | how a sandbox's traffic is filtered, and how a credential is injected without ever entering the sandbox | | [Cross-cluster routing](/docs/designs/cross-cluster) | how a create lands on another cluster, and how every later connection follows it | They are written for someone deciding whether to trust the platform with something — a benchmark run, an untrusted model's code, a multi-cluster rollout — rather than for someone about to change the code. Where a design has an open question, it says so. --- # What is here (/docs/examples) Every page in this section is a file from [`config/samples`](https://github.com/scitix/Agent-Sandbox/tree/develop/config/samples), rendered as it ships: the manifest's own comments become the introduction, and the page links to the same file on `develop`. | Group | What it holds | |---|---| | [Templates](/docs/examples/templates) | three `SandboxTemplate`s to start from: [E2B Basic](/docs/examples/templates/e2b), [E2B Docker](/docs/examples/templates/e2b-docker), [E2B Kata](/docs/examples/templates/e2b-kata) | | [Environments](/docs/examples/envs) | a `SandboxEnv` with a warm pool, and the one line that points it at another template | Apply a template first, then an environment — an environment with no template to bind has nothing to render. [Installation](/docs/installation#first-sandbox) walks the two commands in order. --- # abx-common (/docs/skills/abx-common) **Generated from** > [`plugin/skills/abx-common/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-common/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## abx: the shared half `abx` addresses the **platform**: the environments, warm pools and autoscaling groups that sandboxes are claimed from. The sandboxes themselves — creating one, running commands in it, reading its filesystem — are the **E2B SDK's** job, and that line is the product, not an omission. If the task is "run something in a sandbox", reach for E2B. If it is "there is nowhere to run it yet", or "it will not scale", you are in the right place. ## Do this first ```bash abx agent-context # the whole tool as one JSON document ``` Every resource, every filter, every column, every write, and the grammar that assembles them. It is generated from the same registry the CLI dispatches on, so it cannot describe a command that does not exist. Reading it once costs less than three `--help` calls and answers more. ## The grammar ``` abx list abx get abx sub-list abx sub-get abx create -f FILE make one; it must not exist yet abx update -f FILE change one; it must exist abx delete abx scale --replicas N pools only ``` **A verb is read in the first position and nowhere else.** Everything after it is an address, which is why `abx envs apply` is the env *called* `apply` and why a name can never collide with a write. Reads have no verb at all. The address matches the console URL segment for segment, so `/clusters/c/envs/e/pools/p` and `abx envs e pools p --cluster c` are the same thing said twice. ## Driving an Env over E2B, and the one read that comes first Sandboxes belong to the E2B SDK, and what that SDK needs and you cannot guess is **where the endpoints are**: the E2B-compatible API URL, the data-plane domain, and whether the data plane is http or https. Every Env carries its template's documentation, already rendered for the cluster that Env lives on: ```bash abx envs docs ``` **Read it before creating a sandbox against an Env you have not used.** It is the only place that knows the endpoints of the cluster you are actually talking to, and the failure it prevents is the quiet one: the SDK pointed at the wrong host, or at e2b.dev, never connects and never says why. When the document offers more than one way in — the same cluster, another cluster, the public one — **take the public one unless you are told otherwise** or you can see you are already inside that cluster's network. It is the only path that does not depend on where the caller happens to be running. The key in that document stays written as `${AGBX_API_KEY}`. That is not a missing value: it is your own credential, the one `abx` is already authenticating with, and the SDK reads the same key from `E2B_API_KEY`. No rendered document ever carries a live token — for a person's key or an agent's — which is what makes these pages safe to read, print and relay. The console fills the placeholder in for a person reading it there; nothing else should. ## Finding a shape — never from memory, never from a document Anything a command takes is answerable *by that command*, and asking is the only way to get the answer for the build you are actually talking to. Prose — this file included — is a copy, and a copy of a schema goes stale the first time the schema moves: ```bash abx create envs --help # the file create takes: example + every field abx create envs --schema # …the same body as JSON, refs resolved, no key needed abx update envs --help # the same file, plus how to obtain one abx update envs --schema # …and the same body as JSON abx envs pools --help # …and for a member pool, addressed the same way abx --help # columns, filters, sub-resources, writes abx envs --editable # the current values, in exactly that shape abx agent-context # all of it as one JSON document ``` `abx create|update
--help` is generated from the API schema at build time, so its field list cannot drift from the server you are writing to. When you are about to write a file, that page *is* the specification; when you are about to change one, `--editable` is the file to start from. The image carries the contracts themselves, read-only, for the shapes `abx` does not own — a sandbox's own create call, for instance: ``` /opt/agentbox/source/pkg/openapi/native/openapi.yaml the platform API /opt/agentbox/source/pkg/openapi/e2b/openapi.yaml the E2B surface a sandbox speaks /opt/agentbox/source/sdk/ this platform's own SDKs /opt/agentbox/source/cli/src/ this CLI's source ``` That is where to look for anything the CLI only *uses*: the exact fields a sandbox create call accepts are in the E2B spec there, not in a skill. The Python SDK is a third-party package installed in the sandbox, so ask it directly — its signature is the version-matched answer: ```bash python -c "import inspect, e2b; print(inspect.signature(e2b.Sandbox.create))" ``` ## Authentication Two settings are required, resolved as flag → environment → `~/.config/abx/config.json`: | Setting | Flag | Environment | |---|---|---| | Console address | `--endpoint` | `AGENTBOX_ENDPOINT` | | Credential | `--api-key` | `AGENTBOX_API_KEY` | That is the whole setup. The endpoint is the console's own address — the one a person types in a browser — and **one address reaches every cluster** the platform has. Which header the key travels in follows from that, so there is no scheme to set: the console takes `Authorization: Bearer`, a cluster API takes `AGENTBOX-API-KEY`. `--auth-scheme` is an override for a deployment answering to neither, and is otherwise unnecessary. Console links in output are derived from the endpoint too. **The exception**: a sandbox with no network route to the console is configured with `--cluster-api` (`AGENTBOX_CLUSTER_API`) pointing at one cluster's own API instead. Everything below about choosing a cluster then does not apply — that address answers for one cluster and refuses any other. You will not choose this; whoever deployed the platform did, and it shows up already set in the environment. Under the Claude Code plugin the key is in the OS keychain and a hook writes the config file. **Do not read that file, echo the key, or pass it on a command line** — it is deliberately kept out of the conversation, and putting it back in defeats the arrangement. `abx whoami` answers who the key acts as, and — the part worth checking before a write — whether it is an `agent` key. `abx` speaks as one tenant and has no flag to act as another. That is deliberate: acting as somebody means holding their key, not asking yours to pretend. Administrative work across tenants belongs in the console. ## Clusters Management calls are per cluster, and the cluster is chosen **per command** with `--cluster` (`AGENTBOX_CLUSTER`). It is not part of the context: a context names a platform, and that platform may have several clusters, so a default would quietly answer for whichever one happened to be set. Through a console — the normal case — `--cluster` reaches any cluster the platform has, and when there is exactly one it is filled in for you. In the `--cluster-api` exception above, the address answers for its own cluster and **refuses** a `--cluster` naming a different one. That refusal is deliberate: returning the local cluster's rows under another cluster's name is data that is confidently mislabelled, and a reader cannot tell. Treat the refusal as the truth about that sandbox, not as something to work around — there is no route to the other cluster from there. ```bash abx clusters # what this deployment can reach abx envs --cluster # choose one for this command ``` ## Output Three registers, and the middle one is the default: - **table** — a header line, a count, a `view:` link to the same page in the console where the console has a page for it, and `hint:` lines naming what to do next. Truncated at 200 rows with the filters that would narrow it. - `--json` — the raw API shape. No header, no hints. Use it when piping. - `--csv` — flat, for a spreadsheet or `cut`. `--wide` adds the columns held back by default. `--filter key=value` narrows a list, where **key is a column heading** — the heading, the filter key and the CSV column are one name on purpose. You have a shell. Use it: `abx envs pools --json | jq`, loops, aggregation. That is the point of a CLI over a tool-per-operation surface, and the sandbox you are probably running in has no real credentials in it anyway — the egress sidecar substitutes them on the way out. ## Writes, and approval **Do not write the file from memory.** Ask the CLI for it: ```bash abx create envs --help # create's file: address, example, fields abx create envs --schema # the same body as JSON, every ref resolved abx update envs --help # the same file, plus how to get one abx create envs pools --help # …and the same for a member pool ``` That page is generated from the API schema, so it lists every field, which are required, and which are **fixed at create** — it cannot be out of date in the way a copy in a document can. `abx agent-context` carries the same data as JSON, if you want to read it once rather than per command; `--schema` is the same body for one address, and it works with no deployment or key, because the shape comes from the build, not the server. The two verbs are separate words now: `create` makes one (it must not exist yet), `update` changes one (it must exist). They take the **same file** — the file a create wrote is the file an update takes. `update` is a **PUT of the desired state**: a field the file leaves out is a field you are asking to **remove**. That is why the safe edit is to start from what is there rather than from a blank page: ```bash abx envs demo pools --json # find the pool's name abx envs demo pools demo-1c2gi --editable > pool.json # edit pool.json abx update envs demo pools demo-1c2gi -f pool.json ``` The file is NOT the object `--json` prints: those keys are nested under `spec.*`/`status.*`, the write reads top-level ones, and the ones it does not find are the ones it clears. Piping `--json` straight into a write is the fastest way to lose state. `abx scale envs pools --replicas N` is the exception: one field, no clearing, and it re-sends the current bounds unchanged. Use it when size is all you are changing — and not at all when the pool's scaling group has autoscaling enabled, where the autoscaler owns the number and the API says so. **An `agent` key's writes are held for a person to release.** The command comes back saying so, with a link. That is not an error to work around: re-run the command after the person has acted. **Three writes are refused outright, with no approval to wait for**: issuing an API key, promoting one, and deciding an approval. An agent that could mint a credential could mint one without the agent restriction and leave the gate entirely, and no approval dialog conveys that. The answer is a console link for the person to act on themselves. Do not queue, retry, or look for a flag. For the same reason, an agent key gets key **metadata** without key **material**: `abx api-keys` lists what exists and which keys are gated, and the token field is absent. A rendered document — an Env's docs, a Template's docs, the setup guide in the console — comes back with `${AGBX_API_KEY}` intact rather than a live token, for every caller and not only for you. Relay the document and say where the person's own key goes; there is nothing missing to hunt for. The refusal and the wait are different answers and the error codes say which: `APPROVAL_REQUIRED` means a person is about to decide, so re-run it shortly. `FORBIDDEN_FOR_AGENT` means no approval exists or ever will — stop, and hand over the link. ## Errors carry the recovery A rejection names the valid set and the next command. Read it before reformulating — it usually contains the answer, and a guess costs a round trip. ## Where to go next | You want to | Skill | |---|---| | Run RL rollouts at scale | `abx-reinforcement-learning` | | Put sandboxes behind your own agent | `abx-managed-agent` | | Size pools, fix autoscaling, read quota | `abx-resource-capacity` | | Run a benchmark suite | `abx-harbor-framework` | | Run Docker inside a sandbox | `abx-sandbox-docker` | | Stop a sandbox reaching the internet | `abx-sandbox-network` | | Give a sandbox a credential it cannot read | `abx-sandbox-secrets` | | Work out why something is broken or slow | `abx-observe` | ## Read more [The object model](https://scitix.github.io/Agent-Sandbox/docs/concepts/index.md) --- # abx-harbor-framework (/docs/skills/abx-harbor-framework) **Generated from** > [`plugin/skills/abx-harbor-framework/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-harbor-framework/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## Running evaluations on AgentBox [Harbor](https://github.com/harbor-framework/harbor) already knows how to drive a benchmark. `agent-sandbox-harbor` is an environment plugin that makes it claim sandboxes from a warm pool instead of building one per task — which is where the time goes in a normal Harbor run. **No Harbor fork, and no Template Build step.** The plugin attaches through Harbor's official `--environment-import-path`, and because AgentBox pools swap the image in place on an already-running Pod, a task starts with one API call rather than an image build. ## The shape of a run ```bash pip install 'harbor[e2b]' agent-sandbox-harbor cat > agentbox.env <<'EOF' E2B_API_KEY=agbx_… E2B_API_URL=https:///agent-sandbox/api/e2b E2B_DOMAIN=/agent-sandbox/api/data AGBX_POOL_NAME=terminal-bench-pool AGBX_CLUSTER_ID=cluster-a AGBX_IMAGE_PREFIX=registry.internal/agent-sandbox EOF harbor run \ -d terminal-bench@2.0 -a oracle -n 16 -y \ --environment-import-path agent_sandbox_harbor:AgentSandboxEnvironment \ --env-file agentbox.env ``` The two endpoint lines are the env's, not a sketch: `abx envs docs` prints them for the cluster you are actually reaching, including whether the data plane is http or https, and `E2B_API_KEY` is the key you authenticate `abx` with (what that document writes as `${AGBX_API_KEY}`). Where it offers several ways in, use the public one unless you are already inside the cluster. `-n 16` is concurrency, and it is the number that has to exist as **idle Pods** before the run starts moving. Size the pool for it first: ```bash abx envs pools # idleReplicas is the real answer abx scale envs pools --replicas 16 ``` See `abx-resource-capacity` if it will not grow. ## Images: the part that actually bites Every task needs a **pre-built** image. This environment does not build from a Dockerfile and does not mutate a running sandbox, so an image is chosen in exactly this order: 1. **`AGBX_IMAGE_MAP`** — a ` ` file. Used verbatim, no rewriting. This is how you run a dataset whose `task.toml` has no `docker_image` at all, which is the case for **SWE-bench**, where the task *is* a Dockerfile. 2. **`task.toml`'s `docker_image`** — Terminal-Bench's case. Rewritten by `AGBX_IMAGE_PREFIX` (with `docker.io/` stripped first) and `AGBX_IMAGE_TAG`. 3. Neither → **the task is rejected**, deliberately and loudly. So a SWE-bench run is really two jobs: mirror or build the images once and write the map file; then run Harbor against it. Budget for the first. ## Settings that matter under load | Variable | Why you would touch it | |---|---| | `AGBX_STARTUP_TIMEOUT` | default 300s; raise for heavy images | | `AGBX_READY_TIMEOUT` | default 600s; a cold SWE-bench image can exceed it | | `AGBX_IMAGE_PREFIX` | point every `docker.io/…` at an internal mirror | | `AGBX_HTTPS` | `false` when the data plane is plain HTTP — a mismatch shows up as "never connects", never as a scheme error | One version note worth checking before blaming the platform: e2b SDK ≥ 2.24 rejects non-`e2b_` keys client-side. `agent-sandbox-e2b >= 0.0.4` neutralises that so `agbx_` keys work, and `harbor >= 0.13` pulls a new enough e2b to need it. ## When tasks fail rather than the run ```bash abx sandboxes --filter status=Failed abx sandboxes logs abx envs events ``` A whole dataset failing the same way is almost always the image map or the registry; individual tasks failing is usually the task. ## Related - `abx-resource-capacity` — making the concurrency you asked for exist - `abx-observe` — reading what failed - Reference: `sdk/python/harbor/README.md` and `INTEGRATION.md` in the agent-sandbox repository ## Read more [Sandbox pools](https://scitix.github.io/Agent-Sandbox/docs/concepts/pools.md) --- # abx-managed-agent (/docs/skills/abx-managed-agent) **Generated from** > [`plugin/skills/abx-managed-agent/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-managed-agent/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## Sandboxes as your agent's hands Your agent keeps running where it runs. Its `bash`, `read`, `write`, `edit`, `grep`, `glob` and `apply_patch` stop touching that machine and start acting on a sandbox bound to the conversation — so the work survives a process restart, can be browsed from a file UI, and is reclaimed on a timer instead of accumulating in someone's home directory. ``` your agent process AgentBox ┌──────────────────────────┐ ┌────────────────────────┐ │ harness │ │ sandbox for this │ │ ↓ tool call │ │ session │ │ hands binding │ │ │ │ ↓ HTTP │ │ │ │ hands daemon ───────────┼── E2B API ─┼─→ bash / files │ └──────────────────────────┘ └────────────────────────┘ ``` ## Read this before describing it to anyone **Confinement is not isolation.** The agent process still runs on your machine, with your files, your environment and your credentials. What moves into the sandbox is where the agent's *tools* act. That is genuinely worth having, but an agent that can install a package or load a plugin can reach the host again. For real isolation, run the harness itself in a container — orthogonal, and they compose. Saying otherwise to a user is the one mistake here that matters. ## Three bindings, one behaviour | Harness | Binding | |---|---| | Claude Agent SDK | `sdk/hands/typescript/src/harness/claude-code/` | | OpenCode | `sdk/hands/typescript/src/harness/opencode/` | | anything else | generic MCP: `sdk/hands/typescript/src/harness/mcp/` | `core/` decides what the tools do; a binding only says "replace these built-ins with these tools" in one harness's vocabulary. **A binding that reimplements behaviour from `core` is a bug** — two copies drift, and the drift is invisible from the signatures. The session-binding daemon and workspace file API are the Python half (`sdk/hands/python/agentbox_hands/`). ## Ask two things before configuring 1. **Which harness** — Claude Agent SDK, OpenCode, or something else that needs the generic MCP binding. The binding decides the whole integration and there is no useful generic answer. 2. **How many concurrent conversations**, not how many users. Ten people with one session each is ten; the pool is sized against that number. Then say the confinement sentence below **before** they build anything on it. ## What has to exist on the platform first 1. **An env** whose template is what your agent's tools need — an E2B-compatible image with a shell and the language runtimes the work requires. 2. **A pool with idle Pods**, sized to concurrent *conversations*, not total users. Ten people chatting with one session each is ten. 3. **A key for the right identity.** Which identity a sandbox is created as decides whose namespace and whose quota it lands in. ```bash abx envs abx envs pools abx scale envs pools --replicas 10 abx whoami # which identity, and whether this key's writes need approval abx envs docs # where the SDK points: E2B API URL, data domain, scheme ``` That last one is not optional reading before wiring a binding up. The agent's tools reach the sandbox through the E2B API, and which URL that is depends on the cluster the env lives on — `abx envs docs` is where the platform has already worked it out. Prefer the public entry it prints unless the agent process itself runs inside the cluster. ## Identity: the decision to make deliberately Two coherent models, and mixing them is where the confusion comes from: - **Per-person** — each session is created with that person's own credential. Sandboxes land in their namespace and count against their quota, and their sandbox list shows their own work. - **Service-owned** — every session is created as the env's owner, and *who the conversation is for* is recorded in sandbox `metadata`. One quota, one namespace, and the list needs the metadata column to tell sessions apart. Pick one. Under the second, all usage bills to one tenant — which is a choice, not a bug, but only if it was chosen. Whichever you pick, **a tenant key, never an admin key.** A service holding an admin credential means every conversation runs as admin, and nothing in the behaviour reveals it until something goes wrong. ## Credentials inside the sandbox The sandbox is a remote execution environment that holds **no real credentials** — decoy values sit in the environment and the egress sidecar substitutes the real ones on the way out, per host and per header. So an agent can run arbitrary shell there without that shell being a way to exfiltrate a key. When you need a third-party token available to the agent's code, that is what the vault and injection rules are for, not an environment variable. ## Related - `abx-resource-capacity` — sizing the pool for concurrent sessions - `abx-observe` — why a session's sandbox failed - `abx-common` — endpoint, key, cluster, approval - Reference: `sdk/hands/README.md` in the agent-sandbox repository ## Read more [Sandbox environments](https://scitix.github.io/Agent-Sandbox/docs/concepts/envs.md) --- # abx-observe (/docs/skills/abx-observe) **Generated from** > [`plugin/skills/abx-observe/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-observe/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## Observing: what broke, and what it is costing ## Start from the object, not the symptom ```bash abx envs demo # is the env itself ready abx envs demo pools # desired vs idle vs running, per shape abx envs demo events # what Kubernetes said, most recent first abx sandboxes --filter status=Failed abx sandboxes logs ``` `events` is the highest-yield of these and the most often skipped. A pool that will not grow almost always has a Warning event saying why, in Kubernetes' words rather than the platform's. ## Metrics come from E2B, not from abx The platform serves the **E2B-compatible** metrics endpoints, so the way to read resource usage is the E2B SDK you already have in the sandbox — not a separate `abx` command, and not a Prometheus query you have to construct: Which API that is depends on the env's cluster: `abx envs docs` prints the E2B API URL and data-plane domain the SDK should be pointed at. ```python from e2b import Sandbox sbx = Sandbox.connect(sandbox_id) metrics = sbx.get_metrics() # cpu, memory, disk over the sandbox's life ``` Two endpoints, both E2B's own: | | | |---|---| | `GET /sandboxes/{id}/metrics` | a time series for one sandbox | | `GET /sandboxes/metrics?sandbox_ids=…` | the latest point for several | **`/teams/{id}/metrics` is not implemented** — it answers `501`, deliberately: team and cluster administration is not exposed through the E2B surface. For usage across a team, read `abx statistics` or the console. Deliberately no new convention for the two that do exist: a caller who already speaks E2B needs nothing extra, and one who does not is better served learning E2B than a bespoke metrics dialect. **Units, because the two surfaces differ and it has caught people out:** - `cpuUsedPct` from the API is a **percentage of the sandbox's own cores** (`cores_used / cpuCount * 100`). A 16-core sandbox with 8 busy cores reads `50.08`, not `8` and not `5008`. - The console's charts plot **cores**, not a percentage. The same moment reads `50%` in one place and `8` in the other, and both are right. - `diskTotal` / `diskUsed` are `0` wherever the backend does not collect `container_fs_*`. Zero here means "not collected", not "empty disk". Time-series **charts** live in the console. `abx` prints a `view:` link on every command, and for an env or pool that link lands on the page with the charts. When asked for a trend rather than a number, hand over the link. ## The failures that look like something else | Symptom | Usually | |---|---| | Sandbox created, commands fail to connect | the runtime never came up inside the Pod — check `abx sandboxes logs` before anything else | | Pool stuck at 0 available | reservation or quota refused the Pods; see `abx-resource-capacity` — reduce the target, then grow | | Env lists a running sandbox the sandbox list does not show | two different counts: the env's is cluster-wide, the list is filtered to your tenant | | Everything empty but nothing errors | the credential's namespace does not exist on THIS cluster — `abx whoami` shows which namespace it resolved to | | `--cluster` refused | that endpoint serves one cluster; use an endpoint whose path carries `{cluster}` | ## Reading a pool's status honestly ```bash abx envs demo pools demo-1c2gi --json | jq '.status' ``` `idleReplicas` is the only number that answers "can I claim one right now". `replicas` is intent, `updatedReplicas` is rollout progress, and a pool can report `Ready` while having nothing claimable. ## Related - `abx-resource-capacity` — the fix, once you know it is capacity - `abx-common` — endpoint, key, cluster, approval ## Read more [Sandbox pools](https://scitix.github.io/Agent-Sandbox/docs/concepts/pools.md) --- # abx-reinforcement-learning (/docs/skills/abx-reinforcement-learning) **Generated from** > [`plugin/skills/abx-reinforcement-learning/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-reinforcement-learning/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## RL rollouts on AgentBox ## Before configuring anything, ask what concurrency they need It is the only number that decides the whole shape of the answer, it is never in the question, and guessing it wastes the conversation: a pool sized for 8 when they wanted 200 looks like it worked right up until the run stalls. Ask for **peak concurrent sandboxes**, not total episodes. People usually know the second and have to be walked to the first: 10,000 episodes at 64 in flight needs 64. Two more defaults worth stating rather than deciding silently: - **Leave autoscaling on**, with `maxReplicas` at their peak. A fixed pool holds capacity between runs and bills for it; a group with a ceiling drains and comes back. - **Set `minReplicas` to the steady-state floor** when the run ramps faster than the scale-up cooldown, so the autoscaler only handles the tail. ## The division of labour **`abx` provisions, the E2B SDK executes.** You use `abx` once to make sure there is capacity, and then your trainer talks E2B for the rest of the run. ``` trainer process AgentBox ├─ abx scale envs … pools … provision the pool once, up front └─ e2b.Sandbox(...) × N claim / run / discard, per episode ``` Sandboxes are **claimed from a warm pool**, not built. That is why a rollout gets an environment in about a second instead of a minute, and it is also why the pool has to be the right size before the trainer starts. ## Driving sandboxes Standard E2B — the platform serves the E2B-compatible API, so the SDK you would already reach for works unchanged. The call's fields belong to that SDK, not to this CLI: read them from the SDK source the image ships (`/opt/agentbox/source/sdk/`) or the E2B spec beside it, rather than from a sketch in a document. Before the first sandbox, read where to point it. `abx envs docs` prints the env's documentation, rendered for its cluster: the E2B API URL, the data-plane domain, and the scheme — the three things the SDK cannot guess and will not complain about. Where more than one way in is offered, use the public one unless the trainer runs inside the cluster. The key stays `${AGBX_API_KEY}` in that document; the SDK reads the same one from `E2B_API_KEY`. ```python from e2b import Sandbox sbx = Sandbox.create("", …) # the env name; the rest is the SDK's result = sbx.commands.run("python solve.py") sbx.kill() ``` Two things to carry over whatever the signature says: the `template` is the **env name** (`abx envs` lists them), and tagging each sandbox is worth the trouble — it is what tells two of them apart later, and `abx sandboxes --wide` shows it. Both are the SDK's fields, so take their names from the SDK and not from this line. > SWE ReX is **deprecated** — do not reach for it. E2B is the interface. ## Sizing the pool ```bash abx envs # what exists abx envs pools # idleReplicas is what you can claim now abx instancetypes # shapes and relative cost abx quotas # your ceiling abx scale envs pools --replicas 64 ``` Then leave autoscaling on with a ceiling at your peak, so the pool drains between runs rather than holding capacity idle: ```bash abx envs scaling-groups --editable > g.json # edit it — the fields, and what each one does, are in: # abx update envs scaling-groups --help abx update envs scaling-groups -f g.json ``` `update` is a whole-object PUT, and `--editable` is the file it takes — the object `--json` prints is a different document (nested under `spec.*`), and a file built from it clears whatever it does not carry. ## The three stalls, in the order they happen 1. **Claims queue and the pool does not grow.** Demand-anchored scaling grows toward what is actually being asked for; a wide `maxReplicas` is a ceiling, not a request. If the pool is stuck, **reduce the replica target and grow again** — a pool asking for more than the cluster can place stays stuck asking. See `abx-resource-capacity`. 2. **Quota is the ceiling.** `abx quotas`. No pool setting fixes this. 3. **Sandboxes are created but commands fail.** The runtime inside the Pod did not come up. `abx sandboxes logs` first; see `abx-observe`. ## Running a rollout loop against a pool that is also being scaled Claims are served from idle Pods, so a trainer at steady state and an autoscaler adjusting the pool do not fight — but a trainer that ramps faster than the scale-up cooldown will see queueing. If the run's concurrency is known up front, set the floor with `minReplicas` on the group and let the autoscaler only handle the tail. ## Related - `abx-resource-capacity` — the full autoscaler model and the recovery procedure - `abx-harbor-framework` — running an actual benchmark rather than free-form rollouts - `abx-observe` — logs, events and E2B-native metrics - `abx-common` — endpoint, key, cluster, approval ## Read more [In-place update](https://scitix.github.io/Agent-Sandbox/docs/concepts/inplace-update.md) --- # abx-resource-capacity (/docs/skills/abx-resource-capacity) **Generated from** > [`plugin/skills/abx-resource-capacity/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-resource-capacity/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## Capacity: replicas, autoscaling, quota ## The objects, in the order they matter ``` SandboxEnv what a sandbox is made from (one template) └─ SandboxPool a warm pool of one resource shape — this is what has a size └─ group the autoscaling policy several pools share ``` A pool holds **idle** Pods. Claiming one is fast precisely because it already exists; that is the whole design. `replicas` is how many Pods the pool keeps, `idleReplicas` how many are claimable right now, `runningReplicas` how many are in use. ```bash abx envs abx envs demo pools abx envs demo scaling-groups ``` ## Setting a size If the pool's group has autoscaling **off**, the size is yours: ```bash abx scale envs demo pools demo-1c2gi --replicas 40 ``` If the group is **on**, `replicas` belongs to the autoscaler and setting it is refused — the lever is the group's bounds instead: ```bash abx envs demo scaling-groups 1c2gi --editable > g.json # edit it — `abx update envs demo scaling-groups 1c2gi --help` lists the fields abx update envs demo scaling-groups 1c2gi -f g.json ``` Remember `update` is the whole object: a bound you delete from the file is a bound you are removing. That is how you take a ceiling **off**, which is worth knowing because it is the one thing an "update just this field" API could never express. ## How the autoscaler decides Scale-up is **demand-anchored**, not step-anchored. It looks at what is actually being asked for — running sandboxes plus queued requests — and grows toward that, plus a buffer, capped by a per-mode ceiling relative to demand. The mode sets the aggressiveness of both directions: | mode | scale-up buffer | scale-down step | |---|---|---| | `Conservative` | smallest | 1 replica per window | | `Default` | proportional | a quarter of the pool | | `Aggressive` | largest | half the pool | Two consequences worth carrying: - **A wide `maxReplicas` is not an instruction.** Demand is the anchor; the ceiling only stops growth. Setting 0–2560 does not ask for 2560. - **Scale-down is proportional**, so a pool that over-grew drains in minutes rather than one replica per stabilisation window. ## When a pool will not grow Check in this order — the first two are most of the cases: ```bash abx envs demo pools demo-1c2gi # phase, and desired vs actual abx envs demo pools demo-1c2gi --json | jq '.status' abx quotas # is the team's ceiling in the way abx envs demo events # what Kubernetes said about it ``` **The recovery is counter-intuitive and worth stating plainly: reduce the replica count, then grow again.** A pool asking for more than the cluster can place stays stuck asking; nothing retries it into existence. Dropping the target below what is available lets the reservation succeed, and you climb from there. Doubling down on the number that already failed does nothing. If quota is the limit, no amount of pool configuration helps — that is a request to whoever owns the quota, not a setting. Read the `ceiling` column the way `abx quotas` means it: a number is the cap, `0` is a hard zero (nothing allocated — do not submit, ask for an allocation), and `unlimited` means the deployment skips the quota check for that pool, so nothing caps you and the pool's own stock decides. `unlimited` is worth retrying; `0` is not. ## Defaults worth stating out loud - **Autoscaling on, with a ceiling.** A fixed pool holds capacity nobody is using between runs. A group with `maxReplicas` at the expected peak drains and comes back, and the ceiling is what stops a runaway — not a substitute for asking how much they need. - **A wide ceiling is not a request.** Scale-up is anchored on demand; setting 0–2560 does not ask for 2560, and someone who read it as a target has the wrong model of the autoscaler. - **Never raise a target that just failed.** Reduce it, let the reservation succeed, then climb. This is the one procedure people reliably get backwards. ## Planning for N concurrent sandboxes You need `N` claimable Pods at peak, not `N` over the run: 1. `abx instancetypes` — what shapes exist and what each costs per unit. 2. `abx quotas` — what is committed and what the ceiling is. 3. Size the pool for peak concurrency, not total work. A rollout that runs 1000 episodes 50 at a time needs 50. 4. Leave the group enabled with a `maxReplicas` at your peak, so the pool drains between runs instead of holding capacity nobody is using. Once the Pods are there, the caller has to reach them: `abx envs docs` prints the E2B endpoints of that env's cluster, which is what the SDK is pointed at before the first create. ## Related - `abx-observe` — the metrics and logs behind "it is slow" or "it failed" - `abx-reinforcement-learning` — the rollout-shaped version of this - `abx-common` — endpoint, key, cluster, approval ## Read more [Autoscaling](https://scitix.github.io/Agent-Sandbox/docs/concepts/autoscaling.md) --- # abx-sandbox-docker (/docs/skills/abx-sandbox-docker) **Generated from** > [`plugin/skills/abx-sandbox-docker/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-sandbox-docker/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## Docker inside a sandbox A sandbox can run `docker` and `docker compose`. It is not on by default: the template decides it, because a container runtime changes the Pod spec — a dockerd, a data directory on node disk, and either a microVM or a privileged container around it. ## The choice that matters, and it is not a preference | Runtime | Isolation | Use it | |---|---|---| | **kata** (Firecracker microVM) | dockerd runs inside a guest VM; `privileged` applies to the guest | **This one.** Full functionality, compose included. | | **runc + privileged** | dockerd shares the host kernel and holds every host capability | Only where the cluster has no kata runtime. | Say the second one's consequence plainly whenever you recommend it: **escape means the host**. It is not "slightly less isolated", and a user choosing it should be choosing it knowingly. Rootless dind is not the middle ground people expect — it has been measured on these clusters and does not work under runc. The way to avoid dangerous privilege is kata, where the privilege is confined to a microVM, not a rootless daemon. ## Finding out what this deployment has ```bash abx templates # which templates exist here abx templates # what it carries ``` Look for a template whose description names dind or kata. If there is none, the answer is that this deployment has not published one — not that you should hand the user a template to install, which is an admin action. ## What it looks like from the SDK Nothing special. It is the same E2B create; the template is what differs. The call's fields are the SDK's, not this CLI's — read them off the E2B spec in the image or the installed SDK, the way `abx-common` describes. What matters here is one number: ```python sbx = Sandbox.create("", …) # with a longer timeout than a plain sandbox print(sbx.commands.run("docker version").stdout) sbx.commands.run("docker compose up -d", cwd="/home/user/project") ``` The endpoint this factory talks to is the env's: `abx envs docs` prints it for the cluster the env lives on — the E2B API URL, the data-plane domain and the scheme. Take the public entry from that document unless you are already inside the cluster. Give it a longer timeout than you would a plain sandbox: dockerd starts in the background while envd comes up in front, and the first `docker` call can arrive before the daemon is listening. A short retry around the first command is ordinary, not a symptom. ## Registries Public images may not be reachable — many deployments are on internal networks. Two things to check before concluding the image is broken: - an internal mirror, with a prefix to rewrite `docker.io/...` onto; - registry credentials on the env, which the platform materialises as an image-pull secret. Those cover the **sandbox's own** image, not what `docker` pulls from inside it — inside, `docker login` is the user's to run. ## When `docker` is not found The env is on a template without it. Check which: ```bash abx envs --json | jq '.spec.templateRef' abx templates ``` Moving an env to another template is a template change, not a sandbox one, and it rolls the pool. ## Related - `abx-sandbox-network` — letting the sandbox reach a registry while the agent inside it cannot reach the internet - `abx-resource-capacity` — a dind sandbox is a bigger sandbox; size for it - `abx-common` — endpoint, key, cluster, approval ## Read more [Sandbox templates](https://scitix.github.io/Agent-Sandbox/docs/concepts/templates.md) --- # abx-sandbox-network (/docs/skills/abx-sandbox-network) **Generated from** > [`plugin/skills/abx-sandbox-network/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-sandbox-network/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## What a sandbox can reach Two separate decisions, and conflating them is the usual mistake: ``` SandboxEnv the env carries it does this environment HAVE a gateway create call the sandbox carries it what THIS sandbox may reach ``` The env carries a switch and no rules. Rules belong to the individual sandbox and arrive with the create call, in the E2B SDK's own vocabulary — there is no AgentBox dialect to learn. The env's side is `abx update envs --help` (one field, and that page names it); the sandbox's side is the SDK, whose shape is in the image rather than in this file. ## Why the switch is on the env and the rules are not Installing the proxy sidecar changes the Pod spec, so it rolls the pool. That genuinely is an environment-level decision. Rules are per sandbox because an environment is shared: an env-level allowlist would be a default that every create overrides anyway, so it would buy nothing and cost a second configuration surface. **Fail-closed, deliberately:** a create that carries filtering rules against an env with no gateway is **refused with 400**, not accepted-and-ignored. A Pod without the sidecar has no redirection either, so accepting it would mean the rules silently did nothing — which for an evaluation is the worst possible outcome, because the run completes and the numbers are wrong. ## Cutting an evaluation off from the internet This is the common case: the agent under test must not fetch the answer, and must not install its way around a missing dependency. The create call takes a network policy; **ask for its shape rather than recalling it** — the fields are in the E2B spec the image ships (`/opt/agentbox/source/pkg/openapi/e2b/openapi.yaml`), and the installed SDK will tell you its own signature. `abx-common` has the whole recipe; the short version is `abx update envs --help` for anything the env owns and the spec above for anything the sandbox owns. The call itself needs somewhere to go first: `abx envs docs` prints the E2B API URL and data-plane domain for the env's cluster, which is what the create is sent to. Take the public entry it offers unless the caller runs inside the cluster. What matters, and does not change with the field names: **deny-all-then-allow**, never allow-all-then-deny. A denylist is a list of the routes you thought of. Three things worth checking before declaring a run isolated: 1. **The env has a gateway.** Without it the create is refused — verify you saw a sandbox, not a 400. 2. **The package index is not on the allowlist** unless the task needs it. It is the most common accidental hole: the agent cannot search, but it can `pip install` something that can. 3. **Test it from inside.** `sbx.commands.run("curl -sS -m 5 https://example.com")` should fail. An isolation you did not observe failing is an isolation you are assuming. ## Allowing one thing and nothing else An evaluation that needs a model API but nothing else is the same shape: deny everything, then allow that one host. If the credential for it must not be readable inside the sandbox, that is `abx-sandbox-secrets` — the sidecar can inject it on the way out, so the sandbox reaches the API while holding no key. ## When traffic gets through anyway - Only `:80` and `:443` are parsed at layer 7. Traffic on another port passes through without inspection, so a rule written for `:3000` does nothing. - Check the env actually has the gateway on, and that the sandbox you are testing came from that env. ## Related - `abx-sandbox-secrets` — reaching a service without holding its credential - `abx-harbor-framework` — running an evaluation suite on top of this - `abx-common` — endpoint, key, cluster, approval ## Read more [Egress and secrets](https://scitix.github.io/Agent-Sandbox/docs/concepts/egress-and-secrets.md) --- # abx-sandbox-secrets (/docs/skills/abx-sandbox-secrets) **Generated from** > [`plugin/skills/abx-sandbox-secrets/SKILL.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/plugin/skills/abx-sandbox-secrets/SKILL.md) — the same file the installer > writes to `~/.agents/skills`. Edit it there; this page follows. ## Secrets a sandbox can use but cannot read The mechanism, and the reason it is shaped this way: ``` vault (write-only) → operator memory → sidecar tmpfs → outbound header ``` The value never enters the sandbox. The sandbox holds a **decoy** — a placeholder that looks like a token — and the egress sidecar substitutes the real one as the request leaves, for the hosts and headers the rule names. So an agent that can run arbitrary shell in there still cannot print the credential, because it is not in there. That is the property worth protecting, and it is why "just put it in envVars" is the wrong answer even when it works. ## Storing one The vault is the E2B `/secrets` surface, so the SDK you already have speaks it. Values are write-only: you can list names and overwrite, never read back. Both surfaces are reached through the same E2B API, whose address and data-plane domain are the env's: `abx envs docs` prints them for the cluster the env lives on. Use the public entry it gives unless the caller is inside the cluster. ```python sbx_secrets.set("OPENAI_KEY", "sk-…") # stored; not readable afterwards ``` In the console it is the **Vault** page. ## Using one Reference it by name on the create call. **A plaintext value in `network.rules` is refused with 400** — deliberately, because accepting it would put the credential in the request body, the access log, and the caller's source, which is the exposure the whole feature exists to remove. The create call carries the sandbox's environment and the injection rules. **Look its shape up rather than recalling it**: it belongs to the E2B surface, so it is in `/opt/agentbox/source/pkg/openapi/e2b/openapi.yaml`, and the installed SDK will print its own signature. `abx-common` has the general recipe for finding any shape this way. Two things about it are not about the field names, and are the part to get right: the value the sandbox's code reads is a **decoy**, and the real one is referenced by *name* — the sidecar substitutes it on the way out. A literal credential in the request is a 400, because it would put the secret in the request body, the access log and the caller's source. The code inside runs unmodified: it reads `OPENAI_API_KEY`, sends it, and the sidecar replaces it. Library code that has never heard of AgentBox works. ## Prerequisites, each of which fails silently The sandbox runs, the request goes out, and the credential is simply not substituted. Check in this order: 1. **The env has the gateway on** (`overrides.gateway.enabled`). Without it, a create carrying rules is refused — so if you have a running sandbox, this one is satisfied. 2. **The host matches the rule exactly.** The rule keys on the host it was written for; a redirect elsewhere is not covered. 3. **The port is 80 or 443.** Only those are parsed at layer 7. A rule for a service on another port never fires, and nothing says so. 4. **The secret name exists in the vault of the acting user.** A name that resolves to nothing leaves the decoy in place, and the upstream returns 401 — which reads as a bad key rather than a missing one. `abx whoami` tells you which identity the vault is being read as; a secret stored by one user is not visible to another. ## For an agent doing an integration Writing vault secrets **is** allowed for an agent credential, under the normal approval gate. This is deliberate and worth knowing: the credential lands in the acting person's own vault and widens nobody's authority, and finishing an integration end to end is exactly what people want an agent for. Minting an AgentBox API key is the thing that is refused outright — different act, different answer. See `abx-common`. ## Related - `abx-sandbox-network` — deny everything, then allow the one host this rule targets - `abx-managed-agent` — the same mechanism is how the platform's own assistant holds no real credential - `abx-common` — endpoint, key, cluster, approval ## Read more [Egress and secrets](https://scitix.github.io/Agent-Sandbox/docs/concepts/egress-and-secrets.md) --- # Using the skills (/docs/skills) A **skill** is a Markdown file an agent reads before it touches the platform. It carries the part a `--help` page cannot: when to reach for a command at all, and what the objects mean once you have. `abx` ships nine of them. They are not documentation *about* the product — they are the vocabulary the platform hands an agent, which is why this site renders them here: what you read on these pages is the file the installer writes to `~/.agents/skills`, generated from that same file at build time. ## Install The skills arrive with the CLI. There is nothing to install separately. **Installer** `abx`'s installer writes the binary to `~/.local/bin` and the nine skills to `~/.agents/skills`: ```bash curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh | sh ``` Both paths are overridable, which is what an image build wants: ```bash AGBX_BIN_DIR=/usr/local/bin AGBX_SKILL_DIR=/opt/skills sh install.sh ``` **Claude Code** The plugin installs the skills and keeps the API key in the OS keychain, where the agent never sees it: ```bash /plugin marketplace add scitix/agent-sandbox /plugin install agentbox ``` **Sandbox image** A sandbox that is meant to drive the platform itself has the same two things to do — the binary on `PATH`, the skills somewhere readable. `/opt/skills` is the predictable place, and any harness told to read that directory will find them: ```dockerfile RUN curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh \ | AGBX_BIN_DIR=/usr/local/bin AGBX_SKILL_DIR=/opt/skills sh ``` ## The nine | Skill | Read it when | |---|---| | [`abx-common`](/docs/skills/abx-common) | reaching a platform at all: the endpoint, the key, which cluster, and what happens to a write that needs a person's approval. Every other skill defers here. | | [`abx-resource-capacity`](/docs/skills/abx-resource-capacity) | sizing: how many sandboxes fit, why a pool will not grow, how autoscaling and quota interact. | | [`abx-observe`](/docs/skills/abx-observe) | something is failing or slow — events, pool status, sandbox logs, metrics. | | [`abx-harbor-framework`](/docs/skills/abx-harbor-framework) | running a benchmark suite (Terminal-Bench, SWE-bench, a custom dataset) against pre-warmed pools. | | [`abx-reinforcement-learning`](/docs/skills/abx-reinforcement-learning) | rollouts for training: concurrency, environment resets, collecting trajectories. | | [`abx-managed-agent`](/docs/skills/abx-managed-agent) | giving your own agent a sandbox for its file and shell tools, durably across restarts. | | [`abx-sandbox-docker`](/docs/skills/abx-sandbox-docker) | the workload needs Docker inside the sandbox. | | [`abx-sandbox-network`](/docs/skills/abx-sandbox-network) | cutting a sandbox off the network, or allowing exactly one host. | | [`abx-sandbox-secrets`](/docs/skills/abx-sandbox-secrets) | a sandbox needs a credential it must not be able to read. | ## How an agent uses them A skill describes a task; the tool it reaches for is the CLI, whose shape is one document: ```bash abx agent-context # every resource, filter, column and write body, as JSON ``` Two habits are worth engineering into whatever agent you run, and both are in `abx-common`: have it read that document once rather than guessing at commands, and have it read the environment's own documentation (`abx envs YOUR_ENV docs --cluster YOUR_CLUSTER`) rather than carrying endpoints in its prompt. Give an agent an **agent**-mode key ([section 6 of the CLI guide](/docs/tutorials/cli#6-api-key-permissions)), never an unrestricted one: the platform's writes then wait for your approval, and sandbox work is unaffected. ## See also - [CLI guide](/docs/tutorials/cli) — installing `abx` itself, and using it from an agent - [E2B SDK guide](/docs/tutorials/e2b) — the sandbox half these skills assume --- # Agent Sandbox CLI Guide (/docs/tutorials/cli) Agent Sandbox runs on the command line, which is what makes it usable by an agent as well as by a person: everything the console does, `abx` does, and everything a sandbox does, the E2B SDK does. Hand this page to an agent and it can create, use and reclaim its own sandboxes without a browser. Two tools, and the split between them is the product: | Tool | Covers | |---|---| | `abx` | the platform — environments, warm pools, autoscaling, quotas, templates | | E2B SDK | the sandboxes themselves — create, exec, files, network | The objects those commands address — templates, envs, pools — are described in [Concepts](/docs/concepts). This page is how to drive them. Values written in `UPPER_CASE` are placeholders. The console's own copy of this guide — the floating **CLI** button on any page — arrives with the address, the cluster and the E2B endpoints already filled in for your deployment. ## 1. Get an API key Create one in the console, under **API Keys**. The plaintext is shown once at creation, and the platform keeps a copy you can retrieve from that page at any time. If the key is going to an unattended agent, issue it in **agent** mode — see [section 6](#6-api-key-permissions). ## 2. Install `abx` `abx` is a single binary with nothing behind it: no Node.js, no virtualenv. One script installs it on either platform and writes nine skills to `~/.agents/skills` — see [section 4](#4-using-it-from-an-agent). It needs a platform to talk to; if you have not installed one yet, start at [Installation](/docs/installation). **macOS** ```bash curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh | sh ``` The binary lands in `~/.local/bin`. New shells find it through `~/.zprofile`; for the one you already have open: ```bash export PATH="$HOME/.local/bin:$PATH" ``` **Linux** ```bash curl -fsSL https://oss-ap-southeast.scitix.ai/scitix/packages/agentbox/cli/latest/install.sh | sh ``` The binary lands in `~/.local/bin`. Not every distribution puts that on `PATH`; the script tells you when yours does not, and this fixes the shell you are in: ```bash export PATH="$HOME/.local/bin:$PATH" ``` ## 3. Initialise `abx` ```bash abx context set YOUR_DEPLOYMENT \ --endpoint 'https://YOUR_CONSOLE/agentbox' \ --api-key agbx_... abx clusters ``` That is the entire configuration: the console address and the key. - The **address** reaches the platform, not one cluster: `abx clusters` lists every cluster it has, and any command takes `--cluster `; a single-cluster platform fills it in for you. - The **first command** is therefore `abx clusters`, not `abx envs`: listing clusters needs no `--cluster`, so it is the one command that cannot fail for want of an id — and it prints the ids every other command takes. - The **auth header** follows from the address, so there is nothing else to configure. - To work with several Agent Sandbox platforms, save each as its own context and switch with `abx context use `. The context is which platform; `--cluster` is which of its clusters, per command. ```bash abx agent-context # the whole CLI as JSON — hand this to an agent abx envs YOUR_ENV pools --cluster YOUR_CLUSTER ``` In CI or other unattended settings, `AGENTBOX_ENDPOINT` and `AGENTBOX_API_KEY` stand in for a context. ## 4. Using it from an agent The install above is the whole prerequisite: the binary, and nine skills that give an agent the platform's vocabulary — rollouts, capacity, evaluations, Docker-in-sandbox, egress isolation, vault secrets, approvals. What differs is where your agent looks for them. **Claude Code** The plugin installs the skills and keeps your key in the OS keychain, where the agent never sees it: ```bash /plugin marketplace add scitix/agent-sandbox /plugin install agentbox ``` Then set the endpoint and key in the plugin's settings; a hook writes them to `~/.config/abx/config.json` for the CLI to read. **Codex** Codex reads `~/.agents/skills` directly, so the install already did it: ```bash ls ~/.agents/skills # abx-common, abx-resource-capacity, abx-observe, … ``` Give it a context ([section 3](#3-initialise-abx)) and the skills are live; the only other thing worth telling it is to read `abx agent-context` rather than guess at commands. **OpenCode** The same directory, named from OpenCode's own instructions (an `AGENTS.md`, or whatever file your project uses for agent guidance): ```markdown Platform access is documented in ~/.agents/skills/abx-common/SKILL.md. Run `abx agent-context` before composing an abx command. ``` Nothing about the platform changes with the tool: same binary, same key, same commands. **Anything else** Any agent that can run a shell needs two things: the binary on its `PATH`, and one of these two commands, whose output is written to be read by a model rather than parsed by a person. ```bash abx agent-context # the whole CLI as JSON — the entry point abx envs YOUR_ENV docs --cluster YOUR_CLUSTER # this cluster's endpoints ``` Point it at `~/.agents/skills` if it reads a skills directory, and at `abx agent-context` if it does not. Whichever it is, give the agent an **agent**-mode key ([section 6](#6-api-key-permissions)), never an unrestricted one — the platform's writes then wait for your approval, while sandbox work is unaffected. And have it read the environment's own documentation ([section 5](#5-run-code-in-a-sandbox)) rather than carrying endpoints in its prompt: that document is rendered for the cluster it belongs to, and the prompt is not. All nine are also rendered on this site, one page each, for reading rather than installing: [Skills](/docs/skills). ## 5. Run code in a sandbox Every environment carries the documentation its template was written with, rendered for the cluster it belongs to — the E2B API URL, the data-plane domain, and the scheme to use: ```bash abx envs YOUR_ENV docs --cluster YOUR_CLUSTER ``` Read it before driving an environment through E2B: it is the only place that knows which endpoints your cluster answers on, and it covers what this page does not — the two access paths, pool and scaling-group selection, cross-cluster routing, and how images are rewritten between regions. The [E2B SDK guide](/docs/tutorials/e2b) goes through those in full. Sandboxes are E2B's surface, so the official SDK works unchanged — one call before the import points it at Agent Sandbox instead of `e2b.dev`. ```bash uv venv && source .venv/bin/activate uv pip install e2b agent-sandbox-e2b ``` ```python from agent_sandbox_e2b import patch_e2b patch_e2b( api_url="https://YOUR_GATEWAY/agent-sandbox/api/e2b", domain="YOUR_GATEWAY/agent-sandbox/api/data", ) # patch_e2b() must run BEFORE this import, or the SDK talks to e2b.dev. from e2b import Sandbox sbx = Sandbox.create("YOUR_ENV", timeout=3600, secure=False) print(sbx.commands.run("python -V").stdout) sbx.kill() ``` ## 6. API key permissions An API key acts as **you** — your team, your namespace, your quota. It cannot list or reach another user's sandboxes. Delete it in the console and it stops working everywhere, including in any agent you handed it to. A key is issued in one of two modes: | Mode | What it may do | |---|---| | **Unrestricted** | everything you can do, with no further confirmation. The right mode on your own machine. | | **Agent** | the same, except that the platform's writes — creating an environment, scaling or deleting anything — wait for your approval in the console. | Sandbox work is identical in both modes: an agent key starts sandboxes, runs commands and moves files exactly as an unrestricted one does. The gate is on the platform's write surface, which is the part that outlives the sandbox. An agent key also cannot issue itself another key, and cannot read key material — not from `abx api-keys`, and not from a rendered document such as this one. --- # E2B Python SDK (/docs/tutorials/e2b) Agent Sandbox is a sandbox service for agentic workloads — reasoning evaluation, training rollouts, managed agents — with secure isolation and flexible deployment. [E2B](https://e2b.dev/docs) is the sandbox provider behind Manus, and `envd` is the runtime E2B provides; Agent Sandbox serves the E2B-compatible API, so the official SDK reaches it unchanged. This page is the SDK side of the platform: how to point the SDK at a cluster, how to get a sandbox out of a warm pool, and what the create call takes. The platform side — environments, pools, autoscaling, quotas — is `abx`; see the [CLI guide](/docs/tutorials/cli), and [Concepts](/docs/concepts) is the object model behind both. Every deployment renders its own copy of this material with the real addresses filled in, and that copy is authoritative for the cluster you are using: ```bash abx envs YOUR_ENV docs --cluster YOUR_CLUSTER ``` On the console it is the **Env Docs** panel on the environment's page, key included. This page keeps placeholders where that one has values. ## 1. Install ```bash uv pip install 'agent-sandbox-e2b[e2b]' ``` The `[e2b]` extra pulls in the official `e2b` package. Two version notes worth knowing: **`agent-sandbox-e2b >= 0.0.6` is required for public access** (older builds do not assemble the gateway path, so sandbox ports do not connect), and the server rejects clients below its minimum supported version with `426 Upgrade Required`. ## 2. Quick start `patch_e2b()` must run **before** `from e2b import Sandbox`; otherwise the SDK connects to the official E2B service instead of Agent Sandbox. ```python import os os.environ["E2B_API_KEY"] = "agbx_..." # your API key os.environ["E2B_API_URL"] = "https://YOUR_GATEWAY/agent-sandbox/api/e2b" os.environ["E2B_DOMAIN"] = "YOUR_GATEWAY/agent-sandbox/api/data" # no scheme os.environ["E2B_HTTPS"] = "true" # "false" for a plain-http data plane from agent_sandbox_e2b import patch_e2b patch_e2b() from e2b import Sandbox sbx = Sandbox.create("YOUR_ENV", timeout=3000, secure=False) print(sbx.is_running()) sbx.kill() ``` The same three values can be passed as arguments instead, which take precedence over the environment: ```python patch_e2b( api_url="https://YOUR_GATEWAY/agent-sandbox/api/e2b", domain="YOUR_GATEWAY/agent-sandbox/api/data", https=True, ) ``` `secure=False` skips E2B's signed handshake: Agent Sandbox authenticates with an API key, so it stays `False`. ## 3. Access paths There are three ways in, and they differ only in how much of the network the request crosses. The environment's own documentation names the values for the cluster you are on. | Path | Who it is for | What changes | |---|---|---| | **Public** | your laptop, CI, anything outside the cluster | the full set of values above: `E2B_API_URL`, `E2B_DOMAIN`, `E2B_HTTPS` | | **Internal network** | a machine on the same private network as the cluster | point the gateway hostname at the cluster's internal IP — no code change | | **In-cluster** | code running inside the cluster | nothing: `patch_e2b()` falls back to the cluster's own services | **Public** is the default and the one to reach for unless you know you are inside. It is the longest path, but it depends on nothing about where the caller runs. **Internal** skips the public hop by resolving the gateway host to the cluster's internal IP. For a single machine, add the alias to `/etc/hosts`: ```bash echo "YOUR_INNER_IP YOUR_GATEWAY_HOST" | sudo tee -a /etc/hosts ``` For a workload that is itself a Pod, `hostAliases` does the same without touching a shared file: ```yaml spec: hostAliases: - ip: "YOUR_INNER_IP" hostnames: - "YOUR_GATEWAY_HOST" ``` **In-cluster** is the shortest path: leave `E2B_API_URL`, `E2B_DOMAIN` and `E2B_HTTPS` unset — `patch_e2b()` falls back to `agentbox-e2b-api.agentbox-system.svc.cluster.local` and `agentbox-data-plane.agentbox-system.svc.cluster.local` — and keep only `E2B_API_KEY`. ## 4. What a sandbox comes from: template, env, pool A sandbox is not built from scratch at create time. The platform keeps a set of **pre-warmed Pods**, and `Sandbox.create()` claims one and swaps in your image, which is why it takes seconds rather than minutes. | Object | What it is | Who makes it | |---|---|---| | `SandboxTemplate` | the Pod shape, the runtime, this documentation | platform administrators | | `SandboxEnv` | your entry name (`YOUR_ENV`), bound to one template | you | | `SandboxPool` | a group of pre-warmed Pods under an env, one per resource shape | you | Create them in the console: **Sandbox Envs → new environment** (name, template, optional overrides, autoscaling, image pull secrets, network policy), then **add pool** on the environment (quota, resource mode — instance type × multiplier, or explicit CPU and memory — and replica count). Replicas are the concurrency ceiling: tasks beyond the number of idle Pods queue. The environment's page shows `idle` / `running` / `desired` per member pool. `idle > 0` is the point at which `Sandbox.create()` returns immediately. ## 5. `Sandbox.create()` parameters The first argument (E2B calls it `template`) has four spellings here: | Form | Meaning | |---|---| | `"YOUR_ENV"` | **recommended** — the environment in the current cluster; the platform schedules across its member pools | | `"YOUR_ENV//IMAGE"` | the same, with the main container image replaced | | `"CLUSTER_ID::YOUR_ENV"` | a named cluster plus environment (cross-cluster — see section 6) | | `"CLUSTER_ID::POOL_NAME"` | a named cluster plus one specific pool, skipping env scheduling | The `//IMAGE` suffix composes with any of them, e.g. `"CLUSTER_ID::YOUR_ENV//docker.io/library/ubuntu:24.04"`. Other parameters: | Parameter | Meaning | |---|---| | `timeout=3000` | **idle** timeout in seconds. A sandbox with no activity for that long is reclaimed; extend a live one with `sandbox.set_timeout(1800)` | | `secure=False` | keep it `False` — authentication is the API key | | `metadata={...}` | your own labels, readable from the sandbox object. Three keys are reserved by the platform (below) | Reserved metadata keys: | Key | Effect | Example | |---|---|---| | `agentbox.scitix.ai/image` | overrides the main container image (same as `//IMAGE`, and wins over it) | `"registry.example.com/my/img:v1"` | | `agentbox.scitix.ai/startup-timeout` | **startup** timeout in seconds — how long create waits for Ready; raise it for large images | `"900"` | | `agentbox.scitix.ai/scaling-group` | **routes by scaling group**: only member pools in that group are eligible | `"1c16gi"` | ```python sbx = Sandbox.create( "YOUR_ENV", timeout=3000, secure=False, metadata={ "agentbox.scitix.ai/scaling-group": "1c16gi", # only 1c16gi pools "agentbox.scitix.ai/startup-timeout": "900", # large image, give it time "run_id": "eval-2026-07-31", # ordinary label }, ) ``` A scaling group the environment has no pool for **fails the create (503)** rather than falling back to another shape. That is deliberate: during an evaluation, a size that silently drifts is harder to find than a request that refuses. ## 6. Cross-cluster An environment belongs to one cluster; same-named environments in different clusters federate, so the platform can route a create to whichever side has idle Pods. | Target | Spelling | Behaviour | |---|---|---| | a specific pool in a specific cluster | `"other-cluster::POOL_NAME"` | straight to that pool; cluster and size both pinned | | a specific cluster's environment | `"other-cluster::YOUR_ENV"` | forwarded to that cluster, which schedules across the environment's member pools (**the recommended cross-cluster form** — pool names can change) | | either side, whichever is free | `"YOUR_ENV"` | local first; when the local side has no idle Pods and cannot scale, the request is forwarded to a cluster that has them | The third form needs the environment to exist in every participating cluster: in the console, **new environment → extend another cluster's environment**, pick the existing one, and add member pools on the new side too. Two prerequisites for cross-cluster scheduling: 1. **The image must exist in every region.** A sandbox lands where there is capacity, and it has to be able to pull from there — see section 7. 2. **Team-to-namespace mapping must match** across clusters, or federation cannot pair `(namespace, env)` and a bare name will not spread. Spell the target out (`"other-cluster::YOUR_ENV"`) when that is the case. ## 7. Images and regions Each cluster declares its own region's image registry. When the image you name belongs to **another cluster's** private registry, the platform rewrites the registry host to this cluster's equivalent, so a sandbox pulls from its own region instead of across the world: ```text you write: registry-region-a.example.com/team/swebench:260328 actually pulled: YOUR_REGISTRY_HOST/team/swebench:260328 ``` The rules, in short: - only the **host** is rewritten — the path and tag survive, so the same image has to exist under the same path in each region's registry; - the rewrite happens only between registries of the **same type** (the platform configuration's `type`); public registries such as `docker.io` are never rewritten; - when no registry of that type exists in this cluster, the image is pulled from the address you wrote — possibly across regions, and possibly slowly. So before a cross-cluster evaluation, confirm the image has been synced to every region you expect to run in. ## 8. Common operations ```python sbx.is_running() # run commands — as the standard user, or as root sbx.commands.run("echo hello && python --version") sbx.commands.run("id", user="root") # files sbx.files.write("/tmp/script.py", b"print('hello from AgentBox')\n") content = sbx.files.read("/tmp/script.py") # extend the idle timeout sbx.set_timeout(1800) # always destroy it: a running sandbox holds a warm Pod sbx.kill() ``` ## 9. Troubleshooting | Symptom | Cause and fix | |---|---| | `400 sandbox env or pool "x" not found` | the name is wrong, or the environment is not in this cluster — for another cluster write `CLUSTER_ID::x` | | `503 ... has no eligible members` | the environment has no live member pool here: add one, and if you set `scaling-group`, check that the group exists | | a placeholder appears instead of an API key | the server never renders a key. The console fills it in from the key you select; a CLI or API caller replaces it with its own | | `426 Upgrade Required` | the client is below the server's minimum version: upgrade `agent-sandbox-e2b` | | the sandbox is created but its ports do not connect | usually `agent-sandbox-e2b < 0.0.6`, or `E2B_DOMAIN` missing the gateway path (it should be `YOUR_GATEWAY/agent-sandbox/api/data`) | | creates keep timing out | the image is large or being pulled for the first time: raise `agentbox.scitix.ai/startup-timeout`, and check the image is present in this region's registry | | `Sandbox.create()` hangs | it is waiting for a Pod to become Ready, from the warm pool or a cold start. Set the startup timeout so it fails with a reason instead of waiting | | a sandbox disappeared | the idle `timeout` elapsed during a quiet stretch; call `set_timeout()` before that stretch begins | --- # A warm pool (/docs/examples/envs/e2b) **Source** > [`config/samples/e2b_sandboxenv.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/config/samples/e2b_sandboxenv.yaml) on the `develop` > branch — the file this page is rendered from, and the one to apply. An Env binds one template and fans out to member pools. This one keeps two Pods warm on the cluster you name, so `Sandbox.create("e2b")` is served by a Pod that already exists rather than by a cold start. `kubectl apply -f config/samples/e2b_sandboxenv.yaml` Two things have to be true first: - the template exists — `kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml` - clusterID below is this cluster's id (the worker's `localClusterId`), and the namespace is one your team owns The other two templates take this file unchanged except for one line: point templateRef at e2b-envd-docker or e2b-envd-kata, and give the member a name and a size that match what it runs — a container engine wants more memory than a bare sandbox, and a microVM carries its own kernel inside the same numbers. A member's `spec` is the pool's frozen snapshot; leaving it out of the file is deliberate — the Env renders it from the template on the next reconcile, so what is written here stays the part a person decides: shape, size, replicas. On a deployment that bills by quota (`abx envs ` prints poolSizing), a pool must name a quota and an instance type instead of inlineResources. ```yaml apiVersion: agents.navix.sh/v1alpha1 kind: SandboxEnv metadata: name: e2b spec: templateRef: name: e2b-envd mode: WarmPool autoscaling: groups: - name: 1c2gi enabled: true minReplicas: 0 maxReplicas: 4 clusters: - clusterID: YOUR_CLUSTER_ID members: - name: e2b-1c2gi config: scalingGroup: 1c2gi # The Pod's size. A billed deployment names instanceType and # multiplier here instead. inlineResources: requests: cpu: "1" memory: 2Gi limits: cpu: "1" memory: 2Gi spec: # Warm capacity: how many sandboxes this pool can hand out at once # before anyone has to wait for one to come back. replicas: 2 ``` ```bash kubectl apply -f config/samples/e2b_sandboxenv.yaml ``` --- # Using the environments (/docs/examples/envs) A [SandboxEnv](/docs/concepts/envs) binds one template and fans out to member pools; the pool is what holds the Pods a sandbox is claimed from. One environment is enough to show how they are put together — the page is the manifest from [`config/samples`](https://github.com/scitix/Agent-Sandbox/tree/develop/config/samples), and the other two templates take it with a one-line change. | Environment | On which template | |---|---| | [A warm pool](/docs/examples/envs/e2b) | [E2B Basic](/docs/examples/templates/e2b) — two Pods kept warm | Switching it to [E2B Docker](/docs/examples/templates/e2b-docker) or [E2B Kata](/docs/examples/templates/e2b-kata) is `templateRef.name`, plus a member name and a size that match what the sandbox actually runs: a container engine needs more memory than a bare sandbox, and a microVM carries its own kernel inside the same numbers. ## Order of application The template has to exist first, and the cluster id has to be the one the worker was installed with (`controller.localClusterId` and `spec.clusters[].clusterID` are the same value): ```bash kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml kubectl apply -f config/samples/e2b_sandboxenv.yaml ``` ## What the sample decides, and what the platform does The member declares the shape (`inlineResources`), how it scales (`config.scalingGroup`, matched by an entry in `spec.autoscaling.groups`), and how many Pods to keep warm (`spec.replicas`). It does **not** repeat the Pod spec: the environment renders each member from the template on the next reconcile, so the file stays the part a person decides. On a deployment that bills by quota (`abx envs ` prints `poolSizing`), a member must name an instance type and a quota label instead of `inlineResources` — the CLI guide's [pools section](/docs/concepts/pools) has the three shapes. ## Check it ```bash kubectl get sandboxenv e2b abx envs --cluster YOUR_CLUSTER # the env, and its pools abx envs e2b --cluster YOUR_CLUSTER # one env, member by member ``` `idleReplicas` above zero means a claim will be served instantly; until then the Pods are still starting. --- # E2B Docker (/docs/examples/templates/e2b-docker) **Source** > [`config/samples/e2b-docker_sandboxtemplate.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/config/samples/e2b-docker_sandboxtemplate.yaml) on the `develop` > branch — the file this page is rendered from, and the one to apply. Same envd runtime as e2b-basic_sandboxtemplate.yaml, but the Pod runs a container engine of its own: dockerd starts in the background, and code inside the sandbox can `docker build`, `docker run` and `docker compose up`. That is what a workload needs when the task is "build this image", "bring up this compose file" or "run a database". `kubectl apply -f config/samples/e2b-docker_sandboxtemplate.yaml` WARNING: privileged. A container engine needs more of the kernel than a container is usually given, so this template runs the sandbox Pod privileged (and starts dockerd inside it). A container escape from here reaches the node. It is the right trade for a trusted workload, and the wrong one for untrusted code. On a cluster with a microVM runtime class — kata, or anything that can run a Pod in its own kernel — add `runtimeClassName` to the Pod spec below and the same privileged dockerd is confined to a guest VM instead of sharing the node's kernel. e2b-kata_sandboxtemplate.yaml is that template with the field already set. Every image is public: the container engine comes from Docker Hub's own dind image, and the runtime pieces from this project's GHCR packages. ```yaml apiVersion: agents.navix.sh/v1alpha1 kind: SandboxTemplate metadata: name: e2b-envd-docker spec: version: 0.0.1 description: E2B-compatible sandbox with Docker inside — dockerd and the compose plugin in the sandbox. idleImage: ghcr.io/scitix/agent-sandbox-idle:0.0.10 runtimes: - name: envd port: 49983 protocol: TCP description: E2B envd runtime readinessProbe: httpGet: port: 49983 path: /health template: spec: automountServiceAccountToken: false enableServiceLinks: false shareProcessNamespace: false containers: - name: sandbox # A container engine, not a workload image: dockerd, the docker CLI # and the compose plugin are all in this one. image: docker:29-dind imagePullPolicy: IfNotPresent command: - /mnt/agentbox/tini - -g - -- - /bin/sh - -c - | # An idle Pod wears the idle image, which has no dockerd to start. if [ "$AGENTBOX_IS_IDLE_IMAGE" = "true" ] || [ -f /etc/agentbox_is_idle_image ]; then echo "[dind] idle Pod; waiting for a claim." exec sleep infinity fi # Plain unix socket: the socket is inside this Pod, and nothing # outside it needs to reach the daemon. export DOCKER_TLS_CERTDIR="" echo "[dind] starting dockerd" dockerd --host=unix:///var/run/docker.sock >/var/log/dockerd.log 2>&1 & # Best effort: the first `docker` call should not race the daemon. # Not fatal — a failure here is visible in the log above. for _ in $(seq 1 30); do if docker version >/dev/null 2>&1; then echo "[dind] dockerd is ready" break fi sleep 1 done # Hand the foreground to envd, which is what the E2B SDK talks to. exec /bin/sh /mnt/agentbox/agentbox-entrypoint.sh env: - name: AGENTBOX_DIR value: /mnt/agentbox - name: PORT value: "49999" # Scrub the service environment Kubernetes injects, so a command # inside the sandbox cannot see the cluster's addresses. - name: KUBERNETES_SERVICE_HOST - name: KUBERNETES_SERVICE_PORT - name: KUBERNETES_SERVICE_PORT_HTTPS - name: KUBERNETES_PORT - name: KUBERNETES_PORT_443_TCP - name: KUBERNETES_PORT_443_TCP_ADDR - name: KUBERNETES_PORT_443_TCP_PORT - name: KUBERNETES_PORT_443_TCP_PROTO ports: - containerPort: 49983 name: envd protocol: TCP - containerPort: 49999 name: app protocol: TCP # A container engine plus whatever it builds: size for both. As with # the E2B template, a Pool's own sizing overrides this. resources: limits: cpu: "2" memory: 4Gi requests: cpu: "2" memory: 4Gi securityContext: privileged: true runAsUser: 0 runAsGroup: 0 volumeMounts: - name: shared-bin mountPath: /mnt/agentbox readOnly: true # Node disk, not the container's own overlay root: overlay2 on top # of overlayfs is something dockerd refuses to start with. - name: docker-storage mountPath: /var/lib/docker initContainers: - name: tini-injector image: ghcr.io/scitix/agent-sandbox-tini:v0.19.0-static imagePullPolicy: IfNotPresent command: [sh, -c] args: - | TARGET_DIR="/mnt/agentbox" mkdir -p "$TARGET_DIR" if [ -f "$TARGET_DIR/tini" ]; then echo "[Init] tini already exists at $TARGET_DIR. Skipping copy." else cp /workspace/tini "$TARGET_DIR/tini" chmod +x "$TARGET_DIR/tini" echo "[Init] tini static injected successfully." fi volumeMounts: - name: shared-bin mountPath: /mnt/agentbox - name: envd-injector image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2 imagePullPolicy: IfNotPresent command: [sh, -c] args: - | TARGET_DIR="/mnt/agentbox" mkdir -p "$TARGET_DIR" if [ -f "$TARGET_DIR/envd" ] && [ -f "$TARGET_DIR/agentbox-entrypoint.sh" ]; then echo "[Init] envd and scripts already exist in $TARGET_DIR. Skipping copy." else cp /workspace/envd "$TARGET_DIR/" cp /workspace/agentbox-entrypoint.sh "$TARGET_DIR/" chmod +x "$TARGET_DIR/envd" "$TARGET_DIR/agentbox-entrypoint.sh" echo "[Init] envd & scripts successfully injected." fi volumeMounts: - name: shared-bin mountPath: /mnt/agentbox volumes: - name: shared-bin emptyDir: {} - name: docker-storage emptyDir: {} ``` ```bash kubectl apply -f config/samples/e2b-docker_sandboxtemplate.yaml ``` --- # E2B Kata (/docs/examples/templates/e2b-kata) **Source** > [`config/samples/e2b-kata_sandboxtemplate.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/config/samples/e2b-kata_sandboxtemplate.yaml) on the `develop` > branch — the file this page is rendered from, and the one to apply. The same sandbox as e2b-basic_sandboxtemplate.yaml, with one line that matters: `runtimeClassName`, which asks the cluster to run the Pod in its own kernel (Kata Containers, Firecracker, or whatever your cluster publishes) instead of sharing the node's. A container escape from a shared-kernel Pod lands on the node; from here it lands in a virtual machine. `kubectl apply -f config/samples/e2b-kata_sandboxtemplate.yaml` Two things to check before you use it: - the class exists: `kubectl get runtimeclass`. The name below is the conventional one; use whatever yours is called, or the Pod never schedules. - the runtime has the RAM: a microVM carries its own kernel, so a sandbox that fits in 2Gi on runc may want more here. Cost is the other half of the trade: start-up is slower and each sandbox occupies more than a container would. Pay it where the code is untrusted — evaluations of a model you do not control, agents given a shell and a network. Every image below is public, like the other examples. ```yaml apiVersion: agents.navix.sh/v1alpha1 kind: SandboxTemplate metadata: name: e2b-envd-kata spec: version: 0.0.1 description: E2B-compatible sandbox in a microVM — envd on a runtime class with its own kernel. idleImage: ghcr.io/scitix/agent-sandbox-idle:0.0.10 runtimes: - name: envd port: 49983 protocol: TCP description: E2B envd runtime readinessProbe: httpGet: port: 49983 path: /health template: spec: # The whole difference from the basic template. Set this to the class # `kubectl get runtimeclass` shows on your cluster. runtimeClassName: kata-fc automountServiceAccountToken: false enableServiceLinks: false shareProcessNamespace: false containers: - name: sandbox image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2 imagePullPolicy: IfNotPresent command: - /mnt/agentbox/tini - -g - -- - /bin/sh - /mnt/agentbox/agentbox-entrypoint.sh env: - name: AGENTBOX_DIR value: /mnt/agentbox - name: PORT value: "49999" ports: - containerPort: 49983 name: envd protocol: TCP - containerPort: 49999 name: app protocol: TCP # More headroom than the basic template on purpose: the guest kernel # and its own page cache live inside these numbers. resources: limits: cpu: "1" memory: 4Gi requests: cpu: "1" memory: 4Gi securityContext: runAsGroup: 0 runAsUser: 0 volumeMounts: - mountPath: /mnt/agentbox name: shared-bin readOnly: true initContainers: - name: tini-injector image: ghcr.io/scitix/agent-sandbox-tini:v0.19.0-static imagePullPolicy: IfNotPresent command: [sh, -c] args: - | TARGET_DIR="/mnt/agentbox" mkdir -p "$TARGET_DIR" if [ -f "$TARGET_DIR/tini" ]; then echo "[Init] tini already exists at $TARGET_DIR. Skipping copy." else cp /workspace/tini "$TARGET_DIR/tini" chmod +x "$TARGET_DIR/tini" echo "[Init] tini static injected successfully." fi volumeMounts: - name: shared-bin mountPath: /mnt/agentbox - name: envd-injector image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2 imagePullPolicy: IfNotPresent command: [sh, -c] args: - | TARGET_DIR="/mnt/agentbox" mkdir -p "$TARGET_DIR" if [ -f "$TARGET_DIR/envd" ] && [ -f "$TARGET_DIR/agentbox-entrypoint.sh" ]; then echo "[Init] envd and scripts already exist in $TARGET_DIR. Skipping copy." else cp /workspace/envd "$TARGET_DIR/" cp /workspace/agentbox-entrypoint.sh "$TARGET_DIR/" chmod +x "$TARGET_DIR/envd" "$TARGET_DIR/agentbox-entrypoint.sh" echo "[Init] envd & scripts successfully injected." fi volumeMounts: - name: shared-bin mountPath: /mnt/agentbox volumes: - name: shared-bin emptyDir: {} ``` ```bash kubectl apply -f config/samples/e2b-kata_sandboxtemplate.yaml ``` --- # E2B Basic (/docs/examples/templates/e2b) **Source** > [`config/samples/e2b-basic_sandboxtemplate.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/config/samples/e2b-basic_sandboxtemplate.yaml) on the `develop` > branch — the file this page is rendered from, and the one to apply. The base template: a Pod that runs envd, the agent the E2B SDK talks to. A sandbox created from an Env on this template can run any image the cluster can pull — the image is swapped in on the claim, so there is nothing to rebuild when the workload changes. `kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml` Every image below is public, so this works on a cluster that has never seen this project before. Pinning: the tags are the ones the platform runs; move them together with your platform release. ```yaml apiVersion: agents.navix.sh/v1alpha1 kind: SandboxTemplate metadata: name: e2b-envd spec: version: 0.0.1 description: E2B-compatible sandbox — envd inside a Pod, any image on claim. idleImage: ghcr.io/scitix/agent-sandbox-idle:0.0.10 runtimes: - name: envd port: 49983 protocol: TCP description: E2B envd runtime readinessProbe: httpGet: port: 49983 path: /health template: spec: automountServiceAccountToken: false containers: - command: - /mnt/agentbox/tini - -g - -- - /bin/sh - /mnt/agentbox/agentbox-entrypoint.sh env: - name: AGENTBOX_DIR value: /mnt/agentbox - name: PORT value: "49999" image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2 imagePullPolicy: IfNotPresent name: sandbox ports: - containerPort: 49983 name: envd protocol: TCP - containerPort: 49999 name: app protocol: TCP # The Pod's size when nothing else decides it. A Pool on this Env # usually does: on a billed deployment its instance type (times the # multiplier) is the envelope, and inlineResources is what the Pod # actually asks for. This is the default for a free-form Pool. resources: limits: cpu: "1" memory: 2Gi requests: cpu: "1" memory: 2Gi securityContext: runAsGroup: 0 runAsUser: 0 volumeMounts: - mountPath: /mnt/agentbox name: shared-bin readOnly: true enableServiceLinks: false initContainers: - args: - > TARGET_DIR="/mnt/agentbox" mkdir -p "$TARGET_DIR" if [ -f "$TARGET_DIR/tini" ]; then echo "[Init] tini already exists at $TARGET_DIR. Skipping copy." else cp /workspace/tini "$TARGET_DIR/tini" chmod +x "$TARGET_DIR/tini" echo "[Init] tini static injected successfully." fi command: - sh - -c image: ghcr.io/scitix/agent-sandbox-tini:v0.19.0-static imagePullPolicy: IfNotPresent name: tini-injector resources: {} volumeMounts: - mountPath: /mnt/agentbox name: shared-bin - args: - > TARGET_DIR="/mnt/agentbox" mkdir -p "$TARGET_DIR" if [ -f "$TARGET_DIR/envd" ] && [ -f "$TARGET_DIR/agentbox-entrypoint.sh" ]; then echo "[Init] envd and scripts already exist in $TARGET_DIR. Skipping copy." else cp /workspace/envd "$TARGET_DIR/" cp /workspace/agentbox-entrypoint.sh "$TARGET_DIR/" chmod +x "$TARGET_DIR/envd" "$TARGET_DIR/agentbox-entrypoint.sh" echo "[Init] envd & scripts successfully injected." fi command: - sh - -c image: ghcr.io/scitix/agent-sandbox-envd:0.9.0-2 imagePullPolicy: IfNotPresent name: envd-injector resources: {} volumeMounts: - mountPath: /mnt/agentbox name: shared-bin shareProcessNamespace: false volumes: - emptyDir: {} name: shared-bin ``` ```bash kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml ``` --- # Using the templates (/docs/examples/templates) A [SandboxTemplate](/docs/concepts/templates) is what a sandbox is made from: the Pod shape, the idle image, the runtime, and the defaults a claim inherits. These pages are the real files — each one is the manifest in [`config/samples`](https://github.com/scitix/Agent-Sandbox/tree/develop/config/samples) on the `develop` branch, rendered here so it can be read before it is applied. All of the images are public (this project's GHCR packages, and Docker Hub's own images), so each works on a cluster that has never seen Agent Sandbox before. | Template | What it is for | |---|---| | [E2B Basic](/docs/examples/templates/e2b) | the sandbox itself: envd in a Pod, any image swapped in on claim. Start here. | | [E2B Docker](/docs/examples/templates/e2b-docker) | the same, plus a container engine inside the sandbox — `docker build`, `docker run`, `docker compose up`. | | [E2B Kata](/docs/examples/templates/e2b-kata) | the same sandbox in its own kernel, on a microVM runtime class. For code you do not trust. | ## Apply one ```bash kubectl apply -f config/samples/e2b-basic_sandboxtemplate.yaml ``` Then give it an [environment](/docs/examples/envs) to hold capacity — a template on its own creates nothing. ## Reading a template Three fields decide almost everything, and each example page comments the rest: | Field | What it decides | |---|---| | `idleImage` | what a Pod runs while no sandbox holds it. Keeping this small is what makes a deep warm pool affordable. | | `runtimes` | the runtime inside the Pod, and the ports it answers on. `envd` is the one the E2B SDK speaks to. | | `template.spec` | the Pod: images, resources, security context, volumes. Anything you would write in a Pod spec, plus the `runtimeClassName` that picks the isolation level. | The one field worth understanding before you edit an example: the running image is chosen when a sandbox is *claimed*, not when the template is applied, so a template does not pin your workload image — see [in-place update](/docs/concepts/inplace-update). --- # Brain and Hands (/docs/tutorials/managed-agents/brain-hands) An agent harness has two jobs: decide what to do, and do it. The first is a model loop with your prompt and your tools. The second is `bash`, `read`, `write`, `edit`, `grep`, `glob`, `apply_patch` — and today it happens on whatever machine the harness is running on. **Brain and Hands** separates those two. The brain keeps running wherever you run it; the hands become a sandbox bound to the conversation. ```mermaid flowchart LR subgraph host["wherever your agent runs"] H["harness
Claude Code · OpenCode · your loop"] B["hands binding
the seven tools"] D["hands daemon
session → sandbox"] H -->|tool call| B B -->|HTTP| D end S["sandbox for this session
/home/agents/…"] D -->|E2B API| S ``` ## Why the split is worth it | Without it | With it | |---|---| | the work lands in someone's home directory and stays there | it lands in a sandbox that is reclaimed on a timer | | a restarted process loses its scratch space | the sandbox is bound to the conversation, not the process | | the agent's `rm -rf` is aimed at the machine you develop on | it is aimed at a machine that exists for this task | | the filesystem is invisible unless you go and look | it is inspectable: the same sandbox is a page in the console | **This is confinement, not a security boundary** > The agent process still runs on your machine, with your files, your environment > and your credentials. What moves into the sandbox is where the agent's *tools* > act. An agent that can install a package or load a plugin can reach the host > again — cooperative confinement, not isolation. If you need the stronger thing, > run the harness itself in a container or a sandbox; the two compose, and > [E2B Kata](/docs/examples/templates/e2b-kata) is how you get the sandbox half > to carry its own kernel. ## Three layers, and the rule that holds them together `sdk/hands/` is one package with a deliberate seam: | Layer | Path | What it does | |---|---|---| | **core** | `typescript/src/core/` | the seven tools and their behaviour contract — harness-neutral, knows nothing about harnesses | | **binding** | `typescript/src/harness/{claude-code,opencode,mcp}/` | expresses "replace these built-ins with these tools" in one harness's vocabulary | | **daemon** | `python/agentbox_hands/` | binds a stable session id to a sandbox, and serves the workspace file API | The rule: **a binding must not reimplement behaviour from core.** A binding that does drifts from the contract, and the drift is invisible from the signatures. How completely the built-ins can be *taken away* differs by harness, and it is worth knowing which one you are on before trusting the confinement: | Harness | Mechanism | How complete | |---|---|---| | Claude Code | `tools: []` + `disallowedTools` + `mcpServers` + `toolAliases`, together | complete; subagents inherit it, and a test drives a real session to check | | OpenCode | a tool whose name matches a built-in replaces it | **unverified** — the override appears to register alongside the built-in | | Anything over MCP | the generic binding | MCP can *add* tools; it cannot remove a harness's built-ins | That second row is the reason the console's own assistant runs on Claude Code: an override that registers alongside the built-in leaves the built-in in play, which is the one outcome this package exists to prevent. ## Plugging it in **Claude Agent SDK** ```ts import { sandboxToolOptions } from '@scitix/agentbox-hands/claude-code' for await (const msg of query({ prompt, options: { ...sandboxToolOptions({ sessionKey: threadId }) }, })) { … } ``` `sandboxToolOptions` returns all five tool-related options as one object on purpose: applying four of five loses the guarantee and nothing complains. **OpenCode** Point the config directory's `tools/.ts` at the binding — one line each, and the filename is what does the overriding: ```ts // ~/.config/opencode/tools/bash.ts export { default } from '@scitix/agentbox-hands/opencode/tools/bash' ``` Raise `tool_output` in `opencode.json` at the same time, or the confinement leaks: OpenCode truncates an oversized tool result by writing the full text to a file on the machine running the harness and handing the agent that path. The binding's own offload writes into the sandbox instead, and needs the limit out of the way. **Any agent, over MCP** ```ts import { createServer } from 'node:http' import { handsMcpHttpHandler } from '@scitix/agentbox-hands/mcp/http' createServer(handsMcpHttpHandler()).listen(8766) ``` Every request must carry the conversation's own stable id in `X-Hands-Session`. The MCP transport's session id looks right and is not: it changes on reconnect (the conversation quietly moves to a new sandbox) and is shared when one connection serves several conversations (they end up on one filesystem). ## Where the code is The package is in [`sdk/hands/`](https://github.com/scitix/Agent-Sandbox/tree/develop/sdk/hands), with Claude Agent SDK, OpenCode and MCP bindings and its own README — which is more precise about the open questions than this page can be. **The platform's own assistant is the reference implementation**, and it runs this architecture end to end: the console's floating assistant is a brain in the browser, a gateway, and a sandbox per conversation. | Piece | Where | |---|---| | The brain's instructions | [`brain/AGENTS.md`](https://github.com/scitix/Agent-Sandbox/blob/develop/brain/AGENTS.md) — the prompt a sandbox-backed agent runs with | | The gateway that drives the harness | [`brain/gateway/`](https://github.com/scitix/Agent-Sandbox/tree/develop/brain/gateway) | | The conversation UI and its tool cards | [`dashboard/components/assistant-ui/`](https://github.com/scitix/Agent-Sandbox/tree/develop/dashboard/components/assistant-ui) | | The session → sandbox binding | [`sdk/hands/python/agentbox_hands/`](https://github.com/scitix/Agent-Sandbox/tree/develop/sdk/hands/python/agentbox_hands) | | The pool its sandboxes come from | [`installer/helm/agent-sandbox-hub/values.yaml`](https://github.com/scitix/Agent-Sandbox/blob/develop/installer/helm/agent-sandbox-hub/values.yaml) — `assistant.sandbox` names the Env and the image | Reading those five together is the shortest path from "I have a harness" to "my harness's tools act on a sandbox": the prompt shows what the brain is told, the gateway shows the loop, and the hands package shows the tool layer underneath. ## The patterns behind it The split is not ours to invent — it is where the ecosystem has been heading, and these are the pieces worth reading alongside it: - **OpenAI, [Agents SDK](https://openai.github.io/openai-agents-python/)** — agents as a loop with tools, handoffs and guardrails. - **Anthropic, [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents)** — when a workflow beats an agent, and why tools are the interface. - **Anthropic, [Writing effective tools for agents](https://www.anthropic.com/engineering/writing-tools-for-agents)** — tool shape is prompt engineering; the seven tools here are the result of that thinking applied to a filesystem and a shell. - **Anthropic, [Code execution with MCP](https://www.anthropic.com/engineering/code-execution-with-mcp)** — the argument for putting tool execution *somewhere else* rather than in the agent's own process, which is the premise of this page. - **[Claude Agent SDK](https://docs.claude.com/en/api/agent-sdk/overview)** and **[Model Context Protocol](https://modelcontextprotocol.io/)** — the two integration surfaces the bindings target. ## See also - [`abx-managed-agent`](/docs/skills/abx-managed-agent) — the same material written for an agent - [E2B Python SDK](/docs/tutorials/e2b) — how a sandbox is created, which is what the daemon does - [In-place update](/docs/concepts/inplace-update) — why handing a session its own sandbox is cheap --- # Harbor benchmarks (/docs/tutorials/evals/harbor) [Harbor](https://github.com/harbor-framework/harbor) already knows how to drive a benchmark: datasets, agents, verifiers, results. `agent-sandbox-harbor` is an environment plugin that changes one thing — where the sandbox comes from. Each task claims one from a warm pool instead of building and starting one of its own, and that difference is most of the wall-clock time of a normal run. There is no Harbor fork to maintain and no template build step: Agent Sandbox swaps the workload image into a Pod that already exists, so a task starts with one API call. ## Before you start You need a platform with somewhere to run, which is two objects: an environment and a member pool with idle capacity. The [installation guide](/docs/installation) covers the platform, and [Examples](/docs/examples/templates) has templates and environments to apply if you have none yet. Then read the environment's own documentation — it is the only place that knows which endpoints *your* cluster answers on: ```bash abx envs YOUR_ENV docs --cluster YOUR_CLUSTER ``` Use the public entry it offers; the in-cluster one is for a client running inside the cluster. ## Install ```bash uv pip install 'harbor[e2b]' agent-sandbox-harbor ``` The plugin attaches through Harbor's own `--environment-import-path`, so the Harbor you already have stays the Harbor you have. ## Size the pool for the concurrency `-n` is how many tasks run at once, and every one of them needs a Pod that is **already idle** when the run starts. The number that matters is not the pool's `replicas` but its `idleReplicas`: ```bash abx envs YOUR_ENV pools --cluster YOUR_CLUSTER # idleReplicas is the real answer abx scale envs YOUR_ENV pools YOUR_POOL --replicas 16 --cluster YOUR_CLUSTER ``` A pool that is short does not fail — the run just sits waiting for sandboxes to come back, which looks like a slow model and is not. If the pool will not grow, that is quota or the autoscaler: see [Autoscaling](/docs/concepts/autoscaling). ## Run it The environment file is where the credentials and endpoints go: ```bash cat > agentbox.env <<'EOF' E2B_API_KEY=agbx_… E2B_API_URL=https://YOUR_GATEWAY/agent-sandbox/api/e2b E2B_DOMAIN=YOUR_GATEWAY/agent-sandbox/api/data AGBX_POOL_NAME=YOUR_POOL AGBX_CLUSTER_ID=YOUR_CLUSTER AGBX_IMAGE_PREFIX=registry.example.com/agent-sandbox EOF harbor run \ -d terminal-bench@2.0 -a oracle -n 16 -y \ --environment-import-path agent_sandbox_harbor:AgentSandboxEnvironment \ --env-file agentbox.env ``` `E2B_DOMAIN` carries no scheme — the plugin adds it, and `AGBX_HTTPS=false` is how you say the data plane is plain HTTP. A mismatch there is worth watching for: it shows up as "the sandbox never connects", never as a scheme error. ## Images are the part that bites Every task needs a **pre-built image**. This environment does not build from a Dockerfile and does not mutate a running sandbox, so an image is chosen in exactly this order: 1. **`AGBX_IMAGE_MAP`** — a file of ` ` lines, used verbatim. This is how a dataset whose tasks have no `docker_image` at all gets run, which is the case for **SWE-bench**, where the task *is* a Dockerfile. 2. **the task's own `docker_image`** — Terminal-Bench's case, rewritten by `AGBX_IMAGE_PREFIX` (with `docker.io/` stripped first) and `AGBX_IMAGE_TAG`. 3. neither → **the task is rejected**, loudly and on purpose. So a SWE-bench run is really two jobs: mirror or build the images once and write the map file; then run Harbor against that map. Budget for the first — it is minutes to hours, and it is done once per dataset version. **A dataset with no images is not a slow run, it is a rejected one** > Every task is rejected before Harbor starts it if neither the map nor the task > names an image. `AGBX_IMAGE_MAP` is what makes SWE-bench runnable at all, and > writing it is the step people skip because the run otherwise looks like it is > starting fine. ## Settings worth knowing | Variable | Why you would touch it | |---|---| | `AGBX_STARTUP_TIMEOUT` | default 300s. Raise it for heavy images: this is the wait for a claimed Pod to become ready. | | `AGBX_READY_TIMEOUT` | default 600s. A cold SWE-bench image can exceed it on a first pull. | | `AGBX_IMAGE_PREFIX` | points every `docker.io/…` image at a mirror, which is what keeps a 500-task run from being rate-limited. | | `AGBX_HTTPS` | `false` when the data plane is plain HTTP. | One version note, so you do not blame the platform: E2B's SDK ≥ 2.24 rejects keys that do not start with `e2b_` on the client side. `agent-sandbox-e2b >= 0.0.4` neutralises that so an `agbx_` key works, and `harbor >= 0.13` pulls a new enough E2B SDK to need it. ## When tasks fail, and not the run ```bash abx sandboxes --filter status=Failed --cluster YOUR_CLUSTER abx sandboxes YOUR_SANDBOX_ID logs --cluster YOUR_CLUSTER abx envs YOUR_ENV events --cluster YOUR_CLUSTER ``` A whole dataset failing the same way is almost always the image map or the registry; individual tasks failing is usually the task itself. ## See also - [E2B Python SDK](/docs/tutorials/e2b) — the client the plugin uses underneath - [Autoscaling](/docs/concepts/autoscaling) — making the concurrency exist - [In-place update](/docs/concepts/inplace-update) — why a claim is fast, and what it costs --- # mini SWE Agent (/docs/tutorials/evals/mini-swe-agent) Scitix AgentBox is a sandbox service designed for Agentic scenarios. It provides secure isolation and flexible deployment capabilities, catering to requirements such as inference evaluation and training rollouts. For SWE-Agent inference evaluation/streaming scenarios, we provide a pre-packaged **mini-SWE-Agent** that connects directly to the AgentBox sandbox resource pool, eliminating the need for manual sandbox lifecycle management. --- ## Installation Install the mini-SWE-Agent compatible with Scitix AgentBox. --- ## Quick Start Refer to the aforementioned documentation to apply for an API Key and create a sandbox warm-up pool based on the SWE template. Then, set the following environment variables: ```bash export SCITIX_API_KEY="${AGBX_API_KEY}" # Please apply for the API Key via the platform export SCITIX_POOL_NAME="${AGBX_ENV_NAME}" # The sandbox env to allocate from (a pool name also works) ``` Select the corresponding SWE-Bench image registry based on your current cluster: ```bash export SWEBENCH_REGISTRY="docker.io/swebench" export SWEBENCH_IMAGE_TAG="latest" ``` Then, run `mini-extra swebench`: ```bash mini-extra swebench \ --subset verified \ --split test \ -m openai/zai-org/GLM-4.7 \ -c swebench \ -c swebench_scitix \ -c "environment.idle_timeout=30m" \ -w 10 \ -c "model.model_kwargs.api_base=XXXXXXXX" \ -c "model.model_kwargs.api_key=XXXXXXXX" \ -c "model.cost_tracking=ignore_errors" \ -c "agent.step_limit=1000000" \ -c "agent.cost_limit=1000000" ``` Once running, logs related to Scitix Sandbox creation should appear, indicating successful execution. --- ## Parameter Description ### `-c` Configuration Options | Parameter | Meaning | |---|---| | `-c swebench` | Uses official SWE-Agent default configuration; **must be included** | | `-c swebench_scitix` | Uses Scitix custom initialization configuration (sets Idle Timeout, Startup Timeout, etc.) | | `-c "environment.idle_timeout=30m"` | Overrides the default Idle Timeout (default is 5m); set to an appropriate value | **Order matters** > `-c swebench_scitix` must be added *after* `-c swebench`. Do not replace the > original `-c swebench`. ### Worker Count The `-w` parameter specifies the number of concurrent workers. It is recommended to keep this consistent with the size of the warm-up pool to avoid frequent cold starts. --- ## Cross-Cluster Usage If the sandbox env is in a different cluster, prefix `SCITIX_POOL_NAME` with the cluster ID in the format `clusterId::name`, where `name` is an env name (preferred — the receiving cluster then picks a member pool for you) or a concrete pool name: ```bash export SCITIX_POOL_NAME="${AGBX_CLUSTER_ID}::${AGBX_ENV_NAME}" ``` **Cross-cluster** > Cross-cluster requests are forwarded by the AgentBox control plane through the > gateway. Authentication is the same as within the local cluster, with no extra > configuration; if you get an authentication error, re-issue the API key on the > platform. --- ## FAQ **Q: No Scitix Sandbox creation logs appear after running?** Check if `-c swebench_scitix` has been added and ensure that the three environment variables `SCITIX_ENDPOINT`, `SCITIX_API_KEY`, and `SCITIX_POOL_NAME` are set correctly. **Q: Sandboxes are frequently timing out or being reclaimed?** The default `idle_timeout` is 5 minutes. If a single episode in your streaming task exceeds this duration, the sandbox will be reclaimed prematurely. It is recommended to adjust this using `-c "environment.idle_timeout=30m"`. **Q: Receiving "no idle sandbox" errors during concurrent tasks?** There are insufficient available sandboxes in the warm-up pool. Check if the number of replicas configured for the warm-up pool is equal to or greater than the number of workers specified by `-w`. Consider expanding the warm-up pool or reducing the concurrency.