Harbor benchmarks
Run Terminal-Bench, SWE-bench or a custom dataset on warm pools with Harbor and the agent-sandbox-harbor environment plugin.
Harbor already knows how to drive
a benchmark: datasets, agents, verifiers, results. agent-sandbox-harbor is an
environment plugin that changes one thing — where the sandbox comes from. Each
task claims one from a warm pool instead of building and starting one of its
own, and that difference is most of the wall-clock time of a normal run.
There is no Harbor fork to maintain and no template build step: Agent Sandbox swaps the workload image into a Pod that already exists, so a task starts with one API call.
Before you start
You need a platform with somewhere to run, which is two objects: an environment and a member pool with idle capacity. The installation guide covers the platform, and Examples has templates and environments to apply if you have none yet.
Then read the environment's own documentation — it is the only place that knows which endpoints your cluster answers on:
abx envs YOUR_ENV docs --cluster YOUR_CLUSTERUse the public entry it offers; the in-cluster one is for a client running inside the cluster.
Install
uv pip install 'harbor[e2b]' agent-sandbox-harborThe plugin attaches through Harbor's own --environment-import-path, so the
Harbor you already have stays the Harbor you have.
Size the pool for the concurrency
-n is how many tasks run at once, and every one of them needs a Pod that is
already idle when the run starts. The number that matters is not the pool's
replicas but its idleReplicas:
abx envs YOUR_ENV pools --cluster YOUR_CLUSTER # idleReplicas is the real answer
abx scale envs YOUR_ENV pools YOUR_POOL --replicas 16 --cluster YOUR_CLUSTERA pool that is short does not fail — the run just sits waiting for sandboxes to come back, which looks like a slow model and is not. If the pool will not grow, that is quota or the autoscaler: see Autoscaling.
Run it
The environment file is where the credentials and endpoints go:
cat > agentbox.env <<'EOF'
E2B_API_KEY=agbx_…
E2B_API_URL=https://YOUR_GATEWAY/agent-sandbox/api/e2b
E2B_DOMAIN=YOUR_GATEWAY/agent-sandbox/api/data
AGBX_POOL_NAME=YOUR_POOL
AGBX_CLUSTER_ID=YOUR_CLUSTER
AGBX_IMAGE_PREFIX=registry.example.com/agent-sandbox
EOF
harbor run \
-d terminal-bench@2.0 -a oracle -n 16 -y \
--environment-import-path agent_sandbox_harbor:AgentSandboxEnvironment \
--env-file agentbox.envE2B_DOMAIN carries no scheme — the plugin adds it, and AGBX_HTTPS=false is
how you say the data plane is plain HTTP. A mismatch there is worth watching
for: it shows up as "the sandbox never connects", never as a scheme error.
Images are the part that bites
Every task needs a pre-built image. This environment does not build from a Dockerfile and does not mutate a running sandbox, so an image is chosen in exactly this order:
AGBX_IMAGE_MAP— a file of<task-name> <image>lines, used verbatim. This is how a dataset whose tasks have nodocker_imageat all gets run, which is the case for SWE-bench, where the task is a Dockerfile.- the task's own
docker_image— Terminal-Bench's case, rewritten byAGBX_IMAGE_PREFIX(withdocker.io/stripped first) andAGBX_IMAGE_TAG. - neither → the task is rejected, loudly and on purpose.
So a SWE-bench run is really two jobs: mirror or build the images once and write the map file; then run Harbor against that map. Budget for the first — it is minutes to hours, and it is done once per dataset version.
A dataset with no images is not a slow run, it is a rejected one
Every task is rejected before Harbor starts it if neither the map nor the task
names an image. AGBX_IMAGE_MAP is what makes SWE-bench runnable at all, and
writing it is the step people skip because the run otherwise looks like it is
starting fine.
Settings worth knowing
| Variable | Why you would touch it |
|---|---|
AGBX_STARTUP_TIMEOUT | default 300s. Raise it for heavy images: this is the wait for a claimed Pod to become ready. |
AGBX_READY_TIMEOUT | default 600s. A cold SWE-bench image can exceed it on a first pull. |
AGBX_IMAGE_PREFIX | points every docker.io/… image at a mirror, which is what keeps a 500-task run from being rate-limited. |
AGBX_HTTPS | false when the data plane is plain HTTP. |
One version note, so you do not blame the platform: E2B's SDK ≥ 2.24 rejects
keys that do not start with e2b_ on the client side. agent-sandbox-e2b >= 0.0.4 neutralises that so an agbx_ key works, and harbor >= 0.13 pulls a
new enough E2B SDK to need it.
When tasks fail, and not the run
abx sandboxes --filter status=Failed --cluster YOUR_CLUSTER
abx sandboxes YOUR_SANDBOX_ID logs --cluster YOUR_CLUSTER
abx envs YOUR_ENV events --cluster YOUR_CLUSTERA whole dataset failing the same way is almost always the image map or the registry; individual tasks failing is usually the task itself.
See also
- E2B Python SDK — the client the plugin uses underneath
- Autoscaling — making the concurrency exist
- In-place update — why a claim is fast, and what it costs