abx-harbor-framework
Run Harbor benchmarks (Terminal-Bench, SWE-bench, custom datasets) on AgentBox pre-warmed pools via the agent-sandbox-harbor environment plugin. Use when asked to evaluate an agent, run a benchmark suite, reproduce a leaderboard number, or work out why benchmark tasks are being rejected or timing out. Pool sizing lives in abx-resource-capacity; endpoint and key in abx-common.
Generated from
plugin/skills/abx-harbor-framework/SKILL.md — the same file the installer
writes to ~/.agents/skills. Edit it there; this page follows.
Running evaluations on AgentBox
Harbor already knows how to drive
a benchmark. agent-sandbox-harbor is an environment plugin that makes it claim
sandboxes from a warm pool instead of building one per task — which is where the
time goes in a normal Harbor run.
No Harbor fork, and no Template Build step. The plugin attaches through
Harbor's official --environment-import-path, and because AgentBox pools swap
the image in place on an already-running Pod, a task starts with one API call
rather than an image build.
The shape of a run
pip install 'harbor[e2b]' agent-sandbox-harbor
cat > agentbox.env <<'EOF'
E2B_API_KEY=agbx_…
E2B_API_URL=https://<data-plane host>/agent-sandbox/api/e2b
E2B_DOMAIN=<data-plane host>/agent-sandbox/api/data
AGBX_POOL_NAME=terminal-bench-pool
AGBX_CLUSTER_ID=cluster-a
AGBX_IMAGE_PREFIX=registry.internal/agent-sandbox
EOF
harbor run \
-d terminal-bench@2.0 -a oracle -n 16 -y \
--environment-import-path agent_sandbox_harbor:AgentSandboxEnvironment \
--env-file agentbox.envThe two endpoint lines are the env's, not a sketch: abx envs <env> docs prints
them for the cluster you are actually reaching, including whether the data plane
is http or https, and E2B_API_KEY is the key you authenticate abx with (what
that document writes as ${AGBX_API_KEY}). Where it offers several ways in, use
the public one unless you are already inside the cluster.
-n 16 is concurrency, and it is the number that has to exist as idle Pods
before the run starts moving. Size the pool for it first:
abx envs <env> pools # idleReplicas is the real answer
abx scale envs <env> pools <pool> --replicas 16See abx-resource-capacity if it will not grow.
Images: the part that actually bites
Every task needs a pre-built image. This environment does not build from a Dockerfile and does not mutate a running sandbox, so an image is chosen in exactly this order:
AGBX_IMAGE_MAP— a<task-name> <image>file. Used verbatim, no rewriting. This is how you run a dataset whosetask.tomlhas nodocker_imageat all, which is the case for SWE-bench, where the task is a Dockerfile.task.toml'sdocker_image— Terminal-Bench's case. Rewritten byAGBX_IMAGE_PREFIX(withdocker.io/stripped first) andAGBX_IMAGE_TAG.- Neither → the task is rejected, deliberately and loudly.
So a SWE-bench run is really two jobs: mirror or build the images once and write the map file; then run Harbor against it. Budget for the first.
Settings that matter under load
| Variable | Why you would touch it |
|---|---|
AGBX_STARTUP_TIMEOUT | default 300s; raise for heavy images |
AGBX_READY_TIMEOUT | default 600s; a cold SWE-bench image can exceed it |
AGBX_IMAGE_PREFIX | point every docker.io/… at an internal mirror |
AGBX_HTTPS | false when the data plane is plain HTTP — a mismatch shows up as "never connects", never as a scheme error |
One version note worth checking before blaming the platform: e2b SDK ≥ 2.24
rejects non-e2b_ keys client-side. agent-sandbox-e2b >= 0.0.4 neutralises
that so agbx_ keys work, and harbor >= 0.13 pulls a new enough e2b to need it.
When tasks fail rather than the run
abx sandboxes --filter status=Failed
abx sandboxes <id> logs
abx envs <env> eventsA whole dataset failing the same way is almost always the image map or the registry; individual tasks failing is usually the task.
Related
abx-resource-capacity— making the concurrency you asked for existabx-observe— reading what failed- Reference:
sdk/python/harbor/README.mdandINTEGRATION.mdin the agent-sandbox repository
Read more
abx-common
How to reach an AgentBox platform with `abx` — endpoint and key, choosing a cluster, machine-readable output, and what happens to a write that needs someone's approval. Use when any abx command fails on authentication, addresses the wrong cluster, or comes back asking for approval. Every other abx skill defers here for those four things.
abx-managed-agent
Give your own agent a sandbox for its file and shell tools — the AgentBox hands bindings for Claude Agent SDK, OpenCode and generic MCP, plus the platform side that has to exist first. Use when asked to connect sandboxes to an existing agent or harness, to make an agent's work durable across restarts, or to integrate AgentBox into a product. Pool sizing lives in abx-resource-capacity.