Polyaxon v3 is coming →

Benchmark coding-agent sandboxes with Polyaxon

Benchmark coding-agent sandboxes with controlled workloads, reproducible Polyaxon runs, latency distributions, recovery scenarios, and quality-adjusted results.

January 27, 2026by Polyaxon
SANDBOX BENCHMARKS: three evenly spaced silver candidate tiles one selected with amber outline

A coding-agent sandbox benchmark should answer whether your workflow performs reliably on a particular configuration. A fast empty-container startup does not tell you how long it takes to clone a repository, execute a check, recover work, and deliver a reviewable patch.

Use Polyaxon to organize a repeatable benchmark suite. Compare candidate execution backends through the same application-owned harness rather than through unrelated vendor demos.

Define the unit of useful work

Choose representative tasks before measuring infrastructure. Examples include a small Python fix, a dependency-heavy repository analysis, and a multi-step session that creates and revisits files.

Each fixture should specify the starting commit, input bundle, expected output, allowed dependencies, and completion criterion. A backend that finishes quickly but produces the wrong result should not win.

Keep model behavior constant where possible. To isolate execution infrastructure, replay a fixed command sequence first. Evaluate a live agent separately, because changing tool choices can obscure the infrastructure effect.

Measure the whole path

Record distinct intervals:

  • Request accepted to workload scheduled.
  • Scheduling to application readiness.
  • Input transfer and repository preparation.
  • Command execution.
  • Output collection and verification.
  • Cleanup or checkpoint completion.

Use Polyaxon run timelines for platform lifecycle context and instrument the harness for application-level intervals. Do not subtract timestamps from unsynchronized machines without accounting for clock differences.

Compare distributions and failure rates, not just a single average. Label warm and cold conditions explicitly, including whether images and data were already cached.

Keep the experiment matrix honest

Vary one relevant dimension at a time: backend, resource profile, input size, concurrency, or workspace reuse policy. Record all others as run metadata.

A practical result table includes successful tasks, timeout rate, startup percentiles, command percentiles, bytes transferred, and retained-state correctness. Document the number of repetitions and the observation window.

Avoid claiming a universal winner from one repository or one region. A configuration that performs well for short Python calculations may behave differently for large builds or GPU-dependent tasks.

Start with a fixed fixture and retained trial records

The companion provides a small execution benchmark with no model calls. A fixed Python program sums [2, 3, 5], echoes a unique trial ID, and increments a workspace counter. The host checks the total and counter independently. A fresh service should report counter 1; repeated successful commands in the same service should report 1, 2, 3, ....

Clone the examples repository, open the benchmark directory, and copy the shared lifecycle helper beside the harness:

git clone https://github.com/polyaxon/polyaxon-examples.git
cd polyaxon-examples/blog/sandbox-benchmark
cp ../sandbox-lifecycle/lifecycle.py .

The directory contains benchmark.py, task.py, and sandbox.yaml; lifecycle.py is maintained in the shared lifecycle directory. You need a configured Polyaxon client, an existing evaluation project, permission to launch services, and sandbox support enabled on the compute agent. Replace the illustrative Python image tag with a reviewed digest and retain the same component file when comparing workspace policies.

These commands create real service runs. They are source-reviewed instructions; no measurements are supplied as if they had been executed:

python benchmark.py --project agent-evaluations --mode fresh --repetitions 5 --output-dir results-fresh
python benchmark.py --project agent-evaluations --mode reuse --repetitions 5 --output-dir results-reuse

fresh creates a service per trial. reuse keeps one service across successful trials and replaces it after an error, because an interrupted command may have changed its state. The first reused trial includes service setup; later trials do not. Neither mode guarantees a cold image or node cache. Record those infrastructure conditions separately and alternate the policy order across repeated batches to reduce time-of-day effects.

Companion outputWhat it lets you inspect
manifest.json, component.yaml, task.pyFixture and workload configuration, hashes, policy, and repetition count
trials.jsonlEvery attempted trial's phase, outcome, run UUID, timings, command status, timeout/truncation fields, and expected counter
Per-trial inputs and reportsThe actual values checked and whether state persisted correctly
cleanup-<run UUID>.jsonWhether termination was observed or still needs review
summary.jsonAttempted, failed, and unattempted counts; successful command latency; total wall time including cleanup

The summary reports command latency for successful trials only, alongside failures and command timeouts. Its p95 uses the nearest-rank definition; with five trials it is only a rough smoke-check statistic. Collect enough repetitions for a useful distribution before making a capacity decision. wall_seconds_per_success includes setup, failed attempts, and cleanup, while command_seconds measures the host's command API round trip. Keep these units separate in the comparison.

A lost submission response stops the batch and records an uncertain submission with its correlation tag. Reconcile that run before launching more work. A service timeout remains a backstop, and an unconfirmed cleanup remains visible even when the calculation succeeded. The readiness helper checks deadlines between HTTP calls, not as a hard real-time guarantee.

To run the harness as a tracked job, use controller.yaml after building its documented image, including the copied lifecycle.py beside benchmark.py. It exposes workspace policy, repetitions, a profile label, and target project as inputs. The profile label identifies the resource configuration in the bundled sandbox.yaml; changing that label alone does not change resources. Build separate reviewed component/image revisions when comparing resource profiles. The controller's scoped credentials must be able to create child services in the target project. With --track, the harness writes its files to the run's synced outputs directory and logs the comparison counts and total elapsed time.

Include interruption and isolation checks

Exercise authorized scenarios such as a client disconnect, process timeout, and service replacement. Determine whether the harness can recover the task outcome without blindly repeating side effects.

For session reuse, verify that expected files remain within the session and that a new session does not receive another session's data. Use synthetic fixtures, not production secrets.

The sandbox process interface returns exit status, timing, and truncation information. Treat those fields as part of the benchmark result rather than assuming stdout alone establishes success.

Store enough evidence to rerun the decision

Persist the harness revision, component configuration, image reference, input fixture version, and raw measurements through Polyaxon artifacts. Build summaries from those measurements, preserving unsuccessful trials.

Use run comparison to inspect alternatives under equivalent conditions. Keep cost estimates tied to measured resource use and your actual provider terms instead of mixing incompatible public pricing units.

The benchmark becomes a reusable platform asset. When a runtime, image, or scheduling policy changes, rerun the same suite and compare the evidence before changing the default for coding agents.