Load-test ML services on Kubernetes
Design repeatable Kubernetes load tests for model services using realistic arrivals, tail latency, queueing, accelerator metrics, and recovery criteria.

A load test should answer a capacity or reliability question. Sending as many requests as possible and reporting an average latency rarely does that for an ML service. Model size, input shape, batching, queueing, accelerator memory, cold starts, and autoscaling all influence the result.
Design the experiment so another engineer can repeat it, explain the bottleneck, and decide what to change.
Write the question before the script
Useful questions are specific:
- How many requests per second can this model sustain while keeping p99 latency below the service objective?
- How does a scale-from-zero deployment behave during a sudden burst?
- Does dynamic batching improve throughput without exceeding the latency budget?
- How much traffic can the service absorb when one node or zone becomes unavailable?
- Does a new runtime reduce GPU time per accepted request?
Each question implies a workload shape, duration, metrics, and acceptance criteria. A single test configuration cannot answer all of them.
Model arrivals and inputs
Choose between an open and closed workload model deliberately. A closed model starts a fixed number of virtual users who wait for responses before sending more work. As the service slows, their request rate falls. An open model schedules arrivals independently, which can reveal queue growth and overload more realistically.
Represent production input diversity as well:
- prompt or payload size distribution;
- response length or compute variability;
- model and endpoint mix;
- streaming versus non-streaming requests;
- cache-hit distribution;
- authenticated tenant or priority classes;
- valid error and cancellation behavior.
Synthetic inputs should be safe, but they must exercise comparable tokenization, preprocessing, and execution paths.
Separate test phases
Run explicit phases instead of one undifferentiated traffic block:
- Warmup: initialize clients, connections, caches, and model paths without recording the main result.
- Ramp: increase arrivals in controlled steps to locate the operating range.
- Steady state: hold a target long enough to observe queueing, autoscaling, and thermal or memory behavior.
- Overload: exceed planned capacity to verify rejection, backpressure, and recovery.
- Recovery: reduce traffic and confirm queues, latency, and replicas return to an acceptable state.
Record the exact timestamps for every phase so service and cluster telemetry can be aligned with the generator.
A bounded first experiment
For a first steady-state check, ask: can a staging classifier complete 10 requests per second for one minute with k6 request-duration p99 below 500 ms, fewer than 1% failed or invalid responses, and no unsent scheduled iterations? Those are example acceptance criteria, not measured capacity or universal service objectives.
The companion example uses k6's constant-arrival-rate executor. Clone the examples repository and enter the load-test directory:
git clone https://github.com/polyaxon/polyaxon-examples.git
cd polyaxon-examples/blog/ml-service-load-testIf you already have the repository, use this directory in your existing checkout. It contains:
- load-test.js: bounded arrival rate, response checks, thresholds, and JSON reports.
- fixtures.json: two small synthetic request bodies to replace with a representative, approved input set.
- Dockerfile: an image containing k6 1.8.1 and the versioned example files.
- polyaxonfile.yaml: the same warmup and measurement phases as a Polyaxon Job.
You supply a reachable staging endpoint that accepts POST JSON such as {"inputs":[0.1,0.2,0.3,0.4]} and returns HTTP 200 with a nonempty string field such as {"prediction":"class-a"}. Adapt the fixtures and response check to your API before running. The check validates the response contract; assessing prediction quality requires expected labels or a separate evaluation. This example does not measure streaming responses or time to first token.
With k6 installed on a separate generator machine, set the complete endpoint and run two separate processes:
export TARGET=https://YOUR_STAGING_HOST/infer
mkdir -p results
k6 version
k6 run -e PHASE=warmup -e RATE=2 -e DURATION_SECONDS=10 -e REPORT_DIR=results load-test.jsConfirm that warmup succeeded before starting the measured phase:
k6 run -e PHASE=measurement -e RATE=10 -e DURATION_SECONDS=60 -e REPORT_DIR=results load-test.jsEach iteration sends one request. The script caps the configured rate at 50 requests per second, duration at five minutes, concurrent virtual users at 50, and individual requests at five seconds. It disables redirects and aborts on a cumulative HTTP failure-rate threshold after the initial evaluation delay. These limits bound this example; they are not a replacement for target-specific traffic and resource limits. Supply authentication through your approved secret mechanism if required, and keep credentials out of the target URL and saved reports.
The warmup initializes server-side paths, but a separate measurement process opens its own client connections. For a test of fully warmed client connections, define separate tagged scenarios in one process and exclude warmup samples from every measurement threshold. For a cold-start test, deliberately remove warmup and state that choice.
The script's custom summary writes results/warmup-summary.json and results/measurement-summary.json. Inspect the measurement report and k6's exit status together:
| Evidence | Interpretation |
|---|---|
dropped_iterations is nonzero | The requested arrival pattern was not fully generated; investigate the VU limit, generator resources, and slow responses |
| HTTP failures or invalid responses exceed the threshold | Fast failures or malformed responses do not count as useful capacity |
http_req_duration p99 exceeds 500 ms | The example request-duration objective was missed, even if all requests completed |
| Thresholds pass | This configuration met this short experiment's criteria; repeat over a longer representative window before claiming sustained capacity |
The threshold uses k6's http_req_duration: request sending, server wait, and response receiving time. It excludes initial DNS lookup and connection setup, so it is not a complete end-to-end user latency measurement. Inspect connection and TLS timings separately or instrument the full operation when they belong to your objective. See k6's metric definitions.
Do not read the overall request-rate average as the scheduled arrival rate: graceful completion of in-flight requests can extend the total run duration. Preserve the requested rate, started and completed work, and dropped iterations together. Repeating a tiny fixture set can also inflate cache-hit rates; change the fixture profile when cache behavior matters.
Measure useful work
Request rate alone can reward fast failures. Track accepted and successfully completed work together with quality and latency:
| Layer | Example measurements |
|---|---|
| Client | Arrival rate, completion rate, timeout and cancellation rate |
| Service | p50/p95/p99 latency, time to first token, tokens per second, error class |
| Queue | Depth, wait time, rejection, priority behavior |
| Runtime | Batch size, preprocessing time, inference time, cache hits |
| Accelerator | Activity, memory, power, throttling, communication time |
| Kubernetes | Ready replicas, restarts, scheduling delay, CPU, memory, network |
| Cost | Resource-seconds or GPU-seconds per accepted unit of work |
Our GPU utilization metrics guide explains why allocation, device activity, and throughput must be read together. A busy GPU is not automatically producing useful results.
Isolate the load generator
The generator must not become the hidden bottleneck. Run it on separate capacity from the system under test, especially when testing cluster saturation. Measure its CPU, memory, network, open connections, event-loop lag, and error rate.
For distributed generators, verify that their clocks and phase control are consistent. Keep request identifiers so a client observation can be correlated with gateway, service, and model-runtime telemetry.
Store the generator version, test definition, data profile, and target configuration with the result. A chart without those inputs is not a reproducible benchmark.
Test autoscaling as a control loop
Kubernetes autoscaling is not instantaneous. Metrics are sampled, recommendations are computed, Pods are scheduled, images are pulled, models are loaded, and readiness must succeed before capacity serves traffic. The Horizontal Pod Autoscaler is one part of that sequence.
Measure:
- time from increased arrivals to the first scaling decision;
- time from decision to scheduled Pod;
- model initialization and readiness time;
- queue growth during the gap;
- overshoot and oscillation;
- scale-down behavior and cost after the burst.
If Pods cannot schedule because accelerator capacity is absent, a perfect HPA signal will not help. Include node autoscaling and quota behavior when they are part of production capacity.
Exercise failure during load
Capacity without recovery evidence is incomplete. During a controlled test, remove a replica, drain eligible capacity according to policy, delay a dependency, or introduce a known error. Confirm that traffic shifts, retries remain bounded, and the service recovers without duplicating unsafe work.
Define stopping conditions before the test. Protect shared clusters with traffic ceilings, namespace quotas, cost limits, and an owner who can terminate the run. Never point an unbounded load generator at production by accident; make the target environment explicit in both configuration and review.
Track load tests as experiments
A load test has the same reproducibility needs as an ML experiment: versioned code, parameters, environment, input data, metrics, and artifacts. Run generators as controlled Polyaxon operations, attach their configuration and reports, and compare changes under equivalent conditions.
To use the companion Polyaxonfile, build and push its Dockerfile through your normal image workflow, using blog/ml-service-load-test as the build context. Then replace the image placeholder with that image's immutable tag or digest. The image must be accessible to the execution cluster, the run must be able to reach the target, and its container user must be able to write to the configured artifacts volume. Select a scheduling preset or node pool that keeps generator capacity separate from the service under test.
polyaxon run -f polyaxonfile.yaml -P target=https://YOUR_STAGING_HOST/inferThe Job uses sh -ec, so a failed warmup stops the measured phase and a failed measurement retains k6's nonzero exit status. Both summaries are written under the run's outputs path for artifact collection. Record the target model and deployment revision with the run, and retain the fixture file and image digest alongside the report. These JSON files are artifacts; populating comparable scalar metrics in Polyaxon requires logging the selected values through the tracking client.
The companion files have been reviewed against the documented interfaces, but have not been executed as a benchmark. They include no claimed performance results.
Finish with a decision, not merely a dashboard. State the highest validated operating point, the limiting resource, the behavior beyond capacity, recovery time, and cost per useful unit. That turns load testing from a one-time demonstration into evidence for capacity planning and release control.