Choose the execution target before calling the model
Start a service with the image, working directory, resources, and minimum connections required for the task. Wait for it to reach a running state, then bind the sandbox client to that project and run. Keep the target in host configuration rather than accepting it from model-generated arguments.
The agent loop can run outside the sandbox. The documented examples keep model-provider credentials in the host application, so those keys do not need to be mounted into the service that executes generated code.
Give the agent only the tools the task needs
Define tools around the operations your application supports. For unattended tasks, prefer a narrow action such as processing an approved input file over unrestricted shell access. If you expose general command execution, decide when a person must approve the proposed command.
The linked example implements a host-side approval prompt and validates tool arguments before execution. Those checks belong to the application: using a sandbox client or passing commands as argument lists does not automatically make generated code safe.
Also restrict the service's credentials, mounts, container privileges, and outbound access. Code that runs in the container can use whatever those settings permit.
Tool execution example · Scope secrets · Configure network access
Return useful results without an unbounded loop
Set a command timeout, cap the output returned to the model, and bound the number of tool calls. Inspect the exit code and timeout state before treating a command as successful. Treat command output and file contents as data, not new instructions for the agent.
Process calls return execution results; file APIs let the application supply inputs and retrieve generated reports. Separate commands can reuse files in the same running service, but a new Python process does not inherit another process's variables or imports.
Save the output and end the compute session
Download the report or write it to the run's synced output storage before cleanup. Record which input, code, and configuration produced it if you need to inspect the task later. Closing SandboxClient releases its connection but leaves the service running; stop the run explicitly when the session ends.
This fits applications that need programmable execution on Kubernetes infrastructure they operate. It requires an application-level tool policy and a container security model appropriate to the code. It is not a default hostile-code isolation guarantee or an agent reasoning framework.