FluxSandbox Sandbox Scheduling Engine
Feature Introduction
FluxSandbox is a high-performance Kubernetes sandbox scheduling engine. Serving as the flux runtime backend for OpenSandbox, it provides low-latency, high-throughput sandbox lifecycle management capabilities for AI Agent workloads.
FluxSandbox does not expose APIs directly; OpenSandbox serves as the entry point. Users use the OpenSandbox SDK to create sandboxes, and the OpenSandbox Server forwards requests to FluxSandbox via gRPC to complete scheduling and lifecycle management. The overall flow is:
User (SDK/REST API) → OpenSandbox Server (flux runtime) → FluxSandbox Controller → Scheduler → Agent Pod → Sandbox RuntimeThe sandbox runtime (i.e., the actual carrier of a sandbox instance) has two modes:
- E2B Runtime (default): The sandbox is a microVM instance, created by the E2B Sandbox Runtime and Orchestrator on the node. Supports advanced capabilities such as pause/resume/snapshot, and sandboxes created from E2B templates start in seconds.
- containerd Runtime: The sandbox is a bare container on the node's containerd, suitable for standard Kubernetes clusters. Does not support pause/resume/snapshot.
Application Scenarios
- AI Agent Runtime Environment: Provides isolated runtime sandboxes that can be obtained in seconds and reclaimed on demand, carrying tasks such as command execution and code execution for Agents.
- Large-Scale Sandbox Pool: Declares capacity and resource specifications through SandboxGroup, and the system automatically maintains warm buffers to handle high-concurrency creation requests.
- Unified Resource Management: Sandbox resource specifications are declared uniformly by SandboxGroup, avoiding customization by callers.
- Environment Save and Reuse: Creates snapshots of running sandboxes and quickly creates new sandboxes from snapshots, enabling environment template cloning and fault rollback (E2B runtime only).
- Snapshot Pre-warming: Actively distributes and caches snapshots to multiple SuperPods, accelerating sandbox creation from snapshots (E2B runtime only).
Capability Scope
- Declarative management of sandbox group capacity (SandboxGroup CRD), with automatic scaling of Agent Pods.
- Automatic sandbox scheduling and lifecycle management (create, run, pause, resume, reclaim).
- Warm buffer (sandboxBuffer), reducing sandbox acquisition latency.
- Snapshot management: create, query, delete snapshots, and restore sandboxes from snapshots (E2B runtime only).
- Snapshot pre-warming: Actively distributes and caches snapshots to multiple SuperPods, accelerating sandbox creation from snapshots (E2B runtime only).
- Multi-replica sharded scheduling, supporting horizontal scaling.
- Node-level orphan sandbox automatic cleanup.
Runtime Mode Selection
Table 1 Comparison of the two runtime modes
| Dimension | E2B Runtime (default) | containerd Runtime |
|---|---|---|
| Sandbox form | Firecracker microVM. | Bare container. |
| Sandbox image semantics | E2B template name (e.g., ubuntu-22-04-custom); template must be built in advance. | Standard container image reference (e.g., python:3.11). |
entrypoint | Not injected into sandbox; microVM process determined by template. | Acts as the container main process entry, takes effect. |
env | Not injected into sandbox. | Injected into container, takes effect. |
| Resource specification | Declared by SandboxGroup's sandboxResources (Agent Pod reserves by slot). | Declared by SandboxGroup's sandboxResources, enforced by containerd via cgroup limits. |
| Pause/Resume | Supported. | Not supported. |
| Snapshot (create/restore/pre-warm) | Supported. | Not supported. |
| Command execution carrier | envd daemon inside the sandbox. | Injected execd daemon inside the sandbox. |
| Node prerequisites | Deploy E2B Sandbox Runtime and Orchestrator. | Install containerd. |
| Control plane prerequisites | Deploy E2B Template Manager. | None. |
| SandboxGroup configuration | Default, no additional settings required. | Requires explicit setting of agentEnv.RUNTIME_TYPE=containerd and execdImage. |
Note:
The subsequent chapters of this manual use the E2B runtime by default. If using the containerd runtime, please pay attention to the containerd runtime-specific notes in each chapter.
Basic Concepts
Table 2 Basic concepts
| Concept | Description |
|---|---|
| SandboxGroup | Sandbox group, a CRD resource created directly by the user, defining the resource specifications and capacity of a set of sandboxes. Once the status is Ready, sandboxes can be created. |
| SubSandboxGroup | Automatic shard of a SandboxGroup, generated by the Controller and assigned to each scheduler instance; users do not need to create them manually. |
| Agent Pod | The workload Pod carrying sandbox instances, automatically started and reclaimed by the scheduler based on SandboxGroup capacity. |
| Sandbox | A single sandbox instance, running on the node where the Agent Pod resides, used via the OpenSandbox SDK. |
| OpenSandbox Server | User request entry point, forwarding SDK requests to FluxSandbox via the flux runtime. |
| E2B Template | The sandbox image for the E2B runtime, with the built-in envd daemon; the image in the creation request is the template name. |
| Snapshot (Snapshot) | A persistent record of the complete state of a sandbox at a given point in time, which can be used to create new sandboxes carrying the same state. Only supported by the E2B runtime. |
| SuperPod (SuperPod) | The target node pool for snapshot pre-warming, caching snapshots to accelerate sandbox creation from snapshots; storage availability is reported via MoonCake. |
Implementation Principle
When a user creates a sandbox via the SDK, the OpenSandbox Server routes the request to the corresponding SandboxGroup based on the extensions.sandboxGroup field in the request, forwards it through the FluxSandbox Controller to the owning scheduler shard, and the scheduler selects an Agent Pod from the group to create the sandbox instance. The capacity fields of the SandboxGroup (buffer, max sandboxes per Pod, etc.) drive automatic scaling of Agent Pods, and sandboxResources determines the resource specification occupied by each sandbox.
Under the E2B runtime, before creation, the scheduler resolves the template metadata from the E2B Template Manager based on the image (template name), and then the E2B Sandbox Runtime on the node starts the microVM sandbox; under the containerd runtime, the Agent directly creates a bare container through the node's containerd and injects execd.
Installation
Prerequisites
- The Kubernetes cluster is v1.28 or above and accessible via
kubectl. - E2B runtime (default):
- Worker nodes have deployed the E2B Sandbox Runtime (one instance per node, listening on
unix:///run/cri-multiplex.sock) and the E2B Orchestrator; - The control plane has deployed the E2B Template Manager and prepared available E2B templates (see Prepare E2B Templates).
- Worker nodes have deployed the E2B Sandbox Runtime (one instance per node, listening on
- Helm 3.14 or above is installed (Helm deployment is recommended).
- The OpenSandbox Server image and Chart can be obtained directly from the official source (see Deploy OpenSandbox Server); the four FluxSandbox component images (controller, scheduler, watcher, agent) have been pushed to a registry accessible by the cluster (execute
make docker-buildandmake docker-pushwhen building from source). - When creating sandboxes using the SDK, Python 3.10+ must be installed locally.
[containerd Runtime]
When using the containerd runtime, Worker nodes only need containerd installed; no E2B components are required, but E2B-related capabilities (pause/resume/snapshot/snapshot pre-warming) are unavailable. The relevant chapters can be skipped. For Chart parameter differences during deployment, see the containerd notes in Deploy FluxSandbox.
Deploy OpenSandbox Server
Sandboxes use OpenSandbox as the entry point; you need to deploy the OpenSandbox Server and configure its runtime as flux. Both the Chart and the Server image can be obtained directly from the official repository.
Create the configuration file
config.toml, set the runtime type tofluxand point it to the FluxSandbox Controller:toml[server] host = "0.0.0.0" port = 80 max_sandbox_timeout_seconds = 86400 workers = 15 thread_pool_size = 200 timeout_keep_alive = 120 limit_concurrency = 0 backlog = 65535 loop = "uvloop" http = "httptools" [log] level = "INFO" [runtime] type = "flux" execd_image = "opensandbox/execd:v1.0.20" [storage] allowed_host_paths = [] volume_default_size = "1Gi" [flux_sandbox] endpoint = "flux-sandbox-controller.flux-system.svc:9091" timeout_seconds = 60 use_tls = falseNote:
execd_imageis a required configuration item for the Server; the[ingress]block is auto-generated by the Chart (direct mode by default), do not include it again in the configuration file. The Chart's Service port is fixed at 80,portmust remain 80.- If you need to build the image from source (alternative method): clone the openFuyao/opensandbox repository and execute
docker build -f server/Dockerfile.flux -t opensandbox-server:flux ., then pointserver.image.repository/tagbelow to this image. Note thatDockerfile.fluxdepends on BuildKit features; if the node does not have docker buildx, using the official image directly is recommended.
Deploy the Server using the Chart:
Online deployment (cluster nodes can access the image repository):
helm install opensandbox-server \
oci://cr.openfuyao.cn/charts/opensandbox-server --version 0.2.1-of.1 \
-n flux-system --create-namespace \
--set namespaceOverride=flux-system \
--set server.image.repository=openfuyao-0qfqnk.swr-pro.myhuaweicloud.com/openfuyao/opensandbox/flux-server \
--set server.image.tag=0.2.1-of.1 \
--set server.replicaCount=1 \
--set 'server.env[0].name=OPENSANDBOX_INSECURE_SERVER' \
--set 'server.env[0].value=YES' \
--set-file configToml=config.tomlNote:
helm only downloads the Chart package (templates); the Server image is automatically pulled from the image repository by the node's kubelet when creating the Pod according toimagePullPolicy(defaultIfNotPresent), no manual import required.
Offline deployment (cluster nodes cannot access the image repository):
On a machine with internet access, download the Chart package and pull/export the Server image:
# 1. Download the Chart package
helm pull oci://cr.openfuyao.cn/charts/opensandbox-server --version 0.2.1-of.1
# → opensandbox-server-0.2.1-of.1.tgz
# 2. Pull the Server image
docker pull openfuyao-0qfqnk.swr-pro.myhuaweicloud.com/openfuyao/opensandbox/flux-server:0.2.1-of.1
# 3. Export as an offline image package
docker save -o opensandbox-server-image.tar \
openfuyao-0qfqnk.swr-pro.myhuaweicloud.com/openfuyao/opensandbox/flux-server:0.2.1-of.1
# 4. Transfer to the target cluster
scp opensandbox-server-0.2.1-of.1.tgz opensandbox-server-image.tar root@<node-ip>:/root/Import the image on the target node:
ctr -n k8s.io images import /root/opensandbox-server-image.tar
crictl images | grep flux-serverInstall on the control node. Compared to online deployment, only an additional server.image.pullPolicy=Never is appended: on offline nodes, kubelet only checks local images; the image name must exactly match the one imported on the node, otherwise ErrImageNeverPull is reported:
helm install opensandbox-server /root/opensandbox-server-0.2.1-of.1.tgz \
-n flux-system --create-namespace \
--set namespaceOverride=flux-system \
--set server.image.repository=openfuyao-0qfqnk.swr-pro.myhuaweicloud.com/openfuyao/opensandbox/flux-server \
--set server.image.tag=0.2.1-of.1 \
--set server.image.pullPolicy=Never \
--set server.replicaCount=1 \
--set 'server.env[0].name=OPENSANDBOX_INSECURE_SERVER' \
--set 'server.env[0].value=YES' \
--set-file configToml=config.tomlNote: (applies to both online/offline deployment)
-n flux-systemmust be explicitly included: without-n, the release record lands in the default namespace of the current context.- After uninstalling or deleting the Deployment and Service,
serviceaccount/opensandbox-serverandconfigmap/opensandbox-server-configwill remain (with helm ownership annotations); they must be deleted together before reinstalling.OPENSANDBOX_INSECURE_SERVER=YES: must be set; its absence causes the server worker to repeatedly exit.
Deploy FluxSandbox
Method 1: Helm Deployment (Recommended)
CRDs are installed automatically with the Chart. When using official images, only the global.imageRegistry prefix needs to be set.
Online deployment (cluster nodes can access the image repository):
helm install flux-sandbox \
oci://cr.openfuyao.cn/charts/flux-sandbox --version 26.9.0 \
--namespace flux-system --create-namespace \
--set global.imageRegistry=cr.openfuyao.cn/openfuyao \
--set watcher.args.criSocket="/run/cri-multiplex.sock"Note:
helm only downloads the Chart package (templates); container images are automatically pulled from the image repository by each node's kubelet when creating Pods according toimagePullPolicy(defaultIfNotPresent), no manual import required. For private repositories, executehelm registry loginbefore installation.
Offline deployment (cluster nodes cannot access the image repository):
On a machine with internet access, download the Chart package and pull/export the four component images:
# 1. Download the Chart package
helm pull oci://cr.openfuyao.cn/charts/flux-sandbox --version 26.9.0
# → flux-sandbox-26.9.0.tgz
# 2. Pull the four component images
IMAGES="
cr.openfuyao.cn/openfuyao/flux-sandbox/controller:26.9.0
cr.openfuyao.cn/openfuyao/flux-sandbox/scheduler:26.9.0
cr.openfuyao.cn/openfuyao/flux-sandbox/agent:26.9.0
cr.openfuyao.cn/openfuyao/flux-sandbox/watcher:26.9.0
"
for img in $IMAGES; do docker pull "$img"; done
# 3. Export as an offline image package
docker save -o flux-sandbox-images.tar $IMAGES
# 4. Transfer to the target cluster
scp flux-sandbox-26.9.0.tgz flux-sandbox-images.tar user@your-cluster-ip:/root/Import the image on each worker node (controller/scheduler may be scheduled to any worker node, watcher is a DaemonSet with one per node, and Agent Pods may also land on any worker node):
# When the cluster runtime is containerd, import into the k8s.io namespace
ctr -n k8s.io images import /root/flux-sandbox-images.tar
# Verify successful import
crictl images | grep flux-sandbox
# For docker runtime nodes, use: docker load -i /root/flux-sandbox-images.tarInstall on the control node. Compared to online deployment, only three additional pullPolicy=Never are appended, and the image tag uses the default latest (--set *.image.tag must match the actual image tag imported on the node):
helm install flux-sandbox /root/flux-sandbox-26.9.0.tgz \
--namespace flux-system --create-namespace \
--set global.imageRegistry=cr.openfuyao.cn/openfuyao \
--set controller.image.pullPolicy=Never \
--set scheduler.image.pullPolicy=Never \
--set watcher.image.pullPolicy=Never \
--set watcher.args.criSocket="/run/cri-multiplex.sock"Note:
Offline deployment must retainglobal.imageRegistryand appendpullPolicy=Never: when the image tag islatest, kubelet may treat it asAlwaysand access the registry; offline nodes will reportImagePullBackOff; under theNeverpolicy, kubelet only checks local images, and the image name must exactly match the one imported on the node, otherwiseErrImageNeverPullis reported.
For single-node deployment when you need to reduce the number of scheduler replicas (the Chart deploys 3 replicas by default), append --set scheduler.replicaCount=1, which also drives the controller's --expected-scheduler-replicas. See Table 3 for other common Helm values.
containerd Runtime: For a pure containerd environment (without E2B components), it is recommended to disable E2B configuration injection and adjust the watcher's scan socket:
helm install flux-sandbox ./charts/flux-sandbox \
--namespace flux-system --create-namespace \
--set global.imageRegistry=cr.openfuyao.cn/openfuyao \
--set e2b.enabled=false \
--set watcher.args.containerdSocket=/run/containerd/containerd.sockNote:
Aftere2b.enabled=false, components no longer mount the E2B configuration volume; the watcher targets the E2B cluster by default (criSocket="", containerdSocket=""), and containerd clusters need to be adjusted to the corresponding socket path as shown above.
Common customization options:
Table 3 Common Helm values
| Parameter | Default | Description |
|---|---|---|
global.imageRegistry | "" | Registry prefix for all images |
global.imagePullSecrets | [] | Image pull credentials |
scheduler.replicaCount | 3 | Number of scheduler shards |
e2b.enabled | true | Whether to inject E2B Template Manager configuration; set to false for pure containerd environments |
e2b.apiEndpoint | http://api.e2b.svc.cluster.local:3000 | E2B Template Manager address |
e2b.existingSecret | "" | Name of the Secret holding the E2B config.json; using a Secret to manage E2B credentials is recommended |
e2b.hostPath | /root/.e2b | Legacy method: directory on the host holding config.json |
watcher.args.criSocket | /run/cri-multiplex.sock | E2B CRI socket scanned by watcher; set to empty to disable E2B orphan scanning |
watcher.args.containerdSocket | "" | containerd socket scanned by watcher; set to /run/containerd/containerd.sock for containerd clusters |
*.nodeSelector / *.tolerations | {} | Scheduling constraints for each component |
For complete values documentation, see the flux-sandbox Chart README.
Method 2: Raw Manifests
# Create namespace
kubectl create namespace flux-system
# Install CRDs
make crd-apply
# Deploy each component
kubectl apply -f config/controller/
kubectl apply -f config/scheduler/
kubectl apply -f config/sandbox-watcher/Verify Deployment
kubectl wait --for=condition=Ready pod -n flux-system -l app=flux-sandbox-controller --timeout=120s
kubectl wait --for=condition=Ready pod -n flux-system -l app=flux-sandbox-scheduler --timeout=120sFor the E2B runtime, you should also confirm that the Template Manager configuration is ready: the component logs should show a message indicating successful Template Manager initialization. If E2B credentials are not configured, the components will start in a degraded mode (with a log warning); after the components start, creating sandboxes through an E2B-type SandboxGroup will fail with the error "E2B runtime requires template manager but it is not configured".
Preparation Before Use
After completing deployment, you need to create a SandboxGroup and prepare sandbox templates (E2B runtime) or container images (containerd runtime), after which you can use sandboxes through the OpenSandbox SDK.
Create SandboxGroup
A SandboxGroup defines the resource specifications and capacity of a set of sandboxes and is the only CRD resource users need to create. The Controller splits it into SubSandboxGroups (shards) and assigns them to each scheduler instance, which then starts Agent Pods. Once the sandbox group status becomes Ready, sandboxes can be created.
Execute
kubectl apply -f -to create a SandboxGroup (no additional configuration required for the E2B runtime):bashkubectl apply -f - <<EOF apiVersion: sandbox.flux.io/v1alpha1 kind: SandboxGroup metadata: name: sg-1u1g namespace: flux-system spec: capacity: agentPodMin: 1 agentPodMax: 1 sandboxBufferMin: 1 maxSandboxesPerPod: 5 sandboxResources: requests: { cpu: "1", memory: "2Gi" } limits: { cpu: "1", memory: "2Gi" } EOFCapacity field descriptions:
Table 4 capacity field descriptions
Field Required Description agentPodMin/agentPodMaxYes Lower/upper limit of Agent Pod count; the system auto-scales within this range. sandboxBufferMin/sandboxBufferMaxNo Lower/upper limit of warm buffer idle sandbox count, used to reduce sandbox acquisition latency. maxSandboxesPerPodYes Maximum number of sandboxes allowed to coexist on a single Agent Pod. Sharding mechanism example: Suppose the SandboxGroup is configured with
agentPodMax: 200andmaxSandboxesPerPod: 5, and the Controller startup parameter--max-pod-per-sub-sandbox-groupuses the default value100. The system will automatically split this SandboxGroup into 2 SubSandboxGroup shards (named like<SandboxGroup name>-0,<SandboxGroup name>-1) and assign them to two scheduler instances:- Each shard manages at most 100 Agent Pods, with a maximum capacity of 100 × 5 = 500 Sandboxes;
- The entire SandboxGroup has a maximum capacity of 200 × 5 = 1000 Sandboxes;
- When
agentPodMaxis not evenly divisible by--max-pod-per-sub-sandbox-group, round up; the last shard takes the remainder (e.g.,agentPodMax: 250splits into 3 shards, with the last shard managing 50 Agent Pods).
To adjust the upper limit of creatable Sandbox count:
- Adjust
agentPodMaxormaxSandboxesPerPodin the SandboxGroup spec; total capacity = agentPodMax × maxSandboxesPerPod. WhenmaxSandboxesPerPodis increased, the Agent Pod will reserve resources for more slots based onsandboxResources, and the per-Pod resource request increases linearly; generally, adjustagentPodMaxfirst. - The startup parameter
--max-pod-per-sub-sandbox-group(Helm deployment valuecontroller.args.maxPodPerSubSandboxGroup, default100) does not change total capacity, only controls sharding granularity: increasing it reduces the number of shards, decreasing it spreads shards across more scheduler instances for load balancing. If you do not want sharding, set it to no less thanagentPodMax(e.g., 200 in the above example).
Wait for the SandboxGroup to become ready:
bashkubectl wait --for=condition=Ready sandboxgroup sg-1u1g -n flux-system --timeout=120s
containerd Runtime: When using the containerd runtime, the SandboxGroup must explicitly declare the runtime type and execd image (execd carries the SDK's command execution, required):
spec:
agentEnv:
RUNTIME_TYPE: "containerd"
execdImage: "opensandbox/execd:v1.0.20"Note:
The runtime type is controlled byagentEnv.RUNTIME_TYPE(default is E2B);execdImageonly takes effect for the containerd runtime and is invalid under the E2B runtime.
Prepare E2B Templates
Under the E2B runtime, the image passed when creating a sandbox is the E2B template name (not a container image reference). The template determines the operating system environment, pre-installed software, and processes inside the microVM. The template must meet:
- The template has been built and registered in the E2B environment accessible to the Template Manager through the E2B template mechanism, and the template name is globally available (e.g.,
ubuntu-22-04-custom). - The template has a built-in envd daemon, carrying capabilities such as command execution and file operations for the SDK (official base templates already include this).
- The actual processes and environment running inside the sandbox are determined by the template definition (the
entrypointandenvin the creation request are not injected into the microVM).
When the template name in the creation request does not exist, the Template Manager resolution fails, and the creation request returns an error.
containerd Runtime: image is a standard container image reference; ensure the image can be pulled on all Worker nodes:
# Sandbox runtime image (SDK example uses python:3.11)
docker pull python:3.11
# execd image (specified by execdImage in SandboxGroup)
docker pull opensandbox/execd:v1.0.20Note:
Images must be pushed to a registry accessible by the cluster, or pre-pulled on all Worker nodes.
Using Sandboxes
Prerequisites
- FluxSandbox deployment, SandboxGroup creation, and OpenSandbox Server integration are complete.
- The OpenSandbox Python SDK (adapted version, installation method below) is installed locally, and the Server access address is obtained.
Install OpenSandbox Python SDK (Adapted Version)
This repository (openFuyao/opensandbox) adapts the official SDK (including snapshot APIs aligned with the FluxSandbox Server, etc.) and is not published to PyPI. Executing pip install opensandbox directly will install the PyPI official package, which lacks the above adaptations; you must install using one of the two methods below.
Method 1: With internet access, install directly from the GitCode repository
pip install "git+https://gitcode.com/openFuyao/opensandbox.git@of-dev/v0.2.1#subdirectory=sdks/sandbox/python"Note:
#subdirectory=sdks/sandbox/pythoncannot be omitted; the SDK source code is located in a subdirectory of the repository.
Method 2: Offline environment, install from offline wheel package
Build a wheelhouse on a machine with internet access. The Python minor version and system architecture used for building must match the intranet machine (e.g., both Python 3.12 + linux x86_64), because the dependency pydantic-core is a compiled wheel bound to the Python version and platform:
git clone -b of-dev/v0.2.1 https://gitcode.com/openFuyao/opensandbox.git
cd opensandbox
# Build the SDK itself and all runtime dependencies into the wheelhouse directory
pip wheel -w wheelhouse ./sdks/sandbox/python
# Package and transfer to the intranet machine
tar czf opensandbox-sdk-wheelhouse.tar.gz wheelhouse
scp opensandbox-sdk-wheelhouse.tar.gz root@<intranet-machine-IP>:/root/Install offline on the intranet machine:
tar xzf opensandbox-sdk-wheelhouse.tar.gz
pip install --no-index --find-links=wheelhouse wheelhouse/opensandbox-*.whlNote:
--no-indexforces pip to only take packages from the wheelhouse directory, without accessing any package sources.- When building on a machine with internet access, the
.gitdirectory must be retained, otherwise the version number cannot be derived from the git tag.
Verify Installation
pip show opensandboxA version number with a dev and +g<commit-id> suffix (e.g., 0.1.14.dev26+g6066bc24) indicates the adapted version; a clean 0.1.x version number means the PyPI official package was installed, and you must reinstall.
python -c "
from opensandbox import Sandbox # Async SDK entry
from opensandbox.sync.sandbox import SandboxSync # Sync SDK entry
print(hasattr(Sandbox, 'create_snapshot')) # Should output True
print(hasattr(SandboxSync, 'create_snapshot')) # Should output True
"Both outputs being True indicates the snapshot API adaptation is in place.
Warning:
The adapted version number (0.1.14.devNN+g…) is lower than the PyPI official package version (0.1.16); do not executepip install -U opensandboxto upgrade this package, otherwise it will be overwritten by the official package, resulting in the loss of adapted capabilities such as snapshots; to upgrade the adapted version, rebuild and install using the above method.
Create Sandbox Parameter Description
Creating a sandbox is uniformly done through the OpenSandbox SDK/REST API; FluxSandbox does not expose a separate interface. The required/optional parameter rules are as follows (validation is performed by the OpenSandbox Server; image mode and snapshot_id mode have different rules):
Table 5 Create sandbox parameters (image mode)
| Parameter | Required | Default | Description |
|---|---|---|---|
image | Required (choose one with snapshot_id) | - | E2B template name (container image reference for the containerd runtime). |
entrypoint | Required | SDK default ["tail", "-f", "/dev/null"] | Sandbox main process entry. REST direct call returns 400 if not passed; SDK uses default if not passed. Not injected into microVM for E2B runtime; only takes effect for containerd runtime. |
resource_limits (REST field name resourceLimits) | Required | SDK default {"cpu": "1", "memory": "2Gi"} | Resource specification. REST direct call returns 400 if not passed; SDK uses default if not passed. The actual effective specification is subject to SandboxGroup's sandboxResources (see below). |
extensions.sandboxGroup | Required | - | Target SandboxGroup name; the group must be Ready. Returns 400 if missing. |
timeout | Optional | SDK default 10 minutes; REST does not auto-expire if not passed | Sandbox auto-expiration duration; automatically reclaimed upon expiration. Must not be less than 60 seconds and must not exceed the Server-configured upper limit (max_sandbox_timeout_seconds). |
env | Optional | {} | Environment variables. Only injected into sandbox for containerd runtime; not injected for E2B runtime. |
metadata | Optional | {} | User-defined labels; keys and values must conform to Kubernetes label specifications. |
resource_requests (REST field name resourceRequests) | Optional | Same as resource_limits | Resource request value; defaults to resource_limits when not specified. The actual effective specification is subject to SandboxGroup. |
platform / network_policy / volumes / secure_access / credential_proxy | Optional | - | Platform constraints, egress policies, storage mounts, and other advanced configurations; use according to OpenSandbox general semantics. |
Table 6 Create sandbox parameters (snapshot_id mode, E2B runtime only)
| Parameter | Required | Description |
|---|---|---|
snapshot_id | Required (choose one with image) | Source snapshot ID; the snapshot must be in Ready state, otherwise returns 404/409. |
entrypoint | Optional | When not passed, the server automatically uses ["tail", "-f", "/dev/null"]. |
image / resource_limits | Optional | image is automatically injected from the snapshot record; the actual effective specification of resource_limits is subject to SandboxGroup. |
extensions.sandboxGroup / timeout / other parameters | Same as image mode | - |
Note:
- Required description for
entrypointandresource_limits: REST requires these for Server parameter validation, but FluxSandbox uses the SandboxGroup configuration as the actual effective value; the parameters in the request do not determine the actual specification. To adjust the specification, modify the corresponding SandboxGroup.- A single request must provide exactly one of
imageandsnapshot_id, otherwise returns 400.
Usage Limitations
- When creating a sandbox, the target SandboxGroup must be specified via
extensions.sandboxGroup, otherwise the request is rejected (400). - Sandbox resource specifications are uniformly determined by SandboxGroup's
sandboxResources; resource parameters passed during creation do not take effect (but must still pass the Server's parameter validation, see Table 5). - Under the E2B runtime,
entrypointandenvare not injected into the sandbox; the processes and environment inside the sandbox are determined by the E2B template. - Pause, resume, and snapshot capabilities are only supported by the E2B runtime; calling the relevant interfaces under the containerd runtime returns an unsupported error.
- Only sandboxes in the
Runningstate can have snapshots created; a single sandbox allows only one snapshot operation at a time.
Quick Start
Connect to the OpenSandbox Server and create a sandbox:
import asyncio
from datetime import timedelta
from opensandbox import Sandbox
from opensandbox.config import ConnectionConfig
async def main():
config = ConnectionConfig(
domain="<server-address>:80", # OpenSandbox Server address (Chart default Service port 80)
)
sandbox = await Sandbox.create(
"ubuntu-22-04-custom", # E2B template name (for containerd runtime, pass image reference, e.g., "python:3.11")
connection_config=config,
timeout=timedelta(minutes=30),
extensions={"sandboxGroup": "sg-1u1g"}, # Required: route to SandboxGroup
skip_health_check=True,
)
print("created:", sandbox.id)
info = await sandbox.get_info()
print("state:", info.status.state)
asyncio.run(main())Delete the sandbox:
python3 kill_sandbox.py <sandbox-id>import asyncio, sys
from datetime import timedelta
from opensandbox.config import ConnectionConfig
from opensandbox.sandbox import Sandbox
async def main(sandbox_id):
config = ConnectionConfig(domain="<server-address>", api_key="")
sb = await Sandbox.get(sandbox_id, connection_config=config)
await sb.kill()
await sb.close()
print(f"Sandbox {sandbox_id} destroyed")
asyncio.run(main(sys.argv[1]))For complete execution commands, see envd_manual_verify.md.
Create Sandbox
Create from template/image:
sandbox = await Sandbox.create(
"ubuntu-22-04-custom", # image: E2B template name (or container image reference)
connection_config=config,
entrypoint=["tail", "-f", "/dev/null"], # Optional, SDK default is this value
env={"FOO": "bar"}, # Environment variables (only injected for containerd runtime)
timeout=timedelta(minutes=30), # Auto-expiration time; REST does not auto-expire if not passed
extensions={"sandboxGroup": "sg-1u1g"},
)Create from snapshot (restores the complete state of the snapshot, ready in seconds; E2B runtime only):
sandbox = await Sandbox.create(
connection_config=config,
snapshot_id="<snapshot-id>", # Choose one with image
timeout=timedelta(minutes=30),
extensions={"sandboxGroup": "sg-1u1g"},
)Note:
You must pass exactly one ofimageandsnapshot_id. The snapshot must be inReadystate; a non-existent or not-ready snapshot returns 404/409. When creating from a snapshot, you do not need to passimageandentrypoint; the server automatically uses the image from the snapshot record and the default entrypoint.
containerd Runtime Usage Example
A complete usage flow using the containerd runtime as an example: create sandbox → execute command → stream output → file read/write → query status. Example requirements:
- SandboxGroup is for the containerd runtime (see the containerd notes in Create SandboxGroup), and the execd image is configured;
- The sandbox image
python:3.11and the execd imageopensandbox/execd:v1.0.20can be pulled on Worker nodes.
Access the OpenSandbox Server locally via port-forward:
kubectl port-forward -n flux-system svc/opensandbox-server 8080:80Complete example code:
import asyncio
from datetime import timedelta
from opensandbox import Sandbox
from opensandbox.config import ConnectionConfig
from opensandbox.models.execd import ExecutionHandlers
from opensandbox.models.filesystem import WriteEntry, SearchEntry
async def main():
config = ConnectionConfig(
domain="localhost:8080",
api_key="your-secret-key",
use_server_proxy=True,
)
# ---------- Complete flow for a brand-new sandbox ----------
async with await Sandbox.create(
"python:3.11",
connection_config=config,
timeout=timedelta(minutes=30), # Note: SDK 0.1.16 this parameter not passed to server
extensions={"sandboxGroup": "sg-1u1g"},
) as sandbox:
print(f"sandbox id: {sandbox.id}")
# 1. Execute command + exit code
result = await sandbox.commands.run("python -c 'print(1 + 1)'")
print(f"exit_code={result.exit_code}, out={result.logs.stdout[0].text.strip()}")
# 2. Streaming output (SDK 0.1.16 requires async callbacks, both on_stdout/on_stderr required)
async def on_stdout(msg):
print(f"STDOUT: {msg.text}")
async def on_stderr(msg):
print(f"STDERR: {msg.text}")
handlers = ExecutionHandlers(on_stdout=on_stdout, on_stderr=on_stderr)
await sandbox.commands.run("for i in 1 2 3; do echo $i; done", handlers=handlers)
# 3. File operations: write / read / search
await sandbox.files.write_files([
WriteEntry(path="/tmp/hello.txt", data="Hello World", mode=644)
])
content = await sandbox.files.read_file("/tmp/hello.txt")
print(f"file content: {content}")
files = await sandbox.files.search(SearchEntry(path="/tmp", pattern="*.txt"))
print(f"search: {[f.path for f in files]}")
# 4. Query status
info = await sandbox.get_info()
print(f"state: {info.status.state}, expires: {info.expires_at}")
asyncio.run(main())Note:
- The
timeoutparameter of SDK 0.1.16 is not passed through to the server; the sandbox will not auto-expire based on this value. To enable auto-expiration, control it via thetimeoutfield (in seconds, minimum 60) of the REST request body.- The streaming callbacks of SDK 0.1.16 require
on_stdout/on_stderrto both be async functions and both are mandatory, otherwise no output is received.api_keyis only required when the Server deployment uses authentication;use_server_proxy=Trueindicates accessing the sandbox through the Server proxy.
Manage Snapshots
Snapshots capture the complete state of a sandbox at a given point in time and can be used for environment preservation, batch cloning, and fault rollback. Snapshot creation is a synchronous operation that returns with the final state (Ready or Failed); during creation, the sandbox remains Running, and on failure, the sandbox automatically resumes running.
[containerd Runtime] Snapshot capabilities are only supported by the E2B runtime; this section does not apply.
Manage snapshots (create, query, list, delete) via SandboxManager, and restore sandboxes from snapshots via Sandbox.create(snapshot_id=...):
import asyncio
from datetime import timedelta
from opensandbox import Sandbox, SandboxManager
from opensandbox.config import ConnectionConfig
from opensandbox.models.sandboxes import SnapshotFilter
config = ConnectionConfig(domain="<server-address>:80") # OpenSandbox Server address
async def main():
# ---------- Create snapshot ----------
# Sandbox must be in Running state; creation is a synchronous operation that returns with the final state (Ready or Failed)
async with await SandboxManager.create(connection_config=config) as manager:
snap = await manager.create_snapshot(
sandbox_id="<sandbox-id>", # Required: source sandbox ID, sandbox must be Running
name="my-snapshot", # Optional: snapshot name
)
print(f"snapshot: {snap.id}, state: {snap.status.state}")
# ---------- Query a single snapshot ----------
info = await manager.get_snapshot(snap.id) # Required: snapshot ID
print(f"state: {info.status.state}, reason: {info.status.reason}")
# ---------- List snapshots (supports filtering by source sandbox, state, and pagination) ----------
snaps = await manager.list_snapshots(
SnapshotFilter(
sandbox_id="<sandbox-id>", # Optional: filter by source sandbox ID
states=["READY"], # Optional: filter by state (READY/FAILED)
page_size=20, # Optional: items per page, default 20
page=1, # Optional: page number, starting from 1
)
)
for s in snaps.snapshot_infos:
print(f"{s.id}: {s.status.state}")
# ---------- Delete snapshot ----------
await manager.delete_snapshot(snap.id) # Required: snapshot ID
print(f"deleted: {snap.id}")
# ---------- Restore sandbox from snapshot ----------
# Snapshot must be in Ready state; E2B sandbox has no execd process, SDK automatically skips endpoint resolution and health check
# snapshot_id and image are mutually exclusive; when restoring, only pass snapshot_id, image is automatically injected from the snapshot record
sandbox = await Sandbox.create(
snapshot_id=snap.id, # Required: source snapshot ID (choose one with image)
connection_config=config, # Required: Server connection configuration
timeout=timedelta(minutes=30), # Optional: auto-expiration time; no auto-expiration if not passed
resource={"cpu": "1", "memory": "2Gi"}, # Optional: resource specification (actual value subject to SandboxGroup)
extensions={"sandboxGroup": "sg-1u1g"}, # Required: target SandboxGroup name
)
print(f"restored sandbox: {sandbox.id}")
info = await sandbox.get_info()
print(f"state: {info.status.state}")
await sandbox.close()
asyncio.run(main())Note:
Snapshot records are persistently stored and can still be queried and used after the OpenSandbox Server restarts. Deleted snapshots can no longer be used to create sandboxes. After creation, snapshots can be pre-warmed and distributed to SuperPod caches to further reduce the latency of creating sandboxes from snapshots; see Snapshot Pre-warming.
Query and Manage Existing Sandboxes
from opensandbox.manager import SandboxManager
from opensandbox.models.sandboxes import SandboxFilter
async with await SandboxManager.create(connection_config=config) as manager:
# List running sandboxes
sandboxes = await manager.list_sandbox_infos(
SandboxFilter(states=["RUNNING"], page_size=10)
)
for info in sandboxes.sandbox_infos:
print(f"Found sandbox: {info.id}")
# Terminate the specified sandbox
await manager.kill_sandbox(info.id)Sandbox State Description
Table 7 Sandbox state transitions
| State | Description |
|---|---|
| PENDING | Being created, not yet ready. |
| RUNNING | Running; commands can be executed, files operated, and snapshots created. |
| PAUSING | Pausing (transient); does not accept new pause/resume requests during this period. |
| PAUSED | Paused; processes suspended, resource placeholders released. |
| RESUMING | Resuming (transient). |
| TERMINATED | Terminated. |
Core transition: PENDING → RUNNING ⇄ (pause/resume) PAUSED → TERMINATED.
Table 8 Snapshot states
| State | Description |
|---|---|
| Ready | Snapshot is available and can be used to create sandboxes. |
| Failed | Snapshot creation failed; the source sandbox has automatically resumed running. |
Common Error Descriptions
Table 9 Common error scenarios
| Scenario | Error Code | Suggested Handling |
|---|---|---|
Creating sandbox without passing extensions.sandboxGroup | 400 | Add the sandboxGroup field with the target SandboxGroup name. |
Creating sandbox without passing image/snapshot_id, or passing both | 400 | Provide exactly one of the two. |
Not passing entrypoint or resource_limits in image mode (REST direct call) | 400 | Add the required parameters; or use the SDK (has defaults). |
timeout less than 60 seconds or exceeding the Server-configured upper limit | 400 | Adjust the timeout value. |
| Specified E2B template does not exist (E2B runtime) | 500 | Confirm the template has been built and registered in the E2B environment, and the template name matches the request. |
| Specified container image does not exist (containerd runtime) | 500 | Confirm the image has been pushed to a registry accessible by the cluster or pre-pulled on nodes. |
| Querying/deleting a non-existent sandbox or snapshot | 404 | Verify the ID is correct and whether the resource has been deleted. |
| Creating a snapshot from a non-Running sandbox | 409 | Restore the sandbox to Running state first. |
| Repeatedly initiating a snapshot while one is in progress | 409 | Wait for the current snapshot operation to complete. |
| Performing a resume on a running sandbox | 409 | Only paused (PAUSED) sandboxes need to be resumed. |
| Repeated operations during pause/resume | 409 | The state is in the PAUSING/RESUMING transient; wait for the transition to complete. |
| Pause/resume/snapshot operations return unsupported | - | The SandboxGroup uses the containerd runtime; the relevant capabilities are only supported by the E2B runtime. |
| Pause failed (PauseFailed) | - | Check whether the target Agent Pod is ready, then retry. |
Snapshot Pre-warming
[containerd Runtime] Snapshot pre-warming depends on E2B snapshot capabilities and is only available for the E2B runtime; this section does not apply.
Snapshot pre-warming actively distributes and caches existing snapshots to multiple SuperPods, so that when creating sandboxes from the snapshot, no on-demand fetching is required, reducing cold start latency. The pre-warming subsystem is embedded in the Controller process and is invoked via the Controller's HTTP interface (default :8080, API prefix /api/v1), suitable for integration with external cluster management systems.
Table 10 Pre-warming interface overview
| Method | Path | Purpose |
|---|---|---|
| POST | /api/v1/snapshot-warm-tasks | Create a pre-warming task (executed asynchronously) |
| GET | /api/v1/snapshot-warm-tasks/{task_id} | Query task pre-warming progress (paginated) |
| GET | /api/v1/snapshots/cache-details | Query snapshots cached on each SuperPod (paginated) |
| GET | /api/v1/superpods/storage-details | Query storage availability of each SuperPod |
All responses carry the X-Request-ID response header (passed through from the request's same-named header, auto-generated if not present), which can be used for issue tracking. Error responses uniformly use the {"error": {"code", "message", "retryable"}} structure.
Enable Pre-warming
The pre-warming subsystem is disabled by default and must be explicitly enabled in the Controller startup parameters. All the following parameters are required; missing any one will cause startup failure:
--warm-enabled=true
--warm-postgres-dsn=<PostgreSQL connection string> # Task and pre-warming detail persistence
--warm-etcd-endpoints=<MoonCake etcd address> # Storage availability query
--warm-superpod-ids=<comma-separated SuperPod ID list> # Pre-warming target pool
--warm-e2b-api-endpoint=<E2B API-Server address> # Required when pre-warming backend is e2bThe pre-warming backend is specified via --warm-backend-type, default e2b, optional node (requires --warm-node-port, default 44772). For the database schema script, see flux-sandbox repository warm-migrations/0001_warm.sql.
Create Pre-warming Task
curl -X POST http://<controller>:8080/api/v1/snapshot-warm-tasks \
-H 'Content-Type: application/json' \
-d '{"snapshot_ids": ["snap-001", "snap-002", "snap-003"]}'Response 201 Created:
{
"task_id": "warm-task-20260908-1786230000000000000-1",
"state": "WAIT",
"accepted_snapshot_count": 3,
"created_at": "2026-09-08T02:00:00Z"
}The task executes asynchronously with state transitions: WAIT (enqueued) → WARMING (executing) → SUCCEED / FAILED. Each snapshot is load-balanced and assigned to several SuperPods (default 5); SuperPods that fail pre-warming are automatically retried 3 times (with a 10-second interval), retrying only the failed SuperPods; the overall task timeout is 30 minutes.
snapshot_ids is required and must not contain empty strings or duplicates; the task executes on the full array as submitted, and batching is not supported. The request body is limited to 64 MiB (approximately 100,000 IDs); exceeding the limit returns 400.
Query Pre-warming Progress
curl 'http://<controller>:8080/api/v1/snapshot-warm-tasks/<task_id>?page_size=100&page_num=1'The summary in the response reflects overall task progress:
{
"total": 3,
"success_count": 2,
"failed_count": 0,
"in_progress_count": 1
}- Whether the task has ended is determined by
success_count + failed_count == total. - The
snapshotsarray returns the pre-warming results for each snapshot in submission order with pagination:warm_success=trueindicates that all targets for this snapshot have completed pre-warming;warmed_superpod_idslists the SuperPods that have actually cached this snapshot (snapshots being pre-warmed or failed will also list the successfully completed portions). page_sizedefaults to 1000 with an upper limit of 10000 (out-of-range returns 400);page_numstarts from 1. A non-existent task ID returns 404.
Query Cache Distribution and Storage Availability
# Snapshots cached on each SuperPod (only counts SuperPods that were actually pre-warmed successfully)
curl 'http://<controller>:8080/api/v1/snapshots/cache-details?page_size=1000&page_num=1'
# MoonCake storage availability of each SuperPod, can be used for capacity assessment before pre-warming
curl http://<controller>:8080/api/v1/superpods/storage-details- Cache details are returned grouped by SuperPod, including fields such as
snapshot_id,category,cached_at. - Storage availability data is read directly from MoonCake etcd reports: if
observed_atis more than 60 seconds ago, it is considered stale and that SuperPod'shealth=UNAVAILABLE; configured SuperPods always appear in the response (those without reports or with stale data show asUNAVAILABLE).
Typical Invocation Flow
BASE=http://<controller>:8080/api/v1
# 1. Submit pre-warming task
TASK_ID=$(curl -s -X POST $BASE/snapshot-warm-tasks \
-H 'Content-Type: application/json' \
-d '{"snapshot_ids": ["snap-001", "snap-002"]}' | jq -r .task_id)
# 2. Poll progress until all snapshots have results
curl -s "$BASE/snapshot-warm-tasks/$TASK_ID?page_size=1000" | jq .summary
# 3. View cache distribution and SuperPod storage levels
curl -s "$BASE/snapshots/cache-details" | jq .
curl -s "$BASE/superpods/storage-details" | jq .Pre-warming Common Errors
Table 11 Pre-warming common errors
| Error Code | Trigger Scenario | Retryable |
|---|---|---|
| 400 INVALID_ARGUMENT | Request body is invalid or exceeds limit; snapshot_ids is empty, contains empty strings, or contains duplicates; pagination parameters out of range | No |
| 404 NOT_FOUND | Task ID does not exist | No |
| 503 UNAVAILABLE | PostgreSQL / MoonCake etcd temporarily unavailable | Yes |
Errors with retryable=true can be retried as-is after a short wait.
Usage Limitations
- The pre-warming interface has no authentication; network policies must be used to restrict access at deployment time.
- Tasks cannot be canceled or deleted after creation; failed snapshots need a new task submission (idempotent: successfully pre-warmed snapshots will not be re-distributed).
- Snapshots currently submitted via API are all processed as the
REGULARcategory, with each snapshot pre-warmed to 5 SuperPods by default.
View Running Status
# View SandboxGroup status and capacity
kubectl get sandboxgroup -n flux-system -o wide
# View sandbox instances
kubectl get sandbox -n flux-systemAfter the SandboxGroup is ready, status.phase is Ready; you can use kubectl describe sandboxgroup to view Conditions to locate issues. Sandbox running status is perceived via the SDK's get_info or management interfaces.
Upgrade and Uninstall
# Upgrade
helm upgrade flux-sandbox ./charts/flux-sandbox --namespace flux-system -f my-values.yaml
# Uninstall
helm uninstall flux-sandbox --namespace flux-systemNote:
CRDs are not deleted withhelm uninstall(established Helm behavior). To delete, executekubectl delete -f charts/flux-sandbox/crds/; note this will cascade-delete all sandbox-related CRs.
Appendix
SandboxGroup Spec Field Reference
Table 12 SandboxGroup Spec fields
| Field | Type | Required | Description |
|---|---|---|---|
capacity | object | Yes | Capacity declaration; fields are in Table 4 (agentPodMin, agentPodMax, maxSandboxesPerPod are required). |
sandboxResources | object | Yes | Resource specification per sandbox (requests/limits); the sole authoritative source for resource specifications. |
agentEnv | map | No | Environment variables injected into the Agent Pod, e.g., RUNTIME_TYPE: "containerd". |
execdImage | string | No (required under containerd runtime) | execd image address. Required under the containerd runtime (when set, an execd init container is automatically injected); not needed under the E2B runtime. |
agentPodScheduling | object | No | Node scheduling constraints for Agent Pods (tolerations/nodeSelector/affinity), which can direct the sandbox group to specific node pools. |
schedulingPolicy | object | No | Sandbox placement policy across Agent Pods (strategy, weights, topologyAffinity), optional advanced configuration. |
SandboxGroup Status Field Reference
Table 13 SandboxGroup Status fields
| Field | Description |
|---|---|
phase | Current lifecycle phase; Ready once ready. |
observedGeneration | The generation observed by the Controller, which can be used to determine whether a change has taken effect. |
subSandboxGroupsCount | Number of shards split. |
totalAgentPods | Current total number of Agent Pods. |
totalCapacity | Current total sandbox capacity. |
conditions | Conditions describing the group status, used for troubleshooting. |
SDK Operations and REST API Mapping
Non-Python users or gateway integration scenarios can directly call the OpenSandbox Server REST API:
Table 14 SDK operations and REST API mapping
| Operation | REST API |
|---|---|
| Create sandbox | POST /v1/sandboxes (202) |
| Query sandbox | GET /v1/sandboxes/{id} |
| List sandboxes | GET /v1/sandboxes |
| Delete sandbox | DELETE /v1/sandboxes/{id} |
| Update sandbox metadata | PATCH /v1/sandboxes/{id}/metadata |
| Renew sandbox | POST /v1/sandboxes/{id}/renew-expiration |
| Pause sandbox | POST /v1/sandboxes/{id}/pause |
| Resume sandbox | POST /v1/sandboxes/{id}/resume |
| Get port access address | GET /v1/sandboxes/{id}/endpoints/{port} |
| Create snapshot | POST /v1/sandboxes/{id}/snapshots |
| Query snapshot | GET /v1/snapshots/{id} |
| List snapshots | GET /v1/snapshots |
| Delete snapshot | DELETE /v1/snapshots/{id} |
Create sandbox request example (fields correspond to Table 5/Table 6):
curl -s -X POST "http://${SERVER}/v1/sandboxes" \
-H "Content-Type: application/json" \
-d "{
\"image\":{\"uri\":\"${TEMPLATE_OR_IMAGE}\"},
\"entrypoint\":[\"tail\",\"-f\",\"/dev/null\"],
\"resourceLimits\":{\"cpu\":\"1\",\"memory\":\"2Gi\"},
\"extensions\":{\"sandboxGroup\":\"${SG}\"},
\"timeout\":3600
}"Note:
- When calling REST directly,
entrypointandresourceLimitsare required inimagemode (the SDK auto-fills defaults);timeoutnot being passed means the sandbox does not auto-expire, in seconds, minimum 60.- SDK parameters use snake_case naming (e.g.,
resource_limits), while REST request body fields use camelCase naming (e.g.,resourceLimits).
FAQ
SandboxGroup is not ready after creation
Possible causes
- Controller or Scheduler components are not running properly.
- Agent Pod cannot be scheduled (insufficient node resources, taints not tolerated).
Solution
Execute
kubectl describe sandboxgroup <name> -n flux-systemto view Conditions to locate the cause, and check the status of components and Agent Pods:kubectl get pods -n flux-system.Creating a sandbox returns 400, prompting
extensions['sandboxGroup'] is requiredThe creation request must carry the
extensions.sandboxGroupfield specifying the target SandboxGroup, and the group must be Ready.Creating a sandbox returns 400, prompting that entrypoint or resourceLimits is required
When calling REST directly in
imagemode,entrypointandresourceLimitsare required parameters (the SDK has defaults and will not trigger this); add them per Table 5 and retry. Note that these two values do not affect the actual effective specification; the actual specification is determined by SandboxGroup.Creating a sandbox under the E2B runtime fails with "E2B runtime requires template manager but it is not configured"
The E2B runtime depends on the E2B Template Manager to resolve templates. Check:
- chart's
e2b.enabledistrue(default); - E2B
config.jsoncontainingteamApiKeyhas been provided viae2b.existingSecret(ore2b.hostPath); - the Template Manager service pointed to by
e2b.apiEndpointis reachable.
- chart's
Creating a sandbox under the E2B runtime fails, prompting that the template does not exist
The
imagein the creation request must be a built and registered E2B template name. Confirm the template exists in the E2B environment and the name matches exactly.Creating a sandbox fails under the containerd runtime
To use the containerd runtime in a standard K8s environment, you must: set
agentEnv.RUNTIME_TYPE=containerdandexecdImageon the SandboxGroup; ensure the sandbox image and execd image can be pulled. See the containerd runtime notes in each chapter.Resource parameters passed when creating a sandbox via the SDK do not take effect
Sandbox resource specifications are uniformly determined by SandboxGroup's
sandboxResources, the sole authoritative source; resource parameters passed by the caller are ignored (but must still pass the Server's parameter validation). To adjust the specification, modify the corresponding SandboxGroup.The entrypoint or environment variables passed in are not visible inside the E2B sandbox
This is expected behavior. Under the E2B runtime,
entrypointandenvare not injected into the microVM; the processes and environment inside the sandbox are determined by the E2B template; to customize the environment, build it through the template. Under the containerd runtime, both take effect.Image pull failure when creating a sandbox (containerd runtime)
Confirm that the sandbox runtime image (e.g.,
python:3.11) and the execd image have been pushed to a registry accessible by the cluster, or pre-pulled on all Worker nodes. The E2B runtime does not involve image pulling; the template is solidified during the build phase.Creating a sandbox from a snapshot returns 409
The snapshot must be in
Readystate. Query the snapshot status, confirm, and retry; snapshots inFailedstate are not available.Pause/resume/snapshot interfaces return unsupported
These capabilities are only supported by the E2B runtime. If the SandboxGroup uses the containerd runtime (
RUNTIME_TYPE=containerd), switch to a SandboxGroup using the E2B runtime.Are resources released after pausing a sandbox
Yes. After a sandbox enters the
PAUSEDstate, resource placeholders are released, and new sandboxes of the same capacity can be created immediately; when resumed, the running state is rebuilt based on the snapshot generated at pause time.Does snapshot creation failure affect the original sandbox
No. During snapshot creation, the sandbox remains
Running; on creation failure, the sandbox automatically resumes running, and the snapshot status is set toFailed.How to access ports listened to inside the sandbox from outside
Obtain the access address via
sandbox.get_endpoint(port)and connect directly; see Access Sandbox Ports.CRDs still exist after executing
helm uninstallThis is established Helm behavior. To clean up thoroughly, manually execute
kubectl delete -f charts/flux-sandbox/crds/; note this will cascade-delete all sandbox-related CRs.Pre-warming interface returns 503 UNAVAILABLE
The PostgreSQL or MoonCake etcd that the pre-warming subsystem depends on is temporarily unavailable; retry as-is after a short wait; if it persists, check the pre-warming-related startup parameters and dependency service connectivity.
Can pre-warming tasks be canceled
Task cancellation or deletion is not supported. When individual snapshots fail pre-warming, a new task must be submitted; successfully pre-warmed snapshots will not be re-distributed.