Version: v26.09

FluxSandbox Sandbox Scheduling Engine ​

Feature Introduction ​

FluxSandbox is a high-performance Kubernetes sandbox scheduling engine. Serving as the flux runtime backend for OpenSandbox, it provides low-latency, high-throughput sandbox lifecycle management capabilities for AI Agent workloads.

FluxSandbox does not expose APIs directly; OpenSandbox serves as the entry point. Users use the OpenSandbox SDK to create sandboxes, and the OpenSandbox Server forwards requests to FluxSandbox via gRPC to complete scheduling and lifecycle management. The overall flow is:

User (SDK/REST API) → OpenSandbox Server (flux runtime) → FluxSandbox Controller → Scheduler → Agent Pod → Sandbox Runtime

The sandbox runtime (i.e., the actual carrier of a sandbox instance) has two modes:

  • E2B Runtime (default): The sandbox is a microVM instance, created by the E2B Sandbox Runtime and Orchestrator on the node. Supports advanced capabilities such as pause/resume/snapshot, and sandboxes created from E2B templates start in seconds.
  • containerd Runtime: The sandbox is a bare container on the node's containerd, suitable for standard Kubernetes clusters. Does not support pause/resume/snapshot.

Application Scenarios ​

  • AI Agent Runtime Environment: Provides isolated runtime sandboxes that can be obtained in seconds and reclaimed on demand, carrying tasks such as command execution and code execution for Agents.
  • Large-Scale Sandbox Pool: Declares capacity and resource specifications through SandboxGroup, and the system automatically maintains warm buffers to handle high-concurrency creation requests.
  • Unified Resource Management: Sandbox resource specifications are declared uniformly by SandboxGroup, avoiding customization by callers.
  • Environment Save and Reuse: Creates snapshots of running sandboxes and quickly creates new sandboxes from snapshots, enabling environment template cloning and fault rollback (E2B runtime only).
  • Snapshot Pre-warming: Actively distributes and caches snapshots to multiple SuperPods, accelerating sandbox creation from snapshots (E2B runtime only).

Capability Scope ​

  • Declarative management of sandbox group capacity (SandboxGroup CRD), with automatic scaling of Agent Pods.
  • Automatic sandbox scheduling and lifecycle management (create, run, pause, resume, reclaim).
  • Warm buffer (sandboxBuffer), reducing sandbox acquisition latency.
  • Snapshot management: create, query, delete snapshots, and restore sandboxes from snapshots (E2B runtime only).
  • Snapshot pre-warming: Actively distributes and caches snapshots to multiple SuperPods, accelerating sandbox creation from snapshots (E2B runtime only).
  • Multi-replica sharded scheduling, supporting horizontal scaling.
  • Node-level orphan sandbox automatic cleanup.

Runtime Mode Selection ​

Table 1 Comparison of the two runtime modes

DimensionE2B Runtime (default)containerd Runtime
Sandbox formFirecracker microVM.Bare container.
Sandbox image semanticsE2B template name (e.g., ubuntu-22-04-custom); template must be built in advance.Standard container image reference (e.g., python:3.11).
entrypointNot injected into sandbox; microVM process determined by template.Acts as the container main process entry, takes effect.
envNot injected into sandbox.Injected into container, takes effect.
Resource specificationDeclared by SandboxGroup's sandboxResources (Agent Pod reserves by slot).Declared by SandboxGroup's sandboxResources, enforced by containerd via cgroup limits.
Pause/ResumeSupported.Not supported.
Snapshot (create/restore/pre-warm)Supported.Not supported.
Command execution carrierenvd daemon inside the sandbox.Injected execd daemon inside the sandbox.
Node prerequisitesDeploy E2B Sandbox Runtime and Orchestrator.Install containerd.
Control plane prerequisitesDeploy E2B Template Manager.None.
SandboxGroup configurationDefault, no additional settings required.Requires explicit setting of agentEnv.RUNTIME_TYPE=containerd and execdImage.

iconNote:
The subsequent chapters of this manual use the E2B runtime by default. If using the containerd runtime, please pay attention to the containerd runtime-specific notes in each chapter.

Basic Concepts ​

Table 2 Basic concepts

ConceptDescription
SandboxGroupSandbox group, a CRD resource created directly by the user, defining the resource specifications and capacity of a set of sandboxes. Once the status is Ready, sandboxes can be created.
SubSandboxGroupAutomatic shard of a SandboxGroup, generated by the Controller and assigned to each scheduler instance; users do not need to create them manually.
Agent PodThe workload Pod carrying sandbox instances, automatically started and reclaimed by the scheduler based on SandboxGroup capacity.
SandboxA single sandbox instance, running on the node where the Agent Pod resides, used via the OpenSandbox SDK.
OpenSandbox ServerUser request entry point, forwarding SDK requests to FluxSandbox via the flux runtime.
E2B TemplateThe sandbox image for the E2B runtime, with the built-in envd daemon; the image in the creation request is the template name.
Snapshot (Snapshot)A persistent record of the complete state of a sandbox at a given point in time, which can be used to create new sandboxes carrying the same state. Only supported by the E2B runtime.
SuperPod (SuperPod)The target node pool for snapshot pre-warming, caching snapshots to accelerate sandbox creation from snapshots; storage availability is reported via MoonCake.

Implementation Principle ​

When a user creates a sandbox via the SDK, the OpenSandbox Server routes the request to the corresponding SandboxGroup based on the extensions.sandboxGroup field in the request, forwards it through the FluxSandbox Controller to the owning scheduler shard, and the scheduler selects an Agent Pod from the group to create the sandbox instance. The capacity fields of the SandboxGroup (buffer, max sandboxes per Pod, etc.) drive automatic scaling of Agent Pods, and sandboxResources determines the resource specification occupied by each sandbox.

Under the E2B runtime, before creation, the scheduler resolves the template metadata from the E2B Template Manager based on the image (template name), and then the E2B Sandbox Runtime on the node starts the microVM sandbox; under the containerd runtime, the Agent directly creates a bare container through the node's containerd and injects execd.

Installation ​

Prerequisites ​

  1. The Kubernetes cluster is v1.28 or above and accessible via kubectl.
  2. E2B runtime (default):
    • Worker nodes have deployed the E2B Sandbox Runtime (one instance per node, listening on unix:///run/cri-multiplex.sock) and the E2B Orchestrator;
    • The control plane has deployed the E2B Template Manager and prepared available E2B templates (see Prepare E2B Templates).
  3. Helm 3.14 or above is installed (Helm deployment is recommended).
  4. The OpenSandbox Server image and Chart can be obtained directly from the official source (see Deploy OpenSandbox Server); the four FluxSandbox component images (controller, scheduler, watcher, agent) have been pushed to a registry accessible by the cluster (execute make docker-build and make docker-push when building from source).
  5. When creating sandboxes using the SDK, Python 3.10+ must be installed locally.

icon[containerd Runtime]
When using the containerd runtime, Worker nodes only need containerd installed; no E2B components are required, but E2B-related capabilities (pause/resume/snapshot/snapshot pre-warming) are unavailable. The relevant chapters can be skipped. For Chart parameter differences during deployment, see the containerd notes in Deploy FluxSandbox.

Deploy OpenSandbox Server ​

Sandboxes use OpenSandbox as the entry point; you need to deploy the OpenSandbox Server and configure its runtime as flux. Both the Chart and the Server image can be obtained directly from the official repository.

  1. Create the configuration file config.toml, set the runtime type to flux and point it to the FluxSandbox Controller:

    toml
    [server]
    host = "0.0.0.0"
    port = 80
    max_sandbox_timeout_seconds = 86400
    workers = 15
    thread_pool_size = 200
    timeout_keep_alive = 120
    limit_concurrency = 0
    backlog = 65535
    loop = "uvloop"
    http = "httptools"
    
    [log]
    level = "INFO"
    
    [runtime]
    type = "flux"
    execd_image = "opensandbox/execd:v1.0.20"
    
    [storage]
    allowed_host_paths = []
    volume_default_size = "1Gi"
    
    [flux_sandbox]
    endpoint = "flux-sandbox-controller.flux-system.svc:9091"
    timeout_seconds = 60
    use_tls = false

    iconNote:

    • execd_image is a required configuration item for the Server; the [ingress] block is auto-generated by the Chart (direct mode by default), do not include it again in the configuration file. The Chart's Service port is fixed at 80, port must remain 80.
    • If you need to build the image from source (alternative method): clone the openFuyao/opensandbox repository and execute docker build -f server/Dockerfile.flux -t opensandbox-server:flux ., then point server.image.repository/tag below to this image. Note that Dockerfile.flux depends on BuildKit features; if the node does not have docker buildx, using the official image directly is recommended.
  2. Deploy the Server using the Chart:

Online deployment (cluster nodes can access the image repository):

bash
helm install opensandbox-server \
  oci://cr.openfuyao.cn/charts/opensandbox-server --version 0.2.1-of.1 \
  -n flux-system  --create-namespace \
  --set namespaceOverride=flux-system \
  --set server.image.repository=openfuyao-0qfqnk.swr-pro.myhuaweicloud.com/openfuyao/opensandbox/flux-server \
  --set server.image.tag=0.2.1-of.1 \
  --set server.replicaCount=1 \
  --set 'server.env[0].name=OPENSANDBOX_INSECURE_SERVER' \
  --set 'server.env[0].value=YES' \
  --set-file configToml=config.toml

iconNote:
helm only downloads the Chart package (templates); the Server image is automatically pulled from the image repository by the node's kubelet when creating the Pod according to imagePullPolicy (default IfNotPresent), no manual import required.

Offline deployment (cluster nodes cannot access the image repository):

On a machine with internet access, download the Chart package and pull/export the Server image:

bash
# 1. Download the Chart package
helm pull oci://cr.openfuyao.cn/charts/opensandbox-server --version 0.2.1-of.1
# → opensandbox-server-0.2.1-of.1.tgz

# 2. Pull the Server image
docker pull openfuyao-0qfqnk.swr-pro.myhuaweicloud.com/openfuyao/opensandbox/flux-server:0.2.1-of.1

# 3. Export as an offline image package
docker save -o opensandbox-server-image.tar \
  openfuyao-0qfqnk.swr-pro.myhuaweicloud.com/openfuyao/opensandbox/flux-server:0.2.1-of.1

# 4. Transfer to the target cluster
scp opensandbox-server-0.2.1-of.1.tgz opensandbox-server-image.tar root@<node-ip>:/root/

Import the image on the target node:

bash
ctr -n k8s.io images import /root/opensandbox-server-image.tar
crictl images | grep flux-server

Install on the control node. Compared to online deployment, only an additional server.image.pullPolicy=Never is appended: on offline nodes, kubelet only checks local images; the image name must exactly match the one imported on the node, otherwise ErrImageNeverPull is reported:

bash
helm install opensandbox-server /root/opensandbox-server-0.2.1-of.1.tgz \
  -n flux-system  --create-namespace \
  --set namespaceOverride=flux-system \
  --set server.image.repository=openfuyao-0qfqnk.swr-pro.myhuaweicloud.com/openfuyao/opensandbox/flux-server \
  --set server.image.tag=0.2.1-of.1 \
  --set server.image.pullPolicy=Never \
  --set server.replicaCount=1 \
  --set 'server.env[0].name=OPENSANDBOX_INSECURE_SERVER' \
  --set 'server.env[0].value=YES' \
  --set-file configToml=config.toml

iconNote: (applies to both online/offline deployment)

  • -n flux-system must be explicitly included: without -n, the release record lands in the default namespace of the current context.
  • After uninstalling or deleting the Deployment and Service, serviceaccount/opensandbox-server and configmap/opensandbox-server-config will remain (with helm ownership annotations); they must be deleted together before reinstalling.
  • OPENSANDBOX_INSECURE_SERVER=YES: must be set; its absence causes the server worker to repeatedly exit.

Deploy FluxSandbox ​

CRDs are installed automatically with the Chart. When using official images, only the global.imageRegistry prefix needs to be set.

Online deployment (cluster nodes can access the image repository):

bash
helm install flux-sandbox \
  oci://cr.openfuyao.cn/charts/flux-sandbox --version 26.9.0 \
  --namespace flux-system --create-namespace \
  --set global.imageRegistry=cr.openfuyao.cn/openfuyao \
  --set watcher.args.criSocket="/run/cri-multiplex.sock"

iconNote:
helm only downloads the Chart package (templates); container images are automatically pulled from the image repository by each node's kubelet when creating Pods according to imagePullPolicy (default IfNotPresent), no manual import required. For private repositories, execute helm registry login before installation.

Offline deployment (cluster nodes cannot access the image repository):

On a machine with internet access, download the Chart package and pull/export the four component images:

bash
# 1. Download the Chart package
helm pull oci://cr.openfuyao.cn/charts/flux-sandbox --version 26.9.0
# → flux-sandbox-26.9.0.tgz

# 2. Pull the four component images
IMAGES="
cr.openfuyao.cn/openfuyao/flux-sandbox/controller:26.9.0
cr.openfuyao.cn/openfuyao/flux-sandbox/scheduler:26.9.0
cr.openfuyao.cn/openfuyao/flux-sandbox/agent:26.9.0
cr.openfuyao.cn/openfuyao/flux-sandbox/watcher:26.9.0
"
for img in $IMAGES; do docker pull "$img"; done

# 3. Export as an offline image package
docker save -o flux-sandbox-images.tar $IMAGES

# 4. Transfer to the target cluster
scp flux-sandbox-26.9.0.tgz flux-sandbox-images.tar user@your-cluster-ip:/root/

Import the image on each worker node (controller/scheduler may be scheduled to any worker node, watcher is a DaemonSet with one per node, and Agent Pods may also land on any worker node):

bash
# When the cluster runtime is containerd, import into the k8s.io namespace
ctr -n k8s.io images import /root/flux-sandbox-images.tar

# Verify successful import
crictl images | grep flux-sandbox

# For docker runtime nodes, use: docker load -i /root/flux-sandbox-images.tar

Install on the control node. Compared to online deployment, only three additional pullPolicy=Never are appended, and the image tag uses the default latest (--set *.image.tag must match the actual image tag imported on the node):

bash
helm install flux-sandbox /root/flux-sandbox-26.9.0.tgz \
  --namespace flux-system --create-namespace \
  --set global.imageRegistry=cr.openfuyao.cn/openfuyao \
  --set controller.image.pullPolicy=Never \
  --set scheduler.image.pullPolicy=Never \
  --set watcher.image.pullPolicy=Never \
  --set watcher.args.criSocket="/run/cri-multiplex.sock"

iconNote:
Offline deployment must retain global.imageRegistry and append pullPolicy=Never: when the image tag is latest, kubelet may treat it as Always and access the registry; offline nodes will report ImagePullBackOff; under the Never policy, kubelet only checks local images, and the image name must exactly match the one imported on the node, otherwise ErrImageNeverPull is reported.

For single-node deployment when you need to reduce the number of scheduler replicas (the Chart deploys 3 replicas by default), append --set scheduler.replicaCount=1, which also drives the controller's --expected-scheduler-replicas. See Table 3 for other common Helm values.

containerd Runtime: For a pure containerd environment (without E2B components), it is recommended to disable E2B configuration injection and adjust the watcher's scan socket:

bash
helm install flux-sandbox ./charts/flux-sandbox \
  --namespace flux-system --create-namespace \
  --set global.imageRegistry=cr.openfuyao.cn/openfuyao \
  --set e2b.enabled=false \
  --set watcher.args.containerdSocket=/run/containerd/containerd.sock

iconNote:
After e2b.enabled=false, components no longer mount the E2B configuration volume; the watcher targets the E2B cluster by default (criSocket="", containerdSocket=""), and containerd clusters need to be adjusted to the corresponding socket path as shown above.

Common customization options:

Table 3 Common Helm values

ParameterDefaultDescription
global.imageRegistry""Registry prefix for all images
global.imagePullSecrets[]Image pull credentials
scheduler.replicaCount3Number of scheduler shards
e2b.enabledtrueWhether to inject E2B Template Manager configuration; set to false for pure containerd environments
e2b.apiEndpointhttp://api.e2b.svc.cluster.local:3000E2B Template Manager address
e2b.existingSecret""Name of the Secret holding the E2B config.json; using a Secret to manage E2B credentials is recommended
e2b.hostPath/root/.e2bLegacy method: directory on the host holding config.json
watcher.args.criSocket/run/cri-multiplex.sockE2B CRI socket scanned by watcher; set to empty to disable E2B orphan scanning
watcher.args.containerdSocket""containerd socket scanned by watcher; set to /run/containerd/containerd.sock for containerd clusters
*.nodeSelector / *.tolerations{}Scheduling constraints for each component

For complete values documentation, see the flux-sandbox Chart README.

Method 2: Raw Manifests ​

bash
# Create namespace
kubectl create namespace flux-system

# Install CRDs
make crd-apply

# Deploy each component
kubectl apply -f config/controller/
kubectl apply -f config/scheduler/
kubectl apply -f config/sandbox-watcher/

Verify Deployment ​

bash
kubectl wait --for=condition=Ready pod -n flux-system -l app=flux-sandbox-controller --timeout=120s
kubectl wait --for=condition=Ready pod -n flux-system -l app=flux-sandbox-scheduler --timeout=120s

For the E2B runtime, you should also confirm that the Template Manager configuration is ready: the component logs should show a message indicating successful Template Manager initialization. If E2B credentials are not configured, the components will start in a degraded mode (with a log warning); after the components start, creating sandboxes through an E2B-type SandboxGroup will fail with the error "E2B runtime requires template manager but it is not configured".

Preparation Before Use ​

After completing deployment, you need to create a SandboxGroup and prepare sandbox templates (E2B runtime) or container images (containerd runtime), after which you can use sandboxes through the OpenSandbox SDK.

Create SandboxGroup ​

A SandboxGroup defines the resource specifications and capacity of a set of sandboxes and is the only CRD resource users need to create. The Controller splits it into SubSandboxGroups (shards) and assigns them to each scheduler instance, which then starts Agent Pods. Once the sandbox group status becomes Ready, sandboxes can be created.

  1. Execute kubectl apply -f - to create a SandboxGroup (no additional configuration required for the E2B runtime):

    bash
    kubectl apply -f - <<EOF
    apiVersion: sandbox.flux.io/v1alpha1
    kind: SandboxGroup
    metadata:
      name: sg-1u1g
      namespace: flux-system
    spec:
      capacity:
        agentPodMin: 1
        agentPodMax: 1
        sandboxBufferMin: 1
        maxSandboxesPerPod: 5
      sandboxResources:
        requests: { cpu: "1", memory: "2Gi" }
        limits:   { cpu: "1", memory: "2Gi" }
    EOF

    Capacity field descriptions:

    Table 4 capacity field descriptions

    FieldRequiredDescription
    agentPodMin / agentPodMaxYesLower/upper limit of Agent Pod count; the system auto-scales within this range.
    sandboxBufferMin / sandboxBufferMaxNoLower/upper limit of warm buffer idle sandbox count, used to reduce sandbox acquisition latency.
    maxSandboxesPerPodYesMaximum number of sandboxes allowed to coexist on a single Agent Pod.

    Sharding mechanism example: Suppose the SandboxGroup is configured with agentPodMax: 200 and maxSandboxesPerPod: 5, and the Controller startup parameter --max-pod-per-sub-sandbox-group uses the default value 100. The system will automatically split this SandboxGroup into 2 SubSandboxGroup shards (named like <SandboxGroup name>-0, <SandboxGroup name>-1) and assign them to two scheduler instances:

    • Each shard manages at most 100 Agent Pods, with a maximum capacity of 100 × 5 = 500 Sandboxes;
    • The entire SandboxGroup has a maximum capacity of 200 × 5 = 1000 Sandboxes;
    • When agentPodMax is not evenly divisible by --max-pod-per-sub-sandbox-group, round up; the last shard takes the remainder (e.g., agentPodMax: 250 splits into 3 shards, with the last shard managing 50 Agent Pods).

    To adjust the upper limit of creatable Sandbox count:

    • Adjust agentPodMax or maxSandboxesPerPod in the SandboxGroup spec; total capacity = agentPodMax × maxSandboxesPerPod. When maxSandboxesPerPod is increased, the Agent Pod will reserve resources for more slots based on sandboxResources, and the per-Pod resource request increases linearly; generally, adjust agentPodMax first.
    • The startup parameter --max-pod-per-sub-sandbox-group (Helm deployment value controller.args.maxPodPerSubSandboxGroup, default 100) does not change total capacity, only controls sharding granularity: increasing it reduces the number of shards, decreasing it spreads shards across more scheduler instances for load balancing. If you do not want sharding, set it to no less than agentPodMax (e.g., 200 in the above example).
  2. Wait for the SandboxGroup to become ready:

    bash
    kubectl wait --for=condition=Ready sandboxgroup sg-1u1g -n flux-system --timeout=120s

containerd Runtime: When using the containerd runtime, the SandboxGroup must explicitly declare the runtime type and execd image (execd carries the SDK's command execution, required):

yaml
spec:
  agentEnv:
    RUNTIME_TYPE: "containerd"
  execdImage: "opensandbox/execd:v1.0.20"

iconNote:
The runtime type is controlled by agentEnv.RUNTIME_TYPE (default is E2B); execdImage only takes effect for the containerd runtime and is invalid under the E2B runtime.

Prepare E2B Templates ​

Under the E2B runtime, the image passed when creating a sandbox is the E2B template name (not a container image reference). The template determines the operating system environment, pre-installed software, and processes inside the microVM. The template must meet:

  1. The template has been built and registered in the E2B environment accessible to the Template Manager through the E2B template mechanism, and the template name is globally available (e.g., ubuntu-22-04-custom).
  2. The template has a built-in envd daemon, carrying capabilities such as command execution and file operations for the SDK (official base templates already include this).
  3. The actual processes and environment running inside the sandbox are determined by the template definition (the entrypoint and env in the creation request are not injected into the microVM).

When the template name in the creation request does not exist, the Template Manager resolution fails, and the creation request returns an error.

containerd Runtime: image is a standard container image reference; ensure the image can be pulled on all Worker nodes:

bash
# Sandbox runtime image (SDK example uses python:3.11)
docker pull python:3.11

# execd image (specified by execdImage in SandboxGroup)
docker pull opensandbox/execd:v1.0.20

iconNote:
Images must be pushed to a registry accessible by the cluster, or pre-pulled on all Worker nodes.

Using Sandboxes ​

Prerequisites ​

  • FluxSandbox deployment, SandboxGroup creation, and OpenSandbox Server integration are complete.
  • The OpenSandbox Python SDK (adapted version, installation method below) is installed locally, and the Server access address is obtained.

Install OpenSandbox Python SDK (Adapted Version) ​

This repository (openFuyao/opensandbox) adapts the official SDK (including snapshot APIs aligned with the FluxSandbox Server, etc.) and is not published to PyPI. Executing pip install opensandbox directly will install the PyPI official package, which lacks the above adaptations; you must install using one of the two methods below.

Method 1: With internet access, install directly from the GitCode repository ​
bash
pip install "git+https://gitcode.com/openFuyao/opensandbox.git@of-dev/v0.2.1#subdirectory=sdks/sandbox/python"

iconNote:
#subdirectory=sdks/sandbox/python cannot be omitted; the SDK source code is located in a subdirectory of the repository.

Method 2: Offline environment, install from offline wheel package ​

Build a wheelhouse on a machine with internet access. The Python minor version and system architecture used for building must match the intranet machine (e.g., both Python 3.12 + linux x86_64), because the dependency pydantic-core is a compiled wheel bound to the Python version and platform:

bash
git clone -b of-dev/v0.2.1 https://gitcode.com/openFuyao/opensandbox.git
cd opensandbox
# Build the SDK itself and all runtime dependencies into the wheelhouse directory
pip wheel -w wheelhouse ./sdks/sandbox/python
# Package and transfer to the intranet machine
tar czf opensandbox-sdk-wheelhouse.tar.gz wheelhouse
scp opensandbox-sdk-wheelhouse.tar.gz root@<intranet-machine-IP>:/root/

Install offline on the intranet machine:

bash
tar xzf opensandbox-sdk-wheelhouse.tar.gz
pip install --no-index --find-links=wheelhouse wheelhouse/opensandbox-*.whl

iconNote:

  • --no-index forces pip to only take packages from the wheelhouse directory, without accessing any package sources.
  • When building on a machine with internet access, the .git directory must be retained, otherwise the version number cannot be derived from the git tag.
Verify Installation ​
bash
pip show opensandbox

A version number with a dev and +g<commit-id> suffix (e.g., 0.1.14.dev26+g6066bc24) indicates the adapted version; a clean 0.1.x version number means the PyPI official package was installed, and you must reinstall.

bash
python -c "
from opensandbox import Sandbox                      # Async SDK entry
from opensandbox.sync.sandbox import SandboxSync     # Sync SDK entry
print(hasattr(Sandbox, 'create_snapshot'))           # Should output True
print(hasattr(SandboxSync, 'create_snapshot'))       # Should output True
"

Both outputs being True indicates the snapshot API adaptation is in place.

iconWarning:
The adapted version number (0.1.14.devNN+g…) is lower than the PyPI official package version (0.1.16); do not execute pip install -U opensandbox to upgrade this package, otherwise it will be overwritten by the official package, resulting in the loss of adapted capabilities such as snapshots; to upgrade the adapted version, rebuild and install using the above method.

Create Sandbox Parameter Description ​

Creating a sandbox is uniformly done through the OpenSandbox SDK/REST API; FluxSandbox does not expose a separate interface. The required/optional parameter rules are as follows (validation is performed by the OpenSandbox Server; image mode and snapshot_id mode have different rules):

Table 5 Create sandbox parameters (image mode)

ParameterRequiredDefaultDescription
imageRequired (choose one with snapshot_id)-E2B template name (container image reference for the containerd runtime).
entrypointRequiredSDK default ["tail", "-f", "/dev/null"]Sandbox main process entry. REST direct call returns 400 if not passed; SDK uses default if not passed. Not injected into microVM for E2B runtime; only takes effect for containerd runtime.
resource_limits (REST field name resourceLimits)RequiredSDK default {"cpu": "1", "memory": "2Gi"}Resource specification. REST direct call returns 400 if not passed; SDK uses default if not passed. The actual effective specification is subject to SandboxGroup's sandboxResources (see below).
extensions.sandboxGroupRequired-Target SandboxGroup name; the group must be Ready. Returns 400 if missing.
timeoutOptionalSDK default 10 minutes; REST does not auto-expire if not passedSandbox auto-expiration duration; automatically reclaimed upon expiration. Must not be less than 60 seconds and must not exceed the Server-configured upper limit (max_sandbox_timeout_seconds).
envOptional{}Environment variables. Only injected into sandbox for containerd runtime; not injected for E2B runtime.
metadataOptional{}User-defined labels; keys and values must conform to Kubernetes label specifications.
resource_requests (REST field name resourceRequests)OptionalSame as resource_limitsResource request value; defaults to resource_limits when not specified. The actual effective specification is subject to SandboxGroup.
platform / network_policy / volumes / secure_access / credential_proxyOptional-Platform constraints, egress policies, storage mounts, and other advanced configurations; use according to OpenSandbox general semantics.

Table 6 Create sandbox parameters (snapshot_id mode, E2B runtime only)

ParameterRequiredDescription
snapshot_idRequired (choose one with image)Source snapshot ID; the snapshot must be in Ready state, otherwise returns 404/409.
entrypointOptionalWhen not passed, the server automatically uses ["tail", "-f", "/dev/null"].
image / resource_limitsOptionalimage is automatically injected from the snapshot record; the actual effective specification of resource_limits is subject to SandboxGroup.
extensions.sandboxGroup / timeout / other parametersSame as image mode-

iconNote:

  • Required description for entrypoint and resource_limits: REST requires these for Server parameter validation, but FluxSandbox uses the SandboxGroup configuration as the actual effective value; the parameters in the request do not determine the actual specification. To adjust the specification, modify the corresponding SandboxGroup.
  • A single request must provide exactly one of image and snapshot_id, otherwise returns 400.

Usage Limitations ​

  • When creating a sandbox, the target SandboxGroup must be specified via extensions.sandboxGroup, otherwise the request is rejected (400).
  • Sandbox resource specifications are uniformly determined by SandboxGroup's sandboxResources; resource parameters passed during creation do not take effect (but must still pass the Server's parameter validation, see Table 5).
  • Under the E2B runtime, entrypoint and env are not injected into the sandbox; the processes and environment inside the sandbox are determined by the E2B template.
  • Pause, resume, and snapshot capabilities are only supported by the E2B runtime; calling the relevant interfaces under the containerd runtime returns an unsupported error.
  • Only sandboxes in the Running state can have snapshots created; a single sandbox allows only one snapshot operation at a time.

Quick Start ​

Connect to the OpenSandbox Server and create a sandbox:

python
import asyncio
from datetime import timedelta

from opensandbox import Sandbox
from opensandbox.config import ConnectionConfig

async def main():
    config = ConnectionConfig(
        domain="<server-address>:80",     # OpenSandbox Server address (Chart default Service port 80)
    )

    sandbox = await Sandbox.create(
        "ubuntu-22-04-custom",                  # E2B template name (for containerd runtime, pass image reference, e.g., "python:3.11")
        connection_config=config,
        timeout=timedelta(minutes=30),
        extensions={"sandboxGroup": "sg-1u1g"},   # Required: route to SandboxGroup
        skip_health_check=True,
    ) 
    
    print("created:", sandbox.id) 
    info = await sandbox.get_info()
    print("state:", info.status.state)  

asyncio.run(main())

Delete the sandbox:

bash
python3 kill_sandbox.py <sandbox-id>
python
import asyncio, sys
from datetime import timedelta
from opensandbox.config import ConnectionConfig
from opensandbox.sandbox import Sandbox
 
async def main(sandbox_id):
    config = ConnectionConfig(domain="<server-address>", api_key="")
    sb = await Sandbox.get(sandbox_id, connection_config=config)
    await sb.kill()
    await sb.close()
    print(f"Sandbox {sandbox_id} destroyed")
 
asyncio.run(main(sys.argv[1]))

For complete execution commands, see envd_manual_verify.md.

Create Sandbox ​

Create from template/image:

python
sandbox = await Sandbox.create(
    "ubuntu-22-04-custom",                # image: E2B template name (or container image reference)
    connection_config=config,
    entrypoint=["tail", "-f", "/dev/null"],  # Optional, SDK default is this value
    env={"FOO": "bar"},                    # Environment variables (only injected for containerd runtime)
    timeout=timedelta(minutes=30),         # Auto-expiration time; REST does not auto-expire if not passed
    extensions={"sandboxGroup": "sg-1u1g"},
)

Create from snapshot (restores the complete state of the snapshot, ready in seconds; E2B runtime only):

python
sandbox = await Sandbox.create(
    connection_config=config,
    snapshot_id="<snapshot-id>",            # Choose one with image
    timeout=timedelta(minutes=30),
    extensions={"sandboxGroup": "sg-1u1g"},
)

iconNote:
You must pass exactly one of image and snapshot_id. The snapshot must be in Ready state; a non-existent or not-ready snapshot returns 404/409. When creating from a snapshot, you do not need to pass image and entrypoint; the server automatically uses the image from the snapshot record and the default entrypoint.

containerd Runtime Usage Example ​

A complete usage flow using the containerd runtime as an example: create sandbox → execute command → stream output → file read/write → query status. Example requirements:

  • SandboxGroup is for the containerd runtime (see the containerd notes in Create SandboxGroup), and the execd image is configured;
  • The sandbox image python:3.11 and the execd image opensandbox/execd:v1.0.20 can be pulled on Worker nodes.

Access the OpenSandbox Server locally via port-forward:

bash
kubectl port-forward -n flux-system svc/opensandbox-server 8080:80

Complete example code:

python
import asyncio
from datetime import timedelta

from opensandbox import Sandbox
from opensandbox.config import ConnectionConfig
from opensandbox.models.execd import ExecutionHandlers
from opensandbox.models.filesystem import WriteEntry, SearchEntry


async def main():
    config = ConnectionConfig(
        domain="localhost:8080",
        api_key="your-secret-key",
        use_server_proxy=True,
    )

    # ---------- Complete flow for a brand-new sandbox ----------
    async with await Sandbox.create(
        "python:3.11",
        connection_config=config,
        timeout=timedelta(minutes=30),   # Note: SDK 0.1.16 this parameter not passed to server
        extensions={"sandboxGroup": "sg-1u1g"},
    ) as sandbox:
        print(f"sandbox id: {sandbox.id}")

        # 1. Execute command + exit code
        result = await sandbox.commands.run("python -c 'print(1 + 1)'")
        print(f"exit_code={result.exit_code}, out={result.logs.stdout[0].text.strip()}")

        # 2. Streaming output (SDK 0.1.16 requires async callbacks, both on_stdout/on_stderr required)
        async def on_stdout(msg):
            print(f"STDOUT: {msg.text}")

        async def on_stderr(msg):
            print(f"STDERR: {msg.text}")

        handlers = ExecutionHandlers(on_stdout=on_stdout, on_stderr=on_stderr)
        await sandbox.commands.run("for i in 1 2 3; do echo $i; done", handlers=handlers)

        # 3. File operations: write / read / search
        await sandbox.files.write_files([
            WriteEntry(path="/tmp/hello.txt", data="Hello World", mode=644)
        ])
        content = await sandbox.files.read_file("/tmp/hello.txt")
        print(f"file content: {content}")
        files = await sandbox.files.search(SearchEntry(path="/tmp", pattern="*.txt"))
        print(f"search: {[f.path for f in files]}")

        # 4. Query status
        info = await sandbox.get_info()
        print(f"state: {info.status.state}, expires: {info.expires_at}")

asyncio.run(main())

iconNote:

  • The timeout parameter of SDK 0.1.16 is not passed through to the server; the sandbox will not auto-expire based on this value. To enable auto-expiration, control it via the timeout field (in seconds, minimum 60) of the REST request body.
  • The streaming callbacks of SDK 0.1.16 require on_stdout/on_stderr to both be async functions and both are mandatory, otherwise no output is received.
  • api_key is only required when the Server deployment uses authentication; use_server_proxy=True indicates accessing the sandbox through the Server proxy.

Manage Snapshots ​

Snapshots capture the complete state of a sandbox at a given point in time and can be used for environment preservation, batch cloning, and fault rollback. Snapshot creation is a synchronous operation that returns with the final state (Ready or Failed); during creation, the sandbox remains Running, and on failure, the sandbox automatically resumes running.

[containerd Runtime] Snapshot capabilities are only supported by the E2B runtime; this section does not apply.

Manage snapshots (create, query, list, delete) via SandboxManager, and restore sandboxes from snapshots via Sandbox.create(snapshot_id=...):

python
import asyncio
from datetime import timedelta

from opensandbox import Sandbox, SandboxManager
from opensandbox.config import ConnectionConfig
from opensandbox.models.sandboxes import SnapshotFilter

config = ConnectionConfig(domain="<server-address>:80")  # OpenSandbox Server address

async def main():
    # ---------- Create snapshot ----------
    # Sandbox must be in Running state; creation is a synchronous operation that returns with the final state (Ready or Failed)
    async with await SandboxManager.create(connection_config=config) as manager:
        snap = await manager.create_snapshot(
            sandbox_id="<sandbox-id>",      # Required: source sandbox ID, sandbox must be Running
            name="my-snapshot",               # Optional: snapshot name
        )
        print(f"snapshot: {snap.id}, state: {snap.status.state}")

        # ---------- Query a single snapshot ----------
        info = await manager.get_snapshot(snap.id)  # Required: snapshot ID
        print(f"state: {info.status.state}, reason: {info.status.reason}")

        # ---------- List snapshots (supports filtering by source sandbox, state, and pagination) ----------
        snaps = await manager.list_snapshots(
            SnapshotFilter(
                sandbox_id="<sandbox-id>",   # Optional: filter by source sandbox ID
                states=["READY"],            # Optional: filter by state (READY/FAILED)
                page_size=20,                # Optional: items per page, default 20
                page=1,                      # Optional: page number, starting from 1
            )
        )
        for s in snaps.snapshot_infos:
            print(f"{s.id}: {s.status.state}")

        # ---------- Delete snapshot ----------
        await manager.delete_snapshot(snap.id)  # Required: snapshot ID
        print(f"deleted: {snap.id}")

    # ---------- Restore sandbox from snapshot ----------
    # Snapshot must be in Ready state; E2B sandbox has no execd process, SDK automatically skips endpoint resolution and health check
    # snapshot_id and image are mutually exclusive; when restoring, only pass snapshot_id, image is automatically injected from the snapshot record
    sandbox = await Sandbox.create(
        snapshot_id=snap.id,                               # Required: source snapshot ID (choose one with image)
        connection_config=config,                          # Required: Server connection configuration
        timeout=timedelta(minutes=30),                     # Optional: auto-expiration time; no auto-expiration if not passed
        resource={"cpu": "1", "memory": "2Gi"},            # Optional: resource specification (actual value subject to SandboxGroup)
        extensions={"sandboxGroup": "sg-1u1g"},           # Required: target SandboxGroup name
    )
    print(f"restored sandbox: {sandbox.id}")
    info = await sandbox.get_info()
    print(f"state: {info.status.state}")
    await sandbox.close()

asyncio.run(main())

iconNote:
Snapshot records are persistently stored and can still be queried and used after the OpenSandbox Server restarts. Deleted snapshots can no longer be used to create sandboxes. After creation, snapshots can be pre-warmed and distributed to SuperPod caches to further reduce the latency of creating sandboxes from snapshots; see Snapshot Pre-warming.

Query and Manage Existing Sandboxes ​

python
from opensandbox.manager import SandboxManager
from opensandbox.models.sandboxes import SandboxFilter

async with await SandboxManager.create(connection_config=config) as manager:
    # List running sandboxes
    sandboxes = await manager.list_sandbox_infos(
        SandboxFilter(states=["RUNNING"], page_size=10)
    )
    for info in sandboxes.sandbox_infos:
        print(f"Found sandbox: {info.id}")

    # Terminate the specified sandbox
    await manager.kill_sandbox(info.id)

Sandbox State Description ​

Table 7 Sandbox state transitions

StateDescription
PENDINGBeing created, not yet ready.
RUNNINGRunning; commands can be executed, files operated, and snapshots created.
PAUSINGPausing (transient); does not accept new pause/resume requests during this period.
PAUSEDPaused; processes suspended, resource placeholders released.
RESUMINGResuming (transient).
TERMINATEDTerminated.

Core transition: PENDING → RUNNING ⇄ (pause/resume) PAUSED → TERMINATED.

Table 8 Snapshot states

StateDescription
ReadySnapshot is available and can be used to create sandboxes.
FailedSnapshot creation failed; the source sandbox has automatically resumed running.

Common Error Descriptions ​

Table 9 Common error scenarios

ScenarioError CodeSuggested Handling
Creating sandbox without passing extensions.sandboxGroup400Add the sandboxGroup field with the target SandboxGroup name.
Creating sandbox without passing image/snapshot_id, or passing both400Provide exactly one of the two.
Not passing entrypoint or resource_limits in image mode (REST direct call)400Add the required parameters; or use the SDK (has defaults).
timeout less than 60 seconds or exceeding the Server-configured upper limit400Adjust the timeout value.
Specified E2B template does not exist (E2B runtime)500Confirm the template has been built and registered in the E2B environment, and the template name matches the request.
Specified container image does not exist (containerd runtime)500Confirm the image has been pushed to a registry accessible by the cluster or pre-pulled on nodes.
Querying/deleting a non-existent sandbox or snapshot404Verify the ID is correct and whether the resource has been deleted.
Creating a snapshot from a non-Running sandbox409Restore the sandbox to Running state first.
Repeatedly initiating a snapshot while one is in progress409Wait for the current snapshot operation to complete.
Performing a resume on a running sandbox409Only paused (PAUSED) sandboxes need to be resumed.
Repeated operations during pause/resume409The state is in the PAUSING/RESUMING transient; wait for the transition to complete.
Pause/resume/snapshot operations return unsupported-The SandboxGroup uses the containerd runtime; the relevant capabilities are only supported by the E2B runtime.
Pause failed (PauseFailed)-Check whether the target Agent Pod is ready, then retry.

Snapshot Pre-warming ​

[containerd Runtime] Snapshot pre-warming depends on E2B snapshot capabilities and is only available for the E2B runtime; this section does not apply.

Snapshot pre-warming actively distributes and caches existing snapshots to multiple SuperPods, so that when creating sandboxes from the snapshot, no on-demand fetching is required, reducing cold start latency. The pre-warming subsystem is embedded in the Controller process and is invoked via the Controller's HTTP interface (default :8080, API prefix /api/v1), suitable for integration with external cluster management systems.

Table 10 Pre-warming interface overview

MethodPathPurpose
POST/api/v1/snapshot-warm-tasksCreate a pre-warming task (executed asynchronously)
GET/api/v1/snapshot-warm-tasks/{task_id}Query task pre-warming progress (paginated)
GET/api/v1/snapshots/cache-detailsQuery snapshots cached on each SuperPod (paginated)
GET/api/v1/superpods/storage-detailsQuery storage availability of each SuperPod

All responses carry the X-Request-ID response header (passed through from the request's same-named header, auto-generated if not present), which can be used for issue tracking. Error responses uniformly use the {"error": {"code", "message", "retryable"}} structure.

Enable Pre-warming ​

The pre-warming subsystem is disabled by default and must be explicitly enabled in the Controller startup parameters. All the following parameters are required; missing any one will cause startup failure:

text
--warm-enabled=true
--warm-postgres-dsn=<PostgreSQL connection string>   # Task and pre-warming detail persistence
--warm-etcd-endpoints=<MoonCake etcd address>          # Storage availability query
--warm-superpod-ids=<comma-separated SuperPod ID list>  # Pre-warming target pool
--warm-e2b-api-endpoint=<E2B API-Server address>      # Required when pre-warming backend is e2b

The pre-warming backend is specified via --warm-backend-type, default e2b, optional node (requires --warm-node-port, default 44772). For the database schema script, see flux-sandbox repository warm-migrations/0001_warm.sql.

Create Pre-warming Task ​

bash
curl -X POST http://<controller>:8080/api/v1/snapshot-warm-tasks \
  -H 'Content-Type: application/json' \
  -d '{"snapshot_ids": ["snap-001", "snap-002", "snap-003"]}'

Response 201 Created:

json
{
  "task_id": "warm-task-20260908-1786230000000000000-1",
  "state": "WAIT",
  "accepted_snapshot_count": 3,
  "created_at": "2026-09-08T02:00:00Z"
}

The task executes asynchronously with state transitions: WAIT (enqueued) → WARMING (executing) → SUCCEED / FAILED. Each snapshot is load-balanced and assigned to several SuperPods (default 5); SuperPods that fail pre-warming are automatically retried 3 times (with a 10-second interval), retrying only the failed SuperPods; the overall task timeout is 30 minutes.

snapshot_ids is required and must not contain empty strings or duplicates; the task executes on the full array as submitted, and batching is not supported. The request body is limited to 64 MiB (approximately 100,000 IDs); exceeding the limit returns 400.

Query Pre-warming Progress ​

bash
curl 'http://<controller>:8080/api/v1/snapshot-warm-tasks/<task_id>?page_size=100&page_num=1'

The summary in the response reflects overall task progress:

json
{
  "total": 3,
  "success_count": 2,
  "failed_count": 0,
  "in_progress_count": 1
}
  • Whether the task has ended is determined by success_count + failed_count == total.
  • The snapshots array returns the pre-warming results for each snapshot in submission order with pagination: warm_success=true indicates that all targets for this snapshot have completed pre-warming; warmed_superpod_ids lists the SuperPods that have actually cached this snapshot (snapshots being pre-warmed or failed will also list the successfully completed portions).
  • page_size defaults to 1000 with an upper limit of 10000 (out-of-range returns 400); page_num starts from 1. A non-existent task ID returns 404.

Query Cache Distribution and Storage Availability ​

bash
# Snapshots cached on each SuperPod (only counts SuperPods that were actually pre-warmed successfully)
curl 'http://<controller>:8080/api/v1/snapshots/cache-details?page_size=1000&page_num=1'

# MoonCake storage availability of each SuperPod, can be used for capacity assessment before pre-warming
curl http://<controller>:8080/api/v1/superpods/storage-details
  • Cache details are returned grouped by SuperPod, including fields such as snapshot_id, category, cached_at.
  • Storage availability data is read directly from MoonCake etcd reports: if observed_at is more than 60 seconds ago, it is considered stale and that SuperPod's health=UNAVAILABLE; configured SuperPods always appear in the response (those without reports or with stale data show as UNAVAILABLE).

Typical Invocation Flow ​

bash
BASE=http://<controller>:8080/api/v1

# 1. Submit pre-warming task
TASK_ID=$(curl -s -X POST $BASE/snapshot-warm-tasks \
  -H 'Content-Type: application/json' \
  -d '{"snapshot_ids": ["snap-001", "snap-002"]}' | jq -r .task_id)

# 2. Poll progress until all snapshots have results
curl -s "$BASE/snapshot-warm-tasks/$TASK_ID?page_size=1000" | jq .summary

# 3. View cache distribution and SuperPod storage levels
curl -s "$BASE/snapshots/cache-details" | jq .
curl -s "$BASE/superpods/storage-details" | jq .

Pre-warming Common Errors ​

Table 11 Pre-warming common errors

Error CodeTrigger ScenarioRetryable
400 INVALID_ARGUMENTRequest body is invalid or exceeds limit; snapshot_ids is empty, contains empty strings, or contains duplicates; pagination parameters out of rangeNo
404 NOT_FOUNDTask ID does not existNo
503 UNAVAILABLEPostgreSQL / MoonCake etcd temporarily unavailableYes

Errors with retryable=true can be retried as-is after a short wait.

Usage Limitations ​

  • The pre-warming interface has no authentication; network policies must be used to restrict access at deployment time.
  • Tasks cannot be canceled or deleted after creation; failed snapshots need a new task submission (idempotent: successfully pre-warmed snapshots will not be re-distributed).
  • Snapshots currently submitted via API are all processed as the REGULAR category, with each snapshot pre-warmed to 5 SuperPods by default.

View Running Status ​

bash
# View SandboxGroup status and capacity
kubectl get sandboxgroup -n flux-system -o wide

# View sandbox instances
kubectl get sandbox -n flux-system

After the SandboxGroup is ready, status.phase is Ready; you can use kubectl describe sandboxgroup to view Conditions to locate issues. Sandbox running status is perceived via the SDK's get_info or management interfaces.

Upgrade and Uninstall ​

bash
# Upgrade
helm upgrade flux-sandbox ./charts/flux-sandbox --namespace flux-system -f my-values.yaml

# Uninstall
helm uninstall flux-sandbox --namespace flux-system

iconNote:
CRDs are not deleted with helm uninstall (established Helm behavior). To delete, execute kubectl delete -f charts/flux-sandbox/crds/; note this will cascade-delete all sandbox-related CRs.

Appendix ​

SandboxGroup Spec Field Reference ​

Table 12 SandboxGroup Spec fields

FieldTypeRequiredDescription
capacityobjectYesCapacity declaration; fields are in Table 4 (agentPodMin, agentPodMax, maxSandboxesPerPod are required).
sandboxResourcesobjectYesResource specification per sandbox (requests/limits); the sole authoritative source for resource specifications.
agentEnvmapNoEnvironment variables injected into the Agent Pod, e.g., RUNTIME_TYPE: "containerd".
execdImagestringNo (required under containerd runtime)execd image address. Required under the containerd runtime (when set, an execd init container is automatically injected); not needed under the E2B runtime.
agentPodSchedulingobjectNoNode scheduling constraints for Agent Pods (tolerations/nodeSelector/affinity), which can direct the sandbox group to specific node pools.
schedulingPolicyobjectNoSandbox placement policy across Agent Pods (strategy, weights, topologyAffinity), optional advanced configuration.

SandboxGroup Status Field Reference ​

Table 13 SandboxGroup Status fields

FieldDescription
phaseCurrent lifecycle phase; Ready once ready.
observedGenerationThe generation observed by the Controller, which can be used to determine whether a change has taken effect.
subSandboxGroupsCountNumber of shards split.
totalAgentPodsCurrent total number of Agent Pods.
totalCapacityCurrent total sandbox capacity.
conditionsConditions describing the group status, used for troubleshooting.

SDK Operations and REST API Mapping ​

Non-Python users or gateway integration scenarios can directly call the OpenSandbox Server REST API:

Table 14 SDK operations and REST API mapping

OperationREST API
Create sandboxPOST /v1/sandboxes (202)
Query sandboxGET /v1/sandboxes/{id}
List sandboxesGET /v1/sandboxes
Delete sandboxDELETE /v1/sandboxes/{id}
Update sandbox metadataPATCH /v1/sandboxes/{id}/metadata
Renew sandboxPOST /v1/sandboxes/{id}/renew-expiration
Pause sandboxPOST /v1/sandboxes/{id}/pause
Resume sandboxPOST /v1/sandboxes/{id}/resume
Get port access addressGET /v1/sandboxes/{id}/endpoints/{port}
Create snapshotPOST /v1/sandboxes/{id}/snapshots
Query snapshotGET /v1/snapshots/{id}
List snapshotsGET /v1/snapshots
Delete snapshotDELETE /v1/snapshots/{id}

Create sandbox request example (fields correspond to Table 5/Table 6):

bash
curl -s -X POST "http://${SERVER}/v1/sandboxes" \
  -H "Content-Type: application/json" \
  -d "{
    \"image\":{\"uri\":\"${TEMPLATE_OR_IMAGE}\"},
    \"entrypoint\":[\"tail\",\"-f\",\"/dev/null\"],
    \"resourceLimits\":{\"cpu\":\"1\",\"memory\":\"2Gi\"},
    \"extensions\":{\"sandboxGroup\":\"${SG}\"},
    \"timeout\":3600
  }"

iconNote:

  • When calling REST directly, entrypoint and resourceLimits are required in image mode (the SDK auto-fills defaults); timeout not being passed means the sandbox does not auto-expire, in seconds, minimum 60.
  • SDK parameters use snake_case naming (e.g., resource_limits), while REST request body fields use camelCase naming (e.g., resourceLimits).

FAQ ​

  1. SandboxGroup is not ready after creation

    Possible causes

    • Controller or Scheduler components are not running properly.
    • Agent Pod cannot be scheduled (insufficient node resources, taints not tolerated).

    Solution

    Execute kubectl describe sandboxgroup <name> -n flux-system to view Conditions to locate the cause, and check the status of components and Agent Pods: kubectl get pods -n flux-system.

  2. Creating a sandbox returns 400, prompting extensions['sandboxGroup'] is required

    The creation request must carry the extensions.sandboxGroup field specifying the target SandboxGroup, and the group must be Ready.

  3. Creating a sandbox returns 400, prompting that entrypoint or resourceLimits is required

    When calling REST directly in image mode, entrypoint and resourceLimits are required parameters (the SDK has defaults and will not trigger this); add them per Table 5 and retry. Note that these two values do not affect the actual effective specification; the actual specification is determined by SandboxGroup.

  4. Creating a sandbox under the E2B runtime fails with "E2B runtime requires template manager but it is not configured"

    The E2B runtime depends on the E2B Template Manager to resolve templates. Check:

    • chart's e2b.enabled is true (default);
    • E2B config.json containing teamApiKey has been provided via e2b.existingSecret (or e2b.hostPath);
    • the Template Manager service pointed to by e2b.apiEndpoint is reachable.
  5. Creating a sandbox under the E2B runtime fails, prompting that the template does not exist

    The image in the creation request must be a built and registered E2B template name. Confirm the template exists in the E2B environment and the name matches exactly.

  6. Creating a sandbox fails under the containerd runtime

    To use the containerd runtime in a standard K8s environment, you must: set agentEnv.RUNTIME_TYPE=containerd and execdImage on the SandboxGroup; ensure the sandbox image and execd image can be pulled. See the containerd runtime notes in each chapter.

  7. Resource parameters passed when creating a sandbox via the SDK do not take effect

    Sandbox resource specifications are uniformly determined by SandboxGroup's sandboxResources, the sole authoritative source; resource parameters passed by the caller are ignored (but must still pass the Server's parameter validation). To adjust the specification, modify the corresponding SandboxGroup.

  8. The entrypoint or environment variables passed in are not visible inside the E2B sandbox

    This is expected behavior. Under the E2B runtime, entrypoint and env are not injected into the microVM; the processes and environment inside the sandbox are determined by the E2B template; to customize the environment, build it through the template. Under the containerd runtime, both take effect.

  9. Image pull failure when creating a sandbox (containerd runtime)

    Confirm that the sandbox runtime image (e.g., python:3.11) and the execd image have been pushed to a registry accessible by the cluster, or pre-pulled on all Worker nodes. The E2B runtime does not involve image pulling; the template is solidified during the build phase.

  10. Creating a sandbox from a snapshot returns 409

    The snapshot must be in Ready state. Query the snapshot status, confirm, and retry; snapshots in Failed state are not available.

  11. Pause/resume/snapshot interfaces return unsupported

    These capabilities are only supported by the E2B runtime. If the SandboxGroup uses the containerd runtime (RUNTIME_TYPE=containerd), switch to a SandboxGroup using the E2B runtime.

  12. Are resources released after pausing a sandbox

    Yes. After a sandbox enters the PAUSED state, resource placeholders are released, and new sandboxes of the same capacity can be created immediately; when resumed, the running state is rebuilt based on the snapshot generated at pause time.

  13. Does snapshot creation failure affect the original sandbox

    No. During snapshot creation, the sandbox remains Running; on creation failure, the sandbox automatically resumes running, and the snapshot status is set to Failed.

  14. How to access ports listened to inside the sandbox from outside

    Obtain the access address via sandbox.get_endpoint(port) and connect directly; see Access Sandbox Ports.

  15. CRDs still exist after executing helm uninstall

    This is established Helm behavior. To clean up thoroughly, manually execute kubectl delete -f charts/flux-sandbox/crds/; note this will cascade-delete all sandbox-related CRs.

  16. Pre-warming interface returns 503 UNAVAILABLE

    The PostgreSQL or MoonCake etcd that the pre-warming subsystem depends on is temporarily unavailable; retry as-is after a short wait; if it persists, check the pre-warming-related startup parameters and dependency service connectivity.

  17. Can pre-warming tasks be canceled

    Task cancellation or deletion is not supported. When individual snapshots fail pre-warming, a new task must be submitted; successfully pre-warmed snapshots will not be re-distributed.