celld: Durable Objects on Your Own Storage Chapter 5 · Running a fleet

Chapter 5celld: Durable Objects on Your Own Storage

Running a fleet

Part II · How celld worksEdition celld v0.6.0Length 2,199 words · 10 min

§01

Install and storage configuration

The installer downloads a signed binary. Replication runs in the celld process, so no external replicator is needed. Pin exact releases with CELLD_VERSION and verify build attestations with gh attestation verify. The immutable releases sit behind a single current pointer, which makes a previous SHA the rollback; there is no automatic update agent. Prebuilt binaries cover Linux x86-64, Linux ARM64, and Apple Silicon; Windows is not supported. On Amazon EKS, celld reads Pod Identity credentials from the injected environment and token file.

bash
# one-time install (pin a release with CELLD_VERSION)
curl -fsSL https://celld.dev/install.sh | sh

# Cloudflare R2 bucket — the standard AWS credential chain works
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=auto
export S3_ENDPOINT=https://ACCOUNT_ID.r2.cloudflarestorage.com
export CELLD_BUCKET=s3://cells

Three storage families, one environment contract:

  • S3-compatible (s3://): standard AWS chain; R2 via the S3 endpoint above; EKS Pod Identity supported.
  • Google Cloud Storage (gs://): Application Default Credentials or a GOOGLE_APPLICATION_CREDENTIALS service-account key; no AWS variables, region ignored.
  • Azure Blob (az://): exactly one credential family, which is an account key, VM managed identity, or AKS workload identity; the bucket name is the container.
§02

Develop locally with celld dev

celld dev opens a local store (a local SQLite object store), deploys the application, and starts one node, with no Docker and no cloud bucket:

bash
cd ./my-wrangler-project
celld dev              # Worker listener on http://127.0.0.1:9876 by default
celld dev --port 3000  # pick the port
celld dev --host 0.0.0.0  # expose the Worker listener (internal stays loopback)
celld dev --logs       # show the node's info/warning logs too
celld dev --clean      # discard .celld/dev first
celld dev --watch-ignore "docs/**"   # extra watcher ignores (repeatable)
celld dev --no-watch   # no automatic builds or restarts

State lives in .celld/dev under the project (add .celld/ to .gitignore), survives normal shutdown, and resets with --clean. The command watches the project and adopts a rebuilt deployment automatically; a failed build leaves the current app serving. The watcher ignores .celld, .wrangler, .git, node_modules, target, and any --watch-ignore globs, and a read is not a change. --no-watch turns automatic builds and restarts off entirely and cannot be combined with --watch-ignore. celld dev also reads a .dev.vars file beside the Wrangler config, as wrangler dev does: NAME=value lines, quotes stripped, overriding same-named vars, reloaded on edit, and never shipped to a fleet. Worker projects need esbuild on PATH; asset-only and no_bundle projects do not. The local store is not selectable by fleet nodes or operator subcommands. Fleets require a qualified cloud bucket.

§03

Deploy an application

celld deploy runs from a Wrangler project and accepts module Workers, Durable Object bindings, static assets, service bindings, D1 databases, KV namespaces, Queues, R2 buckets, Workflows, WebAssembly modules, cron triggers, worker_loaders, and containers. It stops with a named error on any Wrangler key it does not model. esbuild on PATH is needed only for projects with Worker code. Two properties of the deploy itself protect the fleet from a bad build. The manifest records a full SHA-256 digest for every JavaScript and WebAssembly module, and a node verifies each one before it builds the deployment, so changed bytes cannot become active. A prebuilt project (no_bundle: true) has its entry JavaScript preserved byte-for-byte, and celld discovers **/*.wasm modules below the entry's directory using Wrangler's default patterns (symlinks refused, no rules/find_additional_modules).

Fig. 1
Deploy, then adoption in place A Wrangler project is bundled with esbuild and pushed by celld deploy into the fleet bucket. Each node polls deploy slash current.json every thirty seconds, builds the new deployment beside the one it is serving, and switches new requests to it in one step without restarting; a request already in flight finishes on the previous deployment, and a Durable Object moves at a safe point. Wrangler project wrangler.jsonc · src/ migrations/ · assets/ esbuild bundle only if the project has Worker code celld deploy signed with the fleet secret fleet bucket deploy/current.json + deploy-blobs/ polled, not pushed EACH FLEET NODE—adoption in place, no restart reads deploy/current.json every CELLD_DEPLOY_POLL_S (30 s) · POST /reload on the internal listener adopts immediately deployment N—serving a request that already started finishes here deployment N+1—built beside it new requests switch to it in one step switch a Durable Object moves at a safe point—no in-flight request, alarm, pending durability, or regular WebSocket it keeps its storage, its epoch, and its hibernatable sockets · in the adoption window the two deployments must accept each other's calls no safe point within CELLD_DEPLOY_MAX_AGE_S (60 s) → the move is forced, with WebSocket code 1012, matching Cloudflare
Deployment is a push to the bucket and a pull by every node; no control plane tells nodes to reload. A node adopts the pulled deployment in place, without a restart.

A running node adopts a new deployment in place. It reads deploy/current.json every 30 seconds (CELLD_DEPLOY_POLL_S), builds the new deployment beside the one it serves, then switches new requests to it in one step; a request started on the previous deployment finishes on it. POST /reload on the internal listener adopts immediately. A Durable Object that is not resident runs the new deployment at its next activation; a resident one moves at a safe point. In the adoption window a request on one deployment can call a Durable Object on the other, so adjacent versions must accept each other's calls.

§04

Start nodes and grow the fleet

Local development needs only the default listener. A fleet node binds two listeners: a public Worker listener for ingress and an internal listener for the peer protocol and operator API. An explicit advertised address requires an explicit internal-listener address. celld also rejects an explicit non-loopback public listener without an internal one; this rule catches a stale single-listener configuration.

bash
celld \
  --bucket "$CELLD_BUCKET" \
  --listen 0.0.0.0:8080 \
  --internal-listen 10.0.0.12:8081 \
  --advertise node-a.internal:8081

To add a node, point it at the same bucket with a distinct internal address. There is no join command and no fixed membership list: nodes find each other through the leases in the bucket. The bucket supplies discovery and authority; it does not supply network reachability. The peer tunnel carries versioned plain HTTP for cell fetch and RPC traffic, with no content signature, so the private network is the security boundary; the fleet HMAC authenticates tunnel establishment and control requests. celld does not terminate TLS. Put the advertised addresses on a private network or an encrypted overlay such as WireGuard or Tailscale, and never expose the internal listener.

§05

Ownership balancing

A joining node takes hibernated cells from the nodes holding the most, so it carries its share within minutes, and the fleet evens out again after a node leaves. The mechanism keeps the no-coordinator posture. Every node reads a shared fleet capacity sample every 5 s (CELLD_REBALANCE_INTERVAL_MS; 0 disables). One node claims the refresh with a conditional write, reads every lease, and writes fleet/capacity-v1.json, so the cost does not grow with the square of the fleet. Each node's target is the fleet's owned cells divided by weight (CELLD_PLACEMENT_WEIGHT, default the CPU count). The node with the most owned cells per unit weight hands at most 32 hibernated cells per sample to the peer furthest below its share, one ownership-record write and one signed acquire each, and the receiver fills to 2% below target.

Only hibernated cells move. A resident cell hibernates through idle eviction first (CELLD_IDLE_EVICT_S), so a fleet without idle eviction balances only the cells that hibernate on their own. A moved cell's parked WebSockets close with code 1012. A draining node, or one with an activation backlog, receives nothing. The fleet moves nothing while any lease lacks a weight, so a rolling upgrade completes before the first move. POST /rebalance/pause and /resume on any internal listener govern the whole fleet.

§06

Graceful shutdown and upgrades

SIGTERM/SIGINT (what systemctl stop, docker stop, and a Kubernetes pod delete send) triggers a graceful drain: /.well-known/celld/health reports unhealthy, new public requests get a 503, and the node hands resident cells to peers. The handoff is batched and heavily engineered. For each batch the node:

  1. reserves a batch of cells, ordered by local request count with the newest request ID breaking ties;
  2. stops new local routes;
  3. cancels firing alarms, arming a durable wake so the successor runs them at least once, and cancels any active internal fetch/RPC handler;
  4. proves the batch durable in the live ensemble;
  5. publishes a full L9 snapshot and verifies the bucket holds a restore object, so the successor skips replay (a database too large for the 10 s durability budget, roughly 80 MiB, skips the snapshot and hands off through its L0 chain);
  6. releases ownership and asks a compatible peer to acquire, waiting for each acknowledgement before the next batch.

One variable bounds the whole stop. CELLD_SHUTDOWN_TOTAL_MS (default 40,000) derives the fleet drain-token wait (3/4 of it, 30 s) and the no-progress bound (5/8, 25 s). CELLD_RELEASES sets the number of concurrent handoffs (default 128). A fresh process holds its first healthy response until the fleet is settled (CELLD_READY_FLEET_GATE_MS, default 120,000). If the gate expires, the node emits a ready_gate_expired event once and keeps readiness closed until the condition clears, so give the orchestrator a rollout deadline that fails a persistent capacity problem.

Fig. 2
Graceful shutdown, batch by batch SIGTERM or SIGINT starts a drain: the health endpoint reports unhealthy, new public requests get a 503, and resident cells are handed to peers in batches, up to 128 handoffs at once. For each batch the node reserves cells by local request count, stops new local routes, cancels firing alarms behind a durable wake and cancels active internal fetch and RPC handlers, proves the batch durable in the live ensemble, publishes an L9 snapshot and verifies the restore object, then releases ownership to a compatible peer and waits for each acknowledgement before the next batch. CELLD_SHUTDOWN_TOTAL_MS, 40 seconds by default, bounds the stop and derives a 30 second drain-token wait and a 25 second no-progress bound. An orchestrator grace shorter than these bounds, such as a 30 second one, ends in SIGKILL before the handoff finishes. SIGTERM / SIGINT systemctl stop · docker stop Kubernetes pod delete the node drains /.well-known/celld/health → unhealthy · new public requests → 503 resident cells hand off to peers, one batch at a time CELLD_RELEASES = 128 handoffs in flight at once FOR EACH BATCH 1 · reserve a batch ordered by local request count newest request ID breaks ties 2 · stop new local routes no new requests start here for these cells 3 · cancel in-flight work firing alarms → a durable wake active internal fetch / RPC 4 · prove it durable the batch, in the live ensemble 5 · publish an L9 snapshot verify the restore object too big (~80 MiB): L0 chain 6 · release, peer acquires a compatible peer takes over wait for each acknowledgement next batch alarms still fire the durable wake runs them at least once, on the successor nothing moves in a draining node receives no cells from rebalancing bounds (defaults) 0 · SIGTERM no progress · 25 s (5/8) drain-token wait · 30 s (3/4) CELLD_SHUTDOWN_TOTAL_MS · 40 s grace long enough stop grace outlasts every bound grace too short SIGKILL at 30 s (the Kubernetes default): the handoff is cut off Give the orchestrator a stop grace longer than the shutdown bounds, or SIGKILL lands mid-handoff. Raise systemd TimeoutStopSec or Kubernetes terminationGracePeriodSeconds past the 40 s default.
The drain is a loop of six steps per batch, and one variable bounds all of it. The hazard is outside celld: if the orchestrator's stop grace ends first, SIGKILL lands mid-handoff. Kubernetes' default 30 s grace and docker stop's default 10 s are both shorter than celld's 40 s default bound, so raise the grace rather than trusting the default.
Warning
Warning
§07

Release notes

Appendix C records what each release changed, one row per release from v0.4.0 to v0.6.0.

§08

Diagnose a fleet

celld diagnose reads the node leases from the bucket and probes each live peer; it never takes a lease or changes ownership. It reports expired records, unsafe or incorrect advertised addresses, unreachable peers, authentication failures, and version disagreements, shows each node's load sample (owned cells, resident cells, WebSockets, RSS, CPU, file descriptors, pressure, shedding), and runs the storage test. celld cell list lists Durable Object instances, each line Class:ID. D1 databases, KV namespaces, and Workflows appear as reserved __ cells. Paginate with --after; a class-name argument scopes the storage prefix so --limit applies to that class alone, and the listing does not load the namespace into memory. During a rolling update, wait for every node to report restoring=0 before restarting the next.

GET /state on the internal listener is an autoscaler feed. It reports owned_cells, occupied, capacity_waiting, activation_waiting, restoring, and shedding, counters for handed_off, rebalanced, rebalance_failed, and remote_route_refreshes, and a node_load object mirroring the lease's sample (placement_weight, resident_cells, host_websockets, rss_bytes, cpu_percent_x100, open_fds, pressured, memory_headroom). It also carries allocator and (Linux) libc_malloc memory counters and a per-script deployment.isolates block: live, live_empty, retiring, freed, V8 heap and external bytes. A live_empty count that persists past 30 seconds signals a stuck maintenance pass, and a persistent retiring count a stuck request. A positive capacity_waiting is the add-a-node signal; scale down only while every remaining node reports headroom and a small restoring backlog. The health path stays a plain boolean: 503 during drain and before settle, no utilization number. /evict/<cell> waits for the eviction and answers {"ok":true}, or {"ok":false,"error":{"kind":…,"reason":…}} with a kind of refused (409/503: cell_active, alarm_imminent, eviction_limit, …), cancelled (new activity or a fence), or failed (500: a lost reply or a durability failure).

§09

Operate D1, KV, Queues, and R2

celld d1 runs SQL and migrations against a deployed D1 database, routing through the fleet to the database cell. The migration extension is ASCII case-insensitive, and migrations_dir must be a relative path inside the project. celld kv reads and writes a deployed KV namespace. The bulk commands use the Wrangler file format, so wrangler kv bulk get can export data for celld kv bulk put, and celld kv bulk get streams rows rather than holding the namespace in memory (an expiration is exported in Unix seconds, as Wrangler expects). celld kv list caps at 1000 keys per read and reports --after to continue. Every celld command writes data to stdout and messages to stderr, so redirects and pipes carry only data. celld queue info/peek/purge/pause/resume/redrive operates queues (purge needs --force; peek and redrive take --limit 1–100). celld r2 get|head|put|delete|list stands in for wrangler r2 object: it reads the fleet bucket directly with no running node, takes the bucket_name rather than the binding name, streams a get to stdout, and preserves the binding's metadata (--content-type, --cache-control, --metadata JSON, …); --local, --remote, and --jurisdiction are refused with an explanation.

§10

Primary environment variables

Tbl. 1
VariablePurpose
CELLD_BUCKETFleet bucket (+ optional key prefix). Same as --bucket.
S3_ENDPOINT, AWS_REGION, AWS_*S3-compatible endpoint and credentials (standard AWS chain; EKS Pod Identity supported).
GOOGLE_*Google credentials for a gs:// bucket (ADC or service-account key).
AZURE_*Azure account/identity for an az:// bucket (exactly one credential family).
CELLD_DURABILITYDurability mode: fleet (default) or bucket. Fleet proof needs ≥ 2 nodes.
CELLD_ADDR / CELLD_INTERNAL_ADDR / CELLD_ADVERTISEPublic listener, internal peer/operator listener, advertised address.
CELLD_ACTIVATIONSConcurrent cold-cell activations (default 8 per CPU, at least 16 and at most 128; a cold activation mostly waits on the store, so the default sits above the CPU count).
CELLD_MAX_RESIDENT_CELLSHard resident-cell cap, enforced at admission.
CELLD_IDLE_EVICT_SSeconds without work after which an idle resident cell hibernates (unset: only pressure or the cap removes it). Balancing moves hibernated cells only.
CELLD_PLACEMENT_WEIGHT / CELLD_REBALANCE_INTERVAL_MSThis node's ownership share relative to its peers (default: CPU count) and the fleet-sample interval (default 5,000; 0 disables balancing).
CELLD_MAX_CELL_REQUESTSConcurrent fetch limit for one Durable Object (default 64). A Queue broker has its own fixed limits: 256 concurrent producer calls, 64 per transaction, four overlapping proofs.
CELLD_MAX_REQUEST_BODY_BYTESBody limit for a public Worker request or direct DO request (default 1 GiB).
CELLD_MAX_RSS_MBMemory threshold for pressure shedding (default 80% of available memory; accounts for cgroup memory on Linux).
CELLD_TTL_MSNode-lease lifetime (default 10,000 ms).
CELLD_OPERATION_DEADLINE_MSDeadline for a non-restore operation (default 15,000).
CELLD_DEPLOY_POLL_S / CELLD_DEPLOY_MAX_AGE_SDeployment-adoption poll interval (30s) and forced-move age (60s).
CELLD_SHUTDOWN_TOTAL_MS / CELLD_RELEASESTotal stop bound (40s; the drain-token wait and no-progress bound derive from it) and concurrent handoffs (default 128).
CELLD_LTX_COMPACTION1 (default) creates additive L1 objects so a takeover reads tens of objects instead of thousands.
CELLD_LTX_PAGED / CELLD_LTX_PAGED_MIN_MB / CELLD_LTX_HYDRATE_MBPSPaged restore (default on) for chains above the threshold (default 256 MiB), and the background fill rate for the paged file (default 16; 0 keeps it sparse).
CELLD_RECOVERY_RETRY_MS / CELLD_RECOVERY_RETRIESNode-log recovery retry pacing (defaults 1,000 and 240); recovery reads bundles in 512 MiB windows and checkpoints every 32 cell epochs.
CELLD_WAKER_TICK_MSInterval of the fleet waker's alarm cleanup pass (default 60,000).
CELLD_ALARM_RESIDENT_MSHow close to its next alarm a cell (a Workflow instance, say) stays resident rather than hibernating (default one hour).
CELLD_DOCKER / CELLD_CONTAINER_PLATFORM / CELLD_CONTAINER_RUNTIMEContainers: the container CLI used to build and pull (default docker; Podman works), the deploy build platform (default linux/amd64), and a node-wide OCI runtime such as runsc or kata.
CELLD_OTEL0 off, 1 Parquet to the bucket, or an OTLP/HTTP collector base URL (see Chapter 8).
Rejected at startupCELLD_OUTPUT_GATE (the gate is always on; celld always waits for the configured durability proof), CELLD_STORAGE_PROBE, CELLD_SHUTDOWN_DRAIN_MS, CELLD_DRAIN_TOKEN_WAIT_MS, CELLD_WORKER_LOADER, CELLD_MAX_LOADED_WORKERS, CELLD_OTEL_SINK, CELLD_AI_BINDING, CELLD_AI_URL, CELLD_REBALANCE_BATCH_CELLS, and a handful of older tuning knobs. A node with any of them set does not start.
Sources

This chapter draws on: celld: documentation at v0.6.0 (9 entries) · celld: release notes (5 entries) · Cloudflare documentation (4 entries). The full entries are in the Bibliography.