Install and storage configuration
The installer downloads a signed binary. Replication runs in the celld process, so no external replicator is needed. Pin exact releases with CELLD_VERSION and verify build attestations with gh attestation verify. The immutable releases sit behind a single current pointer, which makes a previous SHA the rollback; there is no automatic update agent. Prebuilt binaries cover Linux x86-64, Linux ARM64, and Apple Silicon; Windows is not supported. On Amazon EKS, celld reads Pod Identity credentials from the injected environment and token file.
# one-time install (pin a release with CELLD_VERSION)
curl -fsSL https://celld.dev/install.sh | sh
# Cloudflare R2 bucket — the standard AWS credential chain works
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=auto
export S3_ENDPOINT=https://ACCOUNT_ID.r2.cloudflarestorage.com
export CELLD_BUCKET=s3://cellsThree storage families, one environment contract:
- S3-compatible (
s3://): standard AWS chain; R2 via the S3 endpoint above; EKS Pod Identity supported. - Google Cloud Storage (
gs://): Application Default Credentials or aGOOGLE_APPLICATION_CREDENTIALSservice-account key; no AWS variables, region ignored. - Azure Blob (
az://): exactly one credential family, which is an account key, VM managed identity, or AKS workload identity; the bucket name is the container.
Develop locally with celld dev
celld dev opens a local store (a local SQLite object store), deploys the application, and starts one node, with no Docker and no cloud bucket:
cd ./my-wrangler-project
celld dev # Worker listener on http://127.0.0.1:9876 by default
celld dev --port 3000 # pick the port
celld dev --host 0.0.0.0 # expose the Worker listener (internal stays loopback)
celld dev --logs # show the node's info/warning logs too
celld dev --clean # discard .celld/dev first
celld dev --watch-ignore "docs/**" # extra watcher ignores (repeatable)
celld dev --no-watch # no automatic builds or restartsState lives in .celld/dev under the project (add .celld/ to .gitignore), survives normal shutdown, and resets with --clean. The command watches the project and adopts a rebuilt deployment automatically; a failed build leaves the current app serving. The watcher ignores .celld, .wrangler, .git, node_modules, target, and any --watch-ignore globs, and a read is not a change. --no-watch turns automatic builds and restarts off entirely and cannot be combined with --watch-ignore. celld dev also reads a .dev.vars file beside the Wrangler config, as wrangler dev does: NAME=value lines, quotes stripped, overriding same-named vars, reloaded on edit, and never shipped to a fleet. Worker projects need esbuild on PATH; asset-only and no_bundle projects do not. The local store is not selectable by fleet nodes or operator subcommands. Fleets require a qualified cloud bucket.
Deploy an application
celld deploy runs from a Wrangler project and accepts module Workers, Durable Object bindings, static assets, service bindings, D1 databases, KV namespaces, Queues, R2 buckets, Workflows, WebAssembly modules, cron triggers, worker_loaders, and containers. It stops with a named error on any Wrangler key it does not model. esbuild on PATH is needed only for projects with Worker code. Two properties of the deploy itself protect the fleet from a bad build. The manifest records a full SHA-256 digest for every JavaScript and WebAssembly module, and a node verifies each one before it builds the deployment, so changed bytes cannot become active. A prebuilt project (no_bundle: true) has its entry JavaScript preserved byte-for-byte, and celld discovers **/*.wasm modules below the entry's directory using Wrangler's default patterns (symlinks refused, no rules/find_additional_modules).
A running node adopts a new deployment in place. It reads deploy/current.json every 30 seconds (CELLD_DEPLOY_POLL_S), builds the new deployment beside the one it serves, then switches new requests to it in one step; a request started on the previous deployment finishes on it. POST /reload on the internal listener adopts immediately. A Durable Object that is not resident runs the new deployment at its next activation; a resident one moves at a safe point. In the adoption window a request on one deployment can call a Durable Object on the other, so adjacent versions must accept each other's calls.
Start nodes and grow the fleet
Local development needs only the default listener. A fleet node binds two listeners: a public Worker listener for ingress and an internal listener for the peer protocol and operator API. An explicit advertised address requires an explicit internal-listener address. celld also rejects an explicit non-loopback public listener without an internal one; this rule catches a stale single-listener configuration.
celld \
--bucket "$CELLD_BUCKET" \
--listen 0.0.0.0:8080 \
--internal-listen 10.0.0.12:8081 \
--advertise node-a.internal:8081To add a node, point it at the same bucket with a distinct internal address. There is no join command and no fixed membership list: nodes find each other through the leases in the bucket. The bucket supplies discovery and authority; it does not supply network reachability. The peer tunnel carries versioned plain HTTP for cell fetch and RPC traffic, with no content signature, so the private network is the security boundary; the fleet HMAC authenticates tunnel establishment and control requests. celld does not terminate TLS. Put the advertised addresses on a private network or an encrypted overlay such as WireGuard or Tailscale, and never expose the internal listener.
Ownership balancing
A joining node takes hibernated cells from the nodes holding the most, so it carries its share within minutes, and the fleet evens out again after a node leaves. The mechanism keeps the no-coordinator posture. Every node reads a shared fleet capacity sample every 5 s (CELLD_REBALANCE_INTERVAL_MS; 0 disables). One node claims the refresh with a conditional write, reads every lease, and writes fleet/capacity-v1.json, so the cost does not grow with the square of the fleet. Each node's target is the fleet's owned cells divided by weight (CELLD_PLACEMENT_WEIGHT, default the CPU count). The node with the most owned cells per unit weight hands at most 32 hibernated cells per sample to the peer furthest below its share, one ownership-record write and one signed acquire each, and the receiver fills to 2% below target.
Only hibernated cells move. A resident cell hibernates through idle eviction first (CELLD_IDLE_EVICT_S), so a fleet without idle eviction balances only the cells that hibernate on their own. A moved cell's parked WebSockets close with code 1012. A draining node, or one with an activation backlog, receives nothing. The fleet moves nothing while any lease lacks a weight, so a rolling upgrade completes before the first move. POST /rebalance/pause and /resume on any internal listener govern the whole fleet.
Graceful shutdown and upgrades
SIGTERM/SIGINT (what systemctl stop, docker stop, and a Kubernetes pod delete send) triggers a graceful drain: /.well-known/celld/health reports unhealthy, new public requests get a 503, and the node hands resident cells to peers. The handoff is batched and heavily engineered. For each batch the node:
- reserves a batch of cells, ordered by local request count with the newest request ID breaking ties;
- stops new local routes;
- cancels firing alarms, arming a durable wake so the successor runs them at least once, and cancels any active internal fetch/RPC handler;
- proves the batch durable in the live ensemble;
- publishes a full L9 snapshot and verifies the bucket holds a restore object, so the successor skips replay (a database too large for the 10 s durability budget, roughly 80 MiB, skips the snapshot and hands off through its L0 chain);
- releases ownership and asks a compatible peer to acquire, waiting for each acknowledgement before the next batch.
One variable bounds the whole stop. CELLD_SHUTDOWN_TOTAL_MS (default 40,000) derives the fleet drain-token wait (3/4 of it, 30 s) and the no-progress bound (5/8, 25 s). CELLD_RELEASES sets the number of concurrent handoffs (default 128). A fresh process holds its first healthy response until the fleet is settled (CELLD_READY_FLEET_GATE_MS, default 120,000). If the gate expires, the node emits a ready_gate_expired event once and keeps readiness closed until the condition clears, so give the orchestrator a rollout deadline that fails a persistent capacity problem.
docker stop's default 10 s are both shorter than celld's 40 s default bound, so raise the grace rather than trusting the default.Release notes
Appendix C records what each release changed, one row per release from v0.4.0 to v0.6.0.
Diagnose a fleet
celld diagnose reads the node leases from the bucket and probes each live peer; it never takes a lease or changes ownership. It reports expired records, unsafe or incorrect advertised addresses, unreachable peers, authentication failures, and version disagreements, shows each node's load sample (owned cells, resident cells, WebSockets, RSS, CPU, file descriptors, pressure, shedding), and runs the storage test. celld cell list lists Durable Object instances, each line Class:ID. D1 databases, KV namespaces, and Workflows appear as reserved __ cells. Paginate with --after; a class-name argument scopes the storage prefix so --limit applies to that class alone, and the listing does not load the namespace into memory. During a rolling update, wait for every node to report restoring=0 before restarting the next.
GET /state on the internal listener is an autoscaler feed. It reports owned_cells, occupied, capacity_waiting, activation_waiting, restoring, and shedding, counters for handed_off, rebalanced, rebalance_failed, and remote_route_refreshes, and a node_load object mirroring the lease's sample (placement_weight, resident_cells, host_websockets, rss_bytes, cpu_percent_x100, open_fds, pressured, memory_headroom). It also carries allocator and (Linux) libc_malloc memory counters and a per-script deployment.isolates block: live, live_empty, retiring, freed, V8 heap and external bytes. A live_empty count that persists past 30 seconds signals a stuck maintenance pass, and a persistent retiring count a stuck request. A positive capacity_waiting is the add-a-node signal; scale down only while every remaining node reports headroom and a small restoring backlog. The health path stays a plain boolean: 503 during drain and before settle, no utilization number. /evict/<cell> waits for the eviction and answers {"ok":true}, or {"ok":false,"error":{"kind":…,"reason":…}} with a kind of refused (409/503: cell_active, alarm_imminent, eviction_limit, …), cancelled (new activity or a fence), or failed (500: a lost reply or a durability failure).
Operate D1, KV, Queues, and R2
celld d1 runs SQL and migrations against a deployed D1 database, routing through the fleet to the database cell. The migration extension is ASCII case-insensitive, and migrations_dir must be a relative path inside the project. celld kv reads and writes a deployed KV namespace. The bulk commands use the Wrangler file format, so wrangler kv bulk get can export data for celld kv bulk put, and celld kv bulk get streams rows rather than holding the namespace in memory (an expiration is exported in Unix seconds, as Wrangler expects). celld kv list caps at 1000 keys per read and reports --after to continue. Every celld command writes data to stdout and messages to stderr, so redirects and pipes carry only data. celld queue info/peek/purge/pause/resume/redrive operates queues (purge needs --force; peek and redrive take --limit 1–100). celld r2 get|head|put|delete|list stands in for wrangler r2 object: it reads the fleet bucket directly with no running node, takes the bucket_name rather than the binding name, streams a get to stdout, and preserves the binding's metadata (--content-type, --cache-control, --metadata JSON, …); --local, --remote, and --jurisdiction are refused with an explanation.
Primary environment variables
| Variable | Purpose |
|---|---|
CELLD_BUCKET | Fleet bucket (+ optional key prefix). Same as --bucket. |
S3_ENDPOINT, AWS_REGION, AWS_* | S3-compatible endpoint and credentials (standard AWS chain; EKS Pod Identity supported). |
GOOGLE_* | Google credentials for a gs:// bucket (ADC or service-account key). |
AZURE_* | Azure account/identity for an az:// bucket (exactly one credential family). |
CELLD_DURABILITY | Durability mode: fleet (default) or bucket. Fleet proof needs ≥ 2 nodes. |
CELLD_ADDR / CELLD_INTERNAL_ADDR / CELLD_ADVERTISE | Public listener, internal peer/operator listener, advertised address. |
CELLD_ACTIVATIONS | Concurrent cold-cell activations (default 8 per CPU, at least 16 and at most 128; a cold activation mostly waits on the store, so the default sits above the CPU count). |
CELLD_MAX_RESIDENT_CELLS | Hard resident-cell cap, enforced at admission. |
CELLD_IDLE_EVICT_S | Seconds without work after which an idle resident cell hibernates (unset: only pressure or the cap removes it). Balancing moves hibernated cells only. |
CELLD_PLACEMENT_WEIGHT / CELLD_REBALANCE_INTERVAL_MS | This node's ownership share relative to its peers (default: CPU count) and the fleet-sample interval (default 5,000; 0 disables balancing). |
CELLD_MAX_CELL_REQUESTS | Concurrent fetch limit for one Durable Object (default 64). A Queue broker has its own fixed limits: 256 concurrent producer calls, 64 per transaction, four overlapping proofs. |
CELLD_MAX_REQUEST_BODY_BYTES | Body limit for a public Worker request or direct DO request (default 1 GiB). |
CELLD_MAX_RSS_MB | Memory threshold for pressure shedding (default 80% of available memory; accounts for cgroup memory on Linux). |
CELLD_TTL_MS | Node-lease lifetime (default 10,000 ms). |
CELLD_OPERATION_DEADLINE_MS | Deadline for a non-restore operation (default 15,000). |
CELLD_DEPLOY_POLL_S / CELLD_DEPLOY_MAX_AGE_S | Deployment-adoption poll interval (30s) and forced-move age (60s). |
CELLD_SHUTDOWN_TOTAL_MS / CELLD_RELEASES | Total stop bound (40s; the drain-token wait and no-progress bound derive from it) and concurrent handoffs (default 128). |
CELLD_LTX_COMPACTION | 1 (default) creates additive L1 objects so a takeover reads tens of objects instead of thousands. |
CELLD_LTX_PAGED / CELLD_LTX_PAGED_MIN_MB / CELLD_LTX_HYDRATE_MBPS | Paged restore (default on) for chains above the threshold (default 256 MiB), and the background fill rate for the paged file (default 16; 0 keeps it sparse). |
CELLD_RECOVERY_RETRY_MS / CELLD_RECOVERY_RETRIES | Node-log recovery retry pacing (defaults 1,000 and 240); recovery reads bundles in 512 MiB windows and checkpoints every 32 cell epochs. |
CELLD_WAKER_TICK_MS | Interval of the fleet waker's alarm cleanup pass (default 60,000). |
CELLD_ALARM_RESIDENT_MS | How close to its next alarm a cell (a Workflow instance, say) stays resident rather than hibernating (default one hour). |
CELLD_DOCKER / CELLD_CONTAINER_PLATFORM / CELLD_CONTAINER_RUNTIME | Containers: the container CLI used to build and pull (default docker; Podman works), the deploy build platform (default linux/amd64), and a node-wide OCI runtime such as runsc or kata. |
CELLD_OTEL | 0 off, 1 Parquet to the bucket, or an OTLP/HTTP collector base URL (see Chapter 8). |
| Rejected at startup | CELLD_OUTPUT_GATE (the gate is always on; celld always waits for the configured durability proof), CELLD_STORAGE_PROBE, CELLD_SHUTDOWN_DRAIN_MS, CELLD_DRAIN_TOKEN_WAIT_MS, CELLD_WORKER_LOADER, CELLD_MAX_LOADED_WORKERS, CELLD_OTEL_SINK, CELLD_AI_BINDING, CELLD_AI_URL, CELLD_REBALANCE_BATCH_CELLS, and a handful of older tuning knobs. A node with any of them set does not start. |
This chapter draws on: celld: documentation at v0.6.0 (9 entries) · celld: release notes (5 entries) · Cloudflare documentation (4 entries). The full entries are in the Bibliography.