Before you start
celld runs Cloudflare's Workers and Durable Objects programming model on machines you control, with an S3-compatible bucket as the only coordinator. The unit you build with is the cell: in Cloudflare terms, a Durable Object. A cell is a small named server with its own private SQLite database and one thread. It serves HTTP, holds WebSockets, sets alarms, and calls out. You make one cell per user, per document, per chat room, per agent. Cells share no database, so the application is sharded from the start. Because the JavaScript API is the Workers API, the same code runs on Cloudflare or on your fleet.
This guide takes you from an empty directory to a running fleet. The companion overview (Part II) explains the architecture, the guarantees, and the tradeoffs in depth; this guide points there when it needs that depth. A few terms recur throughout. A node is one celld process. A fleet is the set of nodes sharing one fleet bucket. Each cell has exactly one owner node at a time, recorded in an ownership record in the bucket that names the owner and carries a fencing epoch. Each node keeps a node lease in the bucket, renewed while the node lives. Locally, celld dev replaces the bucket with a local store.
You need three things on the development machine:
- A supported platform. Prebuilt binaries cover Linux x86-64, Linux ARM64, and Apple Silicon. Windows is not supported.
- esbuild on
PATH. celld bundles Worker code with it. Install it withnpm i -g esbuildorbrew install esbuild, or pointCELLD_ESBUILDat the binary. An asset-only project does not need it. - No cloud account yet.
celld devruns the whole system against the local store. You need a real bucket only when you stand up a fleet in Step 08.
Install the binary, and pin the release you tested:
# install (the installer downloads a signed binary to ~/.local/bin)
curl -fsSL https://celld.dev/install.sh | sh
# pin an exact release — rerunning with an older tag is the rollback
# (for the binary only; a v0.5.x or v0.6.x fleet's bucket cannot be served by a pre-v0.5.0 node — see Step 10)
CELLD_VERSION=v0.6.0 sh -c "$(curl -fsSL https://celld.dev/install.sh)"
# optionally verify the GitHub Actions build attestation
gh attestation verify ~/.local/bin/celld --repo denoland/celldScaffold the project
A celld application is a Wrangler project. There is no celld-specific project format. celld dev and celld deploy read wrangler.jsonc (or wrangler.json; not wrangler.toml) and accept the same layout Cloudflare's tooling does. Start with three files:
my-app/
├── wrangler.jsonc
├── src/
│ └── index.ts
├── .dev.vars # local-only secrets for celld dev — gitignore it
└── .gitignore # add .celld/ and .dev.vars hereThe minimal configuration declares the Worker entry point and one Durable Object class:
{
"$schema": "./node_modules/wrangler/config-schema.json",
"name": "my-app",
"main": "src/index.ts",
"compatibility_date": "2026-08-18",
"durable_objects": {
"bindings": [{ "name": "COUNTER", "class_name": "Counter" }]
},
"migrations": [{ "tag": "v1", "new_sqlite_classes": ["Counter"] }]
}Two details matter here. The migrations entry with new_sqlite_classes is what gives the class SQLite-backed storage. On celld every cell is SQLite-backed, and this is the declaration that matches Cloudflare's. The binding name COUNTER is how the Worker reaches the class through env.
The entry point, in the shape of celld's own counter example. The default export is the Worker; the exported class is the cell:
import { DurableObject } from "cloudflare:workers";
export interface Env {
COUNTER: DurableObjectNamespace;
}
export class Counter extends DurableObject {
async fetch(request: Request): Promise<Response> {
const n = ((await this.ctx.storage.get<number>("n")) ?? 0) + 1;
await this.ctx.storage.put("n", n);
return Response.json({ n });
}
}
export default {
async fetch(request: Request, env: Env): Promise<Response> {
const name = new URL(request.url).searchParams.get("name") ?? "default";
const id = env.COUNTER.idFromName(name);
return env.COUNTER.get(id).fetch(request);
},
};idFromName is the whole addressing model: the same name maps to the same cell from any node in the fleet, forever. The Worker is the stateless router; the cell is where state lives. TypeScript types come from @cloudflare/workers-types (npm i -D @cloudflare/workers-types). esbuild strips them at bundle time, so the deploy does not type-check; run tsc --noEmit yourself if you want the check.
Run it locally with celld dev
From the project directory:
celld dev # Worker listener on http://127.0.0.1:9876
celld dev --port 3000 # pick the port
celld dev --host 0.0.0.0 # expose the Worker listener; internal stays loopback
celld dev --logs # show the node's info/warning logs too
celld dev --clean # wipe .celld/dev first — a fresh local store
celld dev --watch-ignore "docs/**" # extra watcher ignores (repeatable)
celld dev --no-watch # no automatic builds or restarts (not with --watch-ignore)celld dev opens the local store (a SQLite-backed object store), deploys the application, and starts one node. No Docker, no cloud bucket, no configuration. It also reads a .dev.vars file beside the Wrangler config, exactly as wrangler dev does: one NAME=value per line (quotes stripped, no other dotenv features). Those values override same-named vars, reload on change, and are never carried to a fleet by celld deploy. Then exercise it:
$ curl "http://127.0.0.1:9876/?name=alpha"
{"n":1}
$ curl "http://127.0.0.1:9876/?name=alpha"
{"n":2}
$ curl "http://127.0.0.1:9876/?name=beta"
{"n":1}Two names, two cells, two independent SQLite databases. That is the model working.
Lab 1, Part 5 runs this loop live: new code adopted over old state, then a syntax error that fails the rebuild while the last good deployment keeps serving.
The rules of the loop:
- State lives in
.celld/devunder the project. It survives a normal shutdown;celld dev --cleanresets it. A configuration change is not migrated into stored state, so a renamed class or binding against an old store can fail confusingly. Clean and start again. Add.celld/and.dev.varsto.gitignore. - The watcher ignores
.celld,.wrangler,.git,node_modules, andtarget, plus any--watch-ignoreglobs. It does not follow sources outside the project directory, and it treats a read as no change, so a tool that only reads the project never triggers a build.--no-watchdisables automatic builds entirely. - The local store is dev-only. Fleet nodes and operator subcommands require a qualified cloud bucket; the local store is not selectable for them.
- The internal (operator) listener stays on loopback even with
--host 0.0.0.0, so another machine can reach your app but not the operator API.
Give the cell real state
The counter used the key-value face of cell storage. Every cell also carries a full SQLite database, synchronous from the cell's point of view: a storage operation never interleaves with another request, so there is no transaction ceremony for ordinary work. A chat room shows all three capabilities at once: SQL state, hibernatable WebSockets, and a durable alarm.
import { DurableObject } from "cloudflare:workers";
export class ChatRoom extends DurableObject {
async fetch(request: Request): Promise<Response> {
const pair = new WebSocketPair();
const [client, server] = Object.values(pair);
this.ctx.acceptWebSocket(server); // hibernatable — the room can sleep
return new Response(null, { status: 101, webSocket: client });
}
async webSocketMessage(ws: WebSocket, message: string | ArrayBuffer) {
this.ctx.storage.sql.exec(
"CREATE TABLE IF NOT EXISTS log (at INTEGER, body TEXT)",
);
this.ctx.storage.sql.exec(
"INSERT INTO log (at, body) VALUES (?, ?)",
Date.now(),
String(message),
);
for (const peer of this.ctx.getWebSockets()) {
if (peer !== ws) peer.send(message); // one room, no message bus
}
}
async alarm() {
// prune history older than a day, then reschedule
this.ctx.storage.sql.exec(
"DELETE FROM log WHERE at < ?",
Date.now() - 86_400_000,
);
await this.ctx.storage.setAlarm(Date.now() + 3_600_000);
}
}A cell is in one of three states. It is resident while it is in memory on its owner, hibernated when celld has evicted it from memory while its hibernatable WebSocket clients stay connected and it stays on its node, and inactive when no node holds it and it is only an object in the bucket. The rules that make this shape work in production follow from those states:
- Keep the constructor cheap. A cell keeps no memory across state transitions. The constructor runs again on every wake, including every message to a hibernated room. A schema-version check inside
blockConcurrencyWhile()belongs there; restoring state does not, so restore from storage inside the handler. - Hibernation is the economics.
acceptWebSocket()(rather than the addEventListener API) lets celld evict an idle room from memory while every client stays connected. A thousand quiet rooms cost almost nothing; a message wakes the one room it addresses. - Alarms are durable and covered by the acknowledgement gate. Lab 1, Part 4 shows an alarm set by a request firing on schedule and clearing its pending time. When a handler sets an alarm, celld does not send the successful response until a durable wake entry covers it, so a node loss after the response cannot lose the alarm. A failed
alarm()is retried with backoff. - Batch WebSocket messages. Every frame costs a context switch through the cell's one thread. Pack small logical messages into one frame with an envelope format. The reverse case works too: the output gate (the mechanism that holds a cell's outbound effects until the writes they depend on are durable) holds each outbound frame only for its own proof, so a
webSocketMessage()handler that sends a frame and then awaits delivers it while still running. You can stream an answer through one handler. - Outbound WebSockets do not survive a move. An outbound socket keeps the cell resident and dies when the cell changes nodes. Keep connection intent in storage and re-dial after activation.
- Finish write cursors before you answer. Outside an explicit transaction, a SQL write cursor (an unconsumed
RETURNING, say) must complete before a response, an outbound effect, orstorage.sync(); celld rejects the output with an error otherwise. A transaction orblockConcurrencyWhile()has a 30-second limit, and a timeout resets the object. Pending I/O after a handler returns stays alive withoutctx.waitUntil().
Add KV and Queues
Cells cover per-entity state. The first two cross-cutting services, both implemented as cells underneath, are Workers KV for shared key-value data and Queues for decoupled work. Both are graded Yes on the compatibility page. Each has its own page under celld.dev/docs/services with a narrative, a worked example, and a short "differences from Cloudflare" list. The differences below are the complete list.
KV: a durable store, not a CDN
"kv_namespaces": [
{ "binding": "SESSIONS", "id": "sessions-prod" }
]// in the Worker or any cell
await env.SESSIONS.put(`session:${token}`, JSON.stringify(claims), {
expirationTtl: 3600,
});
const raw = await env.SESSIONS.get(`session:${token}`);The API is Cloudflare's, with celld's storage model behind it. What changes in practice:
- No edge cache.
cacheTtlhas no effect andcacheStatusisnull. Reads route to the namespace's cell. Treat KV as a durable store with KV's API, not as a CDN. Lab 2, Part 2 showscacheStatus: nullon every read and names the namespace's own cell from the node log. - One writer per namespace. Write capacity scales by adding namespaces, not by writing harder to one. The
idaccepts Cloudflare's hex form or any stable string;sessions-prodis fine. - Values above 1 MiB go to the fleet bucket (small values live in the cell). They are stored under the namespace's ownership epoch, so a superseded owner can never delete the current value. Cloudflare's limits apply: keys up to 512 bytes, values up to 25 MiB, metadata up to 1,024 bytes, a minimum TTL of 60 seconds, and 1,000 keys per
list().put()accepts aReadableStreamvalue and checks the size limit while it reads. - Operate from the CLI:
celld kv get/put/delete/listplusbulkvariants that use the Wrangler file format, sowrangler kv bulk getoutput feedscelld kv bulk putdirectly.celld kv listpages at 1,000 keys and prints the--aftercursor on stderr; data goes to stdout, so pipes carry only data.celld kv bulk getstreams rows instead of holding the namespace in memory. A named output file is swapped in atomically when the export completes, but a failed stdout export leaves a truncated JSON array behind.
Queues: one writer, one consumer, four days
"queues": {
"producers": [{ "binding": "OUTBOX", "queue": "outbox" }],
"consumers": [
{ "queue": "outbox", "max_batch_size": 10, "max_batch_timeout": 5,
"max_retries": 2, "dead_letter_queue": "outbox-dead-letter" },
{ "queue": "outbox-dead-letter" }
]
}// one script can be both the HTTP ingress and the consumer (the examples/queues shape)
export default {
async fetch(request: Request, env: Env): Promise<Response> {
await env.OUTBOX.send({ kind: "email", to: "a@example.com", template: "welcome" });
return new Response("queued", { status: 202 });
},
async queue(batch: MessageBatch, env: Env): Promise<void> {
for (const msg of batch.messages) {
try {
await deliver(msg.body);
msg.ack();
} catch {
msg.retry({ delaySeconds: 1 }); // after max_retries it moves to the dead-letter queue
}
}
},
};The consumer settings are Cloudflare's. max_batch_size defaults to 10 (max 100), max_batch_timeout to 5 seconds (max 60), max_retries to 3, and max_concurrency caps at 250. A batch closes when it fills or when the timeout expires after its oldest ready message. A message is at most 128,000 bytes, a sendBatch() at most 100 messages and 256,000 bytes, and delaySeconds runs to 86,400. A handler that returns without settling acknowledges everything; one that throws returns the batch to the queue. A retried message becomes visible again after the delaySeconds passed to retry() or retryAll(), or the consumer's retry_delay by default. celld adds no exponential backoff; if you want one, compute the delay from msg.attempts. When attempts passes max_retries the message moves to the dead_letter_queue, which is an ordinary queue with its own consumer. Without a dead-letter queue, celld deletes the message. On the producer side, contentType ("v8", "json", "text", or "bytes") sets a message's encoding, with the queues_json_messages compatibility flag choosing the default, and a producer entry's delivery_delay sets a default delay for every message of that binding.
celld's Queues carry real constraints, all of the fail-loudly kind:
- One writer per queue. Scale with more queues. Within one queue, producer calls share transactions and durability rounds: a broker admits up to 256 concurrent producer calls, commits at most 64 per transaction, and overlaps four proofs. One queue sustained 7,357 sends per second in a 300,000-send soak. Past the admission limit the owner refuses the call with an error the producer can retry. A caught producer error contains
cell overload: admission refused, so a Worker can pass the same 503 back to its client. - One consumer script per queue. A deployment in which two scripts consume one queue fails. The consumer script may also export
fetch(): the officialexamples/queuesproject exportsfetchandqueuefrom one script, as above. Lab 2, Part 5 shows a consumer without aqueue()handler taking the local node down. - Retention is four days, not configurable. Pull consumers, the Queues HTTP API, dashboard controls, R2 event notifications, and Queue event subscriptions are not available.
- Overload is explicit. A saturated cell or queue answers 503 with
Retry-After: 1andX-Celld-Overload: cell. Count those as rejected work rather than retrying immediately at a fixed rate. - Operate with
celld queue info/peek/purge/pause/resume/redrive.
Orchestrate with Workflows
Cells model entities: named things whose state persists indefinitely. Workflows are the process primitive built on top of them: a sequence of steps that ends, with each step's result stored durably so the sequence survives crashes and restarts. Declare the class in the configuration:
"workflows": [
{ "binding": "REPORTS", "name": "report-builder", "class_name": "ReportBuilder" }
]A workflow extends WorkflowEntrypoint and does its work in run(). This is celld's own example, extended one step:
import { WorkflowEntrypoint } from "cloudflare:workers";
import type { WorkflowEvent, WorkflowStep } from "cloudflare:workers";
export class ReportBuilder extends WorkflowEntrypoint {
async run(event: WorkflowEvent<{ url: string }>, step: WorkflowStep) {
const fetched = await step.do("fetch source", async () => {
const response = await fetch(event.payload.url);
if (!response.ok) throw new Error(`source answered ${response.status}`);
const text = await response.text();
return { bytes: text.length, lines: text.split("\n").length };
});
await step.sleep("cool off", "30 seconds");
return await step.do("store summary", async () => {
// step.do callbacks are the only safe home for side effects
return { ...fetched, storedAt: Date.now() };
});
}
}
export default {
async fetch(request: Request, env: Env): Promise<Response> {
const url = new URL(request.url);
if (url.pathname === "/create") {
const instance = await env.REPORTS.create({
params: { url: url.searchParams.get("url") ?? "https://example.com" },
});
return Response.json({ id: instance.id });
}
const id = url.searchParams.get("id");
if (url.pathname === "/status" && id) {
const instance = await env.REPORTS.get(id);
return Response.json(await instance.status());
}
return new Response("Use /create?url=URL or /status?id=ID.", { status: 404 });
},
};The one rule that decides whether your workflow is correct is the replay rule. A running workflow is stored as its steps. After a crash, run() is replayed from the start: completed steps return their stored results instantly, and everything outside a step runs again. This is also how Cloudflare Workflows behave, which is why the compatibility page does not list it as a difference. The discipline is the same on either platform. Lab 3, Part 3 counts the re-execution: across one durable sleep, the top of run() ran twice while each step ran once.
run() from the top; stored steps are skipped, everything else re-executes.The celld-specific contract, from the workflows service page:
- Retention is yours to choose. A successful or failed instance is kept for 30 days by default. Each duration in the
retentionoption can be at most 30 days, anddelete()removes a completed run. - Limits. A step result, an event payload, and the workflow parameters are each capped at 1 MiB. Work outside a step cannot stay pending longer than 60 seconds. Pass references (an R2 key, a D1 row) between steps, not payloads.
- Defaults worth knowing.
step.do()retries 5 times with a 10-second delay and exponential backoff, each attempt capped at 10 minutes, andNonRetryableErrorstops the loop.waitForEvent()times out after 24 hours. An instance within an hour of its next alarm stays resident (CELLD_ALARM_RESIDENT_MS). Aworkflowsentry cannot carryschedules,limits, or ascript_namenaming another script. The Workflows REST API andwrangler workflowscommands do not operate against celld; drive instances through the binding. locationHintis accepted with Cloudflare's values, but fleet ownership decides where the cell actually runs.- Lifecycle. The page lists no difference for
pause(),resume(), orrestart(), which means celld intends to match Cloudflare's documented behavior for them. It also lists no caveat forcreate()called with a terminal instance's ID. If you rely on create-once semantics, verify against your installed release rather than assuming either way. - Not available: rollback, sensitive step results, and
ReadableStreamstep results.
Add D1 and R2
D1: a database that is a cell
A D1 database on celld is one cell holding one SQLite database, which means it inherits everything cells have: fencing, replication, durable acknowledgement. Declare it, write migrations, apply them:
"d1_databases": [
{ "binding": "DB", "database_name": "ledger" }
]# migrations are NNNN_description.sql files (.SQL works too),
# applied in numeric order, one transaction per file — exactly Wrangler's rule;
# a custom migrations_dir must be a relative path inside the project
mkdir -p migrations
cat > migrations/0001_init.sql <<'SQL'
CREATE TABLE entries (
id INTEGER PRIMARY KEY,
account TEXT NOT NULL,
amount INTEGER NOT NULL,
at INTEGER NOT NULL
);
SQL
# against a fleet (the only documented path — celld d1 finds a node through
# the bucket, so it cannot target celld dev's local store, and a dev deploy
# does not apply migrations; run the SQL from the Worker locally, as celld's
# own examples/d1 does with CREATE TABLE IF NOT EXISTS):
celld d1 migrations apply ledger --bucket "$CELLD_BUCKET"const { results } = await env.DB
.prepare("SELECT account, SUM(amount) AS balance FROM entries WHERE account = ? GROUP BY account")
.bind(account)
.all();- One database, one writer. More capacity comes from more databases, never a bigger one. Per-entity data belongs in the entity's cell; D1 is for the genuinely shared tables.
- Result caps: 100,000 rows or 32 MiB per binding result.
- Importing from Cloudflare:
wrangler d1 export, thencelld d1 execute ledger --file export.sql. A migration already recorded in the imported history does not run twice. - Not available:
dump(), Time Travel, the D1 REST API.wrangler d1commands do not operate against celld. - Migrations on
celld dev. Lab 2, Part 3 confirms that a dev deploy with a migration file leaves no tables until the Worker runs the SQL itself.
R2: your fleet bucket, wearing the R2 API
"r2_buckets": [
{ "binding": "FILES", "bucket_name": "files" }
]await env.FILES.put(`reports/${id}.json`, JSON.stringify(report));
const obj = await env.FILES.get(`reports/${id}.json`);
if (obj) { const report = await obj.json(); }An R2 binding stores its objects in the fleet bucket under r2/<bucket_name>/, which is the fleet bucket earning its name. Use R2 for step results and artifacts that outgrow the 1 MiB Workflow caps. The differences from Cloudflare:
- No public URL. celld serves a bucket through the binding only: no public bucket URL, no presigned URL, no S3 endpoint. To publish a file, put a Worker in front of it.
- Other tools' objects read fine. An object written by another tool reads through the binding, with its user metadata as
customMetadataand its headers ashttpMetadata. Write withcelld r2 putto get the complete record.delete(keys)takes up to 1,000 keys per call. - Not available:
ssecKeyandjurisdiction. - Size limits. A conditional write cannot stream a body larger than 8 MiB. Multipart uploads accept no checksum, cannot resume on another node or after a restart, and out-of-order parts buffer at most 256 MiB.
- Versions. An object's
versionis its content ETag, so identical content yields the same version on a store that reports no version of its own.celld dev's local store numbers each write instead, so identical bytes get different versions there (Lab 2, Part 4). Never use a version to count writes, and check your store before relying on version equality to de-duplicate;checksums.md5is the content hash everywhere. - Keys. Empty key segments count, as on Cloudflare:
a/b,/a/b,a//b, anda/b/are four objects. A key with non-ASCII or special characters is stored percent-encoded, solist()orders it (and appliesstartAfter) by that encoded form. Page with the returnedcursor, notstartAfter, if the order matters.
The R2 CLI reads the fleet bucket directly and needs no running node. celld r2 get|head|put|delete|list replaces wrangler r2 object, takes the bucket_name (not the binding name), and preserves the binding's object metadata:
celld r2 put files reports/2026-09.json --path report.json \
--content-type application/json --metadata '{"release":"1.2.3"}' \
--bucket "$CELLD_BUCKET"
celld r2 list files --prefix reports/ --bucket "$CELLD_BUCKET" # at most 1000 keys
celld r2 get files reports/2026-09.json --bucket "$CELLD_BUCKET" > report.jsonStand up a fleet
Everything so far ran on the local store. A fleet needs one thing the laptop cannot fake: a bucket that honors conditional writes with read-after-write consistency, because ownership of every cell rests on it. The bucket must also serve exact ranged reads, because a large cell is restored page by page. Qualified: Amazon S3, Cloudflare R2, Google Cloud Storage, Tigris, and Azure Blob Storage. Not qualified: Backblaze B2, Hetzner, DigitalOcean Spaces. MinIO passes the storage test but is not qualified for production. Using R2 as the example:
# an R2 bucket + an S3 API token scoped to it
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=auto
export S3_ENDPOINT=https://ACCOUNT_ID.r2.cloudflarestorage.com
export CELLD_BUCKET=s3://my-fleet # a /prefix lets fleets share a bucketGoogle Cloud Storage (gs://, Application Default Credentials) and Azure Blob (az://, where the bucket name is the container, with exactly one credential family) follow the same shape: set CELLD_BUCKET and the platform's own credentials.
Each node probes the store at startup: conditional writes that must succeed and fail in the right places, plus a ranged read that must return exactly the requested bytes. A clear contract violation stops the node at once rather than risk two owners for one cell. An ambiguous failure gets three attempts and then a warning. The probe cannot be switched off; CELLD_STORAGE_PROBE is rejected at startup.
A fleet node binds two listeners: the public Worker listener for ingress, and an internal listener for the peer protocol and operator API. The internal plane carries plain HTTP with no content signature. The network is the security boundary, so put it on a private network or an encrypted overlay (WireGuard, Tailscale), and never expose it:
celld \
--bucket "$CELLD_BUCKET" \
--listen 0.0.0.0:8080 \
--internal-listen 10.0.0.12:8081 \
--advertise node-a.internal:8081To grow the fleet, start another node against the same bucket with a distinct internal address. There is no join command and no membership list; nodes find each other through the node leases in the bucket. Run under a supervisor with no restart limit. A systemd unit carries all the operational rules at once:
[Unit]
Description=celld node
After=network-online.target
[Service]
EnvironmentFile=/etc/celld/env
ExecStart=/usr/local/bin/celld --listen 0.0.0.0:8080 \
--internal-listen 10.0.0.12:8081 --advertise node-a.internal:8081
Restart=always
RestartSec=10 # at least one lease lifetime (CELLD_TTL_MS, 10 s)
TimeoutStopSec=120 # must exceed the stop bound (CELLD_SHUTDOWN_TOTAL_MS, 40 s)
# plus the drain-token wait derived from it (3/4 → 30 s)
[Install]
WantedBy=multi-user.targetThree of those lines are load-bearing:
Restart=alwayswith no attempt limit. A node that loses bucket contact fences itself and exits with code 3. That exit is the design working, and the restart is the recovery. The fence log line names its cause: an expired lease, a missing lease record, or a record another writer replaced.RestartSec=10waits at least one lease lifetime so the old lease expires first.TimeoutStopSec=120exceeds celld's own shutdown bounds, so the graceful handoff (batched, durable, one node at a time across the fleet) finishes before SIGKILL lands. celld derives the drain-token wait and the no-progress interval fromCELLD_SHUTDOWN_TOTAL_MSalone (3/4 and 5/8 of it). The separateCELLD_SHUTDOWN_DRAIN_MSandCELLD_DRAIN_TOKEN_WAIT_MSvariables are rejected at startup.
Kubernetes users: the same three rules map to a restart policy, a backoff, and terminationGracePeriodSeconds, with /.well-known/celld/health as the readiness probe. Set a rollout deadline too, because a fresh node whose readiness gate never clears stays unhealthy instead of reporting healthy at the gate's timeout.
Run at least two nodes if write latency matters. With one node every durable write waits for a bucket round trip. With two or more, the owner and its followers form an ensemble: celld answers on a follower's fsync and uploads afterwards, for the same guarantee at roughly 10× lower write latency.
The fleet balances itself. A node that joins takes hibernated cells from the nodes holding the most, and the fleet evens out again after a node leaves. One node samples every lease every five seconds (CELLD_REBALANCE_INTERVAL_MS; 0 disables) and writes a shared capacity sample. The most-loaded node per unit of weight (CELLD_PLACEMENT_WEIGHT, default the CPU count) hands at most 32 hibernated cells per sample to the peer furthest below its share. Only hibernated cells move, so a fleet without idle eviction balances only the cells that hibernate on their own; set CELLD_IDLE_EVICT_S if you want resident-but-idle cells to become eligible. A moved cell's parked WebSockets close with code 1012 so clients reconnect to the new owner. POST /rebalance/pause and /resume on any node's internal listener stop and restart every move fleet-wide.
Deploy, operate, observe
Deploying to the fleet is one command from the project directory:
celld deploy . --bucket "$CELLD_BUCKET"It bundles (esbuild), signs with the fleet secret, records a full SHA-256 digest of every JavaScript and WebAssembly module in the manifest, and uploads. A node verifies each module against its digest before it builds the deployment. Nodes poll deploy/current.json every 30 seconds and adopt in place, with no restart, exactly as celld dev rehearsed: the new deployment is built beside the old, new requests switch in one step, in-flight requests finish where they started, and resident Durable Objects move at safe points (forced after 60 seconds with WebSocket code 1012, matching Cloudflare). POST /reload on the internal listener adopts immediately. One consequence deserves respect: during the adoption window a request on one deployment can call a Durable Object on the other, so adjacent versions must accept each other's calls.
The operator's toolbox, all reading the bucket, none taking ownership:
# fleet health: leases, reachability, advertised addresses, versions,
# each node's load sample (owned cells, resident cells, RSS, CPU …),
# and the conditional-write + ranged-read storage test
celld diagnose --bucket "$CELLD_BUCKET"
# list Durable Object instances (Class:ID); reserved __ cells are
# D1 databases, KV namespaces, and Workflows. A class argument scopes
# the storage prefix, so --limit then applies to that class alone
celld cell list --bucket "$CELLD_BUCKET"
celld cell list Counter --limit 50 --bucket "$CELLD_BUCKET"
celld cell list --all --json --bucket "$CELLD_BUCKET" |
jq -r 'select(.reserved | not) | .scope'
# data services
celld d1 migrations apply ledger --bucket "$CELLD_BUCKET"
celld kv list sessions --all --json --bucket "$CELLD_BUCKET" > keys.ndjson
celld queue info outbox --bucket "$CELLD_BUCKET"During a binary rolling update (where the release allows one; see Step 10), stop one node, wait for its replacement to report healthy and for every node to show restoring=0 in celld diagnose, then move on. That lets one restart's cold work finish before the next removes more warm capacity. The first healthy response of a fresh node already waits for the fleet to settle (the readiness gate), so the orchestrator needs no separate fleet-level pause. It does need a rollout deadline, because a node that never settles stays unhealthy rather than passing at the gate's timeout.
For an autoscaler, read the node leases or GET /state on the internal listener. Both publish owned_cells, resident_cells, host_websockets, rss_bytes, cpu_percent_x100, memory_headroom, restoring, and capacity_waiting, plus per-script isolate counts (live, live_empty, retiring, freed, with V8 heap bytes) and allocator counters. A live_empty count that persists past 30 seconds signals a stuck maintenance pass. A positive capacity_waiting is the add-a-node signal; scale down only while every remaining node reports headroom and a small restore backlog. The health path stays a plain boolean. The /evict/<cell> operator route waits for the eviction and reports the outcome as JSON: {"ok":true}, or a refused / cancelled / failed kind with a reason such as cell_active or alarm_imminent.
Observability is one variable away, and the sink choice is folded into it. CELLD_OTEL=1 writes spans as Parquet under telemetry/ in the fleet bucket: every request, cell event, outbound fetch, and console.log, trace-joined and queryable with DuckDB directly, no collector required. CELLD_OTEL=https://collector.internal:4318 sends the same data over OTLP/HTTP to that base URL instead. The separate CELLD_OTEL_SINK variable is rejected at startup.
CREATE VIEW traces AS SELECT * FROM
read_parquet('s3://my-fleet/telemetry/traces/*/*/*/*/*/*.parquet');
SELECT name, duration_us, trace_id FROM traces
ORDER BY duration_us DESC LIMIT 20;Production rules
The rules that keep a celld application healthy, gathered in one place. Most have appeared above; the rest come from the guarantees and limitations pages.
- One writer each: per cell, per D1 database, per KV namespace, per queue. Scale by adding entities, never by growing one.
- Make remote operations idempotent, with stable operation IDs. celld never retries a proxied call after body transmission begins (it keeps no replay copy), and it retries only an attempt that provably never started. An ambiguous attempt is yours to retry: same operation ID, handler tolerant of the repeat. The same rule covers WebSocket reconnects, which never move a transport between owners.
- Keep constructors trivial; batch WebSocket frames; restore state in handlers.
- Leave headroom. Balancing moves only hibernated cells and counts cells by node weight, not by what each cell costs. A fleet at its resident limit still has nowhere to put a lost node's cells.
- Respect the upgrade cliffs. Some version transitions must be full stops; others may roll. The complete list:
- v0.1→v0.2: full stop.
- v0.2.1→v0.3: may roll. Never downgrade without
node-log close: sealed epochin the shutdown log. - v0.3→v0.4: full stop.
- v0.4.0→v0.4.1: may roll, but never start a v0.4.0 binary once the fleet has paged a large cell in. That node can never activate it.
- v0.4.1→v0.5.0: full stop. The procedure from the guarantees page: stop application traffic and deployment writers; stop every old node and its supervisor; wait for every lease to expire; back up the bucket and node data directories; make sure no old binary can restart against the bucket (the format marker cannot stop one from writing); then start v0.5.0 on every node with the same names, addresses, and data directories. Startup migrates the alarm wake format before serving.
- v0.5.0→v0.5.1: rolls. Stop one node, wait for its replacement to report healthy, continue.
- v0.5.1→v0.6.0: full stop under the default
fleetdurability. A v0.6.0 node needs its followers to answer log-tail requests in a ranged format that v0.5.1 nodes do not speak, so stop every v0.5.1 node before starting v0.6.0. A fleet onCELLD_DURABILITY=buckethas no followers and can roll. - Do not roll back by starting an old binary. The only rollback is restoring the stopped-fleet backup, which loses writes made after it.
- Appendix C records what each release changed.
- Drop the removed knobs before you upgrade. celld refuses to start with any of
CELLD_OUTPUT_GATE(the write gate is always on),CELLD_STORAGE_PROBE,CELLD_SHUTDOWN_DRAIN_MS,CELLD_DRAIN_TOKEN_WAIT_MS,CELLD_WORKER_LOADER,CELLD_MAX_LOADED_WORKERS,CELLD_OTEL_SINK,CELLD_AI_BINDING,CELLD_AI_URL, orCELLD_REBALANCE_BATCH_CELLSin the environment. Grep/etc/celld/envfirst. - The security boundary is yours. TLS terminates in your ingress proxy. The internal listener stays on a private network or an encrypted overlay, and its operator API is mostly unauthenticated. celld does not authenticate your application's users; that is application code, plus
CELLD_MAX_REQUEST_BODY_BYTESandCELLD_MAX_CELL_REQUESTSfor limits. - One application per fleet, single-tenant. Hostile multi-tenancy is explicitly out of scope for the beta.
- Pin the release; keep operator tooling and binary together. The operator API can change between releases. Re-read the compatibility page before every version bump.
- Watch the health path. It is
/.well-known/celld/health. A fleet coming from v0.3 must update load balancers and probes to it.
The through-line of the whole guide: divide the application into named entities from the start, put every side effect inside something durable (a cell's storage, a queue message, a workflow step), and let the bucket, the epochs, and the fail-loudly deploys do what they were built to do.
This chapter draws on: celld: documentation at v0.6.0 (7 entries) · celld: release notes (4 entries). The full entries are in the Bibliography.