celld makes two claims that a distributed system must actually prove: one node owns a cell at a time, and a write is durable before it is acknowledged. Both rest on the bucket and on epochs, not on clocks. The docs page What celld guarantees states the two promises up front and shows the mechanism.
The ownership record and the epoch
Each cell has one ownership record in the bucket. It names the owner (the node session that may run the cell) and a fencing epoch. A node acquires a cell with a conditional write: create when no record exists, compare-and-swap when one does. The bucket accepts only one such write, so two nodes cannot acquire the same cell. Every activation advances the epoch: a takeover advances it, and a local wake advances it too. The replicator writes each cell's SQLite data under an epoch prefix, cells/<cell>/ltx/e<epoch>/.
The acknowledgement rule (RPO=0)
The output gate holds each write response until a durability proof covers it. After a bucket proof, celld re-reads the ownership record and acknowledges only if it still names this node at this epoch. A partitioned node can commit locally and replicate into its superseded prefix, but the ownership read reveals the new owner, so the write is not acknowledged. The check reads the record rather than comparing a clock, so a paused process or a skewed clock cannot pass it. A fleet proof requires every follower to fsync the write, and a takeover seals the prior node-log session before restoring, so the stale owner cannot complete another fleet proof. The output gate applies one ordering rule across responses, outbound calls, Queue deliveries, and WebSocket sends, and read-only output waits for earlier request or alarm writes to become durable.
Self-fencing
Each node holds a node lease in the bucket with an expiry (CELLD_TTL_MS, default 10,000 ms), renewed after one third of the lifetime. A node that cannot reach the bucket cannot renew or replicate, so it must not own cells. When its published expiry passes it fences itself: it stops each active cell, fails incomplete requests, logs a line starting SELF-FENCE:, and exits with code 3. The fence names its cause with a distinct event:
node_lease_watchdog_fence: the lease expired.node_lease_record_missing_fence: the record is gone.node_lease_record_mismatch_fence: another writer replaced it. The node cannot prove who, so it names no author.
A failed renewal retries before the authority expires. RUST_LOG=celld=info,store=debug logs every lease read and write with its outcome, at no cost when off. The fenced state is terminal; only a restart returns the node to the fleet. Two requirements follow:
- Run under a supervisor (systemd, Docker restart policy, Kubernetes) with no attempt limit, waiting at least one lease lifetime between attempts. A node that cannot acquire a lease at startup retries rather than exiting. A restarting node first recovers its previous session's log before it takes a lease. If a peer is already recovering that log, it waits behind the peer's heartbeat and takes over only when the heartbeat stops, which is what makes a whole-fleet restart recover every acknowledged write.
- A request is refused before the fence runs. celld compares the current time against the published expiry on every route, so a node with a lapsed lease refuses the request. The dispatch check keeps one owner per cell even while the fence is in flight.
Remote calls and retries
Every proxied cell call (fetch, RPC, and WebSocket) runs through one versioned peer tunnel that streams the request body to the owner. The consequence for application code: celld does not retry a call after transmission begins, because it keeps no replay copy of the body. It retries only a peer attempt that proves the handler did not start; an ambiguous attempt (the handler may have completed without returning) is not retried. A stale route, a call that resolves an owner generation a replacement process now rejects, is handled underneath: celld waits for a different (node, epoch), refreshes the route, and re-attempts within CELLD_OPERATION_DEADLINE_MS. Keep one stable operation ID when you retry an ambiguous fetch, RPC, D1, or service operation, and make operations idempotent at the application layer. An AbortSignal passes through an RPC call on the same node; it does not cross a node boundary.
This chapter draws on: celld: documentation at v0.6.0 (9 entries) · celld: release notes (5 entries) · Cloudflare documentation (4 entries). The full entries are in the Bibliography.