celld: Durable Objects on Your Own Storage Chapter 4 · Ownership, leases, and fencing

Chapter 4celld: Durable Objects on Your Own Storage

Ownership, leases, and fencing

Part II · How celld worksEdition celld v0.6.0Length 742 words · 3 min

celld makes two claims that a distributed system must actually prove: one node owns a cell at a time, and a write is durable before it is acknowledged. Both rest on the bucket and on epochs, not on clocks. The docs page What celld guarantees states the two promises up front and shows the mechanism.

§01

The ownership record and the epoch

Each cell has one ownership record in the bucket. It names the owner (the node session that may run the cell) and a fencing epoch. A node acquires a cell with a conditional write: create when no record exists, compare-and-swap when one does. The bucket accepts only one such write, so two nodes cannot acquire the same cell. Every activation advances the epoch: a takeover advances it, and a local wake advances it too. The replicator writes each cell's SQLite data under an epoch prefix, cells/<cell>/ltx/e<epoch>/.

Fig. 1
A takeover, read as a sequence Three lanes over time. Node A owns chat-1 at epoch e1 and stops renewing its lease. Node B acquires the cell with one conditional write that advances the ownership record to epoch e2, then restores from the e2 lineage. When node A comes back it writes into the superseded e1 prefix, re-reads the ownership record, finds node B at e2, does not acknowledge the write, and self-fences with exit code 3. node A the old owner the bucket the authority node B the new owner A owns chat-1 epoch e1 · lease held renewed each ~3.3 s A stops renewing paused, or cut off from the bucket A comes back PUT → cells/chat-1/ltx/e1/ re-reads the record → B · e2 not acked → SELF-FENCE, exit 3 chat-1 → node A · e1 superseded CAS · one winner chat-1 → node B · e2 the live record cells/chat-1/ltx/e1/ superseded—a restore never reads it cells/chat-1/ltx/e2/ the live lineage B acquires conditional write epoch e1 → e2 B restores reads the e2 lineage then serves chat-1 time The epoch in the object key is the fence, so the data path needs no conditional writes at all. Only the ownership record is a compare-and-swap; every LTX segment is a plain PUT under its own epoch prefix.
Fencing shown as a sequence rather than asserted as a property. The ownership record is the only compare-and-swap in the picture; the data path is plain PUTs, fenced by the epoch in the key. A node that cannot reach the bucket cannot renew its lease or replicate either, so it fences itself rather than serve state it can no longer prove.
§02

The acknowledgement rule (RPO=0)

The output gate holds each write response until a durability proof covers it. After a bucket proof, celld re-reads the ownership record and acknowledges only if it still names this node at this epoch. A partitioned node can commit locally and replicate into its superseded prefix, but the ownership read reveals the new owner, so the write is not acknowledged. The check reads the record rather than comparing a clock, so a paused process or a skewed clock cannot pass it. A fleet proof requires every follower to fsync the write, and a takeover seals the prior node-log session before restoring, so the stale owner cannot complete another fleet proof. The output gate applies one ordering rule across responses, outbound calls, Queue deliveries, and WebSocket sends, and read-only output waits for earlier request or alarm writes to become durable.

§03

Self-fencing

Each node holds a node lease in the bucket with an expiry (CELLD_TTL_MS, default 10,000 ms), renewed after one third of the lifetime. A node that cannot reach the bucket cannot renew or replicate, so it must not own cells. When its published expiry passes it fences itself: it stops each active cell, fails incomplete requests, logs a line starting SELF-FENCE:, and exits with code 3. The fence names its cause with a distinct event:

  • node_lease_watchdog_fence: the lease expired.
  • node_lease_record_missing_fence: the record is gone.
  • node_lease_record_mismatch_fence: another writer replaced it. The node cannot prove who, so it names no author.

A failed renewal retries before the authority expires. RUST_LOG=celld=info,store=debug logs every lease read and write with its outcome, at no cost when off. The fenced state is terminal; only a restart returns the node to the fleet. Two requirements follow:

  • Run under a supervisor (systemd, Docker restart policy, Kubernetes) with no attempt limit, waiting at least one lease lifetime between attempts. A node that cannot acquire a lease at startup retries rather than exiting. A restarting node first recovers its previous session's log before it takes a lease. If a peer is already recovering that log, it waits behind the peer's heartbeat and takes over only when the heartbeat stops, which is what makes a whole-fleet restart recover every acknowledged write.
  • A request is refused before the fence runs. celld compares the current time against the published expiry on every route, so a node with a lapsed lease refuses the request. The dispatch check keeps one owner per cell even while the fence is in flight.
Fig. 2
Self-fencing, read as a sequence Three lanes over time. The node renews its lease in the bucket every 3.3 seconds, a third of the 10,000 millisecond CELLD_TTL_MS. When renewals stop landing, it retries before the expiry; once the published expiry passes, every route refuses requests. The node then fences itself: it stops each active cell, fails incomplete requests, logs SELF-FENCE, and exits with code 3, naming its cause with one of three events: node_lease_watchdog_fence when the lease expired, node_lease_record_missing_fence when the record is gone, and node_lease_record_mismatch_fence when another writer replaced it. A supervisor restarts the node after at least one lease lifetime; it recovers its previous log before taking a new lease. HEALTHY LEASE LAPSES FENCE RESTART the node the leaseholder the bucket the lease record requests every route renews its lease every ~3.3 s, a third of CELLD_TTL_MS renewal fails bucket unreachable retries before expiry fences itself stops each active cell fails incomplete requests logs SELF-FENCE: · exit 3 restarts supervisor, no attempt cap waits ≥ one lease lifetime recovers its log first lease record expiry published TTL 10,000 ms no renewal lands the published expiry passes renew the lease expired node_lease_watchdog_fence the record is gone node_lease_record_missing_fence another writer replaced it node_lease_record_mismatch_fence the fence names its cause with exactly one event a mismatch names no author: it cannot prove who wrote it served the route checks the expiry: still ahead refused expiry has passed before the fence runs back in the fleet only after a restart and a new lease time A node that cannot prove its lease stops serving: first each request, then the whole process. The fenced state is terminal; only a supervisor restart, at least one lease lifetime later, brings it back.
Self-fencing shown as a sequence. Requests are refused the moment the published expiry passes, before the fence itself runs, so a node never serves a cell it can no longer prove it owns. The fence then ends the process with exit code 3 and one of three events that name its cause. Because the fenced state is terminal, run celld under a supervisor with no attempt limit.
Note
§04

Remote calls and retries

Every proxied cell call (fetch, RPC, and WebSocket) runs through one versioned peer tunnel that streams the request body to the owner. The consequence for application code: celld does not retry a call after transmission begins, because it keeps no replay copy of the body. It retries only a peer attempt that proves the handler did not start; an ambiguous attempt (the handler may have completed without returning) is not retried. A stale route, a call that resolves an owner generation a replacement process now rejects, is handled underneath: celld waits for a different (node, epoch), refreshes the route, and re-attempts within CELLD_OPERATION_DEADLINE_MS. Keep one stable operation ID when you retry an ambiguous fetch, RPC, D1, or service operation, and make operations idempotent at the application layer. An AbortSignal passes through an RPC call on the same node; it does not cross a node boundary.

Sources

This chapter draws on: celld: documentation at v0.6.0 (9 entries) · celld: release notes (5 entries) · Cloudflare documentation (4 entries). The full entries are in the Bibliography.