celld: Durable Objects on Your Own Storage Chapter 3 · The bucket is the coordinator

Chapter 3celld: Durable Objects on Your Own Storage

The bucket is the coordinator

Part II · How celld worksEdition celld v0.6.0Length 884 words · 4 min

This is the architectural centerpiece. There is no membership protocol, no failure detector, and no consensus service. A node is one celld process; a fleet is the set of nodes sharing one fleet bucket. Ownership of a cell is a record in that bucket, claimed with one atomic write. celld's built-in replicator continuously ships each cell's SQLite state to the bucket as LTX segments, Litestream's replica format from Ben Johnson. The loss of a node cannot lose an acknowledged write, because celld does not answer a write until the data survives a failure (RPO=0).

Fig. 1
Fleet topology and the ownership path Clients reach any celld node. Each node owns a set of resident cells, one SQLite database per cell. A request for chat-1 that lands on node B is proxied over the peer tunnel to node A, the owner, because only the owner may run the cell. Every node replicates its cells to the fleet bucket as LTX segments and reads leases, ownership records, and restores back out of it. clients—HTTP · WebSockets celld node A chat-1 · sqlite OWNER user-42 · sqlite __d1:ledger · sqlite … one thread each warm path · zero bucket ops p50 ≈ 1.1 ms celld node B chat-1 → owner: A room-9 · sqlite user-7 · sqlite … one thread each any node can ingress a non-owner call is proxied celld node C agent-3 · sqlite user-88 · sqlite free capacity … one thread each balancing hands it hibernated cells from full peers peer tunnel to the owner fetch · RPC · WS versioned no join command no membership list LTX segments continuous · RPO=0 every cell, every write restore on activation lease discovery S3-compatible bucket—the only coordinator ownership records · node leases one conditional write per claim cells/<id>/ltx/e<epoch>/ LTX segments · inactive cells deploy/ · telemetry/ · r2/ · kv large KV values · R2 objects no membership protocol · no failure detector · no consensus service—discovery is the leases in the bucket
The bucket supplies discovery and authority: nodes find each other through the leases in it, a node acquires a cell with one atomic write, and every cell's state is continuously replicated there as LTX segments. What it does not supply is network reachability. Peers talk over a private network or an encrypted overlay. Note the consequence of one-owner-per-cell: any node can take the request, but only the owner can run the cell, so a call that lands elsewhere is proxied to the owner over the versioned peer tunnel. That hop is why routing a cell's traffic to its owner is worth doing when latency matters.
§01

How RPO=0 is earned: bucket proof vs. fleet proof

The durability mechanism depends on fleet size, and it is explicit.

  • With one node, every write waits for the bucket. This is the bucket proof: one storage round trip, which is the minimum latency for a durable write.
  • With two or more nodes, the node serving the cell sends each write to another node and answers as soon as that node has the data on its own disk; the bucket upload finishes afterwards. This is the fleet proof. The owner and the nodes it sends to form the cell's ensemble, and each of those nodes is a follower. A node picks one or two followers (never itself), so a fleet of three or more nodes holds three copies of an acknowledged write, and the ensemble keeps acknowledging while one follower remains.

CELLD_DURABILITY selects the mode (default fleet). A single node requests the fleet posture and does not get it, falling back to bucket proof. Run two or more nodes if write latency matters.

Fig. 2
Where the acknowledgement lands, by fleet size Two timelines sharing a start. In a single-node fleet the write is acknowledged only after the bucket upload completes—one full storage round trip. In a fleet of two or more nodes the write is acknowledged as soon as a follower has fsynced it, and the bucket upload finishes afterwards, off the response path. the write ONE NODE bucket proof bucket upload ack one storage round trip — the floor for a durable write TWO+ NODES fleet proof follower fsync bucket upload—off the response path ack v0.3.0 measured ≈10× lower write latency and 100× fewer Class A bucket operations time—not to scale the ensemble—the owner picks one or two followers, never itself, so three nodes hold three copies of an acknowledged write, and the ensemble keeps acknowledging while one follower remains CELLD_DURABILITY=fleet is the default—a single-node fleet asks for the fleet posture, does not get it, and falls back to bucket proof
RPO=0 is earned differently by fleet size, and the whole difference is where the acknowledgement falls. The bucket is the long-term store in both modes; the fleet proof simply takes the upload off the response path.

Four properties make this possible, and they are non-negotiable requirements on the store:

  • Conditional create: creating an ownership record fails when the object already exists.
  • Conditional overwrite: a compare-and-swap on the prior record fails when the object changed after the read.
  • Read-after-write consistency: a read after a successful write returns that write.
  • Exact ranged reads: a Range request returns precisely the requested bytes. A large cell is restored page by page through a fault-in SQLite VFS that reads each page from the bucket on first use, so a wrong range is a correctness failure, and the startup probe checks it.
Tbl. 1
StoreFleet-qualified?Conditional-write path
Amazon S3YesIf-None-Match: * / If-Match etag CAS
Cloudflare R2Yessame headers; celld's release tests run here
Google Cloud StorageYesXML API x-goog-if-generation-match
TigrisYesdocumented conditional operations
Azure Blob StorageYesIf-None-Match: * / If-Match on Put Blob; qualified 2026-08-18
MinIO (community)Passes test, not qualifiedconditional writes work on RELEASE·2025-09-07 or later (#162 pinned one broken release)
Backblaze B2No—
Hetzner Object StorageNo—
DigitalOcean SpacesNo—
Caution

A bucket value can carry a key prefix (s3://bucket/team-a), so two fleets can share one bucket. A value without a prefix keeps objects at the bucket root, so an existing fleet never moves its data. The bucket also holds the deployments (with container images under deploy/images/), node leases, the shared peer-authentication secret, the fleet capacity sample, the alarm wake entries, large KV values, and all R2 objects: whoever holds the bucket credentials controls the fleet.

Sources

This chapter draws on: celld: documentation at v0.6.0 (9 entries) · celld: release notes (5 entries) · Cloudflare documentation (4 entries). The full entries are in the Bibliography.