celld: Durable Objects on Your Own Storage Chapter 7 · Designing applications around cells

Chapter 7celld: Durable Objects on Your Own Storage

Designing applications around cells

Part II · How celld worksEdition celld v0.6.0Length 677 words · 3 min

The model pays off when you divide the application into named, stateful units from the start: one cell per user, per room, per agent, per document, per device. Because cells share no database, the classic distributed-system problems (locking, message buses, hot shards in a shared table) simply do not arise; the cell is the partition.

§01

The workloads the model fits

  • Real-time applications. A multiplayer game, chat room, or collaborative document is one cell holding both the WebSocket connections and the room's state, with no lock and no external message bus. A single Durable Object can coordinate thousands of clients through the hibernation API.
  • Agents. Each AI agent is one cell holding memory, schedule, and inbox in its own SQLite; an idle agent hibernates to the bucket, so a large agent fleet costs almost nothing between events.
  • Sharded web applications. One cell per user or tenant shards the app from the start; the contention of one shared database never appears because no shared database exists. KV, Queues, D1, and R2 round out the toolbox for the parts that are not per-entity state.
§02

Entities, not processes

A cell models an entity: a named unit with state that persists indefinitely, such as a concert, a user, or a document. A durable-execution engine (Temporal, Restate, Azure Durable Functions) models a process: a sequence of steps that ends, such as an order pipeline. celld's Workflows implementation is the process primitive built on the entity primitive, and its replay semantics (see Chapter 6) are exactly the tradeoff that primitive carries: steps must be idempotent, because a crash re-runs them. Pick the shape that matches the problem rather than the tool you have.

§03

A working cell

This is the whole application shape, a hibernating chat room with durable state:

ts
import { DurableObject } from "cloudflare:workers";

export class ChatRoom extends DurableObject {
  async fetch(req: Request) {
    const pair = new WebSocketPair();
    const [client, server] = Object.values(pair);
    this.ctx.acceptWebSocket(server);          // hibernatable — the room can sleep
    return new Response(null, { status: 101, webSocket: client });
  }

  async webSocketMessage(ws: WebSocket, message: string | ArrayBuffer) {
    for (const peer of this.ctx.getWebSockets()) {
      if (peer !== ws) peer.send(message);      // one room, no message bus
    }
  }

  async alarm() {
    await this.ctx.storage.put("lastAlarm", Date.now());
    this.ctx.storage.setAlarm(Date.now() + 60_000);  // durable self-schedule
  }
}
§04

Operational rules of thumb

  • Batch WebSocket messages. Each frame costs a context switch; pack many small logical messages into one frame with an envelope format. Fewer, larger messages beat many small ones.
  • Keep the constructor cheap. It runs on every wake, including each message to a hibernated cell. A schema-version check inside blockConcurrencyWhile() belongs there; loading the cell's state does not, so restore from storage in the handler.
  • Route a cell's traffic to its owner node when latency matters. The versioned peer tunnel lets any node ingress any cell, but the warm path (zero bucket operations, p50 ≈ 1.1 ms) only exists when the request lands on the owner.
  • Outbound WebSocket connections do not survive a move. An outbound DO socket keeps the cell resident and dies with the node/owner move; keep connection intent in storage and reconnect after activation. A WS transport cannot move to a new owner, so reconnect with a stable operation ID.
  • Make remote operations idempotent. celld does not retry a proxied call after transmission starts. Use a stable operation ID and design handlers to tolerate a retry.
  • One writer per cell, per D1 database, per KV namespace, per queue. Shard by adding entities, never by growing one.
Note
§05

Cloudflare's rules, on celld

Cloudflare's Rules of Durable Objects is the standard design checklist, and most of it carries over unchanged. The table sets each rule against celld's Durable Objects page. Three rows need the most attention: renames do not carry over, the constructor needs a precise rule, and celld has more ways to stop an object than Cloudflare does.

Tbl. 1
Cloudflare's ruleOn celld
Model one object per "atom" of coordination; never route all traffic through a global singletonThe same, and a singleton costs more here: every call to a cell is forwarded to the one node that owns it, so one hot cell loads one machine of your fleet.
Use deterministic IDs (getByName(), idFromName())Both work. The id is an HMAC-SHA-256 of the name under a key derived from the script name and the class name, so one name reaches one cell from any node. A newUniqueId() id cannot be derived again, so keep its string form.
Rename or delete a class through migrationsDoes not carry over. A migrations entry accepts only tag and new_sqlite_classes; a class rename, delete, or transfer stops the deployment. Renaming the Worker script changes every derived id: the renamed script reaches new, empty cells while the old data stays under the old ids. Keep script and class names stable, or migrate the data first.
Give a location hintAccepted with Cloudflare's values, but fleet ownership decides where a cell runs. celld makes no placement, migration, or jurisdiction promise, and the jurisdiction calls throw.
Run schema migrations in the constructor, inside blockConcurrencyWhile(); use it for nothing elseYes, but keep it to a cheap schema-version check: the constructor runs on every wake, and a blockConcurrencyWhile() or transaction longer than 30 seconds resets the object and rolls the transaction back.
Treat in-memory state as a cache; persist what mattersStronger here. Idle eviction drops memory, and ownership moves when a node drains, when rebalancing moves a hibernated cell, and when a node is lost. Only storage and WebSocket attachments survive.
Design for unexpected shutdowns: write progress as you gocelld has more ways to stop an object: idle eviction, a drain handoff that cancels firing alarms and active internal fetch/RPC handlers (Chapter 5), a rebalancing move, a self-fence that exits with code 3 (Chapter 4), and SIGKILL when the orchestrator's stop grace is too short. Persist each step before the next await that could be the last.
Rely on the output gate; don't await writes for safetyThe same promise, with a stronger proof: a response waits until a bucket or fleet proof covers every write it can reveal, and a WebSocket frame waits only for its own proof. transaction() and transactionSync() group writes; a nested transaction that fails discards only its own writes.
Guard against races across non-storage I/OThe same interleaving rule: storage calls are synchronous and never interleave, but an outbound fetch() or RPC await lets another event run. Keep a read-modify-write free of such awaits, or re-check a version after the await before writing.
Make alarm handlers idempotentRequired, and celld helps: alarm() receives retryCount and isRetry, and a drain arms a durable wake so the successor runs a cancelled alarm at least once.
Use hibernatable WebSockets and serializeAttachment()Supported, with attachments and tags. A hibernatable socket survives hibernation on the same node but closes with code 1012 when the cell moves to a new owner, so clients must reconnect.
An object doesn't know its name, so store it in an init() callMostly unnecessary: ctx.id.name carries the name for names up to 1,024 UTF-8 bytes, which is Cloudflare's own limit. A newUniqueId() cell has no name and still needs the record.
Clear an object with deleteAll()Available. celld's docs do not say whether it also clears a scheduled alarm, so call deleteAlarm() as well.
Plan for roughly 500–1,000 requests per second per objectThat is Cloudflare's measurement. celld publishes no per-cell figure; a warm request on the owner does zero bucket operations, but every write waits for its durability proof. Measure on your own fleet and storage.
Test with @cloudflare/vitest-pluginThat pool runs on workerd. celld's differential conformance tests hold celld to workerd's output, but celld's docs name no celld-backed test pool, so test ownership, durability, and moves against celld dev or a fleet.
Prefer RPC methods, and always await themThe same. An RPC stub cannot cross an isolate boundary, an AbortSignal does not cross a node boundary, and a proxied call is not retried once transmission starts, so keep operations idempotent.
Sources

This chapter draws on: celld: documentation at v0.6.0 (9 entries) · celld: release notes (5 entries) · Cloudflare documentation (4 entries). The full entries are in the Bibliography.