Reliability, testing, and telemetry
celld stakes three promises, and its testing is organized around breaking them: an acknowledged write is durable, a cell has one writer at a time, and code written for Cloudflare behaves the same on celld. The engineering is unusually rigorous for a pre-1.0 project, and worth respecting on its merits.
Four test layers
- Differential conformance. Each program runs twice, once on workerd and once on celld, identical bytes, and the outputs must be equal. A test cannot agree with celld's own runtime by accident.
- Exhaustive specification. The coordination protocol is specified in TLA+ and model-checked at small configuration, with pinned expected verdicts, most of them failures that model bugs the protocol once had, kept as a canary. The model found four bugs and a split-brain that lost an acknowledged write, none of which had surfaced in review or testing. The fencing argument itself is checked, not asserted.
- Deterministic simulation. The coordination logic is a pure decision core with no I/O; a seeded scheduler injects latency, CAS races, lost responses, drifting clocks, and crashes at every await point. Safety and liveness properties must survive tens of thousands of seeds; the core protocols have run through millions of schedules. They also test the checkers: deliberately broken protocol variants must be caught.
- Live fleet lab. Real VMs, a real bucket, fault injection between verification passes (SIGKILL mid-write and delete the local DB; freeze an owner and unfreeze it; cut a node off from the bucket; throttle the bucket to 429s; stop a full host). Every scenario's verification sweep found zero lost acknowledged writes.
Numbers with their conditions
| Measurement | Result |
|---|---|
| Epoch fence under contention | 500 claimants, 5,500 attempts, one writer per epoch, zero violations |
| Warm resident request | zero bucket operations; p50 ≈ 1.1 ms, p99 ≈ 7 ms (fixed host) |
| Durable write | one bucket round trip with bucket proof; a fleet proof (≥2 nodes) answers on follower fsync. v0.3.0 measured 10× lower write latency and 100× fewer Class A S3 ops |
| Concurrent writes to one cell | join a single shared upload, so throughput is not one round trip per write |
| Scale (measured) | 10 nodes × 4 vCPU / 8 GB held 10,000 resident cells + 20,000 concurrent WebSockets |
| Node failure | stopping 2 of 10 nodes: every cell's data available again in ~11 s at the tail |
| Queue throughput | one queue sustained 7,357 sends/s over a 300,000-send soak with exact delivery (up from a few hundred before the v0.4.1 producer rebuild) |
| Whole-fleet restart | nodes that restart together recover acknowledged writes: a restarting node serves follower fragments while recovering its predecessor, with checkpoints and a heartbeat so a retry skips finished work |
Telemetry
Telemetry is off by default and costs nothing until CELLD_OTEL is set; that one variable also picks the sink. CELLD_OTEL=1 writes Parquet files to the fleet bucket under the telemetry/ prefix, partitioned by node and hour, so a fleet with a bucket has observability with no other service; DuckDB queries the files directly. CELLD_OTEL=https://collector.internal:4318 (a full HTTP(S) base URL) sends the same data as OTLP/HTTP protobuf instead, with /v1/traces and /v1/logs appended; celld does not read OTEL_EXPORTER_OTLP_ENDPOINT. The OTLP exporter retries a batch up to five times with jittered backoff (honoring Retry-After, 30 s cap) on 408/429/502/503/504, drops it on a permanent refusal, and bounds its buffer at 8,192 events so a collector outage cannot grow memory; new telemetry is dropped and counted rather than blocking requests. Spans cover each request, cell event (fetch, alarm, RPC, WebSocket message), outbound fetch(), and cell start, plus every console.log as a log record joined to its trace. W3C traceparent is read and emitted (a malformed one starts a new trace), so traces join the systems in front of and behind celld.
CREATE VIEW traces AS SELECT * FROM
read_parquet('s3://YOUR-BUCKET/telemetry/traces/*/*/*/*/*/*.parquet');
SELECT name, duration_us, trace_id FROM traces
ORDER BY duration_us DESC LIMIT 20;Defaults: 5-minute / 5 MB flush (the byte threshold is an estimate, so a batch can slightly overshoot it), 30-day retention swept at startup and every six hours (none hands lifecycle to your own rules), OTEL_TRACES_SAMPLER for fractional sampling with a consistent per-trace decision across nodes. For a near-live view, set CELLD_OTEL_FLUSH_MS=5000 and run the one-hour compaction job on a maintenance node; a short flush without compaction makes DuckDB open thousands of small files. There are no metrics yet. That is a named gap, and the span durations cover most of what a metric would answer.
Security boundaries
celld is a beta, and it says so plainly: not safe for hostile multi-tenant use, and security fixes apply to the latest release only. The threat model is single-tenant with a trusted operator. Within that, the boundaries are explicit.
- Two listeners. The public Worker listener (
--listen) is the only thing a load balancer or firewall should expose; it reserves/.well-known/celld/healthand hands every other path to the Worker. The internal listener (--internal-listen) carries the peer protocol and the operator API and must stay on a private network or encrypted overlay. Most of the operator API is unauthenticated: anyone who can reach it can inspect state, evict cells, or stop the process. The one exception is the D1 route, which authenticates with the fleet secret because it runs caller-supplied SQL. - The bucket is the root of authority. It holds deployments, cell state, ownership records, node leases, the shared peer-authentication secret, large KV values, and R2 objects. Scope credentials to one bucket, use a prefix to share a bucket safely, and rotate on any suspicion of disclosure.
- One writer per cell. The epoch fences each cell; a node that loses its lease cannot modify current cell state. Peer requests authenticate with HMAC, body signature, clock limit, and replay protection, but the peer tunnel carries plain HTTP for cell fetch/RPC, so the private network or encrypted overlay is the confidentiality boundary, not the protocol.
- No TLS termination. Put public TLS in your ingress proxy and use WireGuard/Tailscale for the internal plane. By default celld ignores
X-Forwarded-Host/X-Forwarded-Proto; set--trust-forwarded-headersonly behind a trusted proxy that rewrites both, and it reads the last value so a direct client can't spoof it. - Application auth is yours. celld does not authenticate your application's users and does not enforce per-cell quotas; a defective cell can only touch its own database, but it can consume resources on its fleet node. Enforce request limits yourself via
CELLD_MAX_REQUEST_BODY_BYTESandCELLD_MAX_CELL_REQUESTS. - Egress is your network's job. Outbound TCP (
cloudflare:sockets) verifies TLS against a bundled Mozilla root store, but celld does not block the ports Cloudflare blocks. Containers get an nftables-fenced bridge (enableInternet: truereaches the Internet only, never the node, its peers, or private ranges), but that fence is experimental by the docs' own label, and macOScelld devkeeps container egress on with a warning. Internal host functions are not exposed to application code.
Tradeoffs and a verdict
| Cloudflare Durable Objects | celld | |
|---|---|---|
| Placement | vendor scheduler, opaque | your fleet; ownership = a lease in your bucket |
| State & durability | platform DO storage | SQLite replicated as LTX to your bucket, RPO=0 (bucket or fleet proof) |
| Failure domain | shared platform (tenant-coupled) | your nodes + your bucket provider |
| Data services | KV, Queues, D1, R2, Workflows on the platform | all five graded Yes, backed by your bucket / cells |
| Containers & sandboxes | Cloudflare Containers, Sandbox SDK | Experimental: a Docker/Podman engine per node, images in your bucket, an nftables-fenced bridge |
| Placement balance | platform scheduler | weight-proportional balancing of hibernated cells, pausable fleet-wide |
| Multi-tenancy | platform | not yet; one application per fleet |
| Deployments | instant, platform-managed | adopt in place without restart; module digests verified; one app per fleet |
| Observability | Cloudflare dashboard | Parquet in your bucket + DuckDB, or OTLP |
| Write latency | platform-managed | one bucket round trip (single node) or follower fsync (fleet) |
| Ingress / TLS | platform | your proxy; peers over Tailscale/WireGuard |
| Correctness claims | platform SLA | explicit + tested: TLA+, simulation, differential conformance |
What celld buys you is placement and blast radius you choose: no shared Durable Objects scheduler can couple your application to another customer's workload, and when a cell misbehaves the evidence is on your disk (ownership records, SQLite and LTX files, logs), answerable with sqlite3 and grep rather than a status page. It buys RPO=0 as a hard guarantee, which most self-hosted setups cannot claim, and it buys portability: the same Worker/DO code runs on workerd or on your fleet. KV, Queues, Workflows, and R2 close most of the "different primitive" gap; facets, HTMLRewriter, TCP sockets, and full Web Crypto close most of the runtime-API gap; and containers and sandboxes supervised by a Durable Object, on your own nodes, open a gap in celld's favor. The toolbox for building a full application on cells exists, not just the entity primitive.
What it costs is operational ownership. Balancing moves only hibernated cells and counts them by node weight. Updates are manual behind a current pointer. Cold restores touch object storage (paged, for a large cell). A durable write costs at least a follower fsync and eventually a bucket upload. The store must be on the qualified list: a wrong store fails the startup probe, which cannot be disabled, and a store that silently ignores conditional writes or ranged reads fails late and dangerously. The compatibility surface is still a strict subset of Cloudflare's. The platform services (AI, Vectorize, Hyperdrive, Browser Rendering, Email, Python) are absent, the Cache API is an always-miss, and KV/Queues/Workflows/R2 keep real gaps (no edge cache, 4-day fixed queue retention, multipart that cannot resume). Containers are experimental with a self-declared movable security boundary. The beta label changes none of the caveats: one application per fleet, no Windows, an operator API that can change between releases, full-stop upgrades (v0.4.1→v0.5.0 with no binary rollback, and v0.5.1→v0.6.0 under fleet durability), and WS transports that cannot move between owners. One semantic changed under running applications in v0.6.0 and deserves a check in any code that relies on it: a facet write does not commit atomically with the root object's transaction.
The verdict for practical use: the model is the best primitive in distributed systems right now, and celld's engineering discipline (TLA+, seeded simulation, differential conformance, live fault injection) is more rigorous than many production platforms. celld is a credible self-hosted platform for stateful serverless: the data-service toolbox exists, the fleet balances itself, large cells restore page by page, a cell can supervise a container, and celld itself has stopped calling it an alpha. It is ready for a team that wants DO semantics on its own infrastructure and can own the ops: agent fleets (with sandboxes), per-tenant sharding, real-time rooms, self-hosted stateful services. It is still not for hostile multi-tenancy, zero-ops, or anything that needs the full managed platform surface. Treat it as an unusually well-proven beta: pin releases, respect the upgrade cliffs (v0.5.0 is a full stop with a written procedure, and v0.6.0 is one under fleet durability), run under a supervisor, and re-check the compatibility page when you plan a release bump.
This chapter draws on: celld: documentation at v0.6.0 (9 entries) · celld: release notes (5 entries) · Cloudflare documentation (4 entries). The full entries are in the Bibliography.