Loomup Docs
Guide /docs/workload-and-recovery-envelope

Supported workload and recovery envelope

Loomup preserves one SQLite primary writer. “Scales without rewrites” means the resource/sync API and portable project artifacts survive topology changes; it does not mean an unbounded multi-writer database.

Supported topology

LayerSupported v1 envelope
PrimaryOne Loomup process and one writable SQLite file per project
ConcurrencySQLite WAL, bounded Rust connection pool, serialized SQLite writes
ReadsPrimary, dedicated same-file read pool, or round-robin read-only replicas through ReadReplicationProvider
RealtimeRegression gate at 1,000 simultaneous subscribed WebSockets per process
Sync mutation upload1–100 idempotent mutations per request
Event consumersIndependent ordered cursors, at-least-once delivery, 10,000-event maximum internal page
Gateway body64 MiB; generated object upload default is 50 MiB
Local dataOrdinary SQLite; default starter quotas (1 GiB DB, 100 MiB objects) and can be raised or set to zero

Operators should add a read topology when read latency or pool saturation rises while the writer remains healthy. Sustained write contention, recovery growth, or a workload requiring concurrent regional writers is a topology boundary, not a tuning flag; Loomup does not claim multi-writer support.

The current provider is ordinary read-only SQLite files, including files maintained by an external replication agent. The ReadReplicationProvider capability boundary lets a compatible libSQL/managed provider produce the same DbHandle set without changing application calls.

Recovery objectives

FailureRecovery pointExpected operation
Bad application writeAny retained journal sequenceRecord at(...), non-destructive clone, or stopped-project restore
Process crashLast committed SQLite transactionSupervisor restarts after bounded backoff; WAL recovery is SQLite-managed
Primary-file lossLatest verified snapshot available outside the failed volumeRestore/attach the verified SQLite artifact
Stale offline clientLatest authorized server statereset_required, bootstrap, then reapply pending idempotent mutations
Bad deployment/schemaPre-operation verified snapshot plus schema-version recordInspect plan, restore clone, then confirmed managed restore

Snapshots default to a daily check, a 10,000-new-event threshold, and seven retained files. These are policies, not an off-machine durability guarantee: a production operator must copy the snapshot directory to independent storage. RPO for total disk loss is the age of that independently stored verified snapshot. RTO depends on database size and storage bandwidth and is measured by the release recovery drill.

Compaction re-verifies the snapshot checksum and blocks on every active durable consumer and sync cursor. Sync cursors and abandoned consumers inactive for 30 days by default are pruned from the barrier; disabled realtime/push/webhook consumers are retired at process start.

Release evidence

  • tests/prd_bench.rs: CRUD, event append path, realtime latency, and 1,000 sockets.
  • tests/production_qualification.rs: exact release-binary mixed REST/sync load, enforced realtime delivery under load, in-flight process kill, zero acknowledged-write loss, kill-to-ready restart RTO, online backup, and copy-to-ready fresh-instance restore RTO.
  • tests/runtime_supervisor.rs: health wait, crash restart, HTTP/WebSocket gateway, logs, and graceful stop.
  • tests/project_portability.rs: export/run-local, attach, secret isolation, exact restore, and clone.
  • src/recovery.rs: checksummed snapshots, consumer barriers, stale-client reset, compaction, and recovery integrity.
  • tests/sync_v1.rs: authorization, idempotency, cursor reset, and conflicts.

Pull requests run Core CI (format, lint, full Rust tests including recovery/compatibility, and release-binary onboarding). Tag releases and production deployments additionally run .github/workflows/production-qualification.yml: it first rejects any source commit that is not in main history, then builds the release binary, binds evidence to its commit and SHA-256, seeds a 10,000-row sync fixture, applies concurrent REST/sync load, enforces realtime delivery during that load, kills the process with writes in flight, verifies zero acknowledged-write loss and SQLite integrity, enforces kill-to-ready restart RTO, copy-to-ready restore RTO, and write-latency thresholds, boots a fresh instance from an online backup, enforces the 1,000-WebSocket PRD benchmark, and performs a credentialed R2 put/head/get/delete canary. The qualified binary SHA is checked before packaging; a complete package manifest and archive checksum are checked after transfer and again in the immutable release directory. Qualification JSON, onboarding, recovery, and performance artifacts are retained for review.

Qualification evidence and deployment archives are attached to a prerelease named backend-qualification-<run-id>-<attempt>, targeted at the exact qualified source commit. Each attempt uses a new tag, and these prereleases are never marked latest. This handoff uses GitHub release assets rather than the Actions artifact quota. The deploy job receives the archive SHA-256 directly from the qualification job output and checks it before reading the downloaded archive; a replaced asset and its matching sidecar checksum cannot replace the qualified bytes. Failed qualification attempts retain any available evidence but cannot trigger deployment.

The workflow is also scheduled weekly, and every external GitHub Action used by the server workflows is pinned to a full commit SHA. Repository settings must mark the Core CI pull-request check as required; workflow YAML alone does not prevent merging around a failed or missing check. The production environment currently has no reviewer or deployment-branch protection rules, so the workflow itself restricts secret-backed qualification to commits already in main history. GitHub currently returns HTTP 403 for both branch protection and repository rulesets on this private repository's plan, so enforced merge gating requires a plan upgrade, making the repository public, or moving it under an organization policy that provides the feature. SDK matrices and npm publication run independently in bluppco/loomup-js. Limits should be raised only with new regression evidence, not marketing assumptions.