Shared hosted backend
Loomup's hosted mode runs one common HTTP API for every managed project while
preserving one SQLite database and one data directory per project. It replaces
the older model where the Platform process spawned a child loomup serve
process and loopback port for each project.
Request topology
Cloudflare -> Caddy -> active Loomup host generation
|-- root routes -> default project context
| /mnt/loomup_volume/loomup/data/app.sqlite
|-- /platform -> control-plane catalog
|-- /p/<id> -> managed project context
platform/projects/<id>/data/app.sqlite
Each hosted context owns its config, SQLite pools, auth and row rules, storage
root, usage meter, realtime hub, and Axum data-plane router. The outer host
strips /p/<id> and dispatches the request directly to that router. There is no
internal HTTP hop, project PID, watchdog, or assigned serve port.
Projects do not enter an externally visible idle, paused, or cold state. Every
catalog project remains routable for the lifetime of the host generation. The
SQLite pool is deliberately lazy (min_idle = 0): it opens connections on
checkout, with a default maximum of 8 connections per database. Connections
become eligible for cleanup after 60 seconds idle and close on a subsequent
maintenance pass. Releasing a connection is an internal resource decision and
never changes project status.
The database file remains on the attached block volume and can be reopened in
milliseconds. All project pools also share one bounded connection-maintenance
scheduler; r2d2 is not allowed to create its default three maintenance threads
per project.
Realtime-only projects also release their durable realtime lease while they have no connected WebSocket clients, so an unused project does not poll SQLite every 100 ms. A new client reacquires the generation-specific lease and resumes from its persisted cursor. Push continues independently when enabled.
The default project is optional in generic installations. Production supplies
--default-config /etc/loomup/loomup.toml, preserving /api, /auth,
/realtime, /storage, /health, /ready, and other existing root routes.
Self-hosted loomup serve remains a supported single-project command.
Availability and maintenance
All catalog projects load when platform serve starts. A project that cannot
load is recorded as degraded; /p/<id>/* returns 503 project_unavailable
while Platform, the default project, and other managed projects remain live.
Project creation returns success only after the new context is loaded.
Hosted projects do not have Start, Stop, Restart, PID, port, or runtime-log
controls. Project details include a read-only hosting object with state,
health, generation, load time, and a sanitized error.
Schema publication, restore, secret rotation, deletion, attach/promotion, and
other exclusive changes enter project-scoped maintenance. New requests to only
that project receive 503 project_maintenance; its in-flight HTTP requests
drain, its WebSockets are disconnected, the operation and rollback checks run,
and its router and workers are rebuilt. Other projects are never paused.
Maintenance responses include Retry-After and active operation metadata.
Internal management status retains projects in maintenance after their runtime
is unloaded, so quarantined projects remain visible to deployment checks.
Schema applies (including Studio publication) run in a supervised child of the
same release binary. Two operations can execute per host, with one active
operation per project. Only queued operations can be cancelled.
The default legacy path has a 300-second execution budget. With
LOOMUP_SCHEMA_ATOMIC_ENABLED=1, one preparation worker per host performs the
backup, SHA-256 generation, sampled record reads, and rehearsal while project
traffic remains live. Preparation has a five-minute no-progress timeout and a
two-hour absolute ceiling; full integrity scans remain in background recovery
snapshots. Cutover then enters maintenance for a
30-minute budget and applies all database changes in one transaction. Before
commit, recovery rolls back; after commit, it finishes runtime publication.
The earlier recovery point is never restored automatically over newer writes.
Recovery attempts have separate 60-second budgets and never overlap workers that have not stopped. Transient failures retry with persisted backoff. Corrupt or inconsistent recovery evidence requires an operator. Maintenance is released only after runtime validation and publication. Operation polling and catalog status remain accessible while the project is fenced. The shared host and other projects remain available. See recovery for failure codes, retry metadata, and the internal operator endpoint.
Schema applies and recovery are serialized with blue/green deployment. Deployment
closes admission, waits for active schema supervisors, and keeps the fence until
the old slot has stopped. New applies receive 409 deployment_in_progress with
Retry-After. A coordinator crash leaves the deployment fence in place for
operator reconciliation.
Realtime generations
Every host generation uses a distinct durable consumer name in each tenant
database: realtime:<generation-id>. During a deployment, clients already
connected to the old generation continue consuming through its hub while new
connections use the new generation. After the candidate passes live-load and
public verification and the active slot is committed, the old generation sends
1012 service restart; clients reconnect, re-subscribe, and resynchronize
against the candidate. Both cursors advance independently until that handoff,
so each connection receives each event once.
Push, projections, webhooks, snapshots, scheduled operations, and outbound usage export remain singleton work. The workflow transfers their ownership to the settled candidate before admitting it to public traffic. The old generation retains only its generation-specific realtime consumer until its WebSockets close. A clean drain retires those cursors; crash-expired cursors are covered by journal lease pruning.
This work is not fire-and-forget. Mutations commit before the API responds; asynchronous delivery is represented by durable project-local outboxes, consumer cursors, leases, retry counters, and dead letters in SQLite. Process failure therefore leaves work claimable after restart. Only disposable telemetry may use best-effort execution.
The per-project worker handle is an isolation and lifecycle boundary, not a dedicated process or OS thread. All async jobs run on the host's shared Tokio executor and all SQLite pools use the shared bounded maintenance scheduler. The handle supplies the project database, cursor names, and cancellation scope, so maintenance on project A cannot stop delivery for project B. This keeps the executor common while durable queue state remains correctly project-scoped. Worker shutdown is joined before a project context is rebuilt or promoted, so the replacement cannot race the previous task for the same durable lease.
R2 is suitable for uploaded objects, verified backups, and cold archives. It is not mounted or treated as a live SQLite filesystem. Active databases and their WAL/SHM sidecars stay on the DigitalOcean block volume; an R2 copy must be restored locally and verified before it can be opened.
DigitalOcean blue/green deployment
Production uses two permanent systemd slots:
| Slot | Unit | Loopback port |
|---|---|---|
| blue | loomup-host@blue.service | 9798 |
| green | loomup-host@green.service | 9799 |
The active slot is stored at
/mnt/loomup_volume/loomup/platform/active-slot. Slot release symlinks live at
/opt/loomup/slots/{blue,green}; each running process therefore keeps the
exact binary that started it even after /opt/loomup/current changes.
Caddy permanently lists both upstreams and checks /__loomup/slot-health.
For each deployment the workflow promotes and soaks the candidate off-traffic,
then renders it first and the old slot second and reloads Caddy without closing
existing WebSocket tunnels. The settled candidate receives full live load while
the old slot remains automatic health fallback. Only after the rollback window
closes does the old slot signal its sockets to reconnect. /__loomup/* is
blocked at the public proxy.
Management actions additionally require LOOMUP_MANAGEMENT_TOKEN from
/etc/loomup/loomup.env.
The GitHub workflow performs the following sequence:
- Refuse to overwrite an inactive slot that is still draining sockets.
- Back up the catalog and default SQLite database, apply the release's default project migrations, install the release, and start the inactive slot warm.
- Verify the default project, Platform, management status, loaded contexts, and singleton-worker prerequisites such as webhook secrets without acquiring the old generation's durable leases.
- Keep the warm candidate ineligible and give its systemd unit low CPU and I/O weight. Warm loading opens the already-initialized tenant databases without repeating schema bootstraps or starting database-polling workers, so it does not lock the serving generation's shared SQLite files. Require 10 consecutive direct and public health samples. Caddy's constant-time slot probe has a 10-second scheduling cushion, so a brief runtime stall cannot remove the only eligible slot. The host uses 16 Tokio workers so synchronous SQLite waits cannot occupy the entire request executor. Deployment probes gate on HTTP success within a bounded 15-second deadline; latency benchmarking runs outside the release gate and requires repeated samples.
- Relinquish singleton workers on the old slot, restore the candidate’s normal
CPU and I/O weight, promote the still-ineligible candidate, and require 20
consecutive direct readiness samples while worker
startup settles off-traffic. Promotion creates the additive bulk-inbox
selection and idempotency tables in each project's notification database
before starting workers; failure prevents promotion. Warm loading remains
read-only and promotion does not rerun application migrations or rebuild CDC
triggers. Only then put the ineligible candidate first in
Caddy, restore normal resource weight, and make it eligible. Require 30
consecutive target and public readiness samples, demote the old slot, and
require another 10 consecutive samples. These are the routine deployment
windows; dispatching with
full_qualification=trueretains 30/60/60/30 samples. Direct activation defaults to the full windows. Every failed sample restarts its window, and both profiles retain the same rollback checks. - Persist the active slot and tell the old process to drain. The old generation
immediately sends close code
1012with reasonservice restart; clients reconnect to the already-active candidate with jitter. The supervised deployment allows 30 seconds for close handshakes and normal retirement, requests a graceful systemd stop, and sends SIGKILL if the 90-second stop deadline expires. When the old binary predates this close capability, the helper uses the legacy eight-minute drain for first-release compatibility. It resets any failed unit state so a later deployment can reclaim the slot. Drain is a terminal generation transition: the same process cannot be promoted or made traffic-eligible again.
Worker ownership transfer, promotion, and old-slot retirement run inside the guarded deployment. Schema admission stays closed throughout retirement. On failure, rollback stops the candidate and re-promotes the old slot before restoring its persisted eligibility.
If public verification fails before drain, the workflow restores old slot eligibility and stops the new slot. A later deployment reclaims an inactive slot that explicitly reports itself ineligible and draining, using the same bounded graceful-stop and force-stop sequence. It also reclaims the slot when systemd says it is active but its local management endpoint cannot be reached. An active inactive-slot process that is reachable and not safely draining still fails closed for operator review.
loomup-deployment-monitor.timer checks the deployment invariants every minute.
It emits a latched loomup.deployment.alert journald event when both slots have
been active for more than ten minutes, a promotion service runs for more than
ten minutes, public /ready is non-200, or an active slot's loopback port is
closed. A matching resolved event clears each latch. The production OpenTelemetry
collector includes these events in its journald pipeline.
The first migration is a bootstrap exception: it starts the shared blue slot,
stops the legacy loomup and loomup-platform services, installs the permanent
two-upstream Caddy configuration, and reloads Caddy once. Existing WebSockets
may reconnect during this one-time cutover. Later backend deployments reload
Caddy to put that deployment's candidate first while preserving upgraded
connections through verification, then close the retiring generation's sockets
with 1012 service restart after committing the new slot.
Operator checks
cat /mnt/loomup_volume/loomup/platform/active-slot
systemctl status 'loomup-host@*.service'
systemctl status loomup-deployment-monitor.timer
curl http://127.0.0.1:9798/__loomup/slot-health
curl http://127.0.0.1:9799/__loomup/slot-health
findmnt -T /mnt/loomup_volume/loomup/data/app.sqlite
findmnt -T /mnt/loomup_volume/loomup/platform/control.sqlite
Do not restart the active unit for an ordinary deployment. Use the deployment workflow so context warming, health checks, rollback, and WebSocket drain semantics are preserved.
/__loomup/slot-health is deliberately constant-time: it reads only atomic
slot eligibility/drain state and the immutable generation id. Use the
authenticated /__loomup/manage/status endpoint for per-project diagnostics.
Release and host-log retention
After a successful activation and public verification, deployment installs the
local log policy described in observability.md, then runs
loomup-prune-releases --keep 5 --apply. It keeps the five newest release
directories plus /opt/loomup/current, both blue/green slot targets, and releases
whose binaries are still running. A draining generation remains protected.
Only real directories named with a 40-character lowercase Git SHA are eligible;
symlinks and unrelated files are ignored. Missing current or dangling current/slot
references and unreadable process information abort cleanup. Database files,
recovery snapshots, and predeployment database backups are outside its scope.
Preview the exact candidates without deleting anything:
sudo /opt/loomup/current/loomup-prune-releases --keep 5
The helper ships with new releases. Older release artifacts without it can still be deployed using their original workflow, but will not enforce this cleanup policy. Run cleanup only when no separate manual deployment is changing release pointers; GitHub already serializes production deployments.