History, snapshots, and recovery
Every history-enabled resource writes one canonical event in the same SQLite transaction as its application row. That powers record history, point-in-time reads, offline sync, verified recovery points, and non-destructive recovery clones.
Record history
const history = await project.items.history(item.id);
const previous = await project.items.at(item.id, {
sequence: history.data[0].sequence - 1,
});
Current access rules are evaluated independently against each historical before/after state. Inaccessible states are redacted. The Studio's Records tab exposes unredacted history only to an authenticated administrator.
History is newest first, and limit counts visible events. Follow
meta.next_before_sequence as beforeSequence only when it is non-null. The
final page always has a null cursor, including exact multiples of the page size.
Hidden events do not terminate pagination or consume its visible-event limit.
Empty or exhausted history returns 200 with data: [] and a null cursor when
the current record or any retained historical state is readable under current
rules. Deleted records and former owners retain their authorized historical
states. Records with neither current nor historical access remain 404; a
caller-provided cursor does not establish access. Old servers can return 404
for an exhausted page, so clients may keep their existing compatibility handling.
Verified snapshots
New projects enable daily snapshots in loomup.toml:
[events]
snapshots_enabled = true
snapshot_directory = "./data/snapshots"
snapshot_interval_secs = 86400
snapshot_min_events = 10000
snapshot_retain = 7
sync_cursor_ttl_secs = 2592000 # inactive clients safely re-bootstrap after 30 days
consumer_cursor_ttl_secs = 2592000 # abandoned consumers no longer block compaction after 30 days
Create one immediately:
loomup snapshot --config loomup.toml
Loomup uses SQLite's online backup API, runs PRAGMA integrity_check, records the exact copied journal sequence, computes SHA-256, and only then registers the recovery point. The Studio Events view shows snapshots and checksums.
The full database copy has a 60-second cooperative deadline, including lock retries. Copying advances in 100-page steps without a delay after successful progress; only lock contention incurs a retry delay. WAL backups retain a read snapshot so concurrent writes cannot repeatedly restart the copy. Writers can continue, but WAL reclamation may be delayed until the snapshot is released. Non-WAL backups release source read locks between steps to allow writes. Individual filesystem operations cannot be interrupted by this deadline.
Schema migrations still require a verified snapshot before applying any planned
changes, including new tables and columns. A failed snapshot leaves those changes
unapplied. Legacy hosted schema execution retains a 300-second deadline, with project-only
worker termination and recovery as described below.
Backup logs report start, copy progress every five seconds, and completion or
failure, with page counts, page size, journal mode, elapsed time, contention
counts, and retry sleep time. These distinguish slow I/O from lock retries;
unknown page counts before source acquisition are reported as -1.
Hosted schema operation recovery
Hosted schema applies use a separate worker process. The default legacy path
retains its verified rollback snapshot, 300-second execution budget, and
60-second recovery attempt budget. Recovery-format support is deployed before
online preparation is enabled with LOOMUP_SCHEMA_ATOMIC_ENABLED=1 on the host.
With atomic schema execution enabled, reads, writes, and realtime remain live through snapshot copying, bounded record-read checks, SHA-256 generation, and migration/runtime rehearsal on a disposable copy. Preparation stops after five minutes without worker progress, with a two-hour absolute resource ceiling; productive work is no longer cut off at 15 minutes. Copying, hashing, and rehearsal SQL execution report progress. One preparation worker runs per host. The snapshot is private to its operation until the migration transaction registers it, so normal snapshot retention cannot delete it during preparation. Preparation failure stops the worker and fails the operation without fencing the serving project. Readiness checks and Studio schema-history reads inspect metadata without creating tables. Email provider setup also completes during preparation; its resulting configuration is staged for the same commit, with no provider calls during cutover.
Only the final cutover enters project maintenance. Its 30-minute budget includes
draining requests and background writers, applying the migration, and publishing
the runtime. Database changes, including bootstrap metadata and schema revisions,
share one SQLite transaction. Before commit, failure rolls back that transaction
and reopens the old runtime. After commit, a receipt in the same database is the
authoritative decision: recovery finishes installing the staged configuration and
publishing the new runtime. It never restores the earlier online snapshot over
writes acknowledged during preparation. Success is reported only after runtime
publication and fence cleanup. A durable commit may already exist while polling
still reports reloading or recovering; an operator-only recovery failure is
reported as failed without discarding that commit decision. A timeout starts recovery; it does not permit reopening while a
worker or filesystem operation has not stopped.
Atomic results include rollback_strategy: "sqlite_transaction" and a
recovery_point snapshot; rollback_snapshot is absent or null. Atomic execution
requires one hosted SQLite database in WAL mode. Named-database manifests are
rejected before maintenance. Schema/configuration changes invalidate prepared
work and require a fresh operation. Large table rebuilds can exceed the cutover
budget and require a separately planned migration.
Atomic preparation reads the first and last records plus 16 random rowid seeks per ordinary table. Tables without an accessible rowid use bounded reads from both ends of their storage order. These checks avoid full-table scans. The SHA-256 fingerprints the completed snapshot; sampled readability does not certify all pages or index cross-references. Full integrity scans remain in scheduled recovery snapshots and restore validation, outside the atomic deployment gate. Legacy schema workers retain full checks with up to 2 GiB of memory-mapped reads and a 64 MiB SQLite page cache. Migration rehearsal, structural-change validation, transaction rollback, and the 30-minute cutover deadline remain mandatory.
If a legacy operation stopped before its durable migration journal existed, no migration writes began. Recovery validates and reopens the unchanged runtime without repeating a full integrity scan. Unreadable or dangling journals are never treated as absent. Preparation and snapshot cleanup retain recovery files until journal absence is confirmed, including when a journal link is dangling. The database must still exist as a regular file: recovery never creates an empty replacement when it is missing or follows a database symlink. Missing runtime tables or columns require operator recovery instead of automatic retries.
Operation stages include preflight, snapshotting, verifying, rehearsing, draining,
applying, validating, reloading, and recovering. Operation responses may
include phase_deadline_at, database_committed when a commit has been observed,
and recovery metadata (attempt, next_retry_at, last_error_code, retryable).
Poll the operation's Location for its terminal state. Schema-key polling validates credentials without database writes, so a migration writer does not block authentication. Database availability failures return retryable 503 service_unavailable; invalid keys still return 401. A synchronous wait beyond
150 seconds returns 504 operation_timeout with Location and does not cancel
work. Recovered execution timeouts return 504 schema_migration_timeout.
Transient recovery failures retain the fence and retry after approximately 5, 15, 30, then 60 seconds, with jitter and a 60-second cap on the base delay. Each attempt has a separate 60-second budget. Retry state survives restart; attempts never overlap live workers or supervisors, including final retry bookkeeping. Failed attempts are terminal catalog operations so sleeping retries do not block deployment drain. Successful forward recovery completes the original committed operation; rolled-back operations stay failed even after availability is restored.
Corrupt databases, invalid journals, and inconsistent commit receipts report
schema_recovery_unsafe and require an operator. Legacy quarantines without a
journal can automatically validate and reopen; other legacy quarantines retain
the operator recovery procedure. Automatic retries run only in the serving
generation, and pause through deployment. After addressing an operator-only
failure, request recovery on the active slot:
curl --fail-with-body -X POST \
-H "x-loomup-management-token: $LOOMUP_MANAGEMENT_TOKEN" \
"http://127.0.0.1:9798/__loomup/manage/projects/$PROJECT_ID/schema-recover"
Use port 9799 when green is active. This route is blocked by the public proxy,
requires the management token, and accepts no replacement paths/schema. It
returns 202 with a recovery operation and Location; poll with authorized
platform/project credentials. Concurrent work returns 409. Successful recovery
clears the protective marker itself. Never delete journals or quarantine markers
to force traffic back on.
Live worker/supervisor locks prevent another catalog opener or generation from recovering an operation that is still executing. Corrupt project journals degrade that project instead of preventing the shared server from starting.
Recovery clone
Always preview first:
loomup recover --at-sequence 1842 --output ./data/recovered.sqlite --dry-run
Then create the clone:
loomup recover --at-sequence 1842 --output ./data/recovered.sqlite
# Or use an inclusive Unix timestamp:
loomup recover --at-timestamp 1783771200 --output ./data/recovered.sqlite
The source is never modified. Loomup copies it, reverses later application events inside one transaction, trims future internal receipts/cursors in the clone, reinstalls capture triggers, records an admin audit entry, verifies database integrity, and atomically publishes the output file.
Recovery refuses to claim correctness when an application table lacks durable capture, the target is outside retained history, the schema cannot be replayed safely, or a row invariant fails.
Every applied application manifest is hashed into _lb_schema_versions with its journal boundary and plan. Additive schema changes keep the latest compatible shape while historical row data is reconstructed inside it; unsupported manifest/event encodings fail before a recovery file is published. Fixture tests cover replay across an added field.
Managed projects expose the same machinery with automatic rollback points:
loomup platform backup-project --id "$PROJECT_ID"
loomup platform clone-project --id "$PROJECT_ID" --name "Investigation" --at-sequence 1842
loomup platform restore-project --id "$PROJECT_ID" --at-sequence 1842 --yes
Restore automatically places only the target project in maintenance, drains its in-flight requests, disconnects its WebSockets, and rebuilds its hosted context.
Conservative compaction
Compaction is explicit:
loomup compact-events \
--snapshot-id <verified-id> \
--config loomup.toml \
--yes
The snapshot file and checksum are re-verified. Compaction stops if any independent event consumer or active registered sync client cursor is below the recovery point. A sync client remains active for events.sync_cursor_ttl_secs; older registrations are pruned during compaction and receive reset_required plus a safe bootstrap if they return. Event consumers idle longer than events.consumer_cursor_ttl_secs (and without a live lease) are likewise pruned from the barrier. Disabling realtime, push, or a webhook retires that consumer at process start so it cannot block compaction forever; operators can also retire a consumer from Studio or DELETE /admin/api/events/consumers/{name}. The compaction report states how many stale sync clients and consumers were reset.
R2 archives and local cleanup
Hosted recovery can keep one local snapshot and the latest seven verified recovery
bundles in the existing R2 bucket. The latest local copy is also archived. Archives
use _loomup_recovery/v1/{project_id}/{snapshot_id}/, outside application asset
namespaces. Application storage APIs and CDN paths cannot address that namespace.
Set host environment variables LOOMUP_SNAPSHOT_ARCHIVE_ENABLED=true to enable
uploads and archived reads, and LOOMUP_SNAPSHOT_EVICTION_ENABLED=true only after
all serving binaries support archived snapshots and a restore drill has passed.
The archive client uses LOOMUP_R2_BUCKET, LOOMUP_R2_ENDPOINT (or
LOOMUP_R2_ACCOUNT_ID), LOOMUP_R2_ACCESS_KEY_ID, and LOOMUP_R2_SECRET_ACCESS_KEY.
Self-hosted installations without archive configuration keep local retention.
The independent catalog is stored alongside the database in app.recovery/ and
survives replacement of app.sqlite. Snapshot IDs are stable and retention uses
creation order, not journal sequence. Uploads stream multipart data, then read it
back and verify SHA-256 and size before any local eviction. The manifest is also
read back and checked. Failed uploads retain their local files for retry.
Failures during upload initialization also record retry state. Local snapshots
remain restorable without R2 credentials; archived downloads have a deadline,
including time spent waiting for another download to release its lock.
Maintenance runs at startup and every five minutes, including when periodic snapshot creation is disabled. It serializes background archive transfers across hosted projects. Active operations and readers pin their recovery files, so the one-local/seven-archive policy may temporarily retain more files. R2 expiry is managed by the server, not a blanket object-expiration lifecycle. Configure a prefix-scoped incomplete-multipart-upload abort rule of one day in R2, preserving existing bucket lifecycle rules; failed uploads also attempt an immediate abort.
With bucket-administration credentials in the LOOMUP_R2_* environment, run
python3 deploy/loomup-configure-recovery-r2 to check the rule, then add --apply
to configure and verify it. The script preserves unrelated rules and rejects
expiration rules that overlap recovery archives. Object-only credentials may
return AccessDenied; keep eviction disabled until the lifecycle configuration
and restore drill have been verified.
Snapshot operations remove temporary SQLite files and WAL/SHM/journal supporting files on completion or failure. Startup/periodic cleanup removes recognized abandoned temporary snapshots older than one hour only when they have no active lock or recovery reference. Schema preparation keeps committed recovery points. Cleanup checks both the live registry and independent recovery catalog while holding the artifact lock. Schema preparation also preserves catalog-owned snapshots and their SQLite sidecars after a restore changes the live registry. Local retention counts files that are actually present, so a missing newest file does not cause the last available local recovery point to be evicted.
Restore an archived snapshot
Project managers can open Recovery snapshots in Studio, select a recovery point, and confirm the project name. Progress is available through the durable project operation endpoint. A restore replaces the database and saved schema configuration, including schema/manifest files and generated table rules, while preserving current infrastructure configuration and secrets. Object storage bytes and server binaries are not rolled back.
The server CLI supports the same snapshot selection while the project is stopped:
loomup platform restore-project --id "$PROJECT_ID" --snapshot-id "$SNAPSHOT_ID" --yes
GET /platform/api/projects/{id}/snapshots lists recovery points and storage
availability. POST /platform/api/projects/{id}/restore accepts
{"snapshot_id":"<uuid>","confirm":"<project-id>"} and returns 202 with a
Location operation URL. Choose exactly one of snapshot_id or target_sequence;
existing sequence restores retain their current response contract. Snapshot-ID
restores download and verify before maintenance, create a recovery point of the
current state, replace database/configuration under a durable journal, and validate
the replacement runtime. Failure rolls back; unresolved recovery keeps the project
blocked. Cancellation is honored before replacement begins.
Rollback can recover from invalid restored schema files without loading those
files first. Offline CLI restores hold project and deployment ownership through
the final durable restore decision.
Legacy snapshots without saved schema configuration are not assigned guessed configuration. They can restore only when their SQL schema matches the current project. Archived history and compaction fail explicitly when recovery data is unavailable or corrupt rather than silently returning incomplete historical data.
Rollout
Deploy archive-aware readers everywhere with eviction off. Import existing recovery points and allow uploads, verify an archived restore into an isolated project, then enable eviction. Keep the catalog and archive settings when replacing a serving binary; never roll back to a binary that requires local-only snapshots after eviction. Previously deleted recovery points cannot be backfilled.