Production observability
Loomup keeps billing metering and diagnostic telemetry separate. The existing SQLite usage tables and idempotent outbox are authoritative for billing. Logs, host metrics, and sampled traces are best-effort diagnostic signals exported to Managed ClickStack through a loopback-only OpenTelemetry Collector.
Runtime configuration
Add these non-secret values to /etc/loomup/loomup.env:
LOOMUP_LOG_FORMAT=json
LOOMUP_OBSERVABILITY_URL=https://your-managed-clickstack.example
OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:4317
OTEL_SERVICE_NAME=loomup
OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=production
The application never receives the Managed ClickStack credential. Put the
remote endpoint and authorization header in /etc/loomup/otel.env, readable by
only root and the otelcol group. Install deploy/otel-collector.yaml as
/etc/loomup/otel-collector.yaml and the systemd unit as
/etc/systemd/system/loomup-otel-collector.service.
Before starting the service, validate the exact collector version and config:
sudo -u otelcol /usr/local/bin/otelcol-contrib validate --config=/etc/loomup/otel-collector.yaml
sudo systemctl daemon-reload
sudo systemctl enable --now loomup-otel-collector.service
The collector is independent of the blue/green Loomup slots and must not be restarted during an ordinary application promotion. Its persistent retry queue lives on the root disk rather than the SQLite block volume. Batches are capped at 1,024 records and the persistent queue at 256 batches, so a prolonged remote outage cannot create an unbounded collector backlog; once full, diagnostic telemetry is dropped while requests and authoritative billing continue.
Operator bootstrap
Grant the initial browser operator locally on the host:
loomup platform grant-operator \
--email operator@example.com \
--catalog /mnt/loomup_volume/loomup/platform/control.sqlite \
--projects-root /mnt/loomup_volume/loomup/platform/projects
The account then receives an Observability navigation item at
/platform/observability. Detailed investigation remains in ClickStack; the
native page uses only local rolling metrics and durable usage, so it remains
available during a remote telemetry outage.
Request completion logs are emitted once per request at INFO, including hosted
project requests. Successful /health, /ready, and internal slot-health
probes are omitted; unsuccessful probes remain logged. Request lifecycle
diagnostics are DEBUG-level. Trace export and billing metering are unchanged.
Categorizing the usage error total
The usage meter's http_errors counts all HTTP statuses >= 400, including
authentication failures, denied access, missing resources, sync resets, quota
rejections, and server failures. It is not a server outage counter. A sync reset
can be a required permission refresh; verify the following bootstrap using the
same project and client_hash before assessing its user impact.
Use the offline report to group retained logs by status, normalized route, and correlated sync reason:
python3 scripts/report-http-errors.py --project PROJECT_UUID \
--since 2026-09-01 --until 2026-10-01 journal.jsonl syslog.7.gz
Input can be journalctl -o cat output, JSON embedded in archived syslog lines,
gzip files, or standard input (-). Overlapping archives are deduplicated by
request ID within the selected project. --json emits a structured report;
--meter month.json compares against that project's monthly usage report and
includes dates with no retained logs. Ensure both inputs refer to the same
project and cutoff time. Negative differences indicate uncategorized metered
errors; positive differences can indicate a later log cutoff, a pending meter
flush, or inconsistent inputs. Origin request totals exclude CDN receipts.
sync_recovery identifies a correlated reset diagnostic, not proof of successful
recovery. A sync 409 without a diagnostic remains
sync_conflict_reason_unavailable; do not assume every sync conflict is a reset.
HTTP 507 is reported separately as quota_exceeded; other 5xx responses are
server_failure. This report does not change billing totals or API behavior.
Retained timestamp bounds do not prove continuous coverage. Journald's size cap can evict logs much sooner than seven days at high traffic volume. The remote collector and backend must actually be configured and running to provide the longer retention described below; installing a service file alone does not export logs. Do not extrapolate categories from the retained subset to a whole month.
Sync recovery diagnostics
Each reset emits one INFO sync.reset_required event in addition to the ordinary
HTTP completion. Its reason is cursor_expired,
authorization_dependency_changed, schema_or_authorization_changed, or
bootstrap_required. Events include project ID, request ID, normalized route,
status, stable error code, and applicable cursor bounds/dependency/resource names.
changed_columns contains changed field names when valid before/after payloads
are available, and is null otherwise. Resource and changed-column lists are
JSON-encoded log fields and are limited to 32 names, each at most 128 characters.
Successful snapshots emit sync.bootstrap.completed. Both events contain the
same client_hash for a validated client ID within a project: SHA-256 of the
project ID, a zero byte, and client ID. Match project and hash to investigate
recovery rather than infer it from neighboring requests. Invalid client IDs have
no hash. Protocol errors emit sync.protocol_mismatch with requested/supported
versions instead of a reset event.
These diagnostic events do not change usage totals. They contain no raw client IDs, credentials, row contents, permission rules, or query strings. Correlation hashes are log fields, not metric labels.
Local disk retention
Restricted SSH log access
Production has a dedicated loomup-logs SSH account with a fixed project allowlist:
approve (961cdf44-0374-4e83-bd12-4917031f54e0, the default) and
loomlytics (4991a905-8998-40bb-8c46-252126a315a5). It exposes structured HTTP completion
and sync reset/bootstrap/protocol diagnostics from the blue/green journals,
plus fixed aggregate usage reports through the usage command.
Messages, request bodies, SQL, span attributes, and query strings are omitted.
It does not expose archived syslog files, database files, or arbitrary queries.
The local private key is ~/.ssh/loomup_logs (mode 0600); its public key is
~/.ssh/loomup_logs.pub. The private key is passphrase-free for automated log
collection and is not installed on the server or stored in this repository.
# Recent errors, newest first. Each reset can have both a diagnostic and an
# HTTP completion record; count completion events for request totals.
ssh -i ~/.ssh/loomup_logs -o IdentitiesOnly=yes loomup-logs@64.227.150.72 \
'errors --since 2026-09-15 --limit 200'
# Recent HTTP/sync diagnostics, or live following.
ssh -i ~/.ssh/loomup_logs -o IdentitiesOnly=yes loomup-logs@64.227.150.72 \
'logs --limit 100'
ssh -i ~/.ssh/loomup_logs -o IdentitiesOnly=yes loomup-logs@64.227.150.72 follow
# Daily histogram for this month, today's hourly report, or JSON for a month.
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'usage month'
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'usage today'
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'usage month 2026-09 --json'
# The same project flag works for usage, logs, errors, and follow.
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'usage month --project loomlytics'
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'errors --project loomlytics --limit 100'
For log commands, dates are UTC; --since is inclusive and --until is exclusive. The default
interval is the last 24 hours; queries may span at most 31 days. --limit defaults
to 200 records and accepts 1–5000. Historical sessions stop after 60 seconds;
follow sessions show the latest 20 journal matches, then live records, and stop
after ten minutes or the output limit. Reconnect to continue following. Log
retention still determines what is available. For categorization, pipe output to
scripts/report-http-errors.py --project PROJECT_UUID -; a bounded export is
not a full daily meter.
The usage command accepts only today, a UTC date (YYYY-MM-DD), or month
with an optional YYYY-MM, plus --json and --project approve|loomlytics.
With no period it reports today. Each command defaults to approve when the
project flag is omitted. The server resolves only these two aliases to fixed
project IDs; the catalog path is fixed. Callers cannot supply source code, SQL,
unlisted projects, file paths, or a different executable. Both authorized keys
on this account have the same access. Reports have a 55-second collection time
limit; timed-out workers and their journal reader are terminated together.
The approved collection/rendering logic comes from the workspace's
server-monitoring/usage.py, installed as root-owned
/usr/local/libexec/loomup-usage.py (0644). It uses read-only SQLite connections
and PRAGMA query_only=ON. Monthly reports read durable project usage aggregates;
hourly reports combine streamed request completions with aggregate CDN receipts
and retain the warning when they differ from the durable daily meter.
The local wrapper supports the restricted connection without sending Python source to the server:
./server-monitoring/usage.sh --restricted --identity ~/.ssh/loomup_logs month
./server-monitoring/usage.sh --restricted --identity /path/to/bot-private-key today --json
./server-monitoring/usage.sh --restricted --identity ~/.ssh/loomup_logs --project loomlytics month
./server-monitoring/usage-loomlytics.sh --restricted --identity ~/.ssh/loomup_logs month
--restricted defaults to loomup-logs@64.227.150.72. The local wrapper accepts
either alias or its project UUID and sends the approved alias to the server.
The existing administrative
transport remains available without that option. After modifying collection
logic, an administrator must install the reviewed usage.py snapshot to the
fixed server path; the restricted account cannot update its own executable.
Installed controls (all root-owned):
deploy/loomup-read-logs→/usr/local/libexec/loomup-read-logs(0755).deploy/sshd-loomup-logs.conf→/etc/ssh/sshd_config.d/40-loomup-logs.conf(0644).deploy/sudoers-loomup-logs→/etc/sudoers.d/loomup-logs(0440)./etc/ssh/authorized_keys/loomup-logs(0644), with bothrestrictand the forced reader command on the public-key line./var/empty/loomup-logsis the root-owned account home. The account belongs only to its own group, notadm,systemd-journal, orsudo.
The forced command may invoke only the reader through sudo with no arguments.
The reader strictly parses SSH_ORIGINAL_COMMAND without executing it as shell
code, fixes the service/project selection, disables paging, and starts journalctl
with a clean environment. SSH disables terminal allocation, user rc files,
forwarding, tunnels, and password authentication for this account. The forced
command also rejects SFTP and arbitrary commands. The ordinary deploy account
is separate.
To revoke new connections immediately, use the administrative account to empty
/etc/ssh/authorized_keys/loomup-logs; no SSH reload is needed for key-file
changes. To also terminate existing log sessions, run
sudo loginctl terminate-user loomup-logs. Keep an administrative session open
when changing SSH configuration, validate with sudo sshd -t, and validate
sudoers changes with sudo visudo -c before reloading SSH.
Validation on September 15, 2026: positive log/error/follow reads succeeded with the new key and SSH-agent use disabled. Arbitrary commands, environment-file reads, PTY allocation, SFTP, both forwarding directions, and general sudo were rejected. Usage reports were verified in histogram and JSON formats; the closed September daily totals matched the existing meter. Live attempts to select an unlisted project or another catalog were rejected. Six focused reader tests and five usage tests cover command injection, fixed selection, project isolation, output filtering, malformed records, restricted transport, read-only queries, and metering exclusions.
Host retention settings
deploy/loomup-configure-logging installs the production host policy:
- Journald: 512 MiB maximum use, 64 MiB journal files, seven-day maximum age, and 2 GiB free-space reserve. Runtime journals are limited to 64 MiB. Active files can temporarily exceed the target until rotation.
- Rsyslog: discard the duplicate
loomup-hostandloomupmessages after journald has stored them. Other services retain their existing syslog path. - General syslog: rotate daily, retain seven archives, compress immediately, and rotate above 100 MiB when logrotate runs. This is a rotation threshold, not a hard cap between timer runs. Auth and other system logs keep their existing four-week policy.
The deployment workflow applies the policy after successful public verification.
The installer validates rsyslog and logrotate, restarts only the logging daemons
when their files change, and restores the previous configuration on failure.
Original configuration is retained under /var/backups/loomup-logging.
It first compresses any legacy uncompressed syslog.1 to preserve that
history when switching away from delayed compression. It does not force
rotation or vacuum existing logs; normal rotation retires older history. The collector continues reading journald.
Data and retention policy
- Request/application logs: 14 days.
- Traces: 7 days; retain 5% normally, all errors, and all requests over one second.
- Metrics: 30 days at the backend's normal resolution.
- Never export authorization, cookies, request/response bodies, query strings, SQL values, email addresses, or client IP addresses.
Configure these TTLs in Managed ClickStack. The repository collector config implements trace sampling; the backend remains responsible for retention.
Alerts
Create a DigitalOcean Uptime HTTPS check for the public /ready endpoint and
alert on two minutes of downtime, five minutes above one-second latency, and
certificate expiry within 14 days. In ClickStack, begin with alerts for missing
host telemetry, sustained 5xx rate, p95 latency, CPU/memory pressure, low disk,
usage-outbox delay, and collector export failures. Tune thresholds after the
first week of production baselining.
Failure checks
Stopping the collector or blocking its outbound connection must never fail an application request or billing write. Confirm local logs with:
journalctl -f -u loomup-host@blue.service -u loomup-host@green.service
journalctl -f -u loomup-otel-collector.service
Shadow request pricing is intentionally non-billable. Observe it for at least 30 days before creating a versioned customer plan.
Local SQLite query diagnostics
SQL diagnostics work without ClickHouse, ClickStack or an OpenTelemetry collector.
They use a separate diagnostics.sqlite next to the platform catalog (hosted) or
application database (standalone), and do not modify billing counters or quotas.
Collection is disabled unless LOOMUP_SQL_DIAGNOSTICS_ENABLED=true.
| Setting | Default | Purpose |
|---|---|---|
LOOMUP_SQL_DIAGNOSTICS_ENABLED | false | Enable local SQL collection on process startup |
LOOMUP_SQL_DIAGNOSTICS_PATH | adjacent diagnostics.sqlite | Override the persistent local file |
LOOMUP_SQL_SLOW_MS | 100 | Retain events at or above this duration |
LOOMUP_SQL_SAMPLE_PERCENT | 1 | Sample percentage of faster executions (0–100) |
LOOMUP_SQL_MAX_BYTES | 536870912 | Main-file page budget; retention also manages WAL growth |
Restart with collection disabled to stop recording. Existing configurations stay
disabled. Begin with a canary, compare enabled/disabled throughput on the same
workload, and enable broadly only if throughput regresses by no more than 5%.
The checked-in Cargo configuration enables SQLite's native SQL normalizer; builds
must preserve LIBSQLITE3_FLAGS=SQLITE_ENABLE_NORMALIZE.
The collector aggregates observed statements into minute buckets, flushing every five seconds and retaining up to 24 hours. All slow executions and sampled fast executions are retained subject to bounded buffers and storage. The dashboard reports dropped diagnostics, stale persistence and capacity eviction. A crash can lose unflushed data; these are diagnostic counts, never billing records. Background tasks, named databases and read pools are distinguished. Pool acquisition wait is measured once per checkout, separately from statement elapsed time. Statement timing excludes preparation and pool wait, but can include lock waiting and row consumption. Counts and latency histograms describe observed executions; they are not an exact audit trail. p95 is a histogram approximation.
SQLite's native normalizer removes literal values and comments while retaining query structure and identifiers. Bound values, results, authorization headers, request bodies and URL query strings are not stored. Oversized SQL is omitted. Diagnostics queries and plan inspection are not collected recursively.
Views and permissions
- Operators: Observability → SQL query performance, at
/platform/observability/sql. Operators can filter all projects, including the operator-only__system__catalog. - Project Studio: Query performance, at
/platform/projects/{id}/sql. The existing manager rule applies: workspace owner or project creator. - Standalone Admin: Query performance, requiring the project admin role.
The dashboard sorts query patterns by total time, frequency, approximate p95 or maximum duration, shows recent execution samples and refreshes every five seconds while visible. Project administrators cannot access another project's data by changing filters. Normal project members do not receive these diagnostics.
API and plan inspection
Three endpoint families exist under /admin/api,
/platform/api/v1/projects/{id} and /platform/api/v1/operator:
GET /sql/queries: aggregates;GET /sql/events: retained samples.POST /sql/queries/{fingerprint}/explain?database=default: current plan preview.
Read filters: database, source=request|background, fingerprint,
minutes=1..1440 (default 60), limit=1..200 (default 50),
offset=0..10000, sort=total_ms|count|p95_ms|max_ms.
Operators may supply project; it is required for their plan inspection.
Responses follow {data: ...} and {error: {code, message}} envelopes.
Disabled collection returns data.enabled=false. Invalid filters, unavailable
history, unsupported statements and inspection timeouts return a diagnostic error.
Cookie-authenticated inspection requires the existing matching-origin CSRF check.
Inspect plan resolves only a captured fingerprint within the authorized database.
It does not accept arbitrary SQL or paths. It uses EXPLAIN QUERY PLAN on a separate
read-only connection, never executing the original SELECT/INSERT/UPDATE/DELETE.
PRAGMA, DDL, ATTACH and multiple statements are unsupported. Inspection is limited
to one concurrent request per project and a 500 ms execution budget.
The preview uses the current schema and NULL placeholder values. It may differ from the historical plan, particularly for value-dependent or expression indexes. The original execution's full-scan steps, sort count, automatic-index rows and VM steps remain available separately. Scan/index/sorting plan text is advisory, not an "optimized/unoptimized" verdict. No index is created automatically.
Durable row usage
Row accounting (usage.row_metering_enabled, default true) is separate from
sampled SQL diagnostics. It uses a pinned instrumented SQLite engine and an
unsampled attempt context across pooled connections and blocking workers.
A <database-filename>.row-usage.sqlite journal beside the primary database
persists attempt markers, immutable receipts, UTC daily project/user totals,
and export acknowledgements. Each primary database has its own journal.
Keep this file outside application snapshot restores. Customer mutations also
store _usage_row_receipts in their own transaction; recovery reconciles these
before classifying dead owners as incomplete. The journal instance identifies
its receipts, so a cloned application database cannot import source-project usage.
Recovery batches receipt reconciliation and excludes expired receipts restored
from old snapshots, even after their deduplication records have been pruned.
Existing usage responses add rows_read, rows_written, rows_returned, and
row_metering_gaps, plus metering version and the first receipt timestamp in the
requested interval. Only known, durably recorded dimensions enter row totals.
Row retention is at least 90 days, even if general usage retention is shorter;
unacknowledged export evidence is retained. Receipts
contain no row values, SQL parameters, tokens, or email addresses.
The exporter derives /platform/api/usage/rows/ingest from the configured
canonical /platform/api/usage/ingest URL and uses the existing project ingestion
credential, including token-file rotation. The receiver deduplicates by project
and attempt ID across batches. It requires version 1 and rejects receipts older
than 90 days, so a replay cannot become billable after deduplication expiry.
Deploy the receiver before enabling producers.
A missing export token leaves receipts pending. Receipts that age beyond the
receiver's window without acknowledgement are retained with exported = 2 in
the journal for operator review; newer receipts continue exporting. The local
incidents table records recovered outage notices and expired-export counts.
Monitor row_meter.outage, row_meter.gap, and row_meter.receipt_failed events.
An accounting outage does not intentionally block CRUD. Bounded in-memory retries
can themselves be lost in a process crash; outages and incomplete attempts must
remain visible rather than be converted into estimated charges. A process still
alive during a rolling deployment is not classified as crashed.
Hosted warm candidates keep row recovery/export dormant until promotion. Worker
shutdown waits for an in-flight row recovery/export pass before maintenance.
Before enabling production metering, run the engine conformance and server/SDK regression suites and compare enabled/disabled throughput and tail latency. SQLite upgrades must retain metering semantics or introduce a new meter version.
Reproduce the engine comparison against the pristine pinned amalgamation with
python3 vendor/libsqlite3-sys/benchmark.py /path/to/upstream/sqlite3.c. It reports
five-run medians for the upstream engine, the patched engine with accounting off,
and the patched engine with accounting on. For an HTTP/journal latency smoke
benchmark, run cargo test --test row_metering row_metering_latency_benchmark -- --ignored --nocapture.
Use a release build and production-like concurrent load for a rollout decision;
this local benchmark is not a production capacity test.