Loomup Docs
Guide /docs/observability

Production observability

Loomup keeps billing metering and diagnostic telemetry separate. The existing SQLite usage tables and idempotent outbox are authoritative for billing. Logs, host metrics, and sampled traces are best-effort diagnostic signals exported to Managed ClickStack through a loopback-only OpenTelemetry Collector.

Runtime configuration

Add these non-secret values to /etc/loomup/loomup.env:

dotenv
LOOMUP_LOG_FORMAT=json
LOOMUP_OBSERVABILITY_URL=https://your-managed-clickstack.example
OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:4317
OTEL_SERVICE_NAME=loomup
OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=production

The application never receives the Managed ClickStack credential. Put the remote endpoint and authorization header in /etc/loomup/otel.env, readable by only root and the otelcol group. Install deploy/otel-collector.yaml as /etc/loomup/otel-collector.yaml and the systemd unit as /etc/systemd/system/loomup-otel-collector.service.

Before starting the service, validate the exact collector version and config:

bash
sudo -u otelcol /usr/local/bin/otelcol-contrib validate --config=/etc/loomup/otel-collector.yaml
sudo systemctl daemon-reload
sudo systemctl enable --now loomup-otel-collector.service

The collector is independent of the blue/green Loomup slots and must not be restarted during an ordinary application promotion. Its persistent retry queue lives on the root disk rather than the SQLite block volume. Batches are capped at 1,024 records and the persistent queue at 256 batches, so a prolonged remote outage cannot create an unbounded collector backlog; once full, diagnostic telemetry is dropped while requests and authoritative billing continue.

Operator bootstrap

Grant the initial browser operator locally on the host:

bash
loomup platform grant-operator \
  --email operator@example.com \
  --catalog /mnt/loomup_volume/loomup/platform/control.sqlite \
  --projects-root /mnt/loomup_volume/loomup/platform/projects

The account then receives an Observability navigation item at /platform/observability. Detailed investigation remains in ClickStack; the native page uses only local rolling metrics and durable usage, so it remains available during a remote telemetry outage.

Request completion logs are emitted once per request at INFO, including hosted project requests. Successful /health, /ready, and internal slot-health probes are omitted; unsuccessful probes remain logged. Request lifecycle diagnostics are DEBUG-level. Trace export and billing metering are unchanged.

Categorizing the usage error total

The usage meter's http_errors counts all HTTP statuses >= 400, including authentication failures, denied access, missing resources, sync resets, quota rejections, and server failures. It is not a server outage counter. A sync reset can be a required permission refresh; verify the following bootstrap using the same project and client_hash before assessing its user impact.

Use the offline report to group retained logs by status, normalized route, and correlated sync reason:

bash
python3 scripts/report-http-errors.py --project PROJECT_UUID \
  --since 2026-09-01 --until 2026-10-01 journal.jsonl syslog.7.gz

Input can be journalctl -o cat output, JSON embedded in archived syslog lines, gzip files, or standard input (-). Overlapping archives are deduplicated by request ID within the selected project. --json emits a structured report; --meter month.json compares against that project's monthly usage report and includes dates with no retained logs. Ensure both inputs refer to the same project and cutoff time. Negative differences indicate uncategorized metered errors; positive differences can indicate a later log cutoff, a pending meter flush, or inconsistent inputs. Origin request totals exclude CDN receipts.

sync_recovery identifies a correlated reset diagnostic, not proof of successful recovery. A sync 409 without a diagnostic remains sync_conflict_reason_unavailable; do not assume every sync conflict is a reset. HTTP 507 is reported separately as quota_exceeded; other 5xx responses are server_failure. This report does not change billing totals or API behavior.

Retained timestamp bounds do not prove continuous coverage. Journald's size cap can evict logs much sooner than seven days at high traffic volume. The remote collector and backend must actually be configured and running to provide the longer retention described below; installing a service file alone does not export logs. Do not extrapolate categories from the retained subset to a whole month.

Sync recovery diagnostics

Each reset emits one INFO sync.reset_required event in addition to the ordinary HTTP completion. Its reason is cursor_expired, authorization_dependency_changed, schema_or_authorization_changed, or bootstrap_required. Events include project ID, request ID, normalized route, status, stable error code, and applicable cursor bounds/dependency/resource names. changed_columns contains changed field names when valid before/after payloads are available, and is null otherwise. Resource and changed-column lists are JSON-encoded log fields and are limited to 32 names, each at most 128 characters.

Successful snapshots emit sync.bootstrap.completed. Both events contain the same client_hash for a validated client ID within a project: SHA-256 of the project ID, a zero byte, and client ID. Match project and hash to investigate recovery rather than infer it from neighboring requests. Invalid client IDs have no hash. Protocol errors emit sync.protocol_mismatch with requested/supported versions instead of a reset event.

These diagnostic events do not change usage totals. They contain no raw client IDs, credentials, row contents, permission rules, or query strings. Correlation hashes are log fields, not metric labels.

Local disk retention

Restricted SSH log access

Production has a dedicated loomup-logs SSH account with a fixed project allowlist: approve (961cdf44-0374-4e83-bd12-4917031f54e0, the default) and loomlytics (4991a905-8998-40bb-8c46-252126a315a5). It exposes structured HTTP completion and sync reset/bootstrap/protocol diagnostics from the blue/green journals, plus fixed aggregate usage reports through the usage command. Messages, request bodies, SQL, span attributes, and query strings are omitted. It does not expose archived syslog files, database files, or arbitrary queries.

The local private key is ~/.ssh/loomup_logs (mode 0600); its public key is ~/.ssh/loomup_logs.pub. The private key is passphrase-free for automated log collection and is not installed on the server or stored in this repository.

bash
# Recent errors, newest first. Each reset can have both a diagnostic and an
# HTTP completion record; count completion events for request totals.
ssh -i ~/.ssh/loomup_logs -o IdentitiesOnly=yes loomup-logs@64.227.150.72 \
  'errors --since 2026-09-15 --limit 200'

# Recent HTTP/sync diagnostics, or live following.
ssh -i ~/.ssh/loomup_logs -o IdentitiesOnly=yes loomup-logs@64.227.150.72 \
  'logs --limit 100'
ssh -i ~/.ssh/loomup_logs -o IdentitiesOnly=yes loomup-logs@64.227.150.72 follow

# Daily histogram for this month, today's hourly report, or JSON for a month.
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'usage month'
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'usage today'
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'usage month 2026-09 --json'

# The same project flag works for usage, logs, errors, and follow.
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'usage month --project loomlytics'
ssh -i ~/.ssh/loomup_logs loomup-logs@64.227.150.72 'errors --project loomlytics --limit 100'

For log commands, dates are UTC; --since is inclusive and --until is exclusive. The default interval is the last 24 hours; queries may span at most 31 days. --limit defaults to 200 records and accepts 1–5000. Historical sessions stop after 60 seconds; follow sessions show the latest 20 journal matches, then live records, and stop after ten minutes or the output limit. Reconnect to continue following. Log retention still determines what is available. For categorization, pipe output to scripts/report-http-errors.py --project PROJECT_UUID -; a bounded export is not a full daily meter.

The usage command accepts only today, a UTC date (YYYY-MM-DD), or month with an optional YYYY-MM, plus --json and --project approve|loomlytics. With no period it reports today. Each command defaults to approve when the project flag is omitted. The server resolves only these two aliases to fixed project IDs; the catalog path is fixed. Callers cannot supply source code, SQL, unlisted projects, file paths, or a different executable. Both authorized keys on this account have the same access. Reports have a 55-second collection time limit; timed-out workers and their journal reader are terminated together.

The approved collection/rendering logic comes from the workspace's server-monitoring/usage.py, installed as root-owned /usr/local/libexec/loomup-usage.py (0644). It uses read-only SQLite connections and PRAGMA query_only=ON. Monthly reports read durable project usage aggregates; hourly reports combine streamed request completions with aggregate CDN receipts and retain the warning when they differ from the durable daily meter.

The local wrapper supports the restricted connection without sending Python source to the server:

bash
./server-monitoring/usage.sh --restricted --identity ~/.ssh/loomup_logs month
./server-monitoring/usage.sh --restricted --identity /path/to/bot-private-key today --json
./server-monitoring/usage.sh --restricted --identity ~/.ssh/loomup_logs --project loomlytics month
./server-monitoring/usage-loomlytics.sh --restricted --identity ~/.ssh/loomup_logs month

--restricted defaults to loomup-logs@64.227.150.72. The local wrapper accepts either alias or its project UUID and sends the approved alias to the server. The existing administrative transport remains available without that option. After modifying collection logic, an administrator must install the reviewed usage.py snapshot to the fixed server path; the restricted account cannot update its own executable.

Installed controls (all root-owned):

  • deploy/loomup-read-logs → /usr/local/libexec/loomup-read-logs (0755).
  • deploy/sshd-loomup-logs.conf → /etc/ssh/sshd_config.d/40-loomup-logs.conf (0644).
  • deploy/sudoers-loomup-logs → /etc/sudoers.d/loomup-logs (0440).
  • /etc/ssh/authorized_keys/loomup-logs (0644), with both restrict and the forced reader command on the public-key line.
  • /var/empty/loomup-logs is the root-owned account home. The account belongs only to its own group, not adm, systemd-journal, or sudo.

The forced command may invoke only the reader through sudo with no arguments. The reader strictly parses SSH_ORIGINAL_COMMAND without executing it as shell code, fixes the service/project selection, disables paging, and starts journalctl with a clean environment. SSH disables terminal allocation, user rc files, forwarding, tunnels, and password authentication for this account. The forced command also rejects SFTP and arbitrary commands. The ordinary deploy account is separate.

To revoke new connections immediately, use the administrative account to empty /etc/ssh/authorized_keys/loomup-logs; no SSH reload is needed for key-file changes. To also terminate existing log sessions, run sudo loginctl terminate-user loomup-logs. Keep an administrative session open when changing SSH configuration, validate with sudo sshd -t, and validate sudoers changes with sudo visudo -c before reloading SSH.

Validation on September 15, 2026: positive log/error/follow reads succeeded with the new key and SSH-agent use disabled. Arbitrary commands, environment-file reads, PTY allocation, SFTP, both forwarding directions, and general sudo were rejected. Usage reports were verified in histogram and JSON formats; the closed September daily totals matched the existing meter. Live attempts to select an unlisted project or another catalog were rejected. Six focused reader tests and five usage tests cover command injection, fixed selection, project isolation, output filtering, malformed records, restricted transport, read-only queries, and metering exclusions.

Host retention settings

deploy/loomup-configure-logging installs the production host policy:

  • Journald: 512 MiB maximum use, 64 MiB journal files, seven-day maximum age, and 2 GiB free-space reserve. Runtime journals are limited to 64 MiB. Active files can temporarily exceed the target until rotation.
  • Rsyslog: discard the duplicate loomup-host and loomup messages after journald has stored them. Other services retain their existing syslog path.
  • General syslog: rotate daily, retain seven archives, compress immediately, and rotate above 100 MiB when logrotate runs. This is a rotation threshold, not a hard cap between timer runs. Auth and other system logs keep their existing four-week policy.

The deployment workflow applies the policy after successful public verification. The installer validates rsyslog and logrotate, restarts only the logging daemons when their files change, and restores the previous configuration on failure. Original configuration is retained under /var/backups/loomup-logging. It first compresses any legacy uncompressed syslog.1 to preserve that history when switching away from delayed compression. It does not force rotation or vacuum existing logs; normal rotation retires older history. The collector continues reading journald.

Data and retention policy

  • Request/application logs: 14 days.
  • Traces: 7 days; retain 5% normally, all errors, and all requests over one second.
  • Metrics: 30 days at the backend's normal resolution.
  • Never export authorization, cookies, request/response bodies, query strings, SQL values, email addresses, or client IP addresses.

Configure these TTLs in Managed ClickStack. The repository collector config implements trace sampling; the backend remains responsible for retention.

Alerts

Create a DigitalOcean Uptime HTTPS check for the public /ready endpoint and alert on two minutes of downtime, five minutes above one-second latency, and certificate expiry within 14 days. In ClickStack, begin with alerts for missing host telemetry, sustained 5xx rate, p95 latency, CPU/memory pressure, low disk, usage-outbox delay, and collector export failures. Tune thresholds after the first week of production baselining.

Failure checks

Stopping the collector or blocking its outbound connection must never fail an application request or billing write. Confirm local logs with:

bash
journalctl -f -u loomup-host@blue.service -u loomup-host@green.service
journalctl -f -u loomup-otel-collector.service

Shadow request pricing is intentionally non-billable. Observe it for at least 30 days before creating a versioned customer plan.

Local SQLite query diagnostics

SQL diagnostics work without ClickHouse, ClickStack or an OpenTelemetry collector. They use a separate diagnostics.sqlite next to the platform catalog (hosted) or application database (standalone), and do not modify billing counters or quotas. Collection is disabled unless LOOMUP_SQL_DIAGNOSTICS_ENABLED=true.

SettingDefaultPurpose
LOOMUP_SQL_DIAGNOSTICS_ENABLEDfalseEnable local SQL collection on process startup
LOOMUP_SQL_DIAGNOSTICS_PATHadjacent diagnostics.sqliteOverride the persistent local file
LOOMUP_SQL_SLOW_MS100Retain events at or above this duration
LOOMUP_SQL_SAMPLE_PERCENT1Sample percentage of faster executions (0–100)
LOOMUP_SQL_MAX_BYTES536870912Main-file page budget; retention also manages WAL growth

Restart with collection disabled to stop recording. Existing configurations stay disabled. Begin with a canary, compare enabled/disabled throughput on the same workload, and enable broadly only if throughput regresses by no more than 5%. The checked-in Cargo configuration enables SQLite's native SQL normalizer; builds must preserve LIBSQLITE3_FLAGS=SQLITE_ENABLE_NORMALIZE.

The collector aggregates observed statements into minute buckets, flushing every five seconds and retaining up to 24 hours. All slow executions and sampled fast executions are retained subject to bounded buffers and storage. The dashboard reports dropped diagnostics, stale persistence and capacity eviction. A crash can lose unflushed data; these are diagnostic counts, never billing records. Background tasks, named databases and read pools are distinguished. Pool acquisition wait is measured once per checkout, separately from statement elapsed time. Statement timing excludes preparation and pool wait, but can include lock waiting and row consumption. Counts and latency histograms describe observed executions; they are not an exact audit trail. p95 is a histogram approximation.

SQLite's native normalizer removes literal values and comments while retaining query structure and identifiers. Bound values, results, authorization headers, request bodies and URL query strings are not stored. Oversized SQL is omitted. Diagnostics queries and plan inspection are not collected recursively.

Views and permissions

  • Operators: Observability → SQL query performance, at /platform/observability/sql. Operators can filter all projects, including the operator-only __system__ catalog.
  • Project Studio: Query performance, at /platform/projects/{id}/sql. The existing manager rule applies: workspace owner or project creator.
  • Standalone Admin: Query performance, requiring the project admin role.

The dashboard sorts query patterns by total time, frequency, approximate p95 or maximum duration, shows recent execution samples and refreshes every five seconds while visible. Project administrators cannot access another project's data by changing filters. Normal project members do not receive these diagnostics.

API and plan inspection

Three endpoint families exist under /admin/api, /platform/api/v1/projects/{id} and /platform/api/v1/operator:

  • GET /sql/queries: aggregates; GET /sql/events: retained samples.
  • POST /sql/queries/{fingerprint}/explain?database=default: current plan preview.

Read filters: database, source=request|background, fingerprint, minutes=1..1440 (default 60), limit=1..200 (default 50), offset=0..10000, sort=total_ms|count|p95_ms|max_ms. Operators may supply project; it is required for their plan inspection. Responses follow {data: ...} and {error: {code, message}} envelopes. Disabled collection returns data.enabled=false. Invalid filters, unavailable history, unsupported statements and inspection timeouts return a diagnostic error. Cookie-authenticated inspection requires the existing matching-origin CSRF check.

Inspect plan resolves only a captured fingerprint within the authorized database. It does not accept arbitrary SQL or paths. It uses EXPLAIN QUERY PLAN on a separate read-only connection, never executing the original SELECT/INSERT/UPDATE/DELETE. PRAGMA, DDL, ATTACH and multiple statements are unsupported. Inspection is limited to one concurrent request per project and a 500 ms execution budget.

The preview uses the current schema and NULL placeholder values. It may differ from the historical plan, particularly for value-dependent or expression indexes. The original execution's full-scan steps, sort count, automatic-index rows and VM steps remain available separately. Scan/index/sorting plan text is advisory, not an "optimized/unoptimized" verdict. No index is created automatically.

Durable row usage

Row accounting (usage.row_metering_enabled, default true) is separate from sampled SQL diagnostics. It uses a pinned instrumented SQLite engine and an unsampled attempt context across pooled connections and blocking workers.

A <database-filename>.row-usage.sqlite journal beside the primary database persists attempt markers, immutable receipts, UTC daily project/user totals, and export acknowledgements. Each primary database has its own journal. Keep this file outside application snapshot restores. Customer mutations also store _usage_row_receipts in their own transaction; recovery reconciles these before classifying dead owners as incomplete. The journal instance identifies its receipts, so a cloned application database cannot import source-project usage. Recovery batches receipt reconciliation and excludes expired receipts restored from old snapshots, even after their deduplication records have been pruned.

Existing usage responses add rows_read, rows_written, rows_returned, and row_metering_gaps, plus metering version and the first receipt timestamp in the requested interval. Only known, durably recorded dimensions enter row totals. Row retention is at least 90 days, even if general usage retention is shorter; unacknowledged export evidence is retained. Receipts contain no row values, SQL parameters, tokens, or email addresses.

The exporter derives /platform/api/usage/rows/ingest from the configured canonical /platform/api/usage/ingest URL and uses the existing project ingestion credential, including token-file rotation. The receiver deduplicates by project and attempt ID across batches. It requires version 1 and rejects receipts older than 90 days, so a replay cannot become billable after deduplication expiry. Deploy the receiver before enabling producers.

A missing export token leaves receipts pending. Receipts that age beyond the receiver's window without acknowledgement are retained with exported = 2 in the journal for operator review; newer receipts continue exporting. The local incidents table records recovered outage notices and expired-export counts.

Monitor row_meter.outage, row_meter.gap, and row_meter.receipt_failed events. An accounting outage does not intentionally block CRUD. Bounded in-memory retries can themselves be lost in a process crash; outages and incomplete attempts must remain visible rather than be converted into estimated charges. A process still alive during a rolling deployment is not classified as crashed. Hosted warm candidates keep row recovery/export dormant until promotion. Worker shutdown waits for an in-flight row recovery/export pass before maintenance.

Before enabling production metering, run the engine conformance and server/SDK regression suites and compare enabled/disabled throughput and tail latency. SQLite upgrades must retain metering semantics or introduce a new meter version.

Reproduce the engine comparison against the pristine pinned amalgamation with python3 vendor/libsqlite3-sys/benchmark.py /path/to/upstream/sqlite3.c. It reports five-run medians for the upstream engine, the patched engine with accounting off, and the patched engine with accounting on. For an HTTP/journal latency smoke benchmark, run cargo test --test row_metering row_metering_latency_benchmark -- --ignored --nocapture. Use a release build and production-like concurrent load for a rollout decision; this local benchmark is not a production capacity test.