Files
discord_bot/specs/SPEC-012-ops.md
T

2.3 KiB

SPEC-012 — Operations hardening

Runtime + host operability (FDB-012). Backups and host wiring are manual coverage; the in-process alerting is test.

OPS-13 — Consistent DB backups (coverage: test)

deploy/backup_db.py writes a gzipped snapshot of bot.db using the sqlite3 online-backup API — consistent even while the bot writes (WAL-safe) — with 0600 permissions. Restoring a snapshot yields a readable database with the same rows.

OPS-14 — Backups are rotated (coverage: test)

The newest backup-keep (default 14) snapshots are kept; older ones are deleted. Timestamped names sort chronologically so rotation is a pure list operation.

OPS-15 — Backup cron on each host (coverage: manual)

Each host runs backup_db.py daily via cron, writing to ~/backups/<bot>/ (outside ~/fjerkroa_bot, so deploys and service restarts never touch it). Verified by presence of the cron line and a fresh snapshot.

OPS-16 — Repeated API errors alert staff (coverage: test)

The responder counts consecutive OpenAI request failures; at api-error-alert-threshold (default 5) in a row it fires one staff alert (rate-limited like all staff alerts) so a silently-broken bot (cf. the gpt-5.6 tools/reasoning incident) surfaces within minutes instead of hours. A success resets the counter.

OPS-18 — Health monitor watches spend, disk, task-queue (coverage: test)

When enable-monitoring is true, a loop wakes every monitor-interval (default 300 s) and checks three thresholds, alerting the staff channel when one is crossed: daily spend at or above monitor-spend-alert-frac (default 0.8) of daily-budget-usd; free disk below monitor-disk-min-mb (default 500 MB); open task-queue depth at or above monitor-taskqueue-max (default 20). A check with no data to evaluate (no budget set, no store, a failed disk read) is skipped, never fatal. With the flag off the loop does nothing.

OPS-19 — Alerts fire once per crossing and re-arm on recovery (coverage: test)

Each metric alerts only on the rising edge — the first tick that finds it over its threshold — and stays silent while it remains over, so a persistent condition does not repeat every interval. When the metric falls back below the threshold the alert re-arms silently, ready to fire again on the next crossing. All alerts still pass through the rate-limited staff-alert path (OPS-07).