# SPEC-012 — Operations hardening Runtime + host operability (FDB-012). Backups and host wiring are `manual` coverage; the in-process alerting is `test`. ### OPS-13 — Consistent DB backups (coverage: test) `deploy/backup_db.py` writes a gzipped snapshot of `bot.db` using the sqlite3 online-backup API — consistent even while the bot writes (WAL-safe) — with 0600 permissions. Restoring a snapshot yields a readable database with the same rows. ### OPS-14 — Backups are rotated (coverage: test) The newest `backup-keep` (default 14) snapshots are kept; older ones are deleted. Timestamped names sort chronologically so rotation is a pure list operation. ### OPS-15 — Backup cron on each host (coverage: manual) Each host runs `backup_db.py` daily via cron, writing to `~/backups//` (outside `~/fjerkroa_bot`, so deploys and service restarts never touch it). Verified by presence of the cron line and a fresh snapshot. ### OPS-16 — Repeated API errors alert staff (coverage: test) The responder counts consecutive OpenAI request failures; at `api-error-alert-threshold` (default 5) in a row it fires one staff alert (rate-limited like all staff alerts) so a silently-broken bot (cf. the gpt-5.6 tools/reasoning incident) surfaces within minutes instead of hours. A success resets the counter. ### OPS-18 — Health monitor watches spend, disk, task-queue (coverage: test) When `enable-monitoring` is true, a loop wakes every `monitor-interval` (default 300 s) and checks three thresholds, alerting the staff channel when one is crossed: daily spend at or above `monitor-spend-alert-frac` (default 0.8) of `daily-budget-usd`; free disk below `monitor-disk-min-mb` (default 500 MB); open task-queue depth at or above `monitor-taskqueue-max` (default 20). A check with no data to evaluate (no budget set, no store, a failed disk read) is skipped, never fatal. With the flag off the loop does nothing. ### OPS-19 — Alerts fire once per crossing and re-arm on recovery (coverage: test) Each metric alerts only on the rising edge — the first tick that finds it over its threshold — and stays silent while it remains over, so a persistent condition does not repeat every interval. When the metric falls back below the threshold the alert re-arms silently, ready to fire again on the next crossing. All alerts still pass through the rate-limited staff-alert path (OPS-07).