Files
discord_bot/docs/RUNBOOK.md
T

5.1 KiB
Raw Blame History

Operator runbook — Fjærkroa / Luma bot

One page for "something is wrong, what do I do". Two deployments of one codebase, both on uberspace (push-based deploy from the dev machine — there is no git checkout on the hosts).

Fjærkroa (café) Luma (GGG clan)
SSH host ssh fjerkroa (pictor.uberspace.de) ssh ggg
Service kroa luma
Config ~/fjerkroa_bot/kroa.toml ~/fjerkroa_bot/ggg.toml
Staff channel #kassa #mods
Language / persona Norwegian, café host German, "Luma"

Common paths on each host: bot code ~/fjerkroa_bot, venv ~/venv-bot, database ~/fjerkroa_bot/history/bot.db (SQLite, WAL), our snapshots ~/backups/<kroa|luma>/, logs under ~/logs and ~/tmp.

From Discord (staff channel only, prefix !bot)

No SSH needed for day-to-day control. Type !bot help in the staff channel for the full, grouped list. The essentials:

  • !bot pause / !bot resume — stop / start all replies.
  • !bot quiet <minutes> — go silent for a while, then auto-resume.
  • !bot status — replies/images/tasks flags + quiet time left.
  • !bot spend — today's estimated USD spend, tokens, images, budget.
  • !bot images on|off, !bot tasks on|off — kill-switches.

!help works in any channel (for everyone) and lists only what is usable there. !forgetme and !privacy also work everywhere, even while the bot is paused.

Restart / check health (SSH)

ssh <host>
supervisorctl status <kroa|luma>          # RUNNING + uptime
supervisorctl restart <kroa|luma>
tail -n 40 ~/tmp/<kroa|luma>-stderr*.log   # discord login / errors
tail -n 40 ~/logs/supervisord.log          # "We have logged in as ..."

A healthy start shows a fresh connected to Gateway + We have logged in as ... line within ~15 s.

Deploy a release / roll back

From the dev machine (~/Repos/FjerkroaBot), tags only:

git tag -m "<msg>" vX.Y.Z && git push --tags        # cut the release first
bash deploy/deploy.sh ggg vX.Y.Z                    # luma
DEPLOY_FORCE=1 bash deploy/deploy.sh fjerkroa vX.Y.Z # kroa (see window)
  • kroa refuses to deploy 11:0022:00 Europe/Oslo (restaurant hours); DEPLOY_FORCE=1 overrides. Café is closed Mondays.
  • The script backs up bot.dbbot.db.pre-<tag> before restart, then smoke-tests (RUNNING + fresh login) and fails loudly if either misses.
  • Rollback = deploy the previous tag. If the schema version moved between the two tags, restore the matching bot.db.pre-<newtag> first (see below) so the older code meets a schema it understands.

Restore the database

Three independent daily backup layers exist — pick the freshest good one.

ssh <host>
supervisorctl stop <kroa|luma>
DB=~/fjerkroa_bot/history/bot.db

# 1) uberspace nightly backup of the whole home (read-only):
#    /backup = current + daily.0..7 + weekly.1..7  (15 restore points)
cp /backup/daily.1/home/<user>/fjerkroa_bot/history/bot.db "$DB"

# 2) our own rotated gzip snapshot (03:17 UTC cron, keep 14):
gunzip -c ~/backups/<kroa|luma>/bot-YYYYMMDD-HHMMSS.db.gz > "$DB"

# 3) the pre-deploy snapshot for a given release:
cp "$DB".pre-vX.Y.Z "$DB"

rm -f "$DB"-wal "$DB"-shm    # drop stale WAL sidecars after a restore
supervisorctl start <kroa|luma>

<user> is fjerkroa or ggg. The DB holds conversation history, structured memory, usage ledger, image cache index, tasks, and the news store — all regenerable, none critical. That is why there is no off-host backup: uberspace /backup + the on-host snapshots are enough.

Rotate a secret

Secrets live only in the host *.toml (never in the repo). Edit in place and restart:

ssh <host>
# OpenAI: edit  openai-token = "sk-..."   in kroa.toml / ggg.toml
# Discord: edit discord-token = "..."     (get a new token from the
#          Discord developer portal → Bot → Reset Token first)
supervisorctl restart <kroa|luma>

After rotating an OpenAI key, revoke the old one in the OpenAI dashboard. Keep a *.toml backup before editing; a broken TOML crash-loops the service (validate: ~/venv-bot/bin/python -c 'import tomlkit; tomlkit.load(open("kroa.toml"))').

Scheduled jobs (crontab -l)

Host When (server time) Job
both 17 3 * * * backup_db.py~/backups/<bot>/ (keep 14)
kroa 5 * * * * news digest → {news} file + news store
ggg */15 * * * * news poster → #news/#newsjp webhooks + store

Logs: ~/backups/<bot>/backup.log, ~/backups/<bot>/news*.log.

Quick triage

  • Bot silent everywhere!bot status (paused/quiet?), else supervisorctl status; if not RUNNING, restart and read stderr.
  • Bot silent in one channel → check the host config ignore-channels / short-path rules for that channel (a stray short-path rule can archive messages without replying).
  • Repeated API errors → the bot posts a rate-limited alert to the staff channel after 5 consecutive OpenAI failures (OPS-16); check !bot spend (budget hit?) and the OpenAI status/key.
  • Bad deploy → roll back to the previous tag (above).