5.1 KiB
Operator runbook — Fjærkroa / Luma bot
One page for "something is wrong, what do I do". Two deployments of one codebase, both on uberspace (push-based deploy from the dev machine — there is no git checkout on the hosts).
| Fjærkroa (café) | Luma (GGG clan) | |
|---|---|---|
| SSH host | ssh fjerkroa (pictor.uberspace.de) |
ssh ggg |
| Service | kroa |
luma |
| Config | ~/fjerkroa_bot/kroa.toml |
~/fjerkroa_bot/ggg.toml |
| Staff channel | #kassa |
#mods |
| Language / persona | Norwegian, café host | German, "Luma" |
Common paths on each host: bot code ~/fjerkroa_bot, venv ~/venv-bot,
database ~/fjerkroa_bot/history/bot.db (SQLite, WAL), our snapshots
~/backups/<kroa|luma>/, logs under ~/logs and ~/tmp.
From Discord (staff channel only, prefix !bot)
No SSH needed for day-to-day control. Type !bot help in the staff
channel for the full, grouped list. The essentials:
!bot pause/!bot resume— stop / start all replies.!bot quiet <minutes>— go silent for a while, then auto-resume.!bot status— replies/images/tasks flags + quiet time left.!bot spend— today's estimated USD spend, tokens, images, budget.!bot images on|off,!bot tasks on|off— kill-switches.
!help works in any channel (for everyone) and lists only what is
usable there. !forgetme and !privacy also work everywhere, even
while the bot is paused.
Restart / check health (SSH)
ssh <host>
supervisorctl status <kroa|luma> # RUNNING + uptime
supervisorctl restart <kroa|luma>
tail -n 40 ~/tmp/<kroa|luma>-stderr*.log # discord login / errors
tail -n 40 ~/logs/supervisord.log # "We have logged in as ..."
A healthy start shows a fresh connected to Gateway + We have logged in as ... line within ~15 s.
Deploy a release / roll back
From the dev machine (~/Repos/FjerkroaBot), tags only:
git tag -m "<msg>" vX.Y.Z && git push --tags # cut the release first
bash deploy/deploy.sh ggg vX.Y.Z # luma
DEPLOY_FORCE=1 bash deploy/deploy.sh fjerkroa vX.Y.Z # kroa (see window)
- kroa refuses to deploy 11:00–22:00 Europe/Oslo (restaurant hours);
DEPLOY_FORCE=1overrides. Café is closed Mondays. - The script backs up
bot.db→bot.db.pre-<tag>before restart, then smoke-tests (RUNNING + fresh login) and fails loudly if either misses. - Rollback = deploy the previous tag. If the schema version moved
between the two tags, restore the matching
bot.db.pre-<newtag>first (see below) so the older code meets a schema it understands.
Restore the database
Three independent daily backup layers exist — pick the freshest good one.
ssh <host>
supervisorctl stop <kroa|luma>
DB=~/fjerkroa_bot/history/bot.db
# 1) uberspace nightly backup of the whole home (read-only):
# /backup = current + daily.0..7 + weekly.1..7 (15 restore points)
cp /backup/daily.1/home/<user>/fjerkroa_bot/history/bot.db "$DB"
# 2) our own rotated gzip snapshot (03:17 UTC cron, keep 14):
gunzip -c ~/backups/<kroa|luma>/bot-YYYYMMDD-HHMMSS.db.gz > "$DB"
# 3) the pre-deploy snapshot for a given release:
cp "$DB".pre-vX.Y.Z "$DB"
rm -f "$DB"-wal "$DB"-shm # drop stale WAL sidecars after a restore
supervisorctl start <kroa|luma>
<user> is fjerkroa or ggg. The DB holds conversation history,
structured memory, usage ledger, image cache index, tasks, and the news
store — all regenerable, none critical. That is why there is no off-host
backup: uberspace /backup + the on-host snapshots are enough.
Rotate a secret
Secrets live only in the host *.toml (never in the repo). Edit in
place and restart:
ssh <host>
# OpenAI: edit openai-token = "sk-..." in kroa.toml / ggg.toml
# Discord: edit discord-token = "..." (get a new token from the
# Discord developer portal → Bot → Reset Token first)
supervisorctl restart <kroa|luma>
After rotating an OpenAI key, revoke the old one in the OpenAI dashboard.
Keep a *.toml backup before editing; a broken TOML crash-loops the
service (validate: ~/venv-bot/bin/python -c 'import tomlkit; tomlkit.load(open("kroa.toml"))').
Scheduled jobs (crontab -l)
| Host | When (server time) | Job |
|---|---|---|
| both | 17 3 * * * |
backup_db.py → ~/backups/<bot>/ (keep 14) |
| kroa | 5 * * * * |
news digest → {news} file + news store |
| ggg | */15 * * * * |
news poster → #news/#newsjp webhooks + store |
Logs: ~/backups/<bot>/backup.log, ~/backups/<bot>/news*.log.
Quick triage
- Bot silent everywhere →
!bot status(paused/quiet?), elsesupervisorctl status; if not RUNNING,restartand read stderr. - Bot silent in one channel → check the host config
ignore-channels/short-pathrules for that channel (a strayshort-pathrule can archive messages without replying). - Repeated API errors → the bot posts a rate-limited alert to the
staff channel after 5 consecutive OpenAI failures (OPS-16); check
!bot spend(budget hit?) and the OpenAI status/key. - Bad deploy → roll back to the previous tag (above).