131 lines
5.1 KiB
Markdown
131 lines
5.1 KiB
Markdown
# Operator runbook — Fjærkroa / Luma bot
|
||
|
||
One page for "something is wrong, what do I do". Two deployments of one
|
||
codebase, both on **uberspace** (push-based deploy from the dev machine —
|
||
there is no git checkout on the hosts).
|
||
|
||
| | Fjærkroa (café) | Luma (GGG clan) |
|
||
| --- | --- | --- |
|
||
| SSH host | `ssh fjerkroa` (pictor.uberspace.de) | `ssh ggg` |
|
||
| Service | `kroa` | `luma` |
|
||
| Config | `~/fjerkroa_bot/kroa.toml` | `~/fjerkroa_bot/ggg.toml` |
|
||
| Staff channel | `#kassa` | `#mods` |
|
||
| Language / persona | Norwegian, café host | German, "Luma" |
|
||
|
||
Common paths on each host: bot code `~/fjerkroa_bot`, venv `~/venv-bot`,
|
||
database `~/fjerkroa_bot/history/bot.db` (SQLite, WAL), our snapshots
|
||
`~/backups/<kroa|luma>/`, logs under `~/logs` and `~/tmp`.
|
||
|
||
## From Discord (staff channel only, prefix `!bot`)
|
||
|
||
No SSH needed for day-to-day control. Type `!bot help` in the staff
|
||
channel for the full, grouped list. The essentials:
|
||
|
||
- `!bot pause` / `!bot resume` — stop / start all replies.
|
||
- `!bot quiet <minutes>` — go silent for a while, then auto-resume.
|
||
- `!bot status` — replies/images/tasks flags + quiet time left.
|
||
- `!bot spend` — today's estimated USD spend, tokens, images, budget.
|
||
- `!bot images on|off`, `!bot tasks on|off` — kill-switches.
|
||
|
||
`!help` works in **any** channel (for everyone) and lists only what is
|
||
usable there. `!forgetme` and `!privacy` also work everywhere, even
|
||
while the bot is paused.
|
||
|
||
## Restart / check health (SSH)
|
||
|
||
```sh
|
||
ssh <host>
|
||
supervisorctl status <kroa|luma> # RUNNING + uptime
|
||
supervisorctl restart <kroa|luma>
|
||
tail -n 40 ~/tmp/<kroa|luma>-stderr*.log # discord login / errors
|
||
tail -n 40 ~/logs/supervisord.log # "We have logged in as ..."
|
||
```
|
||
|
||
A healthy start shows a fresh `connected to Gateway` + `We have logged
|
||
in as ...` line within ~15 s.
|
||
|
||
## Deploy a release / roll back
|
||
|
||
From the **dev machine** (`~/Repos/FjerkroaBot`), tags only:
|
||
|
||
```sh
|
||
git tag -m "<msg>" vX.Y.Z && git push --tags # cut the release first
|
||
bash deploy/deploy.sh ggg vX.Y.Z # luma
|
||
DEPLOY_FORCE=1 bash deploy/deploy.sh fjerkroa vX.Y.Z # kroa (see window)
|
||
```
|
||
|
||
- kroa refuses to deploy **11:00–22:00 Europe/Oslo** (restaurant hours);
|
||
`DEPLOY_FORCE=1` overrides. Café is closed Mondays.
|
||
- The script backs up `bot.db` → `bot.db.pre-<tag>` before restart, then
|
||
smoke-tests (RUNNING + fresh login) and fails loudly if either misses.
|
||
- **Rollback** = deploy the previous tag. If the schema version moved
|
||
between the two tags, restore the matching `bot.db.pre-<newtag>` first
|
||
(see below) so the older code meets a schema it understands.
|
||
|
||
## Restore the database
|
||
|
||
Three independent daily backup layers exist — pick the freshest good one.
|
||
|
||
```sh
|
||
ssh <host>
|
||
supervisorctl stop <kroa|luma>
|
||
DB=~/fjerkroa_bot/history/bot.db
|
||
|
||
# 1) uberspace nightly backup of the whole home (read-only):
|
||
# /backup = current + daily.0..7 + weekly.1..7 (15 restore points)
|
||
cp /backup/daily.1/home/<user>/fjerkroa_bot/history/bot.db "$DB"
|
||
|
||
# 2) our own rotated gzip snapshot (03:17 UTC cron, keep 14):
|
||
gunzip -c ~/backups/<kroa|luma>/bot-YYYYMMDD-HHMMSS.db.gz > "$DB"
|
||
|
||
# 3) the pre-deploy snapshot for a given release:
|
||
cp "$DB".pre-vX.Y.Z "$DB"
|
||
|
||
rm -f "$DB"-wal "$DB"-shm # drop stale WAL sidecars after a restore
|
||
supervisorctl start <kroa|luma>
|
||
```
|
||
|
||
`<user>` is `fjerkroa` or `ggg`. The DB holds conversation history,
|
||
structured memory, usage ledger, image cache index, tasks, and the news
|
||
store — all regenerable, none critical. That is why there is no off-host
|
||
backup: uberspace `/backup` + the on-host snapshots are enough.
|
||
|
||
## Rotate a secret
|
||
|
||
Secrets live only in the host `*.toml` (never in the repo). Edit in
|
||
place and restart:
|
||
|
||
```sh
|
||
ssh <host>
|
||
# OpenAI: edit openai-token = "sk-..." in kroa.toml / ggg.toml
|
||
# Discord: edit discord-token = "..." (get a new token from the
|
||
# Discord developer portal → Bot → Reset Token first)
|
||
supervisorctl restart <kroa|luma>
|
||
```
|
||
|
||
After rotating an OpenAI key, revoke the old one in the OpenAI dashboard.
|
||
Keep a `*.toml` backup before editing; a broken TOML crash-loops the
|
||
service (validate: `~/venv-bot/bin/python -c 'import tomlkit; tomlkit.load(open("kroa.toml"))'`).
|
||
|
||
## Scheduled jobs (crontab -l)
|
||
|
||
| Host | When (server time) | Job |
|
||
| --- | --- | --- |
|
||
| both | `17 3 * * *` | `backup_db.py` → `~/backups/<bot>/` (keep 14) |
|
||
| kroa | `5 * * * *` | news digest → `{news}` file + news store |
|
||
| ggg | `*/15 * * * *` | news poster → #news/#newsjp webhooks + store |
|
||
|
||
Logs: `~/backups/<bot>/backup.log`, `~/backups/<bot>/news*.log`.
|
||
|
||
## Quick triage
|
||
|
||
- **Bot silent everywhere** → `!bot status` (paused/quiet?), else
|
||
`supervisorctl status`; if not RUNNING, `restart` and read stderr.
|
||
- **Bot silent in one channel** → check the host config `ignore-channels`
|
||
/ `short-path` rules for that channel (a stray `short-path` rule can
|
||
archive messages without replying).
|
||
- **Repeated API errors** → the bot posts a rate-limited alert to the
|
||
staff channel after 5 consecutive OpenAI failures (OPS-16); check
|
||
`!bot spend` (budget hit?) and the OpenAI status/key.
|
||
- **Bad deploy** → roll back to the previous tag (above).
|