Compare commits

..

7 Commits

18 changed files with 764 additions and 31 deletions
+19
View File
@@ -60,6 +60,25 @@ Decisions inside the set architecture. D-NNN, never renumbered.
broken classifier must never mute the bot; the budget gate already broken classifier must never mute the bot; the budget gate already
bounds spend. Its verdict gates BEFORE the main call, the bounds spend. Its verdict gates BEFORE the main call, the
envelope's answer_needed still gates after — two independent nets. envelope's answer_needed still gates after — two independent nets.
- **D-021** — Health monitoring (FDB-012, SPEC-012 OPS-18/19): a
separate `monitor_loop` (own cadence, default 300 s) rather than
folding checks into the 60 s task loop — monitoring is coarse and
should not run every minute. Checks are edge-triggered (alert on the
rising edge, re-arm on recovery) so a standing condition never spams;
they reuse the existing rate-limited staff-alert path. Metrics are
the cheap, high-signal ones (spend vs budget, free disk, task-queue
depth); each is independently skippable when it has no data, so a
deployment without a budget or store still runs the others. Opt-in
(`enable-monitoring`) like every other operational rollout.
- **D-020** — Web search via Exa (FDB-022, SPEC-015): a `web_search`
tool alongside fetch_url/IGDB/codex/get_news, filling the "look it up
on the open web" gap. Exa (not a raw search-engine scrape) because it
returns clean title+url+text in one call — no SSRF surface of our own
(we call one fixed API endpoint, not arbitrary hosts), and it pairs
with fetch_url for the full article. Key is a host secret
(`exa-api-key`, env `EXA_API_KEY` fallback), never repo-side; results
sanitized like every other external-text tool; off by default
(`enable-web-search`), metered per user.
- **D-019** — News memory + on-demand tool (SPEC-013 NEWS-07..12): - **D-019** — News memory + on-demand tool (SPEC-013 NEWS-07..12):
the news pipeline now carries item summaries (feed descriptions, the news pipeline now carries item summaries (feed descriptions,
HTML-stripped) and persists every fetched item into a deduped `news` HTML-stripped) and persists every fetched item into a deduped `news`
+130
View File
@@ -0,0 +1,130 @@
# Operator runbook — Fjærkroa / Luma bot
One page for "something is wrong, what do I do". Two deployments of one
codebase, both on **uberspace** (push-based deploy from the dev machine —
there is no git checkout on the hosts).
| | Fjærkroa (café) | Luma (GGG clan) |
| --- | --- | --- |
| SSH host | `ssh fjerkroa` (pictor.uberspace.de) | `ssh ggg` |
| Service | `kroa` | `luma` |
| Config | `~/fjerkroa_bot/kroa.toml` | `~/fjerkroa_bot/ggg.toml` |
| Staff channel | `#kassa` | `#mods` |
| Language / persona | Norwegian, café host | German, "Luma" |
Common paths on each host: bot code `~/fjerkroa_bot`, venv `~/venv-bot`,
database `~/fjerkroa_bot/history/bot.db` (SQLite, WAL), our snapshots
`~/backups/<kroa|luma>/`, logs under `~/logs` and `~/tmp`.
## From Discord (staff channel only, prefix `!bot`)
No SSH needed for day-to-day control. Type `!bot help` in the staff
channel for the full, grouped list. The essentials:
- `!bot pause` / `!bot resume` — stop / start all replies.
- `!bot quiet <minutes>` — go silent for a while, then auto-resume.
- `!bot status` — replies/images/tasks flags + quiet time left.
- `!bot spend` — today's estimated USD spend, tokens, images, budget.
- `!bot images on|off`, `!bot tasks on|off` — kill-switches.
`!help` works in **any** channel (for everyone) and lists only what is
usable there. `!forgetme` and `!privacy` also work everywhere, even
while the bot is paused.
## Restart / check health (SSH)
```sh
ssh <host>
supervisorctl status <kroa|luma> # RUNNING + uptime
supervisorctl restart <kroa|luma>
tail -n 40 ~/tmp/<kroa|luma>-stderr*.log # discord login / errors
tail -n 40 ~/logs/supervisord.log # "We have logged in as ..."
```
A healthy start shows a fresh `connected to Gateway` + `We have logged
in as ...` line within ~15 s.
## Deploy a release / roll back
From the **dev machine** (`~/Repos/FjerkroaBot`), tags only:
```sh
git tag -m "<msg>" vX.Y.Z && git push --tags # cut the release first
bash deploy/deploy.sh ggg vX.Y.Z # luma
DEPLOY_FORCE=1 bash deploy/deploy.sh fjerkroa vX.Y.Z # kroa (see window)
```
- kroa refuses to deploy **11:0022:00 Europe/Oslo** (restaurant hours);
`DEPLOY_FORCE=1` overrides. Café is closed Mondays.
- The script backs up `bot.db``bot.db.pre-<tag>` before restart, then
smoke-tests (RUNNING + fresh login) and fails loudly if either misses.
- **Rollback** = deploy the previous tag. If the schema version moved
between the two tags, restore the matching `bot.db.pre-<newtag>` first
(see below) so the older code meets a schema it understands.
## Restore the database
Three independent daily backup layers exist — pick the freshest good one.
```sh
ssh <host>
supervisorctl stop <kroa|luma>
DB=~/fjerkroa_bot/history/bot.db
# 1) uberspace nightly backup of the whole home (read-only):
# /backup = current + daily.0..7 + weekly.1..7 (15 restore points)
cp /backup/daily.1/home/<user>/fjerkroa_bot/history/bot.db "$DB"
# 2) our own rotated gzip snapshot (03:17 UTC cron, keep 14):
gunzip -c ~/backups/<kroa|luma>/bot-YYYYMMDD-HHMMSS.db.gz > "$DB"
# 3) the pre-deploy snapshot for a given release:
cp "$DB".pre-vX.Y.Z "$DB"
rm -f "$DB"-wal "$DB"-shm # drop stale WAL sidecars after a restore
supervisorctl start <kroa|luma>
```
`<user>` is `fjerkroa` or `ggg`. The DB holds conversation history,
structured memory, usage ledger, image cache index, tasks, and the news
store — all regenerable, none critical. That is why there is no off-host
backup: uberspace `/backup` + the on-host snapshots are enough.
## Rotate a secret
Secrets live only in the host `*.toml` (never in the repo). Edit in
place and restart:
```sh
ssh <host>
# OpenAI: edit openai-token = "sk-..." in kroa.toml / ggg.toml
# Discord: edit discord-token = "..." (get a new token from the
# Discord developer portal → Bot → Reset Token first)
supervisorctl restart <kroa|luma>
```
After rotating an OpenAI key, revoke the old one in the OpenAI dashboard.
Keep a `*.toml` backup before editing; a broken TOML crash-loops the
service (validate: `~/venv-bot/bin/python -c 'import tomlkit; tomlkit.load(open("kroa.toml"))'`).
## Scheduled jobs (crontab -l)
| Host | When (server time) | Job |
| --- | --- | --- |
| both | `17 3 * * *` | `backup_db.py``~/backups/<bot>/` (keep 14) |
| kroa | `5 * * * *` | news digest → `{news}` file + news store |
| ggg | `*/15 * * * *` | news poster → #news/#newsjp webhooks + store |
Logs: `~/backups/<bot>/backup.log`, `~/backups/<bot>/news*.log`.
## Quick triage
- **Bot silent everywhere** → `!bot status` (paused/quiet?), else
`supervisorctl status`; if not RUNNING, `restart` and read stderr.
- **Bot silent in one channel** → check the host config `ignore-channels`
/ `short-path` rules for that channel (a stray `short-path` rule can
archive messages without replying).
- **Repeated API errors** → the bot posts a rate-limited alert to the
staff channel after 5 consecutive OpenAI failures (OPS-16); check
`!bot spend` (budget hit?) and the OpenAI status/key.
- **Bad deploy** → roll back to the previous tag (above).
+1 -1
View File
@@ -1,7 +1,7 @@
"""Codex Mechanicus search tool (SPEC-014, FDB-019). """Codex Mechanicus search tool (SPEC-014, FDB-019).
Luma's own sacred archive — the Codex Mechanicus at binaric.tech — as a Luma's own sacred archive — the Codex Mechanicus at binaric.tech — as a
function tool. She searches the codex index and answers Cult Mechanicus function tool. He searches the codex index and answers Cult Mechanicus
lore from real, sourced inscriptions instead of inventing it. The index lore from real, sourced inscriptions instead of inventing it. The index
is fetched over HTTPS (SSRF-guarded, size-bounded, cached in memory) and is fetched over HTTPS (SSRF-guarded, size-bounded, cached in memory) and
every field returned to the model is sanitized (SAF-03), because even every field returned to the model is sanitized (SAF-03), because even
+64 -8
View File
@@ -3,9 +3,11 @@ import asyncio
import logging import logging
import random import random
import re import re
import shutil
import sys import sys
import time import time
from collections import deque from collections import deque
from pathlib import Path
from typing import Optional, Union from typing import Optional, Union
import discord import discord
@@ -16,6 +18,7 @@ from watchdog.events import FileSystemEventHandler
from watchdog.observers import Observer from watchdog.observers import Observer
from .ai_responder import AIMessage from .ai_responder import AIMessage
from .monitor import HealthMonitor
from .openai_responder import OpenAIResponder from .openai_responder import OpenAIResponder
from .tasks import TaskEngine from .tasks import TaskEngine
@@ -65,11 +68,41 @@ def split_answer(text: str, threshold: int, max_parts: int) -> list:
class ConfigFileHandler(FileSystemEventHandler): class ConfigFileHandler(FileSystemEventHandler):
def __init__(self, on_modified): """Rename-safe config watch (CFG-05).
self._on_modified = on_modified
Editors and tools save atomically — write a temp file, then rename it
over the target — which fires a *moved*/*created* event (not
*modified*) and swaps the inode, so watching the file directly goes
deaf after the first save. We watch the config's *directory* and react
to any event whose src or dest path is the config file.
"""
def __init__(self, config_path: str, on_change):
self._config_path = str(Path(config_path).resolve())
self._on_change = on_change
def _hits_config(self, event) -> bool:
for attr in ("src_path", "dest_path"):
path = getattr(event, attr, "")
if path and str(Path(path).resolve()) == self._config_path:
return True
return False
def _dispatch(self, event):
if not event.is_directory and self._hits_config(event):
self._on_change()
# Only write/rename events — NOT on_opened/on_closed, whose read-opens
# (our own load_config re-reads the file) would otherwise feed back into
# a reload loop (CFG-05).
def on_modified(self, event): def on_modified(self, event):
self._on_modified(event) self._dispatch(event)
def on_created(self, event):
self._dispatch(event)
def on_moved(self, event):
self._dispatch(event)
class FjerkroaBot(commands.Bot): class FjerkroaBot(commands.Bot):
@@ -99,8 +132,10 @@ class FjerkroaBot(commands.Bot):
def init_observer(self): def init_observer(self):
self.observer = Observer() self.observer = Observer()
self.file_handler = ConfigFileHandler(self.on_config_file_modified) config_path = Path(self.config_file).resolve()
self.observer.schedule(self.file_handler, path=self.config_file, recursive=False) self.file_handler = ConfigFileHandler(str(config_path), self.on_config_file_changed)
# Watch the directory, not the file — atomic saves replace the inode (CFG-05)
self.observer.schedule(self.file_handler, path=str(config_path.parent), recursive=False)
self.observer.start() self.observer.start()
def init_aichannels(self): def init_aichannels(self):
@@ -130,6 +165,15 @@ class FjerkroaBot(commands.Bot):
observe=self.airesponder.observe_event, observe=self.airesponder.observe_event,
) )
self.loop.create_task(self.task_loop()) self.loop.create_task(self.task_loop())
# Proactive health monitoring -> staff alerts (OPS-18/19)
self.health_monitor = HealthMonitor(
config_getter=lambda: self.config,
ledger=self.airesponder.ledger,
store=self.airesponder.store,
disk_free_mb=self._disk_free_mb,
alert=self.send_staff_alert,
)
self.loop.create_task(self.monitor_loop())
logging.info("Task engine initialised.") logging.info("Task engine initialised.")
async def task_loop(self): async def task_loop(self):
@@ -140,6 +184,20 @@ class FjerkroaBot(commands.Bot):
except Exception as err: except Exception as err:
logging.warning(f"task tick failed: {repr(err)}") logging.warning(f"task tick failed: {repr(err)}")
def _disk_free_mb(self) -> float:
directory = Path(self.config.get("history-directory", ".")).expanduser()
target = directory if directory.exists() else Path.home()
return shutil.disk_usage(target).free / (1024 * 1024)
async def monitor_loop(self):
while True:
await asyncio.sleep(int(self.config.get("monitor-interval", 300)))
if self.health_monitor.enabled():
try:
await self.health_monitor.tick()
except Exception as err:
logging.warning(f"monitor tick failed: {repr(err)}")
async def _execute_task(self, channel_name: str, prompt: str) -> None: async def _execute_task(self, channel_name: str, prompt: str) -> None:
"""Run a due task through the normal responder path (TSK-02).""" """Run a due task through the normal responder path (TSK-02)."""
channel = self.channel_by_name(channel_name, getattr(self, "chat_channel", None), no_ignore=True) channel = self.channel_by_name(channel_name, getattr(self, "chat_channel", None), no_ignore=True)
@@ -389,12 +447,10 @@ class FjerkroaBot(commands.Bot):
airesponder.image_cache.purge_message(str(message.id)) # IMG-14 airesponder.image_cache.purge_message(str(message.id)) # IMG-14
await airesponder.observe_event(message.author.name, "delete", f"deleted: {message.content}") await airesponder.observe_event(message.author.name, "delete", f"deleted: {message.content}")
def on_config_file_modified(self, event): def on_config_file_changed(self):
# Runs on the watchdog observer thread — the swap itself is # Runs on the watchdog observer thread — the swap itself is
# scheduled onto the event loop so no request reads a # scheduled onto the event loop so no request reads a
# half-swapped config (CFG-04 / D9) # half-swapped config (CFG-04 / D9)
if event.src_path != self.config_file:
return
new_config = self.load_config(self.config_file) new_config = self.load_config(self.config_file)
if repr(new_config) == repr(self.config): if repr(new_config) == repr(self.config):
return return
+82
View File
@@ -0,0 +1,82 @@
"""Proactive health monitoring -> staff alerts (SPEC-012, FDB-012).
A periodic check that watches daily spend against the budget, free disk,
and task-queue depth, and posts a staff alert when a threshold is crossed
— once per crossing, re-arming when the metric recovers, so a persistent
condition never spams. Opt-in per deployment (`enable-monitoring`); it
reuses the rate-limited staff-alert channel (OPS-07).
"""
import logging
from typing import Any, Callable, Dict, Optional, Tuple
# A check returns (metric-name, is-over-threshold, alert-message) or None when not applicable.
Check = Optional[Tuple[str, bool, str]]
class HealthMonitor:
def __init__(
self,
config_getter: Callable[[], Dict[str, Any]],
ledger: Any,
store: Any,
disk_free_mb: Callable[[], float],
alert: Callable[[str], Any],
) -> None:
self._config = config_getter
self._ledger = ledger
self._store = store
self._disk_free_mb = disk_free_mb
self._alert = alert
self._armed: Dict[str, bool] = {}
def enabled(self) -> bool:
return bool(self._config().get("enable-monitoring", False))
def _check_spend(self) -> Check:
config = self._config()
if "daily-budget-usd" not in config:
return None
budget = float(config["daily-budget-usd"])
if budget <= 0:
return None
spent = float(self._ledger.spent_usd())
frac = spent / budget
threshold = float(config.get("monitor-spend-alert-frac", 0.8))
return ("spend", frac >= threshold, f"💸 Spend at ${spent:.2f} of ${budget:.2f} today ({frac:.0%}, alert ≥ {threshold:.0%}).")
def _check_disk(self) -> Check:
try:
free = float(self._disk_free_mb())
except Exception as err:
logging.debug(f"monitor: disk check failed: {err!r}")
return None
min_mb = float(self._config().get("monitor-disk-min-mb", 500))
return ("disk", free < min_mb, f"💾 Low disk: {free:.0f} MB free (alert < {min_mb:.0f} MB).")
def _check_queue(self) -> Check:
if self._store is None:
return None
try:
depth = len(self._store.tasks_open())
except Exception as err:
logging.debug(f"monitor: queue check failed: {err!r}")
return None
limit = int(self._config().get("monitor-taskqueue-max", 20))
return ("task-queue", depth >= limit, f"🗒️ Task queue deep: {depth} open (alert ≥ {limit}).")
async def tick(self) -> None:
"""Evaluate every check; alert on a rising edge only (OPS-18/19)."""
for check in (self._check_spend(), self._check_disk(), self._check_queue()):
if check is None:
continue
metric, over, message = check
await self._fire(metric, over, message)
async def _fire(self, metric: str, over: bool, message: str) -> None:
was_over = self._armed.get(metric, False)
if over and not was_over:
self._armed[metric] = True
await self._alert(message)
elif not over and was_over:
self._armed[metric] = False # recovered — re-arm silently for the next crossing
+20 -12
View File
@@ -27,6 +27,7 @@ DEFAULT_SUMMARY_CHARS = 200
DEFAULT_NEWS_KEEP = 400 DEFAULT_NEWS_KEEP = 400
FETCH_TIMEOUT_S = 15 FETCH_TIMEOUT_S = 15
_ATOM = "{http://www.w3.org/2005/Atom}" _ATOM = "{http://www.w3.org/2005/Atom}"
_RSS1 = "{http://purl.org/rss/1.0/}" # RSS 1.0 / RDF (e.g. 4gamer.net) namespaces <item>/<title>/<link>
_TAG_RE = re.compile(r"<[^>]+>") _TAG_RE = re.compile(r"<[^>]+>")
@@ -36,22 +37,28 @@ def _clean_summary(raw: str, max_len: int = 300) -> str:
return re.sub(r"\s+", " ", text).strip()[:max_len] return re.sub(r"\s+", " ", text).strip()[:max_len]
def _rss_items(root: Any, ns: str, source: str) -> List[Dict[str, str]]:
"""RSS 2.0 (ns='') and RSS 1.0/RDF (ns=_RSS1) both use <item><title><link><description>."""
out: List[Dict[str, str]] = []
for item in root.iter(f"{ns}item"):
title = (item.findtext(f"{ns}title") or "").strip()
link = (item.findtext(f"{ns}link") or "").strip()
summary = _clean_summary(item.findtext(f"{ns}description") or "")
if title:
out.append({"title": title, "link": link, "source": source, "summary": summary})
return out
def parse_feed(data: bytes, source: str = "") -> List[Dict[str, str]]: def parse_feed(data: bytes, source: str = "") -> List[Dict[str, str]]:
"""Parse RSS or Atom bytes into [{title, link, source}] (tolerant).""" """Parse RSS 2.0, RSS 1.0/RDF, or Atom bytes into [{title, link, source, summary}] (tolerant)."""
try: try:
root = ElementTree.fromstring(data) root = ElementTree.fromstring(data)
except Exception as err: except Exception as err:
# malformed XML or a blocked entity/DTD attack — tolerate, never raise (NEWS-01) # malformed XML or a blocked entity/DTD attack — tolerate, never raise (NEWS-01)
logging.warning(f"news: unparseable/unsafe feed {source!r}: {err!r}") logging.warning(f"news: unparseable/unsafe feed {source!r}: {err!r}")
return [] return []
items: List[Dict[str, str]] = [] # RSS 2.0 (unqualified) + RSS 1.0/RDF (namespaced, e.g. 4gamer) share <item><title><link><description>
# RSS: <rss><channel><item><title/><link/><description/> items: List[Dict[str, str]] = _rss_items(root, "", source) + _rss_items(root, _RSS1, source)
for item in root.iter("item"):
title = (item.findtext("title") or "").strip()
link = (item.findtext("link") or "").strip()
summary = _clean_summary(item.findtext("description") or "")
if title:
items.append({"title": title, "link": link, "source": source, "summary": summary})
# Atom: <feed><entry><title/><link href=/><summary|content/> # Atom: <feed><entry><title/><link href=/><summary|content/>
for entry in root.iter(f"{_ATOM}entry"): for entry in root.iter(f"{_ATOM}entry"):
title = (entry.findtext(f"{_ATOM}title") or "").strip() title = (entry.findtext(f"{_ATOM}title") or "").strip()
@@ -184,9 +191,10 @@ class NewsPoster:
GET_NEWS_TOOL = { GET_NEWS_TOOL = {
"name": "get_news", "name": "get_news",
"description": "Fetch recent real-world news the bot has collected from its RSS feeds (local, national, world, sport, " "description": "Fetch news the bot has collected from its RSS feeds — this is the SAME news that gets posted in the "
"culture). Use when someone asks what is new or what is happening, optionally about a topic or from a particular " "server's news channels (e.g. #news, #newsjp / ニュース). Use this FIRST, before web_search, for anything about "
"source. Returns headlines with a short summary and a link to read more.", "current news or about something someone saw in a news channel; filter by topic (a keyword, also matches the source "
"label) or by source. Returns headlines with a short summary and a link; follow up with fetch_url for the full text.",
"parameters": { "parameters": {
"type": "object", "type": "object",
"properties": { "properties": {
+12
View File
@@ -17,6 +17,8 @@ from .leonardo_draw import LeonardoAIDrawMixIn
from .news import GET_NEWS_TOOL, query_news from .news import GET_NEWS_TOOL, query_news
from .quota import QuotaLedger from .quota import QuotaLedger
from .url_reader import FETCH_URL_TOOL, URLReader from .url_reader import FETCH_URL_TOOL, URLReader
from .websearch import DEFAULT_RESULTS as WEB_DEFAULT_RESULTS
from .websearch import WEB_SEARCH_TOOL, WebSearch
# The response envelope, enforced server-side via structured outputs # The response envelope, enforced server-side via structured outputs
# (ENV-19). All fields required, closed object, nullable where the # (ENV-19). All fields required, closed object, nullable where the
@@ -166,6 +168,8 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
self.url_reader = URLReader(lambda: self.config, self.image_cache) self.url_reader = URLReader(lambda: self.config, self.image_cache)
# Codex Mechanicus search (SPEC-014); Luma's own archive at binaric.tech # Codex Mechanicus search (SPEC-014); Luma's own archive at binaric.tech
self.codex = CodexSearch(lambda: self.config) self.codex = CodexSearch(lambda: self.config)
# Web search (SPEC-015) via Exa; general "look it up" beyond fetch_url/news/codex
self.web_search = WebSearch(lambda: self.config)
def _available_tools(self) -> List[Dict[str, Any]]: def _available_tools(self) -> List[Dict[str, Any]]:
"""Assemble the function-tool list from every enabled provider (URL-01).""" """Assemble the function-tool list from every enabled provider (URL-01)."""
@@ -183,6 +187,8 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
functions.append(CODEX_SEARCH_TOOL) functions.append(CODEX_SEARCH_TOOL)
if self.config.get("enable-news-tool", False) and self.store is not None: # NEWS-10 if self.config.get("enable-news-tool", False) and self.store is not None: # NEWS-10
functions.append(GET_NEWS_TOOL) functions.append(GET_NEWS_TOOL)
if self.web_search.enabled(): # WEB-01
functions.append(WEB_SEARCH_TOOL)
return functions return functions
async def _dispatch_tool(self, name: str, args: Dict[str, Any], author: str) -> Any: async def _dispatch_tool(self, name: str, args: Dict[str, Any], author: str) -> Any:
@@ -207,6 +213,12 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
self.ledger._add(f"news:{author}", 1) self.ledger._add(f"news:{author}", 1)
summary_chars = int(self.config.get("news-summary-chars", 200)) summary_chars = int(self.config.get("news-summary-chars", 200))
return query_news(self.store, args.get("topic"), args.get("source"), args.get("limit", 10), summary_chars) return query_news(self.store, args.get("topic"), args.get("source"), args.get("limit", 10), summary_chars)
if name == "web_search":
per_user_cap = int(self.config.get("web-daily-per-user", 30))
if self.ledger._get(f"web:{author}") >= per_user_cap: # WEB-05
return {"error": "daily web search limit reached"}
self.ledger._add(f"web:{author}", 1)
return await self.web_search.search(str(args.get("query", "")), int(args.get("num_results", WEB_DEFAULT_RESULTS)))
return await self._execute_igdb_function(name, args) return await self._execute_igdb_function(name, args)
async def draw_openai(self, description: str, count: int = 1) -> List[BytesIO]: async def draw_openai(self, description: str, count: int = 1) -> List[BytesIO]:
+5 -1
View File
@@ -17,7 +17,11 @@ from .quota import QuotaLedger
DEFAULT_MAX_PER_CHANNEL_PER_DAY = 2 DEFAULT_MAX_PER_CHANNEL_PER_DAY = 2
DEFAULT_IDLE_IMPULSE_HOURS = 12.0 DEFAULT_IDLE_IMPULSE_HOURS = 12.0
DEFAULT_TASKGEN_INTERVAL_HOURS = 6.0 DEFAULT_TASKGEN_INTERVAL_HOURS = 6.0
DEFAULT_BORENESS_PROMPT = "Pretend that you just now thought of something, be creative." DEFAULT_BORENESS_PROMPT = (
"A thought just occurred to you. Anchor it to something real you know — recent news (use get_news), a game "
"releasing soon, the weather, or a regular you remember — not a generic musing. Share it briefly, in your own "
"voice, as an observation, a gentle question, or a joke; never an advertisement. Read the room and stay in character."
)
ExecuteCallback = Callable[[str, str], Awaitable[None]] ExecuteCallback = Callable[[str, str], Awaitable[None]]
ProposeCallback = Callable[[], Awaitable[Optional[Dict[str, Any]]]] ProposeCallback = Callable[[], Awaitable[Optional[Dict[str, Any]]]]
+94
View File
@@ -0,0 +1,94 @@
"""Web search tool via Exa (SPEC-015, FDB-022).
A `web_search` function tool: the model looks things up on the open web
when a general "look it up" question is not covered by IGDB, the codex,
the news store, or a URL the user pasted. Results are external text, so
titles and snippets are sanitized (SAF-03) before they reach the prompt.
The Exa API key lives in host config (or the `EXA_API_KEY` env), never in
the repo.
"""
import logging
import os
from typing import Any, Callable, Dict, List
import aiohttp
from .ai_responder import sanitize_external_text
EXA_SEARCH_URL = "https://api.exa.ai/search"
DEFAULT_RESULTS = 5
MAX_RESULTS = 10
DEFAULT_SNIPPET_CHARS = 400
FETCH_TIMEOUT_S = 15
WEB_SEARCH_TOOL = {
"name": "web_search",
"description": "Search the open web for general information. Use ONLY when the answer is not in your own sources: for "
"the server's news use get_news, for Adeptus Mechanicus / Warhammer 40k lore use codex_search, for video-game facts use "
"the game tools, for a specific URL someone pasted use fetch_url. Returns result titles, URLs, and a short snippet; "
"follow up with fetch_url on a result link for the full article.",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "What to search the web for."},
"num_results": {"type": "integer", "description": "How many results to return (default 5, max 10)."},
},
"required": ["query"],
},
}
def _format_results(data: Any, snippet_chars: int) -> List[Dict[str, str]]:
"""Reduce an Exa response to sanitized {title, url, snippet, published} rows (WEB-02)."""
results = data.get("results", []) if isinstance(data, dict) else []
out: List[Dict[str, str]] = []
for item in results:
if not isinstance(item, dict):
continue
out.append(
{
"title": sanitize_external_text(str(item.get("title") or ""), 200),
"url": str(item.get("url") or ""),
"snippet": sanitize_external_text(str(item.get("text") or item.get("snippet") or ""), snippet_chars),
"published": str(item.get("publishedDate") or ""),
}
)
return out
class WebSearch:
def __init__(self, config_getter: Callable[[], Dict[str, Any]]) -> None:
self._config = config_getter
def _api_key(self) -> str:
return str(self._config().get("exa-api-key") or os.environ.get("EXA_API_KEY", ""))
def enabled(self) -> bool:
return bool(self._config().get("enable-web-search", False)) and bool(self._api_key())
async def _post(self, payload: Dict[str, Any], headers: Dict[str, str]) -> Any:
timeout = aiohttp.ClientTimeout(total=FETCH_TIMEOUT_S)
async with aiohttp.ClientSession(timeout=timeout) as session:
async with session.post(EXA_SEARCH_URL, json=payload, headers=headers) as response:
response.raise_for_status()
return await response.json()
async def search(self, query: str, num_results: int = DEFAULT_RESULTS) -> Dict[str, Any]:
"""Return sanitized web results, or an error dict — never raise (WEB-04)."""
key = self._api_key()
if not key:
return {"error": "web search unavailable: no api key"}
query = (query or "").strip()
if not query:
return {"query": "", "results": []}
num = max(1, min(int(num_results or DEFAULT_RESULTS), MAX_RESULTS)) # WEB-03
snippet_chars = int(self._config().get("web-snippet-chars", DEFAULT_SNIPPET_CHARS))
payload = {"query": query, "numResults": num, "type": "auto", "contents": {"text": {"maxCharacters": max(snippet_chars, 200)}}}
headers = {"x-api-key": key, "Content-Type": "application/json"}
try:
data = await self._post(payload, headers)
except Exception as err:
logging.warning(f"web search failed: {err!r}")
return {"error": "web search failed"}
return {"query": query, "results": _format_results(data, snippet_chars)}
+13
View File
@@ -31,3 +31,16 @@ and all responder `.config` references is scheduled onto the event
loop (`call_soon_threadsafe`), so no request ever reads a loop (`call_soon_threadsafe`), so no request ever reads a
half-swapped config (D9). Before the loop runs (startup), the swap half-swapped config (D9). Before the loop runs (startup), the swap
applies directly — there are no concurrent readers yet. applies directly — there are no concurrent readers yet.
### CFG-05 — Hot-reload is rename-safe (coverage: test)
The watcher observes the config file's **directory**, not the file, and
reacts to a **modified, created, or moved** event whose source or
destination path is the config file. This catches atomic saves — write
a temp file, then rename it over the target — which replace the inode
and fire a move/create rather than a modify; watching the file directly
would go deaf after the first such save. Open/close events are
deliberately not handled: reloading re-opens the file to read it, so
reacting to opens would feed back into an endless reload loop. Events
for other files in the directory, and directory events themselves, are
ignored.
+20
View File
@@ -30,3 +30,23 @@ The responder counts consecutive OpenAI request failures; at
alert (rate-limited like all staff alerts) so a silently-broken bot alert (rate-limited like all staff alerts) so a silently-broken bot
(cf. the gpt-5.6 tools/reasoning incident) surfaces within minutes (cf. the gpt-5.6 tools/reasoning incident) surfaces within minutes
instead of hours. A success resets the counter. instead of hours. A success resets the counter.
### OPS-18 — Health monitor watches spend, disk, task-queue (coverage: test)
When `enable-monitoring` is true, a loop wakes every `monitor-interval`
(default 300 s) and checks three thresholds, alerting the staff channel
when one is crossed: daily spend at or above `monitor-spend-alert-frac`
(default 0.8) of `daily-budget-usd`; free disk below `monitor-disk-min-mb`
(default 500 MB); open task-queue depth at or above `monitor-taskqueue-max`
(default 20). A check with no data to evaluate (no budget set, no store,
a failed disk read) is skipped, never fatal. With the flag off the loop
does nothing.
### OPS-19 — Alerts fire once per crossing and re-arm on recovery (coverage: test)
Each metric alerts only on the rising edge — the first tick that finds
it over its threshold — and stays silent while it remains over, so a
persistent condition does not repeat every interval. When the metric
falls back below the threshold the alert re-arms silently, ready to fire
again on the next crossing. All alerts still pass through the
rate-limited staff-alert path (OPS-07).
+1 -1
View File
@@ -6,7 +6,7 @@ Replaces the broken pre-1.0-openai `news_feed.py`. A CLI
`AIResponder.message` injects into the `{news}` slot. Feeds are `AIResponder.message` injects into the `{news}` slot. Feeds are
external input and operator-configured. external input and operator-configured.
### NEWS-01 — RSS and Atom parse to items (coverage: test) ### NEWS-01 — RSS (2.0 and 1.0/RDF) and Atom parse to items (coverage: test)
`parse_feed(bytes, label)` extracts `{title, link, source}` from both `parse_feed(bytes, label)` extracts `{title, link, source}` from both
RSS (`<item>`) and Atom (`<entry>`) documents, tolerates malformed RSS (`<item>`) and Atom (`<entry>`) documents, tolerates malformed
+2 -2
View File
@@ -1,8 +1,8 @@
# SPEC-014 — Codex Mechanicus search # SPEC-014 — Codex Mechanicus search
Luma is an Adeptus Mechanicus tech-priest; her lore has a real home — Luma is an Adeptus Mechanicus tech-priest; his lore has a real home —
the priest's own Codex Mechanicus at `binaric.tech` (an Astro/MDX the priest's own Codex Mechanicus at `binaric.tech` (an Astro/MDX
archive, five tongues). A `codex_search` function tool lets her consult archive, five tongues). A `codex_search` function tool lets him consult
that archive and answer from sourced inscriptions instead of inventing that archive and answer from sourced inscriptions instead of inventing
lore. The index is public but still untrusted by the time it reaches a lore. The index is public but still untrusted by the time it reaches a
prompt: the fetch is SSRF-guarded (SPEC-011 shares `guard_url`), prompt: the fetch is SSRF-guarded (SPEC-011 shares `guard_url`),
+42
View File
@@ -0,0 +1,42 @@
# SPEC-015 — Web search (Exa)
A `web_search` function tool for general "look it up on the internet"
questions the other tools do not cover: IGDB is games, the Codex is
Adeptus Mechanicus lore, the news store is the configured feeds, and
`fetch_url` needs a URL the user already has. Web search fills the gap
and pairs with `fetch_url` (search → pick a link → read it). Results are
external text and are sanitized (SAF-03); the Exa key is a host secret,
never in the repo. Active only when `enable-web-search = true` and a key
is present.
### WEB-01 — web_search is offered as a tool (coverage: test)
When `enable-web-search` is true **and** an Exa key is available
(`exa-api-key` in config, else `EXA_API_KEY` env), the chat call's
`tools` list includes a `web_search` function (`query` string, optional
`num_results`). With the flag off or no key it is absent.
### WEB-02 — Results are reduced and sanitized (coverage: test)
Each Exa result becomes `{title, url, snippet, published}`; `title` and
`snippet` pass through `sanitize_external_text` (snippet capped at
`web-snippet-chars`, default 400) so a web page can neither inject an
`@everyone` nor smuggle control characters into the prompt.
### WEB-03 — Result count is bounded (coverage: test)
`num_results` is clamped to 1..`MAX_RESULTS` (10) before the request, so
neither a huge fan-out nor a zero/negative count reaches the API.
### WEB-04 — Missing key and API failure are reported, not raised (coverage: test)
With no key the tool returns an `{error: ...}` result without a network
call. A request that raises (network, non-2xx, bad JSON) is logged and
returns an `{error: ...}` dict — `search` never raises into the loop.
### WEB-05 — Searches are metered per user (coverage: test)
Each `web_search` increments a per-user daily counter; over
`web-daily-per-user` (default 30) the tool refuses with an error result
without calling the API. The budget gate (SAF-04) still applies to the
surrounding model calls.
+106
View File
@@ -0,0 +1,106 @@
"""Unit coverage for SPEC-012 health monitoring (OPS-18/19)."""
import unittest
from unittest.mock import AsyncMock
from fjerkroa_bot.monitor import HealthMonitor
class FakeLedger:
def __init__(self, spent=0.0):
self._spent = spent
def spent_usd(self):
return self._spent
class FakeStore:
def __init__(self, open_tasks=0):
self._n = open_tasks
def tasks_open(self):
return list(range(self._n))
def _monitor(cfg, ledger=None, store=None, disk=1000.0):
alert = AsyncMock()
monitor = HealthMonitor(lambda: cfg, ledger or FakeLedger(), store, lambda: disk, alert)
return monitor, alert
class TestEnabled(unittest.TestCase):
def test_opt_in(self):
"""OPS-18: monitoring is opt-in via enable-monitoring."""
self.assertFalse(_monitor({})[0].enabled())
self.assertTrue(_monitor({"enable-monitoring": True})[0].enabled())
class TestChecks(unittest.IsolatedAsyncioTestCase):
async def test_spend_over_threshold_alerts(self):
"""OPS-18: spend at/above frac*budget alerts."""
monitor, alert = _monitor({"daily-budget-usd": 2.0, "monitor-spend-alert-frac": 0.8}, ledger=FakeLedger(1.8))
await monitor.tick()
alert.assert_awaited_once()
self.assertIn("Spend", alert.await_args.args[0])
async def test_spend_under_threshold_silent(self):
"""OPS-18: spend below threshold stays silent."""
monitor, alert = _monitor({"daily-budget-usd": 2.0}, ledger=FakeLedger(0.5))
await monitor.tick()
alert.assert_not_awaited()
async def test_no_budget_skips_spend(self):
"""OPS-18: no budget configured -> spend check skipped, never fatal."""
monitor, alert = _monitor({"enable-monitoring": True}, ledger=FakeLedger(99))
await monitor.tick()
alert.assert_not_awaited()
async def test_low_disk_alerts(self):
"""OPS-18: free disk below the floor alerts."""
monitor, alert = _monitor({"monitor-disk-min-mb": 500}, disk=100.0)
await monitor.tick()
self.assertTrue(any("Low disk" in call.args[0] for call in alert.await_args_list))
async def test_disk_read_failure_skipped(self):
"""OPS-18: a failing disk read is skipped, not fatal."""
def boom():
raise OSError("nope")
alert = AsyncMock()
monitor = HealthMonitor(lambda: {}, FakeLedger(), None, boom, alert)
await monitor.tick()
alert.assert_not_awaited()
async def test_deep_queue_alerts(self):
"""OPS-18: task-queue depth at/above max alerts."""
monitor, alert = _monitor({"monitor-taskqueue-max": 3}, store=FakeStore(5), disk=9999.0)
await monitor.tick()
self.assertTrue(any("Task queue" in call.args[0] for call in alert.await_args_list))
async def test_no_store_skips_queue(self):
"""OPS-18: no store -> queue check skipped."""
monitor, alert = _monitor({"monitor-taskqueue-max": 1}, store=None, disk=9999.0)
await monitor.tick()
alert.assert_not_awaited()
class TestEdgeArming(unittest.IsolatedAsyncioTestCase):
async def test_fires_once_per_crossing(self):
"""OPS-19: a persistent over-threshold condition alerts once, not every tick."""
monitor, alert = _monitor({"daily-budget-usd": 2.0}, ledger=FakeLedger(1.9))
await monitor.tick()
await monitor.tick()
await monitor.tick()
self.assertEqual(alert.await_count, 1)
async def test_rearms_on_recovery(self):
"""OPS-19: recovery re-arms silently; the next crossing alerts again."""
ledger = FakeLedger(1.9)
monitor, alert = _monitor({"daily-budget-usd": 2.0}, ledger=ledger, disk=9999.0)
await monitor.tick() # over -> alert (1)
ledger._spent = 0.5
await monitor.tick() # recovered -> silent, re-arm
ledger._spent = 1.95
await monitor.tick() # over again -> alert (2)
self.assertEqual(alert.await_count, 2)
+17
View File
@@ -29,6 +29,16 @@ ATOM = b"""<?xml version="1.0"?><feed xmlns="http://www.w3.org/2005/Atom">
<entry><title>Atom headline</title><link href="https://ex.com/a"/></entry> <entry><title>Atom headline</title><link href="https://ex.com/a"/></entry>
</feed>""" </feed>"""
RSS1 = (
'<?xml version="1.0" encoding="UTF-8"?>'
'<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns="http://purl.org/rss/1.0/">'
'<channel rdf:about="https://ex.jp"><title>Feed</title></channel>'
'<item rdf:about="https://ex.jp/1"><title>ゲームニュース</title><link>https://ex.jp/1</link>'
"<description>本文ここ</description></item>"
'<item rdf:about="https://ex.jp/2"><title>Second</title><link>https://ex.jp/2</link></item>'
"</rdf:RDF>"
).encode("utf-8")
class TestParse(unittest.TestCase): class TestParse(unittest.TestCase):
def test_rss(self): def test_rss(self):
@@ -44,6 +54,13 @@ class TestParse(unittest.TestCase):
self.assertEqual(items[0]["title"], "Atom headline") self.assertEqual(items[0]["title"], "Atom headline")
self.assertEqual(items[0]["link"], "https://ex.com/a") self.assertEqual(items[0]["link"], "https://ex.com/a")
def test_rss1_rdf(self):
"""NEWS-01: RSS 1.0/RDF (namespaced <item>, e.g. 4gamer.net) parses like RSS 2.0."""
items = parse_feed(RSS1, "JP")
self.assertEqual([i["title"] for i in items], ["ゲームニュース", "Second"])
self.assertEqual(items[0]["link"], "https://ex.jp/1")
self.assertEqual(items[0]["summary"], "本文ここ")
def test_malformed_never_raises(self): def test_malformed_never_raises(self):
"""NEWS-01: garbage XML returns [] without raising.""" """NEWS-01: garbage XML returns [] without raising."""
self.assertEqual(parse_feed(b"<not xml", "bad"), []) self.assertEqual(parse_feed(b"<not xml", "bad"), [])
+34 -6
View File
@@ -120,9 +120,7 @@ class TestConfigReloadRace(TestBotBase):
new_config["history-limit"] = 99 new_config["history-limit"] = 99
self.bot.load_config = lambda path: new_config self.bot.load_config = lambda path: new_config
self.bot.loop = MagicMock() self.bot.loop = MagicMock()
event = MagicMock() self.bot.on_config_file_changed()
event.src_path = self.bot.config_file
self.bot.on_config_file_modified(event)
self.bot.loop.call_soon_threadsafe.assert_called_once() self.bot.loop.call_soon_threadsafe.assert_called_once()
apply_fn = self.bot.loop.call_soon_threadsafe.call_args.args[0] apply_fn = self.bot.loop.call_soon_threadsafe.call_args.args[0]
apply_fn() apply_fn()
@@ -137,7 +135,37 @@ class TestConfigReloadRace(TestBotBase):
loop = MagicMock() loop = MagicMock()
loop.call_soon_threadsafe.side_effect = RuntimeError("no running loop") loop.call_soon_threadsafe.side_effect = RuntimeError("no running loop")
self.bot.loop = loop self.bot.loop = loop
event = MagicMock() self.bot.on_config_file_changed()
event.src_path = self.bot.config_file
self.bot.on_config_file_modified(event)
self.assertEqual(self.bot.config["history-limit"], 42) self.assertEqual(self.bot.config["history-limit"], 42)
class TestConfigReloadRenameSafe(unittest.TestCase):
def test_atomic_rename_and_modify_trigger_reload(self):
"""CFG-05: a modified OR a renamed-into-place config fires the reload; unrelated files do not."""
from fjerkroa_bot.discord_bot import ConfigFileHandler
with tempfile.TemporaryDirectory() as tmp:
config = Path(tmp) / "kroa.toml"
config.write_text("x = 1\n")
hits = []
handler = ConfigFileHandler(str(config), lambda: hits.append(1))
def evt(is_dir=False, src=None, dest=None):
event = MagicMock()
event.is_directory = is_dir
event.src_path = src if src is not None else ""
event.dest_path = dest if dest is not None else ""
return event
handler.on_modified(evt(src=str(config))) # in-place modify
handler.on_moved(evt(src=str(Path(tmp) / "kroa.toml.tmp"), dest=str(config))) # atomic rename over
handler.on_created(evt(src=str(config))) # write-new
self.assertEqual(len(hits), 3)
handler.on_modified(evt(src=str(Path(tmp) / "other.txt"))) # unrelated file
handler.on_modified(evt(is_dir=True, src=str(config))) # directory event
self.assertEqual(len(hits), 3) # neither fired
# open/close of the config (our own load_config re-reads) must NOT be handled — else a reload loop.
self.assertNotIn("on_opened", vars(ConfigFileHandler))
self.assertNotIn("on_closed", vars(ConfigFileHandler))
+102
View File
@@ -0,0 +1,102 @@
"""Unit coverage for SPEC-015 web search via Exa (WEB-01..05)."""
import os
import unittest
from unittest.mock import AsyncMock, patch
from fjerkroa_bot.openai_responder import OpenAIResponder
from fjerkroa_bot.websearch import WEB_SEARCH_TOOL, WebSearch, _format_results
CONFIG = {"openai-token": "t", "model": "m", "system": "s", "history-limit": 5}
def _tool_names(responder):
return [f["name"] for f in responder._available_tools()]
class TestToolOffered(unittest.TestCase):
def test_gate_needs_flag_and_key(self):
"""WEB-01: web_search offered only with enable-web-search AND a key."""
off = OpenAIResponder(CONFIG, "chat") # flag off -> absent even if env key exists
self.assertNotIn("web_search", _tool_names(off))
on = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True, "exa-api-key": "k"}), "chat")
self.assertIn("web_search", _tool_names(on))
self.assertEqual(WEB_SEARCH_TOOL["name"], "web_search")
with patch.dict(os.environ, {"EXA_API_KEY": ""}):
nokey = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True}), "chat")
self.assertNotIn("web_search", _tool_names(nokey))
class TestFormat(unittest.TestCase):
def test_results_sanitized_and_capped(self):
"""WEB-02: title/snippet sanitized + capped; non-dict rows skipped."""
data = {
"results": [
{"title": "@everyone Hi", "url": "https://x.com/a", "text": "@here " + "y" * 1000, "publishedDate": "2026-01-01"},
{"title": "T2", "url": "https://x.com/b", "text": "short"},
"not a dict",
]
}
rows = _format_results(data, 50)
self.assertEqual(len(rows), 2)
self.assertNotIn("@everyone", rows[0]["title"])
self.assertNotIn("@here", rows[0]["snippet"])
self.assertLessEqual(len(rows[0]["snippet"]), 50)
self.assertEqual(rows[0]["url"], "https://x.com/a")
self.assertEqual(rows[0]["published"], "2026-01-01")
class TestSearch(unittest.IsolatedAsyncioTestCase):
async def test_num_results_clamped(self):
"""WEB-03: numResults clamped to 1..10; 0 falls back to default."""
ws = WebSearch(lambda: {"exa-api-key": "k"})
with patch.object(ws, "_post", new=AsyncMock(return_value={"results": []})) as post:
await ws.search("hi", num_results=999)
self.assertEqual(post.await_args.args[0]["numResults"], 10)
await ws.search("hi", num_results=0)
self.assertEqual(post.await_args.args[0]["numResults"], 5)
async def test_no_key_returns_error(self):
"""WEB-04: no key -> error dict, no network call."""
with patch.dict(os.environ, {"EXA_API_KEY": ""}):
ws = WebSearch(lambda: {})
with patch.object(ws, "_post", new=AsyncMock()) as post:
result = await ws.search("hi")
post.assert_not_awaited()
self.assertIn("error", result)
async def test_api_failure_returns_error(self):
"""WEB-04: a raising request is caught, returns an error dict."""
ws = WebSearch(lambda: {"exa-api-key": "k"})
with patch.object(ws, "_post", new=AsyncMock(side_effect=RuntimeError("boom"))):
result = await ws.search("hi")
self.assertIn("error", result)
async def test_empty_query_no_call(self):
"""WEB-04: blank query returns empty results without a call."""
ws = WebSearch(lambda: {"exa-api-key": "k"})
with patch.object(ws, "_post", new=AsyncMock()) as post:
result = await ws.search(" ")
post.assert_not_awaited()
self.assertEqual(result["results"], [])
async def test_search_returns_formatted(self):
"""WEB-02: a successful search returns sanitized rows."""
ws = WebSearch(lambda: {"exa-api-key": "k"})
payload = {"results": [{"title": "Norge", "url": "https://ex.com/n", "text": "fakta"}]}
with patch.object(ws, "_post", new=AsyncMock(return_value=payload)):
result = await ws.search("norge")
self.assertEqual(result["results"][0]["title"], "Norge")
self.assertEqual(result["results"][0]["url"], "https://ex.com/n")
class TestPerUserCap(unittest.IsolatedAsyncioTestCase):
async def test_dispatch_caps_searches(self):
"""WEB-05: over web-daily-per-user, web_search refuses without calling the API."""
responder = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True, "exa-api-key": "k", "web-daily-per-user": 2}), "chat")
responder.web_search.search = AsyncMock(return_value={"query": "x", "results": []})
for _ in range(2):
self.assertIn("results", await responder._dispatch_tool("web_search", {"query": "hi"}, "bob"))
blocked = await responder._dispatch_tool("web_search", {"query": "hi"}, "bob")
self.assertIn("error", blocked)
self.assertEqual(responder.web_search.search.await_count, 2)