Compare commits
8 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| da50395dc7 | |||
| ef62cb41a5 | |||
| 7753cc4a07 | |||
| df0bc94489 | |||
| 6b61ed6175 | |||
| 80000528d3 | |||
| cdd5a4cd48 | |||
| 5d01400638 |
@@ -60,6 +60,37 @@ Decisions inside the set architecture. D-NNN, never renumbered.
|
|||||||
broken classifier must never mute the bot; the budget gate already
|
broken classifier must never mute the bot; the budget gate already
|
||||||
bounds spend. Its verdict gates BEFORE the main call, the
|
bounds spend. Its verdict gates BEFORE the main call, the
|
||||||
envelope's answer_needed still gates after — two independent nets.
|
envelope's answer_needed still gates after — two independent nets.
|
||||||
|
- **D-021** — Health monitoring (FDB-012, SPEC-012 OPS-18/19): a
|
||||||
|
separate `monitor_loop` (own cadence, default 300 s) rather than
|
||||||
|
folding checks into the 60 s task loop — monitoring is coarse and
|
||||||
|
should not run every minute. Checks are edge-triggered (alert on the
|
||||||
|
rising edge, re-arm on recovery) so a standing condition never spams;
|
||||||
|
they reuse the existing rate-limited staff-alert path. Metrics are
|
||||||
|
the cheap, high-signal ones (spend vs budget, free disk, task-queue
|
||||||
|
depth); each is independently skippable when it has no data, so a
|
||||||
|
deployment without a budget or store still runs the others. Opt-in
|
||||||
|
(`enable-monitoring`) like every other operational rollout.
|
||||||
|
- **D-020** — Web search via Exa (FDB-022, SPEC-015): a `web_search`
|
||||||
|
tool alongside fetch_url/IGDB/codex/get_news, filling the "look it up
|
||||||
|
on the open web" gap. Exa (not a raw search-engine scrape) because it
|
||||||
|
returns clean title+url+text in one call — no SSRF surface of our own
|
||||||
|
(we call one fixed API endpoint, not arbitrary hosts), and it pairs
|
||||||
|
with fetch_url for the full article. Key is a host secret
|
||||||
|
(`exa-api-key`, env `EXA_API_KEY` fallback), never repo-side; results
|
||||||
|
sanitized like every other external-text tool; off by default
|
||||||
|
(`enable-web-search`), metered per user.
|
||||||
|
- **D-019** — News memory + on-demand tool (SPEC-013 NEWS-07..12):
|
||||||
|
the news pipeline now carries item summaries (feed descriptions,
|
||||||
|
HTML-stripped) and persists every fetched item into a deduped `news`
|
||||||
|
table (schema v6), pruned to a rolling window (`news-keep`). Both the
|
||||||
|
kroa digest run and the ggg posting run write to it, so the store is
|
||||||
|
a single searchable source across both models. A `get_news` tool
|
||||||
|
reads that store (topic/source-filtered, metered, sanitized) rather
|
||||||
|
than re-fetching feeds live: the ambient `{news}` digest stays a
|
||||||
|
small always-on snapshot, while the tool gives unbounded on-demand
|
||||||
|
reach without a fresh network round-trip per call. The store is the
|
||||||
|
same `bot.db` (WAL) the bot uses; the cron process opens it
|
||||||
|
independently — concurrent reader/writer is what WAL is for.
|
||||||
- **D-018** — Codex Mechanicus search (FDB-019, SPEC-014): Luma's
|
- **D-018** — Codex Mechanicus search (FDB-019, SPEC-014): Luma's
|
||||||
lore is grounded in the priest's real archive at binaric.tech via a
|
lore is grounded in the priest's real archive at binaric.tech via a
|
||||||
`codex_search` tool over the site's public `search-index.json`, not
|
`codex_search` tool over the site's public `search-index.json`, not
|
||||||
|
|||||||
+130
@@ -0,0 +1,130 @@
|
|||||||
|
# Operator runbook — Fjærkroa / Luma bot
|
||||||
|
|
||||||
|
One page for "something is wrong, what do I do". Two deployments of one
|
||||||
|
codebase, both on **uberspace** (push-based deploy from the dev machine —
|
||||||
|
there is no git checkout on the hosts).
|
||||||
|
|
||||||
|
| | Fjærkroa (café) | Luma (GGG clan) |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| SSH host | `ssh fjerkroa` (pictor.uberspace.de) | `ssh ggg` |
|
||||||
|
| Service | `kroa` | `luma` |
|
||||||
|
| Config | `~/fjerkroa_bot/kroa.toml` | `~/fjerkroa_bot/ggg.toml` |
|
||||||
|
| Staff channel | `#kassa` | `#mods` |
|
||||||
|
| Language / persona | Norwegian, café host | German, "Luma" |
|
||||||
|
|
||||||
|
Common paths on each host: bot code `~/fjerkroa_bot`, venv `~/venv-bot`,
|
||||||
|
database `~/fjerkroa_bot/history/bot.db` (SQLite, WAL), our snapshots
|
||||||
|
`~/backups/<kroa|luma>/`, logs under `~/logs` and `~/tmp`.
|
||||||
|
|
||||||
|
## From Discord (staff channel only, prefix `!bot`)
|
||||||
|
|
||||||
|
No SSH needed for day-to-day control. Type `!bot help` in the staff
|
||||||
|
channel for the full, grouped list. The essentials:
|
||||||
|
|
||||||
|
- `!bot pause` / `!bot resume` — stop / start all replies.
|
||||||
|
- `!bot quiet <minutes>` — go silent for a while, then auto-resume.
|
||||||
|
- `!bot status` — replies/images/tasks flags + quiet time left.
|
||||||
|
- `!bot spend` — today's estimated USD spend, tokens, images, budget.
|
||||||
|
- `!bot images on|off`, `!bot tasks on|off` — kill-switches.
|
||||||
|
|
||||||
|
`!help` works in **any** channel (for everyone) and lists only what is
|
||||||
|
usable there. `!forgetme` and `!privacy` also work everywhere, even
|
||||||
|
while the bot is paused.
|
||||||
|
|
||||||
|
## Restart / check health (SSH)
|
||||||
|
|
||||||
|
```sh
|
||||||
|
ssh <host>
|
||||||
|
supervisorctl status <kroa|luma> # RUNNING + uptime
|
||||||
|
supervisorctl restart <kroa|luma>
|
||||||
|
tail -n 40 ~/tmp/<kroa|luma>-stderr*.log # discord login / errors
|
||||||
|
tail -n 40 ~/logs/supervisord.log # "We have logged in as ..."
|
||||||
|
```
|
||||||
|
|
||||||
|
A healthy start shows a fresh `connected to Gateway` + `We have logged
|
||||||
|
in as ...` line within ~15 s.
|
||||||
|
|
||||||
|
## Deploy a release / roll back
|
||||||
|
|
||||||
|
From the **dev machine** (`~/Repos/FjerkroaBot`), tags only:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git tag -m "<msg>" vX.Y.Z && git push --tags # cut the release first
|
||||||
|
bash deploy/deploy.sh ggg vX.Y.Z # luma
|
||||||
|
DEPLOY_FORCE=1 bash deploy/deploy.sh fjerkroa vX.Y.Z # kroa (see window)
|
||||||
|
```
|
||||||
|
|
||||||
|
- kroa refuses to deploy **11:00–22:00 Europe/Oslo** (restaurant hours);
|
||||||
|
`DEPLOY_FORCE=1` overrides. Café is closed Mondays.
|
||||||
|
- The script backs up `bot.db` → `bot.db.pre-<tag>` before restart, then
|
||||||
|
smoke-tests (RUNNING + fresh login) and fails loudly if either misses.
|
||||||
|
- **Rollback** = deploy the previous tag. If the schema version moved
|
||||||
|
between the two tags, restore the matching `bot.db.pre-<newtag>` first
|
||||||
|
(see below) so the older code meets a schema it understands.
|
||||||
|
|
||||||
|
## Restore the database
|
||||||
|
|
||||||
|
Three independent daily backup layers exist — pick the freshest good one.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
ssh <host>
|
||||||
|
supervisorctl stop <kroa|luma>
|
||||||
|
DB=~/fjerkroa_bot/history/bot.db
|
||||||
|
|
||||||
|
# 1) uberspace nightly backup of the whole home (read-only):
|
||||||
|
# /backup = current + daily.0..7 + weekly.1..7 (15 restore points)
|
||||||
|
cp /backup/daily.1/home/<user>/fjerkroa_bot/history/bot.db "$DB"
|
||||||
|
|
||||||
|
# 2) our own rotated gzip snapshot (03:17 UTC cron, keep 14):
|
||||||
|
gunzip -c ~/backups/<kroa|luma>/bot-YYYYMMDD-HHMMSS.db.gz > "$DB"
|
||||||
|
|
||||||
|
# 3) the pre-deploy snapshot for a given release:
|
||||||
|
cp "$DB".pre-vX.Y.Z "$DB"
|
||||||
|
|
||||||
|
rm -f "$DB"-wal "$DB"-shm # drop stale WAL sidecars after a restore
|
||||||
|
supervisorctl start <kroa|luma>
|
||||||
|
```
|
||||||
|
|
||||||
|
`<user>` is `fjerkroa` or `ggg`. The DB holds conversation history,
|
||||||
|
structured memory, usage ledger, image cache index, tasks, and the news
|
||||||
|
store — all regenerable, none critical. That is why there is no off-host
|
||||||
|
backup: uberspace `/backup` + the on-host snapshots are enough.
|
||||||
|
|
||||||
|
## Rotate a secret
|
||||||
|
|
||||||
|
Secrets live only in the host `*.toml` (never in the repo). Edit in
|
||||||
|
place and restart:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
ssh <host>
|
||||||
|
# OpenAI: edit openai-token = "sk-..." in kroa.toml / ggg.toml
|
||||||
|
# Discord: edit discord-token = "..." (get a new token from the
|
||||||
|
# Discord developer portal → Bot → Reset Token first)
|
||||||
|
supervisorctl restart <kroa|luma>
|
||||||
|
```
|
||||||
|
|
||||||
|
After rotating an OpenAI key, revoke the old one in the OpenAI dashboard.
|
||||||
|
Keep a `*.toml` backup before editing; a broken TOML crash-loops the
|
||||||
|
service (validate: `~/venv-bot/bin/python -c 'import tomlkit; tomlkit.load(open("kroa.toml"))'`).
|
||||||
|
|
||||||
|
## Scheduled jobs (crontab -l)
|
||||||
|
|
||||||
|
| Host | When (server time) | Job |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| both | `17 3 * * *` | `backup_db.py` → `~/backups/<bot>/` (keep 14) |
|
||||||
|
| kroa | `5 * * * *` | news digest → `{news}` file + news store |
|
||||||
|
| ggg | `*/15 * * * *` | news poster → #news/#newsjp webhooks + store |
|
||||||
|
|
||||||
|
Logs: `~/backups/<bot>/backup.log`, `~/backups/<bot>/news*.log`.
|
||||||
|
|
||||||
|
## Quick triage
|
||||||
|
|
||||||
|
- **Bot silent everywhere** → `!bot status` (paused/quiet?), else
|
||||||
|
`supervisorctl status`; if not RUNNING, `restart` and read stderr.
|
||||||
|
- **Bot silent in one channel** → check the host config `ignore-channels`
|
||||||
|
/ `short-path` rules for that channel (a stray `short-path` rule can
|
||||||
|
archive messages without replying).
|
||||||
|
- **Repeated API errors** → the bot posts a rate-limited alert to the
|
||||||
|
staff channel after 5 consecutive OpenAI failures (OPS-16); check
|
||||||
|
`!bot spend` (budget hit?) and the OpenAI status/key.
|
||||||
|
- **Bad deploy** → roll back to the previous tag (above).
|
||||||
@@ -1,7 +1,7 @@
|
|||||||
"""Codex Mechanicus search tool (SPEC-014, FDB-019).
|
"""Codex Mechanicus search tool (SPEC-014, FDB-019).
|
||||||
|
|
||||||
Luma's own sacred archive — the Codex Mechanicus at binaric.tech — as a
|
Luma's own sacred archive — the Codex Mechanicus at binaric.tech — as a
|
||||||
function tool. She searches the codex index and answers Cult Mechanicus
|
function tool. He searches the codex index and answers Cult Mechanicus
|
||||||
lore from real, sourced inscriptions instead of inventing it. The index
|
lore from real, sourced inscriptions instead of inventing it. The index
|
||||||
is fetched over HTTPS (SSRF-guarded, size-bounded, cached in memory) and
|
is fetched over HTTPS (SSRF-guarded, size-bounded, cached in memory) and
|
||||||
every field returned to the model is sanitized (SAF-03), because even
|
every field returned to the model is sanitized (SAF-03), because even
|
||||||
|
|||||||
@@ -3,9 +3,11 @@ import asyncio
|
|||||||
import logging
|
import logging
|
||||||
import random
|
import random
|
||||||
import re
|
import re
|
||||||
|
import shutil
|
||||||
import sys
|
import sys
|
||||||
import time
|
import time
|
||||||
from collections import deque
|
from collections import deque
|
||||||
|
from pathlib import Path
|
||||||
from typing import Optional, Union
|
from typing import Optional, Union
|
||||||
|
|
||||||
import discord
|
import discord
|
||||||
@@ -16,6 +18,7 @@ from watchdog.events import FileSystemEventHandler
|
|||||||
from watchdog.observers import Observer
|
from watchdog.observers import Observer
|
||||||
|
|
||||||
from .ai_responder import AIMessage
|
from .ai_responder import AIMessage
|
||||||
|
from .monitor import HealthMonitor
|
||||||
from .openai_responder import OpenAIResponder
|
from .openai_responder import OpenAIResponder
|
||||||
from .tasks import TaskEngine
|
from .tasks import TaskEngine
|
||||||
|
|
||||||
@@ -65,11 +68,41 @@ def split_answer(text: str, threshold: int, max_parts: int) -> list:
|
|||||||
|
|
||||||
|
|
||||||
class ConfigFileHandler(FileSystemEventHandler):
|
class ConfigFileHandler(FileSystemEventHandler):
|
||||||
def __init__(self, on_modified):
|
"""Rename-safe config watch (CFG-05).
|
||||||
self._on_modified = on_modified
|
|
||||||
|
|
||||||
|
Editors and tools save atomically — write a temp file, then rename it
|
||||||
|
over the target — which fires a *moved*/*created* event (not
|
||||||
|
*modified*) and swaps the inode, so watching the file directly goes
|
||||||
|
deaf after the first save. We watch the config's *directory* and react
|
||||||
|
to any event whose src or dest path is the config file.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, config_path: str, on_change):
|
||||||
|
self._config_path = str(Path(config_path).resolve())
|
||||||
|
self._on_change = on_change
|
||||||
|
|
||||||
|
def _hits_config(self, event) -> bool:
|
||||||
|
for attr in ("src_path", "dest_path"):
|
||||||
|
path = getattr(event, attr, "")
|
||||||
|
if path and str(Path(path).resolve()) == self._config_path:
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
def _dispatch(self, event):
|
||||||
|
if not event.is_directory and self._hits_config(event):
|
||||||
|
self._on_change()
|
||||||
|
|
||||||
|
# Only write/rename events — NOT on_opened/on_closed, whose read-opens
|
||||||
|
# (our own load_config re-reads the file) would otherwise feed back into
|
||||||
|
# a reload loop (CFG-05).
|
||||||
def on_modified(self, event):
|
def on_modified(self, event):
|
||||||
self._on_modified(event)
|
self._dispatch(event)
|
||||||
|
|
||||||
|
def on_created(self, event):
|
||||||
|
self._dispatch(event)
|
||||||
|
|
||||||
|
def on_moved(self, event):
|
||||||
|
self._dispatch(event)
|
||||||
|
|
||||||
|
|
||||||
class FjerkroaBot(commands.Bot):
|
class FjerkroaBot(commands.Bot):
|
||||||
@@ -99,8 +132,10 @@ class FjerkroaBot(commands.Bot):
|
|||||||
|
|
||||||
def init_observer(self):
|
def init_observer(self):
|
||||||
self.observer = Observer()
|
self.observer = Observer()
|
||||||
self.file_handler = ConfigFileHandler(self.on_config_file_modified)
|
config_path = Path(self.config_file).resolve()
|
||||||
self.observer.schedule(self.file_handler, path=self.config_file, recursive=False)
|
self.file_handler = ConfigFileHandler(str(config_path), self.on_config_file_changed)
|
||||||
|
# Watch the directory, not the file — atomic saves replace the inode (CFG-05)
|
||||||
|
self.observer.schedule(self.file_handler, path=str(config_path.parent), recursive=False)
|
||||||
self.observer.start()
|
self.observer.start()
|
||||||
|
|
||||||
def init_aichannels(self):
|
def init_aichannels(self):
|
||||||
@@ -130,6 +165,15 @@ class FjerkroaBot(commands.Bot):
|
|||||||
observe=self.airesponder.observe_event,
|
observe=self.airesponder.observe_event,
|
||||||
)
|
)
|
||||||
self.loop.create_task(self.task_loop())
|
self.loop.create_task(self.task_loop())
|
||||||
|
# Proactive health monitoring -> staff alerts (OPS-18/19)
|
||||||
|
self.health_monitor = HealthMonitor(
|
||||||
|
config_getter=lambda: self.config,
|
||||||
|
ledger=self.airesponder.ledger,
|
||||||
|
store=self.airesponder.store,
|
||||||
|
disk_free_mb=self._disk_free_mb,
|
||||||
|
alert=self.send_staff_alert,
|
||||||
|
)
|
||||||
|
self.loop.create_task(self.monitor_loop())
|
||||||
logging.info("Task engine initialised.")
|
logging.info("Task engine initialised.")
|
||||||
|
|
||||||
async def task_loop(self):
|
async def task_loop(self):
|
||||||
@@ -140,6 +184,20 @@ class FjerkroaBot(commands.Bot):
|
|||||||
except Exception as err:
|
except Exception as err:
|
||||||
logging.warning(f"task tick failed: {repr(err)}")
|
logging.warning(f"task tick failed: {repr(err)}")
|
||||||
|
|
||||||
|
def _disk_free_mb(self) -> float:
|
||||||
|
directory = Path(self.config.get("history-directory", ".")).expanduser()
|
||||||
|
target = directory if directory.exists() else Path.home()
|
||||||
|
return shutil.disk_usage(target).free / (1024 * 1024)
|
||||||
|
|
||||||
|
async def monitor_loop(self):
|
||||||
|
while True:
|
||||||
|
await asyncio.sleep(int(self.config.get("monitor-interval", 300)))
|
||||||
|
if self.health_monitor.enabled():
|
||||||
|
try:
|
||||||
|
await self.health_monitor.tick()
|
||||||
|
except Exception as err:
|
||||||
|
logging.warning(f"monitor tick failed: {repr(err)}")
|
||||||
|
|
||||||
async def _execute_task(self, channel_name: str, prompt: str) -> None:
|
async def _execute_task(self, channel_name: str, prompt: str) -> None:
|
||||||
"""Run a due task through the normal responder path (TSK-02)."""
|
"""Run a due task through the normal responder path (TSK-02)."""
|
||||||
channel = self.channel_by_name(channel_name, getattr(self, "chat_channel", None), no_ignore=True)
|
channel = self.channel_by_name(channel_name, getattr(self, "chat_channel", None), no_ignore=True)
|
||||||
@@ -389,12 +447,10 @@ class FjerkroaBot(commands.Bot):
|
|||||||
airesponder.image_cache.purge_message(str(message.id)) # IMG-14
|
airesponder.image_cache.purge_message(str(message.id)) # IMG-14
|
||||||
await airesponder.observe_event(message.author.name, "delete", f"deleted: {message.content}")
|
await airesponder.observe_event(message.author.name, "delete", f"deleted: {message.content}")
|
||||||
|
|
||||||
def on_config_file_modified(self, event):
|
def on_config_file_changed(self):
|
||||||
# Runs on the watchdog observer thread — the swap itself is
|
# Runs on the watchdog observer thread — the swap itself is
|
||||||
# scheduled onto the event loop so no request reads a
|
# scheduled onto the event loop so no request reads a
|
||||||
# half-swapped config (CFG-04 / D9)
|
# half-swapped config (CFG-04 / D9)
|
||||||
if event.src_path != self.config_file:
|
|
||||||
return
|
|
||||||
new_config = self.load_config(self.config_file)
|
new_config = self.load_config(self.config_file)
|
||||||
if repr(new_config) == repr(self.config):
|
if repr(new_config) == repr(self.config):
|
||||||
return
|
return
|
||||||
|
|||||||
@@ -0,0 +1,82 @@
|
|||||||
|
"""Proactive health monitoring -> staff alerts (SPEC-012, FDB-012).
|
||||||
|
|
||||||
|
A periodic check that watches daily spend against the budget, free disk,
|
||||||
|
and task-queue depth, and posts a staff alert when a threshold is crossed
|
||||||
|
— once per crossing, re-arming when the metric recovers, so a persistent
|
||||||
|
condition never spams. Opt-in per deployment (`enable-monitoring`); it
|
||||||
|
reuses the rate-limited staff-alert channel (OPS-07).
|
||||||
|
"""
|
||||||
|
|
||||||
|
import logging
|
||||||
|
from typing import Any, Callable, Dict, Optional, Tuple
|
||||||
|
|
||||||
|
# A check returns (metric-name, is-over-threshold, alert-message) or None when not applicable.
|
||||||
|
Check = Optional[Tuple[str, bool, str]]
|
||||||
|
|
||||||
|
|
||||||
|
class HealthMonitor:
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
config_getter: Callable[[], Dict[str, Any]],
|
||||||
|
ledger: Any,
|
||||||
|
store: Any,
|
||||||
|
disk_free_mb: Callable[[], float],
|
||||||
|
alert: Callable[[str], Any],
|
||||||
|
) -> None:
|
||||||
|
self._config = config_getter
|
||||||
|
self._ledger = ledger
|
||||||
|
self._store = store
|
||||||
|
self._disk_free_mb = disk_free_mb
|
||||||
|
self._alert = alert
|
||||||
|
self._armed: Dict[str, bool] = {}
|
||||||
|
|
||||||
|
def enabled(self) -> bool:
|
||||||
|
return bool(self._config().get("enable-monitoring", False))
|
||||||
|
|
||||||
|
def _check_spend(self) -> Check:
|
||||||
|
config = self._config()
|
||||||
|
if "daily-budget-usd" not in config:
|
||||||
|
return None
|
||||||
|
budget = float(config["daily-budget-usd"])
|
||||||
|
if budget <= 0:
|
||||||
|
return None
|
||||||
|
spent = float(self._ledger.spent_usd())
|
||||||
|
frac = spent / budget
|
||||||
|
threshold = float(config.get("monitor-spend-alert-frac", 0.8))
|
||||||
|
return ("spend", frac >= threshold, f"💸 Spend at ${spent:.2f} of ${budget:.2f} today ({frac:.0%}, alert ≥ {threshold:.0%}).")
|
||||||
|
|
||||||
|
def _check_disk(self) -> Check:
|
||||||
|
try:
|
||||||
|
free = float(self._disk_free_mb())
|
||||||
|
except Exception as err:
|
||||||
|
logging.debug(f"monitor: disk check failed: {err!r}")
|
||||||
|
return None
|
||||||
|
min_mb = float(self._config().get("monitor-disk-min-mb", 500))
|
||||||
|
return ("disk", free < min_mb, f"💾 Low disk: {free:.0f} MB free (alert < {min_mb:.0f} MB).")
|
||||||
|
|
||||||
|
def _check_queue(self) -> Check:
|
||||||
|
if self._store is None:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
depth = len(self._store.tasks_open())
|
||||||
|
except Exception as err:
|
||||||
|
logging.debug(f"monitor: queue check failed: {err!r}")
|
||||||
|
return None
|
||||||
|
limit = int(self._config().get("monitor-taskqueue-max", 20))
|
||||||
|
return ("task-queue", depth >= limit, f"🗒️ Task queue deep: {depth} open (alert ≥ {limit}).")
|
||||||
|
|
||||||
|
async def tick(self) -> None:
|
||||||
|
"""Evaluate every check; alert on a rising edge only (OPS-18/19)."""
|
||||||
|
for check in (self._check_spend(), self._check_disk(), self._check_queue()):
|
||||||
|
if check is None:
|
||||||
|
continue
|
||||||
|
metric, over, message = check
|
||||||
|
await self._fire(metric, over, message)
|
||||||
|
|
||||||
|
async def _fire(self, metric: str, over: bool, message: str) -> None:
|
||||||
|
was_over = self._armed.get(metric, False)
|
||||||
|
if over and not was_over:
|
||||||
|
self._armed[metric] = True
|
||||||
|
await self._alert(message)
|
||||||
|
elif not over and was_over:
|
||||||
|
self._armed[metric] = False # recovered — re-arm silently for the next crossing
|
||||||
+124
-16
@@ -11,8 +11,10 @@ CLI: python -m fjerkroa_bot.news --config kroa.toml
|
|||||||
|
|
||||||
import argparse
|
import argparse
|
||||||
import logging
|
import logging
|
||||||
|
import re
|
||||||
import sys
|
import sys
|
||||||
import time
|
import time
|
||||||
|
from html import unescape
|
||||||
from typing import Any, Dict, List, Optional, Tuple
|
from typing import Any, Dict, List, Optional, Tuple
|
||||||
|
|
||||||
import defusedxml.ElementTree as ElementTree # hardened XML: feeds are untrusted (XXE/billion-laughs)
|
import defusedxml.ElementTree as ElementTree # hardened XML: feeds are untrusted (XXE/billion-laughs)
|
||||||
@@ -21,44 +23,68 @@ from .ai_responder import sanitize_external_text
|
|||||||
|
|
||||||
DEFAULT_PER_FEED = 3
|
DEFAULT_PER_FEED = 3
|
||||||
DEFAULT_MAX_ITEMS = 15
|
DEFAULT_MAX_ITEMS = 15
|
||||||
|
DEFAULT_SUMMARY_CHARS = 200
|
||||||
|
DEFAULT_NEWS_KEEP = 400
|
||||||
FETCH_TIMEOUT_S = 15
|
FETCH_TIMEOUT_S = 15
|
||||||
_ATOM = "{http://www.w3.org/2005/Atom}"
|
_ATOM = "{http://www.w3.org/2005/Atom}"
|
||||||
|
_RSS1 = "{http://purl.org/rss/1.0/}" # RSS 1.0 / RDF (e.g. 4gamer.net) namespaces <item>/<title>/<link>
|
||||||
|
_TAG_RE = re.compile(r"<[^>]+>")
|
||||||
|
|
||||||
|
|
||||||
|
def _clean_summary(raw: str, max_len: int = 300) -> str:
|
||||||
|
"""Strip HTML, unescape entities, collapse whitespace (feed descriptions are often HTML)."""
|
||||||
|
text = unescape(_TAG_RE.sub(" ", raw or ""))
|
||||||
|
return re.sub(r"\s+", " ", text).strip()[:max_len]
|
||||||
|
|
||||||
|
|
||||||
|
def _rss_items(root: Any, ns: str, source: str) -> List[Dict[str, str]]:
|
||||||
|
"""RSS 2.0 (ns='') and RSS 1.0/RDF (ns=_RSS1) both use <item><title><link><description>."""
|
||||||
|
out: List[Dict[str, str]] = []
|
||||||
|
for item in root.iter(f"{ns}item"):
|
||||||
|
title = (item.findtext(f"{ns}title") or "").strip()
|
||||||
|
link = (item.findtext(f"{ns}link") or "").strip()
|
||||||
|
summary = _clean_summary(item.findtext(f"{ns}description") or "")
|
||||||
|
if title:
|
||||||
|
out.append({"title": title, "link": link, "source": source, "summary": summary})
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
def parse_feed(data: bytes, source: str = "") -> List[Dict[str, str]]:
|
def parse_feed(data: bytes, source: str = "") -> List[Dict[str, str]]:
|
||||||
"""Parse RSS or Atom bytes into [{title, link, source}] (tolerant)."""
|
"""Parse RSS 2.0, RSS 1.0/RDF, or Atom bytes into [{title, link, source, summary}] (tolerant)."""
|
||||||
try:
|
try:
|
||||||
root = ElementTree.fromstring(data)
|
root = ElementTree.fromstring(data)
|
||||||
except Exception as err:
|
except Exception as err:
|
||||||
# malformed XML or a blocked entity/DTD attack — tolerate, never raise (NEWS-01)
|
# malformed XML or a blocked entity/DTD attack — tolerate, never raise (NEWS-01)
|
||||||
logging.warning(f"news: unparseable/unsafe feed {source!r}: {err!r}")
|
logging.warning(f"news: unparseable/unsafe feed {source!r}: {err!r}")
|
||||||
return []
|
return []
|
||||||
items: List[Dict[str, str]] = []
|
# RSS 2.0 (unqualified) + RSS 1.0/RDF (namespaced, e.g. 4gamer) share <item><title><link><description>
|
||||||
# RSS: <rss><channel><item><title/><link/>
|
items: List[Dict[str, str]] = _rss_items(root, "", source) + _rss_items(root, _RSS1, source)
|
||||||
for item in root.iter("item"):
|
# Atom: <feed><entry><title/><link href=/><summary|content/>
|
||||||
title = (item.findtext("title") or "").strip()
|
|
||||||
link = (item.findtext("link") or "").strip()
|
|
||||||
if title:
|
|
||||||
items.append({"title": title, "link": link, "source": source})
|
|
||||||
# Atom: <feed><entry><title/><link href=/>
|
|
||||||
for entry in root.iter(f"{_ATOM}entry"):
|
for entry in root.iter(f"{_ATOM}entry"):
|
||||||
title = (entry.findtext(f"{_ATOM}title") or "").strip()
|
title = (entry.findtext(f"{_ATOM}title") or "").strip()
|
||||||
link_el = entry.find(f"{_ATOM}link")
|
link_el = entry.find(f"{_ATOM}link")
|
||||||
link = link_el.get("href", "") if link_el is not None else ""
|
link = link_el.get("href", "") if link_el is not None else ""
|
||||||
|
summary = _clean_summary(entry.findtext(f"{_ATOM}summary") or entry.findtext(f"{_ATOM}content") or "")
|
||||||
if title:
|
if title:
|
||||||
items.append({"title": title, "link": link, "source": source})
|
items.append({"title": title, "link": link, "source": source, "summary": summary})
|
||||||
return items
|
return items
|
||||||
|
|
||||||
|
|
||||||
def render_digest(items: List[Dict[str, str]], max_items: int = DEFAULT_MAX_ITEMS) -> str:
|
def render_digest(items: List[Dict[str, str]], max_items: int = DEFAULT_MAX_ITEMS, summary_chars: int = DEFAULT_SUMMARY_CHARS) -> str:
|
||||||
"""Compact sanitized digest for the {news} prompt slot."""
|
"""Compact sanitized digest for the {news} prompt slot (title + short summary + link)."""
|
||||||
lines = []
|
lines = []
|
||||||
for item in items[:max_items]:
|
for item in items[:max_items]:
|
||||||
title = sanitize_external_text(item["title"], 200)
|
title = sanitize_external_text(item["title"], 200)
|
||||||
source = item.get("source", "")
|
source = item.get("source", "")
|
||||||
link = item.get("link", "")
|
link = item.get("link", "")
|
||||||
|
summary = sanitize_external_text(item.get("summary", ""), summary_chars) if summary_chars else ""
|
||||||
prefix = f"[{source}] " if source else ""
|
prefix = f"[{source}] " if source else ""
|
||||||
lines.append(f"- {prefix}{title}" + (f" ({link})" if link else ""))
|
line = f"- {prefix}{title}"
|
||||||
|
if summary:
|
||||||
|
line += f" — {summary}"
|
||||||
|
if link:
|
||||||
|
line += f" ({link})"
|
||||||
|
lines.append(line)
|
||||||
return "\n".join(lines)
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
@@ -102,10 +128,11 @@ def item_key(item: Dict[str, str]) -> str:
|
|||||||
class NewsPoster:
|
class NewsPoster:
|
||||||
"""Post NEW feed items to Discord channel webhooks (ggg model, SPEC-013 NEWS-04..06)."""
|
"""Post NEW feed items to Discord channel webhooks (ggg model, SPEC-013 NEWS-04..06)."""
|
||||||
|
|
||||||
def __init__(self, guard, fetch_bytes, post_webhook) -> None:
|
def __init__(self, guard, fetch_bytes, post_webhook, store: Any = None) -> None:
|
||||||
self._guard = guard
|
self._guard = guard
|
||||||
self._fetch_bytes = fetch_bytes
|
self._fetch_bytes = fetch_bytes
|
||||||
self._post_webhook = post_webhook
|
self._post_webhook = post_webhook
|
||||||
|
self._store = store
|
||||||
|
|
||||||
async def run_post(
|
async def run_post(
|
||||||
self,
|
self,
|
||||||
@@ -115,9 +142,11 @@ class NewsPoster:
|
|||||||
per_feed: int,
|
per_feed: int,
|
||||||
max_per_run: int,
|
max_per_run: int,
|
||||||
seed_only: bool,
|
seed_only: bool,
|
||||||
|
keep: int = DEFAULT_NEWS_KEEP,
|
||||||
) -> Tuple[int, set]:
|
) -> Tuple[int, set]:
|
||||||
"""Returns (posted_count, updated_seen). seed_only marks new items seen without posting."""
|
"""Returns (posted_count, updated_seen). seed_only marks new items seen without posting."""
|
||||||
posted = 0
|
posted = 0
|
||||||
|
harvested: List[Dict[str, str]] = []
|
||||||
for url, label, channel in feeds:
|
for url, label, channel in feeds:
|
||||||
reason = self._guard(url)
|
reason = self._guard(url)
|
||||||
if reason:
|
if reason:
|
||||||
@@ -129,6 +158,7 @@ class NewsPoster:
|
|||||||
logging.warning(f"news-post: fetch failed for {label}: {repr(err)}")
|
logging.warning(f"news-post: fetch failed for {label}: {repr(err)}")
|
||||||
continue
|
continue
|
||||||
for item in parse_feed(data, label)[:per_feed]:
|
for item in parse_feed(data, label)[:per_feed]:
|
||||||
|
harvested.append(item) # NEWS-09: everything parsed feeds the searchable store
|
||||||
key = item_key(item)
|
key = item_key(item)
|
||||||
if not key or key in seen:
|
if not key or key in seen:
|
||||||
continue
|
continue
|
||||||
@@ -136,6 +166,9 @@ class NewsPoster:
|
|||||||
may_post = not seed_only and posted < max_per_run
|
may_post = not seed_only and posted < max_per_run
|
||||||
if may_post and await self._deliver(item, label, channel, webhooks):
|
if may_post and await self._deliver(item, label, channel, webhooks):
|
||||||
posted += 1
|
posted += 1
|
||||||
|
if self._store is not None and harvested:
|
||||||
|
self._store.add_news_items(harvested)
|
||||||
|
self._store.prune_news(keep)
|
||||||
return posted, seen
|
return posted, seen
|
||||||
|
|
||||||
async def _deliver(self, item: Dict[str, str], label: str, channel: str, webhooks: Dict[str, str]) -> bool:
|
async def _deliver(self, item: Dict[str, str], label: str, channel: str, webhooks: Dict[str, str]) -> bool:
|
||||||
@@ -154,6 +187,77 @@ class NewsPoster:
|
|||||||
return False
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
# --- news memory + on-demand retrieval tool (SPEC-013 NEWS-09..12) ---
|
||||||
|
|
||||||
|
GET_NEWS_TOOL = {
|
||||||
|
"name": "get_news",
|
||||||
|
"description": "Fetch news the bot has collected from its RSS feeds — this is the SAME news that gets posted in the "
|
||||||
|
"server's news channels (e.g. #news, #newsjp / ニュース). Use this FIRST, before web_search, for anything about "
|
||||||
|
"current news or about something someone saw in a news channel; filter by topic (a keyword, also matches the source "
|
||||||
|
"label) or by source. Returns headlines with a short summary and a link; follow up with fetch_url for the full text.",
|
||||||
|
"parameters": {
|
||||||
|
"type": "object",
|
||||||
|
"properties": {
|
||||||
|
"topic": {"type": "string", "description": "Optional keywords to filter by, e.g. 'Nordland', 'football', 'weather'."},
|
||||||
|
"source": {"type": "string", "description": "Optional source label, e.g. 'NRK', 'Aftenposten', 'Verden', 'Sport'."},
|
||||||
|
"limit": {"type": "integer", "description": "How many items to return (default 10, max 30)."},
|
||||||
|
},
|
||||||
|
"required": [],
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _news_terms(topic: Optional[str]) -> List[str]:
|
||||||
|
return [t for t in re.split(r"\W+", (topic or "").lower()) if len(t) > 1][:5]
|
||||||
|
|
||||||
|
|
||||||
|
def query_news(
|
||||||
|
store: Any, topic: Optional[str] = None, source: Optional[str] = None, limit: int = 10, summary_chars: int = DEFAULT_SUMMARY_CHARS
|
||||||
|
) -> Dict[str, Any]:
|
||||||
|
"""Retrieve stored news for the get_news tool: term/source-filtered, sanitized (NEWS-11)."""
|
||||||
|
if store is None:
|
||||||
|
return {"error": "news store unavailable"}
|
||||||
|
limit = max(1, min(int(limit or 10), 30))
|
||||||
|
src = (str(source).strip() or None) if source else None
|
||||||
|
terms = _news_terms(topic)
|
||||||
|
try:
|
||||||
|
rows = store.search_news(terms, limit, src) if terms else store.recent_news(limit, src)
|
||||||
|
except Exception as err:
|
||||||
|
logging.warning(f"news: query failed: {err!r}")
|
||||||
|
return {"error": "news lookup failed"}
|
||||||
|
results = [
|
||||||
|
{
|
||||||
|
"source": row.get("source", ""),
|
||||||
|
"title": sanitize_external_text(row.get("title", ""), 200),
|
||||||
|
"summary": sanitize_external_text(row.get("summary", ""), summary_chars),
|
||||||
|
"link": row.get("link", ""),
|
||||||
|
}
|
||||||
|
for row in rows
|
||||||
|
]
|
||||||
|
return {"topic": topic or "", "source": src or "", "results": results}
|
||||||
|
|
||||||
|
|
||||||
|
def _open_store(config: Dict[str, Any]) -> Any:
|
||||||
|
directory = config.get("history-directory")
|
||||||
|
if not directory:
|
||||||
|
return None
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from .persistence import PersistentStore
|
||||||
|
|
||||||
|
return PersistentStore(Path(str(directory)).expanduser() / "bot.db")
|
||||||
|
|
||||||
|
|
||||||
|
def persist_news(config: Dict[str, Any], items: List[Dict[str, str]]) -> int:
|
||||||
|
"""Upsert fetched items into the news store, prune to the rolling window (NEWS-09)."""
|
||||||
|
store = _open_store(config)
|
||||||
|
if store is None or not items:
|
||||||
|
return 0
|
||||||
|
added = store.add_news_items(items)
|
||||||
|
store.prune_news(int(config.get("news-keep", DEFAULT_NEWS_KEEP)))
|
||||||
|
return added
|
||||||
|
|
||||||
|
|
||||||
def load_seen(path: str) -> Tuple[set, bool]:
|
def load_seen(path: str) -> Tuple[set, bool]:
|
||||||
"""(seen-set, existed). Missing/broken state -> empty set, existed=False (seed run)."""
|
"""(seen-set, existed). Missing/broken state -> empty set, existed=False (seed run)."""
|
||||||
import json
|
import json
|
||||||
@@ -231,7 +335,7 @@ async def run_post(config: Dict[str, Any]) -> int:
|
|||||||
logging.error("news-post: need news-post-webhooks and news-post-feeds")
|
logging.error("news-post: need news-post-webhooks and news-post-feeds")
|
||||||
return 0
|
return 0
|
||||||
seen, existed = load_seen(state_path)
|
seen, existed = load_seen(state_path)
|
||||||
poster = NewsPoster(guard_url, _aiohttp_fetch, _aiohttp_post)
|
poster = NewsPoster(guard_url, _aiohttp_fetch, _aiohttp_post, store=_open_store(config))
|
||||||
posted, seen = await poster.run_post(
|
posted, seen = await poster.run_post(
|
||||||
feeds,
|
feeds,
|
||||||
webhooks,
|
webhooks,
|
||||||
@@ -239,6 +343,7 @@ async def run_post(config: Dict[str, Any]) -> int:
|
|||||||
int(config.get("news-post-per-feed", DEFAULT_POST_PER_FEED)),
|
int(config.get("news-post-per-feed", DEFAULT_POST_PER_FEED)),
|
||||||
int(config.get("news-post-max-per-run", DEFAULT_POST_MAX_PER_RUN)),
|
int(config.get("news-post-max-per-run", DEFAULT_POST_MAX_PER_RUN)),
|
||||||
seed_only=not existed, # first run seeds without flooding the channels
|
seed_only=not existed, # first run seeds without flooding the channels
|
||||||
|
keep=int(config.get("news-keep", DEFAULT_NEWS_KEEP)),
|
||||||
)
|
)
|
||||||
save_seen(state_path, seen, int(config.get("news-post-seen-cap", DEFAULT_SEEN_CAP)))
|
save_seen(state_path, seen, int(config.get("news-post-seen-cap", DEFAULT_SEEN_CAP)))
|
||||||
logging.info(f"news-post: posted {posted} item(s)" + (" (seed run — nothing posted)" if not existed else ""))
|
logging.info(f"news-post: posted {posted} item(s)" + (" (seed run — nothing posted)" if not existed else ""))
|
||||||
@@ -258,7 +363,10 @@ async def run(config: Dict[str, Any]) -> Optional[str]:
|
|||||||
return None
|
return None
|
||||||
fetcher = NewsFetcher(guard_url, _aiohttp_fetch)
|
fetcher = NewsFetcher(guard_url, _aiohttp_fetch)
|
||||||
items = await fetcher.collect(feeds, int(config.get("news-per-feed", DEFAULT_PER_FEED)))
|
items = await fetcher.collect(feeds, int(config.get("news-per-feed", DEFAULT_PER_FEED)))
|
||||||
digest = render_digest(items, int(config.get("news-max-items", DEFAULT_MAX_ITEMS)))
|
persist_news(config, items) # NEWS-09: feed the searchable rolling store for get_news
|
||||||
|
digest = render_digest(
|
||||||
|
items, int(config.get("news-max-items", DEFAULT_MAX_ITEMS)), int(config.get("news-summary-chars", DEFAULT_SUMMARY_CHARS))
|
||||||
|
)
|
||||||
header = f"News as of {time.strftime('%Y-%m-%d %H:%M UTC', time.gmtime())}:\n"
|
header = f"News as of {time.strftime('%Y-%m-%d %H:%M UTC', time.gmtime())}:\n"
|
||||||
with open(out_path, "w", encoding="utf-8") as fd:
|
with open(out_path, "w", encoding="utf-8") as fd:
|
||||||
fd.write(header + digest + "\n")
|
fd.write(header + digest + "\n")
|
||||||
|
|||||||
@@ -14,8 +14,11 @@ from .codex import DEFAULT_LIMIT as CODEX_DEFAULT_LIMIT
|
|||||||
from .codex import CodexSearch
|
from .codex import CodexSearch
|
||||||
from .igdblib import IGDBQuery
|
from .igdblib import IGDBQuery
|
||||||
from .leonardo_draw import LeonardoAIDrawMixIn
|
from .leonardo_draw import LeonardoAIDrawMixIn
|
||||||
|
from .news import GET_NEWS_TOOL, query_news
|
||||||
from .quota import QuotaLedger
|
from .quota import QuotaLedger
|
||||||
from .url_reader import FETCH_URL_TOOL, URLReader
|
from .url_reader import FETCH_URL_TOOL, URLReader
|
||||||
|
from .websearch import DEFAULT_RESULTS as WEB_DEFAULT_RESULTS
|
||||||
|
from .websearch import WEB_SEARCH_TOOL, WebSearch
|
||||||
|
|
||||||
# The response envelope, enforced server-side via structured outputs
|
# The response envelope, enforced server-side via structured outputs
|
||||||
# (ENV-19). All fields required, closed object, nullable where the
|
# (ENV-19). All fields required, closed object, nullable where the
|
||||||
@@ -165,6 +168,8 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
|
|||||||
self.url_reader = URLReader(lambda: self.config, self.image_cache)
|
self.url_reader = URLReader(lambda: self.config, self.image_cache)
|
||||||
# Codex Mechanicus search (SPEC-014); Luma's own archive at binaric.tech
|
# Codex Mechanicus search (SPEC-014); Luma's own archive at binaric.tech
|
||||||
self.codex = CodexSearch(lambda: self.config)
|
self.codex = CodexSearch(lambda: self.config)
|
||||||
|
# Web search (SPEC-015) via Exa; general "look it up" beyond fetch_url/news/codex
|
||||||
|
self.web_search = WebSearch(lambda: self.config)
|
||||||
|
|
||||||
def _available_tools(self) -> List[Dict[str, Any]]:
|
def _available_tools(self) -> List[Dict[str, Any]]:
|
||||||
"""Assemble the function-tool list from every enabled provider (URL-01)."""
|
"""Assemble the function-tool list from every enabled provider (URL-01)."""
|
||||||
@@ -180,6 +185,10 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
|
|||||||
functions.append(FETCH_URL_TOOL)
|
functions.append(FETCH_URL_TOOL)
|
||||||
if self.codex.enabled(): # CDX-01
|
if self.codex.enabled(): # CDX-01
|
||||||
functions.append(CODEX_SEARCH_TOOL)
|
functions.append(CODEX_SEARCH_TOOL)
|
||||||
|
if self.config.get("enable-news-tool", False) and self.store is not None: # NEWS-10
|
||||||
|
functions.append(GET_NEWS_TOOL)
|
||||||
|
if self.web_search.enabled(): # WEB-01
|
||||||
|
functions.append(WEB_SEARCH_TOOL)
|
||||||
return functions
|
return functions
|
||||||
|
|
||||||
async def _dispatch_tool(self, name: str, args: Dict[str, Any], author: str) -> Any:
|
async def _dispatch_tool(self, name: str, args: Dict[str, Any], author: str) -> Any:
|
||||||
@@ -197,6 +206,19 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
|
|||||||
self.ledger._add(f"codex:{author}", 1)
|
self.ledger._add(f"codex:{author}", 1)
|
||||||
limit = int(self.config.get("codex-limit", CODEX_DEFAULT_LIMIT))
|
limit = int(self.config.get("codex-limit", CODEX_DEFAULT_LIMIT))
|
||||||
return await self.codex.search(str(args.get("query", "")), str(args.get("lang", "en")), limit)
|
return await self.codex.search(str(args.get("query", "")), str(args.get("lang", "en")), limit)
|
||||||
|
if name == "get_news":
|
||||||
|
per_user_cap = int(self.config.get("news-daily-per-user", 30))
|
||||||
|
if self.ledger._get(f"news:{author}") >= per_user_cap: # NEWS-12
|
||||||
|
return {"error": "daily news lookup limit reached"}
|
||||||
|
self.ledger._add(f"news:{author}", 1)
|
||||||
|
summary_chars = int(self.config.get("news-summary-chars", 200))
|
||||||
|
return query_news(self.store, args.get("topic"), args.get("source"), args.get("limit", 10), summary_chars)
|
||||||
|
if name == "web_search":
|
||||||
|
per_user_cap = int(self.config.get("web-daily-per-user", 30))
|
||||||
|
if self.ledger._get(f"web:{author}") >= per_user_cap: # WEB-05
|
||||||
|
return {"error": "daily web search limit reached"}
|
||||||
|
self.ledger._add(f"web:{author}", 1)
|
||||||
|
return await self.web_search.search(str(args.get("query", "")), int(args.get("num_results", WEB_DEFAULT_RESULTS)))
|
||||||
return await self._execute_igdb_function(name, args)
|
return await self._execute_igdb_function(name, args)
|
||||||
|
|
||||||
async def draw_openai(self, description: str, count: int = 1) -> List[BytesIO]:
|
async def draw_openai(self, description: str, count: int = 1) -> List[BytesIO]:
|
||||||
|
|||||||
@@ -13,7 +13,7 @@ from contextlib import closing
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any, Dict, List, Optional
|
from typing import Any, Dict, List, Optional
|
||||||
|
|
||||||
SCHEMA_VERSION = 5
|
SCHEMA_VERSION = 6
|
||||||
|
|
||||||
|
|
||||||
class PersistentStore:
|
class PersistentStore:
|
||||||
@@ -72,6 +72,13 @@ class PersistentStore:
|
|||||||
" due_at TEXT NOT NULL, payload TEXT NOT NULL, state TEXT NOT NULL DEFAULT 'queued',"
|
" due_at TEXT NOT NULL, payload TEXT NOT NULL, state TEXT NOT NULL DEFAULT 'queued',"
|
||||||
" created_at TEXT NOT NULL DEFAULT (datetime('now')), executed_at TEXT)"
|
" created_at TEXT NOT NULL DEFAULT (datetime('now')), executed_at TEXT)"
|
||||||
)
|
)
|
||||||
|
if version < 6:
|
||||||
|
# News memory (SPEC-013 NEWS-09): deduped rolling store of fetched items
|
||||||
|
conn.execute(
|
||||||
|
"CREATE TABLE IF NOT EXISTS news (id INTEGER PRIMARY KEY, dedup_key TEXT UNIQUE NOT NULL,"
|
||||||
|
" source TEXT NOT NULL DEFAULT '', title TEXT NOT NULL, link TEXT NOT NULL DEFAULT '',"
|
||||||
|
" summary TEXT NOT NULL DEFAULT '', first_seen TEXT NOT NULL DEFAULT (datetime('now')))"
|
||||||
|
)
|
||||||
if version < SCHEMA_VERSION:
|
if version < SCHEMA_VERSION:
|
||||||
conn.execute(f"PRAGMA user_version = {SCHEMA_VERSION}")
|
conn.execute(f"PRAGMA user_version = {SCHEMA_VERSION}")
|
||||||
os.chmod(self.db_path, 0o600) # conversation data (PER-04)
|
os.chmod(self.db_path, 0o600) # conversation data (PER-04)
|
||||||
@@ -115,6 +122,64 @@ class PersistentStore:
|
|||||||
row = conn.execute("SELECT value FROM usage WHERE day = ? AND key = ?", (day, key)).fetchone()
|
row = conn.execute("SELECT value FROM usage WHERE day = ? AND key = ?", (day, key)).fetchone()
|
||||||
return float(row[0]) if row else 0.0
|
return float(row[0]) if row else 0.0
|
||||||
|
|
||||||
|
# --- news memory (SPEC-013 NEWS-09..12) ---
|
||||||
|
|
||||||
|
def add_news_items(self, items: List[Dict[str, Any]]) -> int:
|
||||||
|
"""Insert deduped news rows (by link or title); returns how many were new (NEWS-09)."""
|
||||||
|
added = 0
|
||||||
|
with closing(self._connect()) as conn, conn:
|
||||||
|
for item in items:
|
||||||
|
title = str(item.get("title") or "").strip()
|
||||||
|
key = (str(item.get("link") or "").strip()) or title
|
||||||
|
if not title or not key:
|
||||||
|
continue
|
||||||
|
cursor = conn.execute(
|
||||||
|
"INSERT OR IGNORE INTO news (dedup_key, source, title, link, summary) VALUES (?, ?, ?, ?, ?)",
|
||||||
|
(key, str(item.get("source") or ""), title, str(item.get("link") or ""), str(item.get("summary") or "")),
|
||||||
|
)
|
||||||
|
added += cursor.rowcount
|
||||||
|
return added
|
||||||
|
|
||||||
|
def recent_news(self, limit: int = 20, source: Optional[str] = None) -> List[Dict[str, Any]]:
|
||||||
|
sql = "SELECT source, title, link, summary FROM news"
|
||||||
|
params: List[Any] = []
|
||||||
|
if source:
|
||||||
|
sql += " WHERE source = ?"
|
||||||
|
params.append(source)
|
||||||
|
sql += " ORDER BY id DESC LIMIT ?"
|
||||||
|
params.append(int(limit))
|
||||||
|
with closing(self._connect()) as conn:
|
||||||
|
rows = conn.execute(sql, params).fetchall()
|
||||||
|
return [{"source": r[0], "title": r[1], "link": r[2], "summary": r[3]} for r in rows]
|
||||||
|
|
||||||
|
def search_news(self, terms: List[str], limit: int = 20, source: Optional[str] = None) -> List[Dict[str, Any]]:
|
||||||
|
"""Rows where every term appears in title or summary; optional source filter (NEWS-11)."""
|
||||||
|
params: List[Any] = []
|
||||||
|
clauses = []
|
||||||
|
for term in terms:
|
||||||
|
clauses.append("(title LIKE ? OR summary LIKE ? OR source LIKE ?)")
|
||||||
|
like = f"%{term}%"
|
||||||
|
params += [like, like, like]
|
||||||
|
where = " AND ".join(clauses) if clauses else "1=1"
|
||||||
|
if source:
|
||||||
|
where = f"({where}) AND source = ?"
|
||||||
|
params.append(source)
|
||||||
|
params.append(int(limit))
|
||||||
|
sql = f"SELECT source, title, link, summary FROM news WHERE {where} ORDER BY id DESC LIMIT ?" # nosec B608 - fixed templates; values parameterised
|
||||||
|
with closing(self._connect()) as conn:
|
||||||
|
rows = conn.execute(sql, params).fetchall()
|
||||||
|
return [{"source": r[0], "title": r[1], "link": r[2], "summary": r[3]} for r in rows]
|
||||||
|
|
||||||
|
def prune_news(self, keep: int) -> int:
|
||||||
|
"""Keep the newest `keep` rows, delete the rest (rolling window, NEWS-09)."""
|
||||||
|
with closing(self._connect()) as conn, conn:
|
||||||
|
cursor = conn.execute("DELETE FROM news WHERE id NOT IN (SELECT id FROM news ORDER BY id DESC LIMIT ?)", (int(keep),))
|
||||||
|
return cursor.rowcount
|
||||||
|
|
||||||
|
def news_count(self) -> int:
|
||||||
|
with closing(self._connect()) as conn:
|
||||||
|
return int(conn.execute("SELECT COUNT(*) FROM news").fetchone()[0])
|
||||||
|
|
||||||
def delete_history_of_user(self, user: str) -> int:
|
def delete_history_of_user(self, user: str) -> int:
|
||||||
"""Remove persisted rows carrying this user's messages (SAF-08)."""
|
"""Remove persisted rows carrying this user's messages (SAF-08)."""
|
||||||
with closing(self._connect()) as conn, conn:
|
with closing(self._connect()) as conn, conn:
|
||||||
|
|||||||
@@ -17,7 +17,11 @@ from .quota import QuotaLedger
|
|||||||
DEFAULT_MAX_PER_CHANNEL_PER_DAY = 2
|
DEFAULT_MAX_PER_CHANNEL_PER_DAY = 2
|
||||||
DEFAULT_IDLE_IMPULSE_HOURS = 12.0
|
DEFAULT_IDLE_IMPULSE_HOURS = 12.0
|
||||||
DEFAULT_TASKGEN_INTERVAL_HOURS = 6.0
|
DEFAULT_TASKGEN_INTERVAL_HOURS = 6.0
|
||||||
DEFAULT_BORENESS_PROMPT = "Pretend that you just now thought of something, be creative."
|
DEFAULT_BORENESS_PROMPT = (
|
||||||
|
"A thought just occurred to you. Anchor it to something real you know — recent news (use get_news), a game "
|
||||||
|
"releasing soon, the weather, or a regular you remember — not a generic musing. Share it briefly, in your own "
|
||||||
|
"voice, as an observation, a gentle question, or a joke; never an advertisement. Read the room and stay in character."
|
||||||
|
)
|
||||||
|
|
||||||
ExecuteCallback = Callable[[str, str], Awaitable[None]]
|
ExecuteCallback = Callable[[str, str], Awaitable[None]]
|
||||||
ProposeCallback = Callable[[], Awaitable[Optional[Dict[str, Any]]]]
|
ProposeCallback = Callable[[], Awaitable[Optional[Dict[str, Any]]]]
|
||||||
|
|||||||
@@ -0,0 +1,94 @@
|
|||||||
|
"""Web search tool via Exa (SPEC-015, FDB-022).
|
||||||
|
|
||||||
|
A `web_search` function tool: the model looks things up on the open web
|
||||||
|
when a general "look it up" question is not covered by IGDB, the codex,
|
||||||
|
the news store, or a URL the user pasted. Results are external text, so
|
||||||
|
titles and snippets are sanitized (SAF-03) before they reach the prompt.
|
||||||
|
The Exa API key lives in host config (or the `EXA_API_KEY` env), never in
|
||||||
|
the repo.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
from typing import Any, Callable, Dict, List
|
||||||
|
|
||||||
|
import aiohttp
|
||||||
|
|
||||||
|
from .ai_responder import sanitize_external_text
|
||||||
|
|
||||||
|
EXA_SEARCH_URL = "https://api.exa.ai/search"
|
||||||
|
DEFAULT_RESULTS = 5
|
||||||
|
MAX_RESULTS = 10
|
||||||
|
DEFAULT_SNIPPET_CHARS = 400
|
||||||
|
FETCH_TIMEOUT_S = 15
|
||||||
|
|
||||||
|
WEB_SEARCH_TOOL = {
|
||||||
|
"name": "web_search",
|
||||||
|
"description": "Search the open web for general information. Use ONLY when the answer is not in your own sources: for "
|
||||||
|
"the server's news use get_news, for Adeptus Mechanicus / Warhammer 40k lore use codex_search, for video-game facts use "
|
||||||
|
"the game tools, for a specific URL someone pasted use fetch_url. Returns result titles, URLs, and a short snippet; "
|
||||||
|
"follow up with fetch_url on a result link for the full article.",
|
||||||
|
"parameters": {
|
||||||
|
"type": "object",
|
||||||
|
"properties": {
|
||||||
|
"query": {"type": "string", "description": "What to search the web for."},
|
||||||
|
"num_results": {"type": "integer", "description": "How many results to return (default 5, max 10)."},
|
||||||
|
},
|
||||||
|
"required": ["query"],
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _format_results(data: Any, snippet_chars: int) -> List[Dict[str, str]]:
|
||||||
|
"""Reduce an Exa response to sanitized {title, url, snippet, published} rows (WEB-02)."""
|
||||||
|
results = data.get("results", []) if isinstance(data, dict) else []
|
||||||
|
out: List[Dict[str, str]] = []
|
||||||
|
for item in results:
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
continue
|
||||||
|
out.append(
|
||||||
|
{
|
||||||
|
"title": sanitize_external_text(str(item.get("title") or ""), 200),
|
||||||
|
"url": str(item.get("url") or ""),
|
||||||
|
"snippet": sanitize_external_text(str(item.get("text") or item.get("snippet") or ""), snippet_chars),
|
||||||
|
"published": str(item.get("publishedDate") or ""),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
class WebSearch:
|
||||||
|
def __init__(self, config_getter: Callable[[], Dict[str, Any]]) -> None:
|
||||||
|
self._config = config_getter
|
||||||
|
|
||||||
|
def _api_key(self) -> str:
|
||||||
|
return str(self._config().get("exa-api-key") or os.environ.get("EXA_API_KEY", ""))
|
||||||
|
|
||||||
|
def enabled(self) -> bool:
|
||||||
|
return bool(self._config().get("enable-web-search", False)) and bool(self._api_key())
|
||||||
|
|
||||||
|
async def _post(self, payload: Dict[str, Any], headers: Dict[str, str]) -> Any:
|
||||||
|
timeout = aiohttp.ClientTimeout(total=FETCH_TIMEOUT_S)
|
||||||
|
async with aiohttp.ClientSession(timeout=timeout) as session:
|
||||||
|
async with session.post(EXA_SEARCH_URL, json=payload, headers=headers) as response:
|
||||||
|
response.raise_for_status()
|
||||||
|
return await response.json()
|
||||||
|
|
||||||
|
async def search(self, query: str, num_results: int = DEFAULT_RESULTS) -> Dict[str, Any]:
|
||||||
|
"""Return sanitized web results, or an error dict — never raise (WEB-04)."""
|
||||||
|
key = self._api_key()
|
||||||
|
if not key:
|
||||||
|
return {"error": "web search unavailable: no api key"}
|
||||||
|
query = (query or "").strip()
|
||||||
|
if not query:
|
||||||
|
return {"query": "", "results": []}
|
||||||
|
num = max(1, min(int(num_results or DEFAULT_RESULTS), MAX_RESULTS)) # WEB-03
|
||||||
|
snippet_chars = int(self._config().get("web-snippet-chars", DEFAULT_SNIPPET_CHARS))
|
||||||
|
payload = {"query": query, "numResults": num, "type": "auto", "contents": {"text": {"maxCharacters": max(snippet_chars, 200)}}}
|
||||||
|
headers = {"x-api-key": key, "Content-Type": "application/json"}
|
||||||
|
try:
|
||||||
|
data = await self._post(payload, headers)
|
||||||
|
except Exception as err:
|
||||||
|
logging.warning(f"web search failed: {err!r}")
|
||||||
|
return {"error": "web search failed"}
|
||||||
|
return {"query": query, "results": _format_results(data, snippet_chars)}
|
||||||
@@ -31,3 +31,16 @@ and all responder `.config` references is scheduled onto the event
|
|||||||
loop (`call_soon_threadsafe`), so no request ever reads a
|
loop (`call_soon_threadsafe`), so no request ever reads a
|
||||||
half-swapped config (D9). Before the loop runs (startup), the swap
|
half-swapped config (D9). Before the loop runs (startup), the swap
|
||||||
applies directly — there are no concurrent readers yet.
|
applies directly — there are no concurrent readers yet.
|
||||||
|
|
||||||
|
### CFG-05 — Hot-reload is rename-safe (coverage: test)
|
||||||
|
|
||||||
|
The watcher observes the config file's **directory**, not the file, and
|
||||||
|
reacts to a **modified, created, or moved** event whose source or
|
||||||
|
destination path is the config file. This catches atomic saves — write
|
||||||
|
a temp file, then rename it over the target — which replace the inode
|
||||||
|
and fire a move/create rather than a modify; watching the file directly
|
||||||
|
would go deaf after the first such save. Open/close events are
|
||||||
|
deliberately not handled: reloading re-opens the file to read it, so
|
||||||
|
reacting to opens would feed back into an endless reload loop. Events
|
||||||
|
for other files in the directory, and directory events themselves, are
|
||||||
|
ignored.
|
||||||
|
|||||||
@@ -30,3 +30,23 @@ The responder counts consecutive OpenAI request failures; at
|
|||||||
alert (rate-limited like all staff alerts) so a silently-broken bot
|
alert (rate-limited like all staff alerts) so a silently-broken bot
|
||||||
(cf. the gpt-5.6 tools/reasoning incident) surfaces within minutes
|
(cf. the gpt-5.6 tools/reasoning incident) surfaces within minutes
|
||||||
instead of hours. A success resets the counter.
|
instead of hours. A success resets the counter.
|
||||||
|
|
||||||
|
### OPS-18 — Health monitor watches spend, disk, task-queue (coverage: test)
|
||||||
|
|
||||||
|
When `enable-monitoring` is true, a loop wakes every `monitor-interval`
|
||||||
|
(default 300 s) and checks three thresholds, alerting the staff channel
|
||||||
|
when one is crossed: daily spend at or above `monitor-spend-alert-frac`
|
||||||
|
(default 0.8) of `daily-budget-usd`; free disk below `monitor-disk-min-mb`
|
||||||
|
(default 500 MB); open task-queue depth at or above `monitor-taskqueue-max`
|
||||||
|
(default 20). A check with no data to evaluate (no budget set, no store,
|
||||||
|
a failed disk read) is skipped, never fatal. With the flag off the loop
|
||||||
|
does nothing.
|
||||||
|
|
||||||
|
### OPS-19 — Alerts fire once per crossing and re-arm on recovery (coverage: test)
|
||||||
|
|
||||||
|
Each metric alerts only on the rising edge — the first tick that finds
|
||||||
|
it over its threshold — and stays silent while it remains over, so a
|
||||||
|
persistent condition does not repeat every interval. When the metric
|
||||||
|
falls back below the threshold the alert re-arms silently, ready to fire
|
||||||
|
again on the next crossing. All alerts still pass through the
|
||||||
|
rate-limited staff-alert path (OPS-07).
|
||||||
|
|||||||
+47
-1
@@ -6,7 +6,7 @@ Replaces the broken pre-1.0-openai `news_feed.py`. A CLI
|
|||||||
`AIResponder.message` injects into the `{news}` slot. Feeds are
|
`AIResponder.message` injects into the `{news}` slot. Feeds are
|
||||||
external input and operator-configured.
|
external input and operator-configured.
|
||||||
|
|
||||||
### NEWS-01 — RSS and Atom parse to items (coverage: test)
|
### NEWS-01 — RSS (2.0 and 1.0/RDF) and Atom parse to items (coverage: test)
|
||||||
|
|
||||||
`parse_feed(bytes, label)` extracts `{title, link, source}` from both
|
`parse_feed(bytes, label)` extracts `{title, link, source}` from both
|
||||||
RSS (`<item>`) and Atom (`<entry>`) documents, tolerates malformed
|
RSS (`<item>`) and Atom (`<entry>`) documents, tolerates malformed
|
||||||
@@ -52,3 +52,49 @@ whose channel has no configured webhook, and a webhook POST that
|
|||||||
raises are each logged and skipped — one failure never sinks the
|
raises are each logged and skipped — one failure never sinks the
|
||||||
run, and the seen-set still advances for successfully-processed
|
run, and the seen-set still advances for successfully-processed
|
||||||
items.
|
items.
|
||||||
|
|
||||||
|
### NEWS-07 — Item summaries are extracted (coverage: test)
|
||||||
|
|
||||||
|
`parse_feed` also captures each item's short description — RSS
|
||||||
|
`<description>`, Atom `<summary>` or `<content>` — with HTML stripped,
|
||||||
|
entities unescaped, and whitespace collapsed, so an item carries what
|
||||||
|
it is about, not only a headline. Missing descriptions yield an empty
|
||||||
|
summary, never an error.
|
||||||
|
|
||||||
|
### NEWS-08 — The digest carries summaries (coverage: test)
|
||||||
|
|
||||||
|
`render_digest` appends the sanitized, length-capped
|
||||||
|
(`news-summary-chars`, default 200) summary after each headline, so
|
||||||
|
the bot's ambient `{news}` context knows the gist of each story, not
|
||||||
|
just its title. A zero cap restores the title-only digest.
|
||||||
|
|
||||||
|
### NEWS-09 — Fetched news is stored, deduped, and rolled over (coverage: test)
|
||||||
|
|
||||||
|
Both the digest run (kroa) and the posting run (ggg) upsert every
|
||||||
|
fetched item into a `news` table keyed by link (or title), so the same
|
||||||
|
story is stored once. After each run the store is pruned to the newest
|
||||||
|
`news-keep` rows (default 400), a rolling window that bounds growth
|
||||||
|
while keeping recent history searchable.
|
||||||
|
|
||||||
|
### NEWS-10 — get_news is offered as a tool (coverage: test)
|
||||||
|
|
||||||
|
When `enable-news-tool` is true and a store is configured, the chat
|
||||||
|
call's `tools` list includes a `get_news` function (optional `topic`,
|
||||||
|
`source`, `limit`) next to the other tools. Without a store or the
|
||||||
|
flag it is absent.
|
||||||
|
|
||||||
|
### NEWS-11 — get_news retrieves filtered, sanitized items (coverage: test)
|
||||||
|
|
||||||
|
`get_news` returns recent stored items, newest first, optionally
|
||||||
|
narrowed by `topic` (every keyword must appear in the title, summary,
|
||||||
|
or source label — so `topic: "Nordland"` finds items from that source)
|
||||||
|
and/or an exact `source`; `limit` is clamped to 1..30. Each result's title and
|
||||||
|
summary are passed through `sanitize_external_text`. The bot can then
|
||||||
|
`fetch_url` a returned link for the full article.
|
||||||
|
|
||||||
|
### NEWS-12 — get_news is metered per user (coverage: test)
|
||||||
|
|
||||||
|
Each `get_news` call increments a per-user daily counter; over
|
||||||
|
`news-daily-per-user` (default 30) the tool refuses with an error
|
||||||
|
result without touching the store. The budget gate (SAF-04) still
|
||||||
|
applies to the surrounding model calls.
|
||||||
|
|||||||
@@ -1,8 +1,8 @@
|
|||||||
# SPEC-014 — Codex Mechanicus search
|
# SPEC-014 — Codex Mechanicus search
|
||||||
|
|
||||||
Luma is an Adeptus Mechanicus tech-priest; her lore has a real home —
|
Luma is an Adeptus Mechanicus tech-priest; his lore has a real home —
|
||||||
the priest's own Codex Mechanicus at `binaric.tech` (an Astro/MDX
|
the priest's own Codex Mechanicus at `binaric.tech` (an Astro/MDX
|
||||||
archive, five tongues). A `codex_search` function tool lets her consult
|
archive, five tongues). A `codex_search` function tool lets him consult
|
||||||
that archive and answer from sourced inscriptions instead of inventing
|
that archive and answer from sourced inscriptions instead of inventing
|
||||||
lore. The index is public but still untrusted by the time it reaches a
|
lore. The index is public but still untrusted by the time it reaches a
|
||||||
prompt: the fetch is SSRF-guarded (SPEC-011 shares `guard_url`),
|
prompt: the fetch is SSRF-guarded (SPEC-011 shares `guard_url`),
|
||||||
|
|||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# SPEC-015 — Web search (Exa)
|
||||||
|
|
||||||
|
A `web_search` function tool for general "look it up on the internet"
|
||||||
|
questions the other tools do not cover: IGDB is games, the Codex is
|
||||||
|
Adeptus Mechanicus lore, the news store is the configured feeds, and
|
||||||
|
`fetch_url` needs a URL the user already has. Web search fills the gap
|
||||||
|
and pairs with `fetch_url` (search → pick a link → read it). Results are
|
||||||
|
external text and are sanitized (SAF-03); the Exa key is a host secret,
|
||||||
|
never in the repo. Active only when `enable-web-search = true` and a key
|
||||||
|
is present.
|
||||||
|
|
||||||
|
### WEB-01 — web_search is offered as a tool (coverage: test)
|
||||||
|
|
||||||
|
When `enable-web-search` is true **and** an Exa key is available
|
||||||
|
(`exa-api-key` in config, else `EXA_API_KEY` env), the chat call's
|
||||||
|
`tools` list includes a `web_search` function (`query` string, optional
|
||||||
|
`num_results`). With the flag off or no key it is absent.
|
||||||
|
|
||||||
|
### WEB-02 — Results are reduced and sanitized (coverage: test)
|
||||||
|
|
||||||
|
Each Exa result becomes `{title, url, snippet, published}`; `title` and
|
||||||
|
`snippet` pass through `sanitize_external_text` (snippet capped at
|
||||||
|
`web-snippet-chars`, default 400) so a web page can neither inject an
|
||||||
|
`@everyone` nor smuggle control characters into the prompt.
|
||||||
|
|
||||||
|
### WEB-03 — Result count is bounded (coverage: test)
|
||||||
|
|
||||||
|
`num_results` is clamped to 1..`MAX_RESULTS` (10) before the request, so
|
||||||
|
neither a huge fan-out nor a zero/negative count reaches the API.
|
||||||
|
|
||||||
|
### WEB-04 — Missing key and API failure are reported, not raised (coverage: test)
|
||||||
|
|
||||||
|
With no key the tool returns an `{error: ...}` result without a network
|
||||||
|
call. A request that raises (network, non-2xx, bad JSON) is logged and
|
||||||
|
returns an `{error: ...}` dict — `search` never raises into the loop.
|
||||||
|
|
||||||
|
### WEB-05 — Searches are metered per user (coverage: test)
|
||||||
|
|
||||||
|
Each `web_search` increments a per-user daily counter; over
|
||||||
|
`web-daily-per-user` (default 30) the tool refuses with an error result
|
||||||
|
without calling the API. The budget gate (SAF-04) still applies to the
|
||||||
|
surrounding model calls.
|
||||||
@@ -0,0 +1,106 @@
|
|||||||
|
"""Unit coverage for SPEC-012 health monitoring (OPS-18/19)."""
|
||||||
|
|
||||||
|
import unittest
|
||||||
|
from unittest.mock import AsyncMock
|
||||||
|
|
||||||
|
from fjerkroa_bot.monitor import HealthMonitor
|
||||||
|
|
||||||
|
|
||||||
|
class FakeLedger:
|
||||||
|
def __init__(self, spent=0.0):
|
||||||
|
self._spent = spent
|
||||||
|
|
||||||
|
def spent_usd(self):
|
||||||
|
return self._spent
|
||||||
|
|
||||||
|
|
||||||
|
class FakeStore:
|
||||||
|
def __init__(self, open_tasks=0):
|
||||||
|
self._n = open_tasks
|
||||||
|
|
||||||
|
def tasks_open(self):
|
||||||
|
return list(range(self._n))
|
||||||
|
|
||||||
|
|
||||||
|
def _monitor(cfg, ledger=None, store=None, disk=1000.0):
|
||||||
|
alert = AsyncMock()
|
||||||
|
monitor = HealthMonitor(lambda: cfg, ledger or FakeLedger(), store, lambda: disk, alert)
|
||||||
|
return monitor, alert
|
||||||
|
|
||||||
|
|
||||||
|
class TestEnabled(unittest.TestCase):
|
||||||
|
def test_opt_in(self):
|
||||||
|
"""OPS-18: monitoring is opt-in via enable-monitoring."""
|
||||||
|
self.assertFalse(_monitor({})[0].enabled())
|
||||||
|
self.assertTrue(_monitor({"enable-monitoring": True})[0].enabled())
|
||||||
|
|
||||||
|
|
||||||
|
class TestChecks(unittest.IsolatedAsyncioTestCase):
|
||||||
|
async def test_spend_over_threshold_alerts(self):
|
||||||
|
"""OPS-18: spend at/above frac*budget alerts."""
|
||||||
|
monitor, alert = _monitor({"daily-budget-usd": 2.0, "monitor-spend-alert-frac": 0.8}, ledger=FakeLedger(1.8))
|
||||||
|
await monitor.tick()
|
||||||
|
alert.assert_awaited_once()
|
||||||
|
self.assertIn("Spend", alert.await_args.args[0])
|
||||||
|
|
||||||
|
async def test_spend_under_threshold_silent(self):
|
||||||
|
"""OPS-18: spend below threshold stays silent."""
|
||||||
|
monitor, alert = _monitor({"daily-budget-usd": 2.0}, ledger=FakeLedger(0.5))
|
||||||
|
await monitor.tick()
|
||||||
|
alert.assert_not_awaited()
|
||||||
|
|
||||||
|
async def test_no_budget_skips_spend(self):
|
||||||
|
"""OPS-18: no budget configured -> spend check skipped, never fatal."""
|
||||||
|
monitor, alert = _monitor({"enable-monitoring": True}, ledger=FakeLedger(99))
|
||||||
|
await monitor.tick()
|
||||||
|
alert.assert_not_awaited()
|
||||||
|
|
||||||
|
async def test_low_disk_alerts(self):
|
||||||
|
"""OPS-18: free disk below the floor alerts."""
|
||||||
|
monitor, alert = _monitor({"monitor-disk-min-mb": 500}, disk=100.0)
|
||||||
|
await monitor.tick()
|
||||||
|
self.assertTrue(any("Low disk" in call.args[0] for call in alert.await_args_list))
|
||||||
|
|
||||||
|
async def test_disk_read_failure_skipped(self):
|
||||||
|
"""OPS-18: a failing disk read is skipped, not fatal."""
|
||||||
|
|
||||||
|
def boom():
|
||||||
|
raise OSError("nope")
|
||||||
|
|
||||||
|
alert = AsyncMock()
|
||||||
|
monitor = HealthMonitor(lambda: {}, FakeLedger(), None, boom, alert)
|
||||||
|
await monitor.tick()
|
||||||
|
alert.assert_not_awaited()
|
||||||
|
|
||||||
|
async def test_deep_queue_alerts(self):
|
||||||
|
"""OPS-18: task-queue depth at/above max alerts."""
|
||||||
|
monitor, alert = _monitor({"monitor-taskqueue-max": 3}, store=FakeStore(5), disk=9999.0)
|
||||||
|
await monitor.tick()
|
||||||
|
self.assertTrue(any("Task queue" in call.args[0] for call in alert.await_args_list))
|
||||||
|
|
||||||
|
async def test_no_store_skips_queue(self):
|
||||||
|
"""OPS-18: no store -> queue check skipped."""
|
||||||
|
monitor, alert = _monitor({"monitor-taskqueue-max": 1}, store=None, disk=9999.0)
|
||||||
|
await monitor.tick()
|
||||||
|
alert.assert_not_awaited()
|
||||||
|
|
||||||
|
|
||||||
|
class TestEdgeArming(unittest.IsolatedAsyncioTestCase):
|
||||||
|
async def test_fires_once_per_crossing(self):
|
||||||
|
"""OPS-19: a persistent over-threshold condition alerts once, not every tick."""
|
||||||
|
monitor, alert = _monitor({"daily-budget-usd": 2.0}, ledger=FakeLedger(1.9))
|
||||||
|
await monitor.tick()
|
||||||
|
await monitor.tick()
|
||||||
|
await monitor.tick()
|
||||||
|
self.assertEqual(alert.await_count, 1)
|
||||||
|
|
||||||
|
async def test_rearms_on_recovery(self):
|
||||||
|
"""OPS-19: recovery re-arms silently; the next crossing alerts again."""
|
||||||
|
ledger = FakeLedger(1.9)
|
||||||
|
monitor, alert = _monitor({"daily-budget-usd": 2.0}, ledger=ledger, disk=9999.0)
|
||||||
|
await monitor.tick() # over -> alert (1)
|
||||||
|
ledger._spent = 0.5
|
||||||
|
await monitor.tick() # recovered -> silent, re-arm
|
||||||
|
ledger._spent = 1.95
|
||||||
|
await monitor.tick() # over again -> alert (2)
|
||||||
|
self.assertEqual(alert.await_count, 2)
|
||||||
+163
-4
@@ -1,9 +1,24 @@
|
|||||||
"""Unit coverage for SPEC-013 news digest (NEWS-01..03)."""
|
"""Unit coverage for SPEC-013 news digest + memory + tool (NEWS-01..12)."""
|
||||||
|
|
||||||
|
import tempfile
|
||||||
import unittest
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
from unittest.mock import AsyncMock
|
from unittest.mock import AsyncMock
|
||||||
|
|
||||||
from fjerkroa_bot.news import NewsFetcher, NewsPoster, load_seen, parse_feed, render_digest, save_seen
|
from fjerkroa_bot.news import (
|
||||||
|
GET_NEWS_TOOL,
|
||||||
|
NewsFetcher,
|
||||||
|
NewsPoster,
|
||||||
|
load_seen,
|
||||||
|
parse_feed,
|
||||||
|
query_news,
|
||||||
|
render_digest,
|
||||||
|
save_seen,
|
||||||
|
)
|
||||||
|
from fjerkroa_bot.openai_responder import OpenAIResponder
|
||||||
|
from fjerkroa_bot.persistence import PersistentStore
|
||||||
|
|
||||||
|
CONFIG = {"openai-token": "t", "model": "m", "system": "s", "history-limit": 5}
|
||||||
|
|
||||||
RSS = b"""<?xml version="1.0"?><rss><channel>
|
RSS = b"""<?xml version="1.0"?><rss><channel>
|
||||||
<item><title>Game X released</title><link>https://ex.com/x</link></item>
|
<item><title>Game X released</title><link>https://ex.com/x</link></item>
|
||||||
@@ -14,6 +29,16 @@ ATOM = b"""<?xml version="1.0"?><feed xmlns="http://www.w3.org/2005/Atom">
|
|||||||
<entry><title>Atom headline</title><link href="https://ex.com/a"/></entry>
|
<entry><title>Atom headline</title><link href="https://ex.com/a"/></entry>
|
||||||
</feed>"""
|
</feed>"""
|
||||||
|
|
||||||
|
RSS1 = (
|
||||||
|
'<?xml version="1.0" encoding="UTF-8"?>'
|
||||||
|
'<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns="http://purl.org/rss/1.0/">'
|
||||||
|
'<channel rdf:about="https://ex.jp"><title>Feed</title></channel>'
|
||||||
|
'<item rdf:about="https://ex.jp/1"><title>ゲームニュース</title><link>https://ex.jp/1</link>'
|
||||||
|
"<description>本文ここ</description></item>"
|
||||||
|
'<item rdf:about="https://ex.jp/2"><title>Second</title><link>https://ex.jp/2</link></item>'
|
||||||
|
"</rdf:RDF>"
|
||||||
|
).encode("utf-8")
|
||||||
|
|
||||||
|
|
||||||
class TestParse(unittest.TestCase):
|
class TestParse(unittest.TestCase):
|
||||||
def test_rss(self):
|
def test_rss(self):
|
||||||
@@ -29,6 +54,13 @@ class TestParse(unittest.TestCase):
|
|||||||
self.assertEqual(items[0]["title"], "Atom headline")
|
self.assertEqual(items[0]["title"], "Atom headline")
|
||||||
self.assertEqual(items[0]["link"], "https://ex.com/a")
|
self.assertEqual(items[0]["link"], "https://ex.com/a")
|
||||||
|
|
||||||
|
def test_rss1_rdf(self):
|
||||||
|
"""NEWS-01: RSS 1.0/RDF (namespaced <item>, e.g. 4gamer.net) parses like RSS 2.0."""
|
||||||
|
items = parse_feed(RSS1, "JP")
|
||||||
|
self.assertEqual([i["title"] for i in items], ["ゲームニュース", "Second"])
|
||||||
|
self.assertEqual(items[0]["link"], "https://ex.jp/1")
|
||||||
|
self.assertEqual(items[0]["summary"], "本文ここ")
|
||||||
|
|
||||||
def test_malformed_never_raises(self):
|
def test_malformed_never_raises(self):
|
||||||
"""NEWS-01: garbage XML returns [] without raising."""
|
"""NEWS-01: garbage XML returns [] without raising."""
|
||||||
self.assertEqual(parse_feed(b"<not xml", "bad"), [])
|
self.assertEqual(parse_feed(b"<not xml", "bad"), [])
|
||||||
@@ -166,10 +198,137 @@ class TestSeenState(unittest.TestCase):
|
|||||||
def test_cap_bounds_state(self):
|
def test_cap_bounds_state(self):
|
||||||
"""NEWS-05: save keeps at most `cap` keys."""
|
"""NEWS-05: save keeps at most `cap` keys."""
|
||||||
import json
|
import json
|
||||||
import tempfile
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
with tempfile.TemporaryDirectory() as tmp:
|
with tempfile.TemporaryDirectory() as tmp:
|
||||||
path = str(Path(tmp) / "state.json")
|
path = str(Path(tmp) / "state.json")
|
||||||
save_seen(path, {f"k{i}" for i in range(100)}, cap=10)
|
save_seen(path, {f"k{i}" for i in range(100)}, cap=10)
|
||||||
self.assertEqual(len(json.load(open(path))), 10)
|
self.assertEqual(len(json.load(open(path))), 10)
|
||||||
|
|
||||||
|
|
||||||
|
RSS_DESC = b"""<?xml version="1.0"?><rss><channel>
|
||||||
|
<item><title>Storm hits coast</title><link>https://ex.com/s</link>
|
||||||
|
<description><p>Heavy <b>wind</b> expected</p></description></item>
|
||||||
|
</channel></rss>"""
|
||||||
|
|
||||||
|
ATOM_SUM = b"""<?xml version="1.0"?><feed xmlns="http://www.w3.org/2005/Atom">
|
||||||
|
<entry><title>Atom T</title><link href="https://ex.com/a"/><summary>Short gist here</summary></entry>
|
||||||
|
</feed>"""
|
||||||
|
|
||||||
|
|
||||||
|
class TestSummaries(unittest.TestCase):
|
||||||
|
def test_rss_description_stripped(self):
|
||||||
|
"""NEWS-07: RSS description parsed, HTML stripped, entities unescaped, whitespace collapsed."""
|
||||||
|
items = parse_feed(RSS_DESC, "S")
|
||||||
|
self.assertEqual(items[0]["summary"], "Heavy wind expected")
|
||||||
|
|
||||||
|
def test_atom_summary(self):
|
||||||
|
"""NEWS-07: Atom summary collapsed to clean text."""
|
||||||
|
items = parse_feed(ATOM_SUM, "A")
|
||||||
|
self.assertEqual(items[0]["summary"], "Short gist here")
|
||||||
|
|
||||||
|
def test_missing_description_is_empty(self):
|
||||||
|
"""NEWS-07: no description -> empty summary, never an error."""
|
||||||
|
self.assertEqual(parse_feed(RSS, "S")[0]["summary"], "")
|
||||||
|
|
||||||
|
def test_digest_carries_summary(self):
|
||||||
|
"""NEWS-08: digest appends the sanitized capped summary; zero cap = title only."""
|
||||||
|
items = [{"title": "T", "link": "https://ex.com/x", "source": "NRK", "summary": "the gist of it"}]
|
||||||
|
digest = render_digest(items, 10, 100)
|
||||||
|
self.assertIn("[NRK]", digest)
|
||||||
|
self.assertIn("the gist of it", digest)
|
||||||
|
self.assertNotIn("the gist", render_digest(items, 10, 0)) # zero cap -> title only
|
||||||
|
|
||||||
|
|
||||||
|
class NewsStoreBase(unittest.TestCase):
|
||||||
|
def setUp(self):
|
||||||
|
self.tmp = tempfile.TemporaryDirectory()
|
||||||
|
self.addCleanup(self.tmp.cleanup)
|
||||||
|
self.store = PersistentStore(Path(self.tmp.name) / "bot.db")
|
||||||
|
|
||||||
|
|
||||||
|
class TestNewsStore(NewsStoreBase):
|
||||||
|
def test_dedup_and_rolling_window(self):
|
||||||
|
"""NEWS-09: items deduped by link; prune keeps the newest N."""
|
||||||
|
first = [
|
||||||
|
{"title": "A", "link": "L1", "source": "S", "summary": "sa"},
|
||||||
|
{"title": "B", "link": "L2", "source": "S", "summary": "sb"},
|
||||||
|
]
|
||||||
|
self.assertEqual(self.store.add_news_items(first), 2)
|
||||||
|
self.assertEqual(self.store.add_news_items([dict(first[0])]), 0) # dup link ignored
|
||||||
|
self.assertEqual(self.store.news_count(), 2)
|
||||||
|
self.store.prune_news(1)
|
||||||
|
self.assertEqual(self.store.news_count(), 1)
|
||||||
|
self.assertEqual(self.store.recent_news(5)[0]["title"], "B") # newest survives
|
||||||
|
|
||||||
|
def test_dedup_by_title_when_no_link(self):
|
||||||
|
"""NEWS-09: linkless items dedup on title."""
|
||||||
|
self.store.add_news_items([{"title": "Same", "link": "", "source": "S", "summary": ""}])
|
||||||
|
self.store.add_news_items([{"title": "Same", "link": "", "source": "S", "summary": ""}])
|
||||||
|
self.assertEqual(self.store.news_count(), 1)
|
||||||
|
|
||||||
|
|
||||||
|
class TestQueryNews(NewsStoreBase):
|
||||||
|
def seed(self):
|
||||||
|
self.store.add_news_items(
|
||||||
|
[
|
||||||
|
{"title": "Nordland storm", "link": "L1", "source": "Nordland", "summary": "strong wind on the coast"},
|
||||||
|
{"title": "Oslo budget", "link": "L2", "source": "NRK", "summary": "@everyone spending plan"},
|
||||||
|
{"title": "Sport result", "link": "L3", "source": "Sport", "summary": "the match ended"},
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_topic_filter(self):
|
||||||
|
"""NEWS-11: topic keywords must appear in title or summary."""
|
||||||
|
self.seed()
|
||||||
|
res = query_news(self.store, topic="storm")
|
||||||
|
self.assertEqual([r["title"] for r in res["results"]], ["Nordland storm"])
|
||||||
|
|
||||||
|
def test_topic_matches_source_label(self):
|
||||||
|
"""NEWS-11: topic also matches the source label, so 'Nordland' finds regional items."""
|
||||||
|
self.store.add_news_items([{"title": "Ferry delayed", "link": "LX", "source": "Nordland", "summary": "boat late"}])
|
||||||
|
res = query_news(self.store, topic="Nordland")
|
||||||
|
self.assertTrue(any(r["link"] == "LX" for r in res["results"])) # matched via source, not title/summary
|
||||||
|
|
||||||
|
def test_source_filter_and_sanitize(self):
|
||||||
|
"""NEWS-11: source narrows results; title/summary are sanitized."""
|
||||||
|
self.seed()
|
||||||
|
res = query_news(self.store, source="NRK")
|
||||||
|
self.assertTrue(res["results"] and all(r["source"] == "NRK" for r in res["results"]))
|
||||||
|
self.assertNotIn("@everyone", res["results"][0]["summary"])
|
||||||
|
|
||||||
|
def test_limit_clamped_and_no_store(self):
|
||||||
|
"""NEWS-11: limit clamps to 1..30; a missing store returns an error."""
|
||||||
|
self.seed()
|
||||||
|
self.assertLessEqual(len(query_news(self.store, limit=999)["results"]), 30)
|
||||||
|
self.assertGreaterEqual(len(query_news(self.store, limit=0)["results"]), 1)
|
||||||
|
self.assertIn("error", query_news(None))
|
||||||
|
|
||||||
|
|
||||||
|
class TestNewsTool(unittest.IsolatedAsyncioTestCase):
|
||||||
|
def setUp(self):
|
||||||
|
self.tmp = tempfile.TemporaryDirectory()
|
||||||
|
self.addCleanup(self.tmp.cleanup)
|
||||||
|
|
||||||
|
def _responder(self, **extra):
|
||||||
|
cfg = dict(CONFIG, **{"history-directory": self.tmp.name}, **extra)
|
||||||
|
return OpenAIResponder(cfg, "chat")
|
||||||
|
|
||||||
|
def test_tool_offered_needs_flag_and_store(self):
|
||||||
|
"""NEWS-10: get_news offered only with enable-news-tool AND a store."""
|
||||||
|
no_store = OpenAIResponder(dict(CONFIG, **{"enable-news-tool": True}), "chat")
|
||||||
|
self.assertIsNone(no_store.store)
|
||||||
|
self.assertNotIn("get_news", [f["name"] for f in no_store._available_tools()])
|
||||||
|
flag_off = self._responder()
|
||||||
|
self.assertNotIn("get_news", [f["name"] for f in flag_off._available_tools()])
|
||||||
|
on = self._responder(**{"enable-news-tool": True})
|
||||||
|
self.assertIn("get_news", [f["name"] for f in on._available_tools()])
|
||||||
|
self.assertEqual(GET_NEWS_TOOL["name"], "get_news")
|
||||||
|
|
||||||
|
async def test_dispatch_caps_news(self):
|
||||||
|
"""NEWS-12: over news-daily-per-user, get_news refuses without querying."""
|
||||||
|
responder = self._responder(**{"enable-news-tool": True, "news-daily-per-user": 2})
|
||||||
|
responder.store.add_news_items([{"title": "x", "link": "l", "source": "s", "summary": "y"}])
|
||||||
|
for _ in range(2):
|
||||||
|
self.assertIn("results", await responder._dispatch_tool("get_news", {}, "alice"))
|
||||||
|
blocked = await responder._dispatch_tool("get_news", {}, "alice")
|
||||||
|
self.assertIn("error", blocked)
|
||||||
|
|||||||
+34
-6
@@ -120,9 +120,7 @@ class TestConfigReloadRace(TestBotBase):
|
|||||||
new_config["history-limit"] = 99
|
new_config["history-limit"] = 99
|
||||||
self.bot.load_config = lambda path: new_config
|
self.bot.load_config = lambda path: new_config
|
||||||
self.bot.loop = MagicMock()
|
self.bot.loop = MagicMock()
|
||||||
event = MagicMock()
|
self.bot.on_config_file_changed()
|
||||||
event.src_path = self.bot.config_file
|
|
||||||
self.bot.on_config_file_modified(event)
|
|
||||||
self.bot.loop.call_soon_threadsafe.assert_called_once()
|
self.bot.loop.call_soon_threadsafe.assert_called_once()
|
||||||
apply_fn = self.bot.loop.call_soon_threadsafe.call_args.args[0]
|
apply_fn = self.bot.loop.call_soon_threadsafe.call_args.args[0]
|
||||||
apply_fn()
|
apply_fn()
|
||||||
@@ -137,7 +135,37 @@ class TestConfigReloadRace(TestBotBase):
|
|||||||
loop = MagicMock()
|
loop = MagicMock()
|
||||||
loop.call_soon_threadsafe.side_effect = RuntimeError("no running loop")
|
loop.call_soon_threadsafe.side_effect = RuntimeError("no running loop")
|
||||||
self.bot.loop = loop
|
self.bot.loop = loop
|
||||||
event = MagicMock()
|
self.bot.on_config_file_changed()
|
||||||
event.src_path = self.bot.config_file
|
|
||||||
self.bot.on_config_file_modified(event)
|
|
||||||
self.assertEqual(self.bot.config["history-limit"], 42)
|
self.assertEqual(self.bot.config["history-limit"], 42)
|
||||||
|
|
||||||
|
|
||||||
|
class TestConfigReloadRenameSafe(unittest.TestCase):
|
||||||
|
def test_atomic_rename_and_modify_trigger_reload(self):
|
||||||
|
"""CFG-05: a modified OR a renamed-into-place config fires the reload; unrelated files do not."""
|
||||||
|
from fjerkroa_bot.discord_bot import ConfigFileHandler
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory() as tmp:
|
||||||
|
config = Path(tmp) / "kroa.toml"
|
||||||
|
config.write_text("x = 1\n")
|
||||||
|
hits = []
|
||||||
|
handler = ConfigFileHandler(str(config), lambda: hits.append(1))
|
||||||
|
|
||||||
|
def evt(is_dir=False, src=None, dest=None):
|
||||||
|
event = MagicMock()
|
||||||
|
event.is_directory = is_dir
|
||||||
|
event.src_path = src if src is not None else ""
|
||||||
|
event.dest_path = dest if dest is not None else ""
|
||||||
|
return event
|
||||||
|
|
||||||
|
handler.on_modified(evt(src=str(config))) # in-place modify
|
||||||
|
handler.on_moved(evt(src=str(Path(tmp) / "kroa.toml.tmp"), dest=str(config))) # atomic rename over
|
||||||
|
handler.on_created(evt(src=str(config))) # write-new
|
||||||
|
self.assertEqual(len(hits), 3)
|
||||||
|
|
||||||
|
handler.on_modified(evt(src=str(Path(tmp) / "other.txt"))) # unrelated file
|
||||||
|
handler.on_modified(evt(is_dir=True, src=str(config))) # directory event
|
||||||
|
self.assertEqual(len(hits), 3) # neither fired
|
||||||
|
|
||||||
|
# open/close of the config (our own load_config re-reads) must NOT be handled — else a reload loop.
|
||||||
|
self.assertNotIn("on_opened", vars(ConfigFileHandler))
|
||||||
|
self.assertNotIn("on_closed", vars(ConfigFileHandler))
|
||||||
|
|||||||
@@ -0,0 +1,102 @@
|
|||||||
|
"""Unit coverage for SPEC-015 web search via Exa (WEB-01..05)."""
|
||||||
|
|
||||||
|
import os
|
||||||
|
import unittest
|
||||||
|
from unittest.mock import AsyncMock, patch
|
||||||
|
|
||||||
|
from fjerkroa_bot.openai_responder import OpenAIResponder
|
||||||
|
from fjerkroa_bot.websearch import WEB_SEARCH_TOOL, WebSearch, _format_results
|
||||||
|
|
||||||
|
CONFIG = {"openai-token": "t", "model": "m", "system": "s", "history-limit": 5}
|
||||||
|
|
||||||
|
|
||||||
|
def _tool_names(responder):
|
||||||
|
return [f["name"] for f in responder._available_tools()]
|
||||||
|
|
||||||
|
|
||||||
|
class TestToolOffered(unittest.TestCase):
|
||||||
|
def test_gate_needs_flag_and_key(self):
|
||||||
|
"""WEB-01: web_search offered only with enable-web-search AND a key."""
|
||||||
|
off = OpenAIResponder(CONFIG, "chat") # flag off -> absent even if env key exists
|
||||||
|
self.assertNotIn("web_search", _tool_names(off))
|
||||||
|
on = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True, "exa-api-key": "k"}), "chat")
|
||||||
|
self.assertIn("web_search", _tool_names(on))
|
||||||
|
self.assertEqual(WEB_SEARCH_TOOL["name"], "web_search")
|
||||||
|
with patch.dict(os.environ, {"EXA_API_KEY": ""}):
|
||||||
|
nokey = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True}), "chat")
|
||||||
|
self.assertNotIn("web_search", _tool_names(nokey))
|
||||||
|
|
||||||
|
|
||||||
|
class TestFormat(unittest.TestCase):
|
||||||
|
def test_results_sanitized_and_capped(self):
|
||||||
|
"""WEB-02: title/snippet sanitized + capped; non-dict rows skipped."""
|
||||||
|
data = {
|
||||||
|
"results": [
|
||||||
|
{"title": "@everyone Hi", "url": "https://x.com/a", "text": "@here " + "y" * 1000, "publishedDate": "2026-01-01"},
|
||||||
|
{"title": "T2", "url": "https://x.com/b", "text": "short"},
|
||||||
|
"not a dict",
|
||||||
|
]
|
||||||
|
}
|
||||||
|
rows = _format_results(data, 50)
|
||||||
|
self.assertEqual(len(rows), 2)
|
||||||
|
self.assertNotIn("@everyone", rows[0]["title"])
|
||||||
|
self.assertNotIn("@here", rows[0]["snippet"])
|
||||||
|
self.assertLessEqual(len(rows[0]["snippet"]), 50)
|
||||||
|
self.assertEqual(rows[0]["url"], "https://x.com/a")
|
||||||
|
self.assertEqual(rows[0]["published"], "2026-01-01")
|
||||||
|
|
||||||
|
|
||||||
|
class TestSearch(unittest.IsolatedAsyncioTestCase):
|
||||||
|
async def test_num_results_clamped(self):
|
||||||
|
"""WEB-03: numResults clamped to 1..10; 0 falls back to default."""
|
||||||
|
ws = WebSearch(lambda: {"exa-api-key": "k"})
|
||||||
|
with patch.object(ws, "_post", new=AsyncMock(return_value={"results": []})) as post:
|
||||||
|
await ws.search("hi", num_results=999)
|
||||||
|
self.assertEqual(post.await_args.args[0]["numResults"], 10)
|
||||||
|
await ws.search("hi", num_results=0)
|
||||||
|
self.assertEqual(post.await_args.args[0]["numResults"], 5)
|
||||||
|
|
||||||
|
async def test_no_key_returns_error(self):
|
||||||
|
"""WEB-04: no key -> error dict, no network call."""
|
||||||
|
with patch.dict(os.environ, {"EXA_API_KEY": ""}):
|
||||||
|
ws = WebSearch(lambda: {})
|
||||||
|
with patch.object(ws, "_post", new=AsyncMock()) as post:
|
||||||
|
result = await ws.search("hi")
|
||||||
|
post.assert_not_awaited()
|
||||||
|
self.assertIn("error", result)
|
||||||
|
|
||||||
|
async def test_api_failure_returns_error(self):
|
||||||
|
"""WEB-04: a raising request is caught, returns an error dict."""
|
||||||
|
ws = WebSearch(lambda: {"exa-api-key": "k"})
|
||||||
|
with patch.object(ws, "_post", new=AsyncMock(side_effect=RuntimeError("boom"))):
|
||||||
|
result = await ws.search("hi")
|
||||||
|
self.assertIn("error", result)
|
||||||
|
|
||||||
|
async def test_empty_query_no_call(self):
|
||||||
|
"""WEB-04: blank query returns empty results without a call."""
|
||||||
|
ws = WebSearch(lambda: {"exa-api-key": "k"})
|
||||||
|
with patch.object(ws, "_post", new=AsyncMock()) as post:
|
||||||
|
result = await ws.search(" ")
|
||||||
|
post.assert_not_awaited()
|
||||||
|
self.assertEqual(result["results"], [])
|
||||||
|
|
||||||
|
async def test_search_returns_formatted(self):
|
||||||
|
"""WEB-02: a successful search returns sanitized rows."""
|
||||||
|
ws = WebSearch(lambda: {"exa-api-key": "k"})
|
||||||
|
payload = {"results": [{"title": "Norge", "url": "https://ex.com/n", "text": "fakta"}]}
|
||||||
|
with patch.object(ws, "_post", new=AsyncMock(return_value=payload)):
|
||||||
|
result = await ws.search("norge")
|
||||||
|
self.assertEqual(result["results"][0]["title"], "Norge")
|
||||||
|
self.assertEqual(result["results"][0]["url"], "https://ex.com/n")
|
||||||
|
|
||||||
|
|
||||||
|
class TestPerUserCap(unittest.IsolatedAsyncioTestCase):
|
||||||
|
async def test_dispatch_caps_searches(self):
|
||||||
|
"""WEB-05: over web-daily-per-user, web_search refuses without calling the API."""
|
||||||
|
responder = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True, "exa-api-key": "k", "web-daily-per-user": 2}), "chat")
|
||||||
|
responder.web_search.search = AsyncMock(return_value={"query": "x", "results": []})
|
||||||
|
for _ in range(2):
|
||||||
|
self.assertIn("results", await responder._dispatch_tool("web_search", {"query": "hi"}, "bob"))
|
||||||
|
blocked = await responder._dispatch_tool("web_search", {"query": "hi"}, "bob")
|
||||||
|
self.assertIn("error", blocked)
|
||||||
|
self.assertEqual(responder.web_search.search.await_count, 2)
|
||||||
Reference in New Issue
Block a user