Compare commits

..

1 Commits

Author SHA1 Message Date
Oleksandr Kozachuk cdd5a4cd48 web search via exa (spec-015): web_search tool, sanitized results, per-user cap, host-config key 2026-07-14 12:08:48 +02:00
6 changed files with 388 additions and 0 deletions
+9
View File
@@ -60,6 +60,15 @@ Decisions inside the set architecture. D-NNN, never renumbered.
broken classifier must never mute the bot; the budget gate already broken classifier must never mute the bot; the budget gate already
bounds spend. Its verdict gates BEFORE the main call, the bounds spend. Its verdict gates BEFORE the main call, the
envelope's answer_needed still gates after — two independent nets. envelope's answer_needed still gates after — two independent nets.
- **D-020** — Web search via Exa (FDB-022, SPEC-015): a `web_search`
tool alongside fetch_url/IGDB/codex/get_news, filling the "look it up
on the open web" gap. Exa (not a raw search-engine scrape) because it
returns clean title+url+text in one call — no SSRF surface of our own
(we call one fixed API endpoint, not arbitrary hosts), and it pairs
with fetch_url for the full article. Key is a host secret
(`exa-api-key`, env `EXA_API_KEY` fallback), never repo-side; results
sanitized like every other external-text tool; off by default
(`enable-web-search`), metered per user.
- **D-019** — News memory + on-demand tool (SPEC-013 NEWS-07..12): - **D-019** — News memory + on-demand tool (SPEC-013 NEWS-07..12):
the news pipeline now carries item summaries (feed descriptions, the news pipeline now carries item summaries (feed descriptions,
HTML-stripped) and persists every fetched item into a deduped `news` HTML-stripped) and persists every fetched item into a deduped `news`
+130
View File
@@ -0,0 +1,130 @@
# Operator runbook — Fjærkroa / Luma bot
One page for "something is wrong, what do I do". Two deployments of one
codebase, both on **uberspace** (push-based deploy from the dev machine —
there is no git checkout on the hosts).
| | Fjærkroa (café) | Luma (GGG clan) |
| --- | --- | --- |
| SSH host | `ssh fjerkroa` (pictor.uberspace.de) | `ssh ggg` |
| Service | `kroa` | `luma` |
| Config | `~/fjerkroa_bot/kroa.toml` | `~/fjerkroa_bot/ggg.toml` |
| Staff channel | `#kassa` | `#mods` |
| Language / persona | Norwegian, café host | German, "Luma" |
Common paths on each host: bot code `~/fjerkroa_bot`, venv `~/venv-bot`,
database `~/fjerkroa_bot/history/bot.db` (SQLite, WAL), our snapshots
`~/backups/<kroa|luma>/`, logs under `~/logs` and `~/tmp`.
## From Discord (staff channel only, prefix `!bot`)
No SSH needed for day-to-day control. Type `!bot help` in the staff
channel for the full, grouped list. The essentials:
- `!bot pause` / `!bot resume` — stop / start all replies.
- `!bot quiet <minutes>` — go silent for a while, then auto-resume.
- `!bot status` — replies/images/tasks flags + quiet time left.
- `!bot spend` — today's estimated USD spend, tokens, images, budget.
- `!bot images on|off`, `!bot tasks on|off` — kill-switches.
`!help` works in **any** channel (for everyone) and lists only what is
usable there. `!forgetme` and `!privacy` also work everywhere, even
while the bot is paused.
## Restart / check health (SSH)
```sh
ssh <host>
supervisorctl status <kroa|luma> # RUNNING + uptime
supervisorctl restart <kroa|luma>
tail -n 40 ~/tmp/<kroa|luma>-stderr*.log # discord login / errors
tail -n 40 ~/logs/supervisord.log # "We have logged in as ..."
```
A healthy start shows a fresh `connected to Gateway` + `We have logged
in as ...` line within ~15 s.
## Deploy a release / roll back
From the **dev machine** (`~/Repos/FjerkroaBot`), tags only:
```sh
git tag -m "<msg>" vX.Y.Z && git push --tags # cut the release first
bash deploy/deploy.sh ggg vX.Y.Z # luma
DEPLOY_FORCE=1 bash deploy/deploy.sh fjerkroa vX.Y.Z # kroa (see window)
```
- kroa refuses to deploy **11:0022:00 Europe/Oslo** (restaurant hours);
`DEPLOY_FORCE=1` overrides. Café is closed Mondays.
- The script backs up `bot.db``bot.db.pre-<tag>` before restart, then
smoke-tests (RUNNING + fresh login) and fails loudly if either misses.
- **Rollback** = deploy the previous tag. If the schema version moved
between the two tags, restore the matching `bot.db.pre-<newtag>` first
(see below) so the older code meets a schema it understands.
## Restore the database
Three independent daily backup layers exist — pick the freshest good one.
```sh
ssh <host>
supervisorctl stop <kroa|luma>
DB=~/fjerkroa_bot/history/bot.db
# 1) uberspace nightly backup of the whole home (read-only):
# /backup = current + daily.0..7 + weekly.1..7 (15 restore points)
cp /backup/daily.1/home/<user>/fjerkroa_bot/history/bot.db "$DB"
# 2) our own rotated gzip snapshot (03:17 UTC cron, keep 14):
gunzip -c ~/backups/<kroa|luma>/bot-YYYYMMDD-HHMMSS.db.gz > "$DB"
# 3) the pre-deploy snapshot for a given release:
cp "$DB".pre-vX.Y.Z "$DB"
rm -f "$DB"-wal "$DB"-shm # drop stale WAL sidecars after a restore
supervisorctl start <kroa|luma>
```
`<user>` is `fjerkroa` or `ggg`. The DB holds conversation history,
structured memory, usage ledger, image cache index, tasks, and the news
store — all regenerable, none critical. That is why there is no off-host
backup: uberspace `/backup` + the on-host snapshots are enough.
## Rotate a secret
Secrets live only in the host `*.toml` (never in the repo). Edit in
place and restart:
```sh
ssh <host>
# OpenAI: edit openai-token = "sk-..." in kroa.toml / ggg.toml
# Discord: edit discord-token = "..." (get a new token from the
# Discord developer portal → Bot → Reset Token first)
supervisorctl restart <kroa|luma>
```
After rotating an OpenAI key, revoke the old one in the OpenAI dashboard.
Keep a `*.toml` backup before editing; a broken TOML crash-loops the
service (validate: `~/venv-bot/bin/python -c 'import tomlkit; tomlkit.load(open("kroa.toml"))'`).
## Scheduled jobs (crontab -l)
| Host | When (server time) | Job |
| --- | --- | --- |
| both | `17 3 * * *` | `backup_db.py``~/backups/<bot>/` (keep 14) |
| kroa | `5 * * * *` | news digest → `{news}` file + news store |
| ggg | `*/15 * * * *` | news poster → #news/#newsjp webhooks + store |
Logs: `~/backups/<bot>/backup.log`, `~/backups/<bot>/news*.log`.
## Quick triage
- **Bot silent everywhere** → `!bot status` (paused/quiet?), else
`supervisorctl status`; if not RUNNING, `restart` and read stderr.
- **Bot silent in one channel** → check the host config `ignore-channels`
/ `short-path` rules for that channel (a stray `short-path` rule can
archive messages without replying).
- **Repeated API errors** → the bot posts a rate-limited alert to the
staff channel after 5 consecutive OpenAI failures (OPS-16); check
`!bot spend` (budget hit?) and the OpenAI status/key.
- **Bad deploy** → roll back to the previous tag (above).
+12
View File
@@ -17,6 +17,8 @@ from .leonardo_draw import LeonardoAIDrawMixIn
from .news import GET_NEWS_TOOL, query_news from .news import GET_NEWS_TOOL, query_news
from .quota import QuotaLedger from .quota import QuotaLedger
from .url_reader import FETCH_URL_TOOL, URLReader from .url_reader import FETCH_URL_TOOL, URLReader
from .websearch import DEFAULT_RESULTS as WEB_DEFAULT_RESULTS
from .websearch import WEB_SEARCH_TOOL, WebSearch
# The response envelope, enforced server-side via structured outputs # The response envelope, enforced server-side via structured outputs
# (ENV-19). All fields required, closed object, nullable where the # (ENV-19). All fields required, closed object, nullable where the
@@ -166,6 +168,8 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
self.url_reader = URLReader(lambda: self.config, self.image_cache) self.url_reader = URLReader(lambda: self.config, self.image_cache)
# Codex Mechanicus search (SPEC-014); Luma's own archive at binaric.tech # Codex Mechanicus search (SPEC-014); Luma's own archive at binaric.tech
self.codex = CodexSearch(lambda: self.config) self.codex = CodexSearch(lambda: self.config)
# Web search (SPEC-015) via Exa; general "look it up" beyond fetch_url/news/codex
self.web_search = WebSearch(lambda: self.config)
def _available_tools(self) -> List[Dict[str, Any]]: def _available_tools(self) -> List[Dict[str, Any]]:
"""Assemble the function-tool list from every enabled provider (URL-01).""" """Assemble the function-tool list from every enabled provider (URL-01)."""
@@ -183,6 +187,8 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
functions.append(CODEX_SEARCH_TOOL) functions.append(CODEX_SEARCH_TOOL)
if self.config.get("enable-news-tool", False) and self.store is not None: # NEWS-10 if self.config.get("enable-news-tool", False) and self.store is not None: # NEWS-10
functions.append(GET_NEWS_TOOL) functions.append(GET_NEWS_TOOL)
if self.web_search.enabled(): # WEB-01
functions.append(WEB_SEARCH_TOOL)
return functions return functions
async def _dispatch_tool(self, name: str, args: Dict[str, Any], author: str) -> Any: async def _dispatch_tool(self, name: str, args: Dict[str, Any], author: str) -> Any:
@@ -207,6 +213,12 @@ class OpenAIResponder(AIResponder, LeonardoAIDrawMixIn):
self.ledger._add(f"news:{author}", 1) self.ledger._add(f"news:{author}", 1)
summary_chars = int(self.config.get("news-summary-chars", 200)) summary_chars = int(self.config.get("news-summary-chars", 200))
return query_news(self.store, args.get("topic"), args.get("source"), args.get("limit", 10), summary_chars) return query_news(self.store, args.get("topic"), args.get("source"), args.get("limit", 10), summary_chars)
if name == "web_search":
per_user_cap = int(self.config.get("web-daily-per-user", 30))
if self.ledger._get(f"web:{author}") >= per_user_cap: # WEB-05
return {"error": "daily web search limit reached"}
self.ledger._add(f"web:{author}", 1)
return await self.web_search.search(str(args.get("query", "")), int(args.get("num_results", WEB_DEFAULT_RESULTS)))
return await self._execute_igdb_function(name, args) return await self._execute_igdb_function(name, args)
async def draw_openai(self, description: str, count: int = 1) -> List[BytesIO]: async def draw_openai(self, description: str, count: int = 1) -> List[BytesIO]:
+93
View File
@@ -0,0 +1,93 @@
"""Web search tool via Exa (SPEC-015, FDB-022).
A `web_search` function tool: the model looks things up on the open web
when a general "look it up" question is not covered by IGDB, the codex,
the news store, or a URL the user pasted. Results are external text, so
titles and snippets are sanitized (SAF-03) before they reach the prompt.
The Exa API key lives in host config (or the `EXA_API_KEY` env), never in
the repo.
"""
import logging
import os
from typing import Any, Callable, Dict, List
import aiohttp
from .ai_responder import sanitize_external_text
EXA_SEARCH_URL = "https://api.exa.ai/search"
DEFAULT_RESULTS = 5
MAX_RESULTS = 10
DEFAULT_SNIPPET_CHARS = 400
FETCH_TIMEOUT_S = 15
WEB_SEARCH_TOOL = {
"name": "web_search",
"description": "Search the open web for current information when the user asks you to look something up and it is not "
"covered by game info (IGDB), the Codex, the news store, or a URL they pasted. Returns result titles, URLs, and a short "
"snippet; follow up with fetch_url on a result link for the full article.",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "What to search the web for."},
"num_results": {"type": "integer", "description": "How many results to return (default 5, max 10)."},
},
"required": ["query"],
},
}
def _format_results(data: Any, snippet_chars: int) -> List[Dict[str, str]]:
"""Reduce an Exa response to sanitized {title, url, snippet, published} rows (WEB-02)."""
results = data.get("results", []) if isinstance(data, dict) else []
out: List[Dict[str, str]] = []
for item in results:
if not isinstance(item, dict):
continue
out.append(
{
"title": sanitize_external_text(str(item.get("title") or ""), 200),
"url": str(item.get("url") or ""),
"snippet": sanitize_external_text(str(item.get("text") or item.get("snippet") or ""), snippet_chars),
"published": str(item.get("publishedDate") or ""),
}
)
return out
class WebSearch:
def __init__(self, config_getter: Callable[[], Dict[str, Any]]) -> None:
self._config = config_getter
def _api_key(self) -> str:
return str(self._config().get("exa-api-key") or os.environ.get("EXA_API_KEY", ""))
def enabled(self) -> bool:
return bool(self._config().get("enable-web-search", False)) and bool(self._api_key())
async def _post(self, payload: Dict[str, Any], headers: Dict[str, str]) -> Any:
timeout = aiohttp.ClientTimeout(total=FETCH_TIMEOUT_S)
async with aiohttp.ClientSession(timeout=timeout) as session:
async with session.post(EXA_SEARCH_URL, json=payload, headers=headers) as response:
response.raise_for_status()
return await response.json()
async def search(self, query: str, num_results: int = DEFAULT_RESULTS) -> Dict[str, Any]:
"""Return sanitized web results, or an error dict — never raise (WEB-04)."""
key = self._api_key()
if not key:
return {"error": "web search unavailable: no api key"}
query = (query or "").strip()
if not query:
return {"query": "", "results": []}
num = max(1, min(int(num_results or DEFAULT_RESULTS), MAX_RESULTS)) # WEB-03
snippet_chars = int(self._config().get("web-snippet-chars", DEFAULT_SNIPPET_CHARS))
payload = {"query": query, "numResults": num, "type": "auto", "contents": {"text": {"maxCharacters": max(snippet_chars, 200)}}}
headers = {"x-api-key": key, "Content-Type": "application/json"}
try:
data = await self._post(payload, headers)
except Exception as err:
logging.warning(f"web search failed: {err!r}")
return {"error": "web search failed"}
return {"query": query, "results": _format_results(data, snippet_chars)}
+42
View File
@@ -0,0 +1,42 @@
# SPEC-015 — Web search (Exa)
A `web_search` function tool for general "look it up on the internet"
questions the other tools do not cover: IGDB is games, the Codex is
Adeptus Mechanicus lore, the news store is the configured feeds, and
`fetch_url` needs a URL the user already has. Web search fills the gap
and pairs with `fetch_url` (search → pick a link → read it). Results are
external text and are sanitized (SAF-03); the Exa key is a host secret,
never in the repo. Active only when `enable-web-search = true` and a key
is present.
### WEB-01 — web_search is offered as a tool (coverage: test)
When `enable-web-search` is true **and** an Exa key is available
(`exa-api-key` in config, else `EXA_API_KEY` env), the chat call's
`tools` list includes a `web_search` function (`query` string, optional
`num_results`). With the flag off or no key it is absent.
### WEB-02 — Results are reduced and sanitized (coverage: test)
Each Exa result becomes `{title, url, snippet, published}`; `title` and
`snippet` pass through `sanitize_external_text` (snippet capped at
`web-snippet-chars`, default 400) so a web page can neither inject an
`@everyone` nor smuggle control characters into the prompt.
### WEB-03 — Result count is bounded (coverage: test)
`num_results` is clamped to 1..`MAX_RESULTS` (10) before the request, so
neither a huge fan-out nor a zero/negative count reaches the API.
### WEB-04 — Missing key and API failure are reported, not raised (coverage: test)
With no key the tool returns an `{error: ...}` result without a network
call. A request that raises (network, non-2xx, bad JSON) is logged and
returns an `{error: ...}` dict — `search` never raises into the loop.
### WEB-05 — Searches are metered per user (coverage: test)
Each `web_search` increments a per-user daily counter; over
`web-daily-per-user` (default 30) the tool refuses with an error result
without calling the API. The budget gate (SAF-04) still applies to the
surrounding model calls.
+102
View File
@@ -0,0 +1,102 @@
"""Unit coverage for SPEC-015 web search via Exa (WEB-01..05)."""
import os
import unittest
from unittest.mock import AsyncMock, patch
from fjerkroa_bot.openai_responder import OpenAIResponder
from fjerkroa_bot.websearch import WEB_SEARCH_TOOL, WebSearch, _format_results
CONFIG = {"openai-token": "t", "model": "m", "system": "s", "history-limit": 5}
def _tool_names(responder):
return [f["name"] for f in responder._available_tools()]
class TestToolOffered(unittest.TestCase):
def test_gate_needs_flag_and_key(self):
"""WEB-01: web_search offered only with enable-web-search AND a key."""
off = OpenAIResponder(CONFIG, "chat") # flag off -> absent even if env key exists
self.assertNotIn("web_search", _tool_names(off))
on = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True, "exa-api-key": "k"}), "chat")
self.assertIn("web_search", _tool_names(on))
self.assertEqual(WEB_SEARCH_TOOL["name"], "web_search")
with patch.dict(os.environ, {"EXA_API_KEY": ""}):
nokey = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True}), "chat")
self.assertNotIn("web_search", _tool_names(nokey))
class TestFormat(unittest.TestCase):
def test_results_sanitized_and_capped(self):
"""WEB-02: title/snippet sanitized + capped; non-dict rows skipped."""
data = {
"results": [
{"title": "@everyone Hi", "url": "https://x.com/a", "text": "@here " + "y" * 1000, "publishedDate": "2026-01-01"},
{"title": "T2", "url": "https://x.com/b", "text": "short"},
"not a dict",
]
}
rows = _format_results(data, 50)
self.assertEqual(len(rows), 2)
self.assertNotIn("@everyone", rows[0]["title"])
self.assertNotIn("@here", rows[0]["snippet"])
self.assertLessEqual(len(rows[0]["snippet"]), 50)
self.assertEqual(rows[0]["url"], "https://x.com/a")
self.assertEqual(rows[0]["published"], "2026-01-01")
class TestSearch(unittest.IsolatedAsyncioTestCase):
async def test_num_results_clamped(self):
"""WEB-03: numResults clamped to 1..10; 0 falls back to default."""
ws = WebSearch(lambda: {"exa-api-key": "k"})
with patch.object(ws, "_post", new=AsyncMock(return_value={"results": []})) as post:
await ws.search("hi", num_results=999)
self.assertEqual(post.await_args.args[0]["numResults"], 10)
await ws.search("hi", num_results=0)
self.assertEqual(post.await_args.args[0]["numResults"], 5)
async def test_no_key_returns_error(self):
"""WEB-04: no key -> error dict, no network call."""
with patch.dict(os.environ, {"EXA_API_KEY": ""}):
ws = WebSearch(lambda: {})
with patch.object(ws, "_post", new=AsyncMock()) as post:
result = await ws.search("hi")
post.assert_not_awaited()
self.assertIn("error", result)
async def test_api_failure_returns_error(self):
"""WEB-04: a raising request is caught, returns an error dict."""
ws = WebSearch(lambda: {"exa-api-key": "k"})
with patch.object(ws, "_post", new=AsyncMock(side_effect=RuntimeError("boom"))):
result = await ws.search("hi")
self.assertIn("error", result)
async def test_empty_query_no_call(self):
"""WEB-04: blank query returns empty results without a call."""
ws = WebSearch(lambda: {"exa-api-key": "k"})
with patch.object(ws, "_post", new=AsyncMock()) as post:
result = await ws.search(" ")
post.assert_not_awaited()
self.assertEqual(result["results"], [])
async def test_search_returns_formatted(self):
"""WEB-02: a successful search returns sanitized rows."""
ws = WebSearch(lambda: {"exa-api-key": "k"})
payload = {"results": [{"title": "Norge", "url": "https://ex.com/n", "text": "fakta"}]}
with patch.object(ws, "_post", new=AsyncMock(return_value=payload)):
result = await ws.search("norge")
self.assertEqual(result["results"][0]["title"], "Norge")
self.assertEqual(result["results"][0]["url"], "https://ex.com/n")
class TestPerUserCap(unittest.IsolatedAsyncioTestCase):
async def test_dispatch_caps_searches(self):
"""WEB-05: over web-daily-per-user, web_search refuses without calling the API."""
responder = OpenAIResponder(dict(CONFIG, **{"enable-web-search": True, "exa-api-key": "k", "web-daily-per-user": 2}), "chat")
responder.web_search.search = AsyncMock(return_value={"query": "x", "results": []})
for _ in range(2):
self.assertIn("results", await responder._dispatch_tool("web_search", {"query": "hi"}, "bob"))
blocked = await responder._dispatch_tool("web_search", {"query": "hi"}, "bob")
self.assertIn("error", blocked)
self.assertEqual(responder.web_search.search.await_count, 2)