health monitor (ops-18/19): spend/disk/task-queue thresholds -> staff alerts, edge-triggered, opt-in
This commit is contained in:
@@ -30,3 +30,23 @@ The responder counts consecutive OpenAI request failures; at
|
||||
alert (rate-limited like all staff alerts) so a silently-broken bot
|
||||
(cf. the gpt-5.6 tools/reasoning incident) surfaces within minutes
|
||||
instead of hours. A success resets the counter.
|
||||
|
||||
### OPS-18 — Health monitor watches spend, disk, task-queue (coverage: test)
|
||||
|
||||
When `enable-monitoring` is true, a loop wakes every `monitor-interval`
|
||||
(default 300 s) and checks three thresholds, alerting the staff channel
|
||||
when one is crossed: daily spend at or above `monitor-spend-alert-frac`
|
||||
(default 0.8) of `daily-budget-usd`; free disk below `monitor-disk-min-mb`
|
||||
(default 500 MB); open task-queue depth at or above `monitor-taskqueue-max`
|
||||
(default 20). A check with no data to evaluate (no budget set, no store,
|
||||
a failed disk read) is skipped, never fatal. With the flag off the loop
|
||||
does nothing.
|
||||
|
||||
### OPS-19 — Alerts fire once per crossing and re-arm on recovery (coverage: test)
|
||||
|
||||
Each metric alerts only on the rising edge — the first tick that finds
|
||||
it over its threshold — and stays silent while it remains over, so a
|
||||
persistent condition does not repeat every interval. When the metric
|
||||
falls back below the threshold the alert re-arms silently, ready to fire
|
||||
again on the next crossing. All alerts still pass through the
|
||||
rate-limited staff-alert path (OPS-07).
|
||||
|
||||
Reference in New Issue
Block a user