diff --git a/docs/waves/backlog.md b/docs/waves/backlog.md index 4071b05..09dfa45 100644 --- a/docs/waves/backlog.md +++ b/docs/waves/backlog.md @@ -554,3 +554,43 @@ deliberately deferred. entry rather than a drive-by. - **Suggested wave or follow-up:** next housekeeping pass, with the check widened so it cannot recur. + +### BL-026 — No version stamp: "is live current?" cannot be answered from the app + +- **Found during:** the 2026-08-21 outage triage (the question that started it) +- **Where:** `Dockerfile` / build, `server/app.py` `/api/health`, admin console +- **What:** the app carries no record of what code it is running. `/api/health` + returns `{"ok": true}` and nothing identifies the deployed commit, so + answering "is the live site on the latest code?" took fingerprinting + (probing for files/routes that only exist after certain merges) in the + middle of an outage. The fix: bake the git SHA into the image at build time + (`ARG GIT_SHA`), return it from `/api/health` + (`{"ok": true, "version": ""}`), and show it on the admin console's + diagnostics card. Then currency is one glance against `git log -1`. +- **Why not now:** new scope — needs its own item id per the working rules + (D13 is the natural next), and it touches the image build, which deserves a + deploy alongside someone with host access. +- **Suggested wave or follow-up:** next housekeeping pass; ~1 task including a + probe check that /api/health carries a version field. + +### BL-027 — Migrations are rehearsed on SQLite only; production is Postgres + +- **Found during:** the 2026-08-21 production outage (D6's `material_items` + migration crash-looped the api container) +- **Where:** `DEPLOYMENT.md` (the update/deploy steps), `tests/` +- **What:** the migration chain is verified end-to-end on scratch SQLite, but + production runs Postgres, and the dialects disagree exactly where it hurts: + `server_default=sa.text('1')` on a Boolean passed every SQLite rehearsal and + was refused by Postgres at deploy (DatatypeMismatch), taking the API down + until the table was created by hand. The hotfix (64eac0c) fixed that one + instance and pinned the Boolean-default class in `materials_check`; the + CLASS of dialect drift is still unguarded. Two cheap layers: (1) a runbook + step — render `alembic upgrade --sql` for the postgresql dialect and read it + before restarting (offline, needs no live DB; this render would have shown + `DEFAULT 1` on a boolean); (2) better, a probe that renders every migration + for the postgresql dialect on each run and fails on anything the dialect + rejects or on known-bad patterns. +- **Why not now:** the outage is resolved and the one known instance is fixed + and pinned; the systematic guard is its own small task, not a hotfix rider. +- **Suggested wave or follow-up:** next housekeeping pass, paired with BL-026 + (both are "deploys should be boring" work).