Files
Project-SDE-WP-Suite/docs/waves/wave-13.md
Matt Mabrey 358469531c D18: revise T13.1 to idle timeout + absolute ceiling
Matt asked whether idle time would be a better fit than a flat session
length. It is, but idle-alone weakens D18's own purpose - a continuously
active session would never force a fresh Okta recheck on its own. Decided:
both. AUTH_IDLE_MINUTES (new, default 30) slides the session on activity;
AUTH_SESSION_HOURS (existing var, meaning changes to an absolute ceiling,
default 12 -> proposed 8) caps it regardless of activity.

Docs only in this commit - implementation is T13.1, next.
2026-09-23 11:38:37 -07:00

243 lines
10 KiB
Markdown

# Wave 13 — Okta/AD deprovisioning sync
**Items:** `D18`
**Depends on:** wave 10 merged (it is). Not blocked by wave 11 or wave 12, but
shares `server/app.py` account-state surface with wave 12 (`CR-020`) — sequence
commits to avoid an avoidable conflict.
**Decision record:** `docs/waves/decisions-2026-09-17.md`
Six tasks, in build order. **T13.3's fail-closed behavior is the most
important done-when list in this wave — do not relax it to ship faster.**
---
### T13.1 — D18: idle timeout with an absolute ceiling
- **Items:** `D18`
- **Depends on:** nothing (first task, independent of the rest)
- **Blocks:** nothing
- **Surface:** `server/`
- **Files:** `server/auth.py`, `server/app.py` (`auth_gate` middleware),
`server/.env.example`, `DEPLOYMENT.md`
**Revised 2026-09-23** — originally just "shrink `AUTH_SESSION_HOURS`". Matt
asked whether an idle timeout would be a better fit than a flat session
length. It is, but not by itself — see the decision record's reasoning on why
idle-with-no-ceiling is actually a worse fit for this item's own threat model
than a flat expiry would have been. Build both.
**Do:**
- Add `login_at` to the JWT payload in `create_token()` — the original
sign-in time, distinct from `iat`, which becomes "when THIS token was
issued" once tokens start getting reissued. `login_at` never changes across
reissues; it's what the absolute ceiling is measured from.
- Add `AUTH_IDLE_MINUTES` (default 30). `AUTH_SESSION_HOURS` stays the name
for the absolute ceiling, default changing from 12 to a proposed 8 — update
its docstring/comment in `auth.py` and `.env.example`, since its MEANING is
changing (session length -> hard ceiling on top of a sliding idle window),
not just its value.
- In `auth_gate` (`server/app.py`), after confirming a request is
authenticated: if `now < login_at + AUTH_SESSION_HOURS` (the ceiling hasn't
passed), compute `new_exp = min(now + AUTH_IDLE_MINUTES, login_at +
AUTH_SESSION_HOURS)`. If `new_exp` is meaningfully later than the current
token's `exp` (throttle this — do not reissue on every single request, only
when enough time has passed to be worth a new cookie; a few minutes of
slack is fine), mint a refreshed token carrying forward `sub`/`username`/
`role`/`ver`/`login_at` unchanged, and set it on the response.
- If the ceiling HAS passed, do not refresh — let the existing token expire
on its own terms (it may already be invalid, or may tick over within the
idle window; either way, no new one is issued past the ceiling).
- The middleware should not need a DB round trip to do this — everything
needed (`sub`, `username`, `role`, `ver`, `login_at`) is already in the
validated claims. `get_current_user`'s existing per-request `is_active`/
`token_version` check is unaffected and still runs separately.
**Do not:** reissue the cookie on every request unconditionally — that's a
`Set-Cookie` header on every API call for no benefit over a coarser refresh.
Do not let the ceiling check silently vanish for tokens issued before this
change ships — a token with no `login_at` claim should fall back to treating
its own `iat` as `login_at`, not bypass the ceiling entirely.
**Done when:**
- [ ] a session with continuous activity stays alive past 30 minutes but is
cut off at the `AUTH_SESSION_HOURS` ceiling regardless
- [ ] a session with no activity for 30+ minutes is rejected on its next
request
- [ ] the cookie is not rewritten on every single request — verify the
refresh is throttled, not unconditional
- [ ] a pre-existing token with no `login_at` claim (simulating a session
issued before this shipped) still gets a hard ceiling, via the `iat`
fallback
- [ ] `.env.example` and `DEPLOYMENT.md` explain both variables and flag both
defaults as proposed, not confirmed against Okta's own session policy
- [ ] existing session/auth tests updated for the new mechanism, not just a
new number
---
### T13.2 — D18: Okta Management API client
- **Items:** `D18`
- **Depends on:** nothing (independent of T13.1)
- **Blocks:** T13.3
- **Surface:** `server/`
- **Files:** new `server/okta_sync.py` (or extend `server/okta_auth.py`
task's call), `server/.env.example`
**Do:** A small client for Okta's user-list endpoint, authenticated with a new,
separate credential (e.g. `OKTA_API_TOKEN`) — not the OIDC client secret used
for sign-in. Fetch the full user list (paginated per Okta's API) rather than
one-by-one lookups per local user; this app has ~20-odd accounts today, and a
list-and-diff is simpler and cheaper than N calls. Return each Okta user's
identity-claim value (matching `OKTA_IDENTITY_CLAIM`, already confirmed live
per wave 10) and status.
Follow the existing pattern for external credentials in this repo (`MICRON_DB_URL`,
`SMTP_PASSWORD`): env-only, never logged, never returned to the browser in an
error message.
**Do not:** reuse `OKTA_CLIENT_ID`/`OKTA_CLIENT_SECRET` for this. Sign-in and
the management API are different trust boundaries with different scopes;
conflating them means a compromise or rotation of one affects the other
unnecessarily.
**Done when:**
- [ ] the client authenticates with its own credential, distinct from the OIDC
client
- [ ] it fetches the complete Okta user list, handling pagination
- [ ] an auth failure or malformed response raises a clear, specific error
rather than returning an empty list indistinguishable from "everyone was
deprovisioned" — this distinction is what T13.3 depends on
- [ ] the credential is never logged or surfaced in any API response
---
### T13.3 — D18: the sync job
- **Items:** `D18`
- **Depends on:** T13.2
- **Blocks:** T13.4
- **Surface:** `server/`
- **Files:** `server/okta_sync.py`, `server/app.py` (or wherever `log_event`
lives)
**Do:** Compare the Okta user list (T13.2) against local `users` rows. For any
local `is_active=True` user whose Okta identity is missing from the list, or
present with a non-active status, set `is_active=False` and write an
`AuditLog` row (`actor="system:okta_sync"`, action e.g.
`user_deprovisioned_by_sync`, detail naming the Okta status found). Never
touch a user already `is_active=False`. Never re-enable anyone.
**The fail-closed rule, non-negotiable:** if T13.2's client raises an error of
any kind (network, auth, malformed response, timeout), this task takes **no
action on any account** for that run and logs the failure clearly (server log,
at minimum). An error is never treated as "Okta returned zero active users."
Write a test that asserts this directly: feed the sync a failing client and
assert zero rows changed and zero `AuditLog` entries written.
**Done when:**
- [ ] a user missing from Okta's list, or present but not active, is disabled
with a correctly-detailed audit row
- [ ] a user already disabled is left alone (no duplicate audit row each run)
- [ ] an active Okta user already active locally produces no audit row (only
changes are logged, not a clean bill of health every cycle)
- [ ] **a simulated Okta API failure results in zero account changes and zero
audit rows** — this is the one check that must never be skipped or
weakened
- [ ] the job never re-enables an account under any input
---
### T13.4 — D18: run it on a schedule
- **Items:** `D18`
- **Depends on:** T13.3
- **Blocks:** T13.6
- **Surface:** `server/`, `docker-compose.yml`
- **Files:** `server/app.py` (startup hook) or a new sidecar per the `backup`
container's pattern — task's call, per the decision record's noted
alternative
**Do:** Wire T13.3 to run on an interval (proposed 15 minutes, env-overridable
— e.g. `OKTA_SYNC_INTERVAL_SECONDS`, matching `BACKUP_INTERVAL_SECONDS`'s
naming). Default choice is an in-process background task in the `api`
container; if a separate container is chosen instead, follow the `backup`
service's shape (own Dockerfile or script, `internal` network plus whatever
egress reaching Okta requires — check whether `outbound` as currently defined
is sufficient or Okta needs a distinct allowance).
**Done when:**
- [ ] the job runs automatically on the configured interval without manual
invocation
- [ ] interval is env-configurable with a sane default
- [ ] a container restart does not produce a duplicate/overlapping run, and a
slow cycle does not stack with the next one
---
### T13.5 — D18: admin visibility
- **Items:** `D18`
- **Depends on:** T13.3
- **Blocks:** T13.6
- **Surface:** `html/`
- **Files:** `html/users.js` / `html/admin.js` (wherever audit history is
already surfaced)
**Do:** Confirm an auto-disable reads clearly wherever admins already look at
account history — the actor string (`system:okta_sync`) should be
self-explanatory in context, not require reading server logs to understand.
Do not build new UI beyond making sure the existing audit surface renders this
actor sensibly. Email/notification to admins on auto-disable is a noted
fast-follow (the decision record flags it as open, not required here) — do not
build it in this task; log it instead if it's tempting to add.
**Done when:**
- [ ] an auto-disabled account's audit entry is visible and legible in the
existing admin UI without special-casing
- [ ] nothing here silently assumes `CR-019`'s activity view exists yet — this
must work standalone
---
### T13.6 — D18: verification
- **Items:** `D18`
- **Depends on:** T13.4, T13.5
- **Blocks:** nothing
- **Surface:** `server/`
- **Files:** as touched above
**Do:** Full verification per `CLAUDE.md`. Beyond the usual run: specifically
re-run T13.3's fail-closed test in isolation and confirm it still passes after
T13.4's scheduling wrapper is in place — the scheduling layer must not
introduce a path that swallows the client's error and proceeds anyway.
**Done when:**
- [ ] all `D18` acceptance criteria in `decisions-2026-09-17.md` are met or a
failure is stated with a reason
- [ ] the fail-closed behavior is verified end to end through the scheduled
wrapper, not just the bare sync function
- [ ] smoke test and `seed_demo.py` both still pass
- [ ] full test suite passes
---
## Wave 13 exit criteria
- [ ] `AUTH_SESSION_HOURS` defaults to 2, documented
- [ ] the sync job runs on a schedule and correctly disables accounts Okta no
longer shows as active
- [ ] every auto-disable is individually audited with a clearly non-human actor
- [ ] an Okta API failure of any kind changes zero accounts — verified, not
assumed
- [ ] the sync never re-enables an account
- [ ] `D18` fully accounted for, no open acceptance criteria