Files
Project-SDE-WP-Suite/docs/waves/wave-13.md
Matt Mabrey 75ac930d0c waves 11-13: planning docs for CR-019, CR-020, D18
CR-019 - usage/activity metrics (admin console), wave 11
CR-020 - bulk editing of users, wave 12
D18    - Okta/AD deprovisioning detection and auto-disable, wave 13

Raised by Matt Mabrey 2026-09-17. Decision record and task breakdowns
only in this commit - no feature code yet.
2026-09-18 09:21:29 -07:00

7.9 KiB

Wave 13 — Okta/AD deprovisioning sync

Items: D18 Depends on: wave 10 merged (it is). Not blocked by wave 11 or wave 12, but shares server/app.py account-state surface with wave 12 (CR-020) — sequence commits to avoid an avoidable conflict. Decision record: docs/waves/decisions-2026-09-17.md

Six tasks, in build order. T13.3's fail-closed behavior is the most important done-when list in this wave — do not relax it to ship faster.


T13.1 — D18: shrink AUTH_SESSION_HOURS

  • Items: D18
  • Depends on: nothing (first task, independent of the rest)
  • Blocks: nothing
  • Surface: server/
  • Files: server/auth.py, server/.env.example, DEPLOYMENT.md

Do: Change the default from 12 to 2 hours. Document the change and the reasoning (defense in depth against an already-live session outliving a deprovisioning event) in server/.env.example next to AUTH_SESSION_HOURS, and in DEPLOYMENT.md's auth section. Note in the PR that this is proposed, not confirmed against the tenant's actual Okta SSO session policy — flag it for Nick/IT rather than treating 2 as unquestionable.

Done when:

  • default is 2 hours; still overridable via env, unchanged mechanism
  • .env.example and DEPLOYMENT.md explain why
  • existing session/auth tests updated for the new default where they assumed 12

T13.2 — D18: Okta Management API client

  • Items: D18
  • Depends on: nothing (independent of T13.1)
  • Blocks: T13.3
  • Surface: server/
  • Files: new server/okta_sync.py (or extend server/okta_auth.py — task's call), server/.env.example

Do: A small client for Okta's user-list endpoint, authenticated with a new, separate credential (e.g. OKTA_API_TOKEN) — not the OIDC client secret used for sign-in. Fetch the full user list (paginated per Okta's API) rather than one-by-one lookups per local user; this app has ~20-odd accounts today, and a list-and-diff is simpler and cheaper than N calls. Return each Okta user's identity-claim value (matching OKTA_IDENTITY_CLAIM, already confirmed live per wave 10) and status.

Follow the existing pattern for external credentials in this repo (MICRON_DB_URL, SMTP_PASSWORD): env-only, never logged, never returned to the browser in an error message.

Do not: reuse OKTA_CLIENT_ID/OKTA_CLIENT_SECRET for this. Sign-in and the management API are different trust boundaries with different scopes; conflating them means a compromise or rotation of one affects the other unnecessarily.

Done when:

  • the client authenticates with its own credential, distinct from the OIDC client
  • it fetches the complete Okta user list, handling pagination
  • an auth failure or malformed response raises a clear, specific error rather than returning an empty list indistinguishable from "everyone was deprovisioned" — this distinction is what T13.3 depends on
  • the credential is never logged or surfaced in any API response

T13.3 — D18: the sync job

  • Items: D18
  • Depends on: T13.2
  • Blocks: T13.4
  • Surface: server/
  • Files: server/okta_sync.py, server/app.py (or wherever log_event lives)

Do: Compare the Okta user list (T13.2) against local users rows. For any local is_active=True user whose Okta identity is missing from the list, or present with a non-active status, set is_active=False and write an AuditLog row (actor="system:okta_sync", action e.g. user_deprovisioned_by_sync, detail naming the Okta status found). Never touch a user already is_active=False. Never re-enable anyone.

The fail-closed rule, non-negotiable: if T13.2's client raises an error of any kind (network, auth, malformed response, timeout), this task takes no action on any account for that run and logs the failure clearly (server log, at minimum). An error is never treated as "Okta returned zero active users." Write a test that asserts this directly: feed the sync a failing client and assert zero rows changed and zero AuditLog entries written.

Done when:

  • a user missing from Okta's list, or present but not active, is disabled with a correctly-detailed audit row
  • a user already disabled is left alone (no duplicate audit row each run)
  • an active Okta user already active locally produces no audit row (only changes are logged, not a clean bill of health every cycle)
  • a simulated Okta API failure results in zero account changes and zero audit rows — this is the one check that must never be skipped or weakened
  • the job never re-enables an account under any input

T13.4 — D18: run it on a schedule

  • Items: D18
  • Depends on: T13.3
  • Blocks: T13.6
  • Surface: server/, docker-compose.yml
  • Files: server/app.py (startup hook) or a new sidecar per the backup container's pattern — task's call, per the decision record's noted alternative

Do: Wire T13.3 to run on an interval (proposed 15 minutes, env-overridable — e.g. OKTA_SYNC_INTERVAL_SECONDS, matching BACKUP_INTERVAL_SECONDS's naming). Default choice is an in-process background task in the api container; if a separate container is chosen instead, follow the backup service's shape (own Dockerfile or script, internal network plus whatever egress reaching Okta requires — check whether outbound as currently defined is sufficient or Okta needs a distinct allowance).

Done when:

  • the job runs automatically on the configured interval without manual invocation
  • interval is env-configurable with a sane default
  • a container restart does not produce a duplicate/overlapping run, and a slow cycle does not stack with the next one

T13.5 — D18: admin visibility

  • Items: D18
  • Depends on: T13.3
  • Blocks: T13.6
  • Surface: html/
  • Files: html/users.js / html/admin.js (wherever audit history is already surfaced)

Do: Confirm an auto-disable reads clearly wherever admins already look at account history — the actor string (system:okta_sync) should be self-explanatory in context, not require reading server logs to understand. Do not build new UI beyond making sure the existing audit surface renders this actor sensibly. Email/notification to admins on auto-disable is a noted fast-follow (the decision record flags it as open, not required here) — do not build it in this task; log it instead if it's tempting to add.

Done when:

  • an auto-disabled account's audit entry is visible and legible in the existing admin UI without special-casing
  • nothing here silently assumes CR-019's activity view exists yet — this must work standalone

T13.6 — D18: verification

  • Items: D18
  • Depends on: T13.4, T13.5
  • Blocks: nothing
  • Surface: server/
  • Files: as touched above

Do: Full verification per CLAUDE.md. Beyond the usual run: specifically re-run T13.3's fail-closed test in isolation and confirm it still passes after T13.4's scheduling wrapper is in place — the scheduling layer must not introduce a path that swallows the client's error and proceeds anyway.

Done when:

  • all D18 acceptance criteria in decisions-2026-09-17.md are met or a failure is stated with a reason
  • the fail-closed behavior is verified end to end through the scheduled wrapper, not just the bare sync function
  • smoke test and seed_demo.py both still pass
  • full test suite passes

Wave 13 exit criteria

  • AUTH_SESSION_HOURS defaults to 2, documented
  • the sync job runs on a schedule and correctly disables accounts Okta no longer shows as active
  • every auto-disable is individually audited with a clearly non-human actor
  • an Okta API failure of any kind changes zero accounts — verified, not assumed
  • the sync never re-enables an account
  • D18 fully accounted for, no open acceptance criteria