Files
Project-SDE-WP-Suite/docs/waves/wave-10.md
Cody Schaefer eecd724e02 T10.7 - close four more done-when boxes offline; token_version was mis-specified
I had said all six remaining boxes needed the live environment. Four did not,
and saying so was lazy scoping. ldap_auth_check now covers them (32/32):

  a just-provisioned account appears in GET /api/auth/users, the request the
  Admin console makes - without which nobody could grant a JIT account access

  granting admin through POST /api/auth/users/{id}/role, the request users.js
  sends, with the role verified to have actually changed. D13 criterion 4 was
  until now only asserted for role PRESERVATION, never for role GRANTING.

  no AD error-49 sub-code appears in any response body, while _err49 does parse
  one out of a real AD message - the log gets the detail, the caller does not

  token_version invalidates a cookie already issued, and leaves other sessions
  alone

That last box was wrong as written. It asked to exercise token_version "by a
role change", and nothing in app.py bumps it on a role change - or on a
deactivation. Neither needs to: get_current_user re-reads the account from the
database every request, so both take effect on the next request regardless.
T10.3's note that "role changes and deactivation should bump it" described an
intention rather than the code, and I repeated it without checking.

What token_version actually is now: a mechanism whose only trigger is the bump
manage_users makes on disable, which is belt-and-braces since is_active already
refuses the request. It works, and it is tested - but nothing much triggers it.
Logged as BL-030 rather than resolved here, because whether to wire it to
something (a "sign out everywhere" control is the usual shape) or remove it is
a session-handling design question, not an auth-wave bug.

Two boxes remain open, and both genuinely need your environment: member_of
against a real NESTED group, and the in-container openssl certificate check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 09:51:58 -05:00

399 lines
24 KiB
Markdown

# Wave 10 — Domain authentication over LDAPS
**Items:** `D13`, `D14`
**Depends on:** wave 9 merged (it is — `a8e28bf`)
One item, eight tasks. The item is stated in `docs/waves/decisions-2026-08-21.md`; read it
before starting, particularly the **Non-negotiables** section, which is where this change
goes wrong if it goes wrong.
**This wave is field-visible in one direction only.** Nobody gains a screen. People sign in
with their Windows password instead of an app password, the "Forgot password?" link
disappears, and admins stop issuing passwords. Everything else is invisible — which means
the done-when checks are the only evidence the wave worked.
> **BEFORE PRODUCTION: `LDAP_REQUIRED_GROUP` must be set.** It is empty by default,
> and empty means *no group gate* — every account in `prime.local` may sign in. The
> successful live sign-in on Aug 24 was made without it, so it proved the bind, the
> certificate chain and JIT provisioning, but **not** the group check: that path never
> executed. The group gate is covered against the fake directory in `ldap_auth_check`
> and has never run against the real directory. The startup line says which state you
> are in — `required group: (none configured — every domain account may sign in)`.
**Build order is task order** for once. `T10.1` is a standalone module with no callers and
should merge first; nothing else can be tested until it exists.
---
### T10.1 — The LDAPS client
- **Items:** `D13` (1)
- **Depends on:** nothing
- **Blocks:** T10.2, T10.4, T10.5, T10.7
- **Surface:** `server/`
- **Files:** new `server/ldap_auth.py`, `server/requirements.txt`, `docker-compose.yml`,
`Dockerfile`, `server/.env.example`
**Problem:** There is no directory client in the repo. `ldap3` is not a dependency, the API
container has no CA bundle, and the `api` service sits on the `internal` network, which has
no default gateway and therefore no route to `192.168.3.x` at all.
**Do:** Add `ldap3` (pinned, per the convention at the top of `requirements.txt`). Write
`server/ldap_auth.py` exposing two functions and nothing else:
- `verify(username, password) -> LdapResult | None` — simple bind as
`f"{username}@{domain}"` against `ldaps://{host}:636`.
- `member_of(conn, group) -> bool` — nested-group aware, via
`LDAP_MATCHING_RULE_IN_CHAIN` (`1.2.840.113556.1.4.1941`). `memberOf` alone is direct
membership only and will wrongly refuse anyone in a nested group.
Configuration by environment: `LDAP_HOST` (default `prime.local`), `LDAP_DOMAIN` (default
`prime.local`), `LDAP_CA_FILE`, `LDAP_REQUIRED_GROUP`, `LDAP_TIMEOUT_SECONDS`. Attach the
`outbound` network to `api` in `docker-compose.yml` — the same reason `assets_db.py` needed
it, and the comment there already explains the gateway-less `internal` network. Mount or
`COPY` the PEM bundle and point `LDAP_CA_FILE` at it.
Connect to **`prime.local`**, never a DC hostname and never an IP. See the decision doc for
why: SAN coverage plus round-robin across six DCs in one move. Wrap the bind in a retry
across resolved addresses, because round-robin will hand out a rebooting DC's address.
**Done when:**
- [x] `ldap3` is pinned to an exact version in `requirements.txt`
- [x] an empty or whitespace-only password returns failure **without calling `bind()`** — proven with `Connection` nulled, so any call to `bind()` would raise
- [x] an empty username returns failure without calling `bind()`
- [x] `Tls` is constructed with `validate=ssl.CERT_REQUIRED` and an explicit `ca_certs_file`
- [x] no code path sets `CERT_NONE`, and none falls back to the system trust store — asserted by an AST walk in `ldap_auth_check`, not a grep: the docstring names it to explain the ban
- [ ] `member_of` returns true for an account in a **nested** child of the required group
- [x] a bind against `192.168.3.37` (raw IP) fails hostname validation rather than silently passing — `selftest()` to the raw IP returns `untrusted`; to `prime.local`, `0 (ok)`
- [ ] `docker compose exec api openssl s_client -connect prime.local:636 -CAfile $LDAP_CA_FILE` reports `Verify return code: 0 (ok)`
- [x] the module imports cleanly with no LDAP env set (unconfigured is a first-class state, as with `MICRON_DB_URL`)
---
### T10.2 — The login path binds instead of hashing
- **Items:** `D13` (1)
- **Depends on:** T10.1
- **Blocks:** T10.3, T10.4
- **Surface:** `server/`
- **Files:** `server/app.py` (`login`), `server/auth.py`
**Problem:** `login()` at `server/app.py:681` calls `auth.verify_password` against
`user.password_hash`. The lockout counter it maintains is about to start counting *domain*
bind failures, which changes what that counter is for.
**Do:** Replace the credential check with `ldap_auth.verify`. Keep the surrounding shape —
the deliberate timing equalisation, the generic `401`, the `403` for a disabled local
account, `last_login_at`, the session cookie. Rework the throttle so the local counter trips
**before** the domain policy is reached and short-circuits without calling the DC: the
failure mode to avoid is this endpoint being usable to lock domain accounts out of Windows.
Log directory error-49 sub-codes for diagnosis; return the same generic message regardless.
**Break-glass: none — decided Aug 21, see the decision doc.** LDAPS is the only way in, so
this task adds no fallback path. What it must add instead is *visibility*: a startup log line
stating whether LDAP is configured and whether the DC answered, because with no fallback a
misconfigured deploy is indistinguishable from a forgotten password at the login box.
**Done when:**
- [x] a correct domain password signs in and sets the session cookie — confirmed against the live domain Aug 24, and in `ldap_auth_check`
- [x] a wrong password is refused with the generic message
- [x] a blank password is refused (guards `T10.1` from the caller's side too)
- [x] a user not in the required group is refused even though the bind succeeded
- [x] `is_active = false` locally still refuses, independent of the directory — a disabled account is refused 403 even though the bind succeeds
- [x] the local throttle trips below the domain lockout threshold and stops calling the DC — proved by presenting a CORRECT password once the budget is spent: a 429 for a credential that would otherwise work is only possible if the throttle runs before the directory is consulted
- [x] no response body distinguishes "no such user" from "wrong password"
- [x] error-49 sub-codes appear in the log and nowhere in any response — asserted both ways: `_err49` parses a real AD message, and no sub-code appears in any response body
---
### T10.3 — Drop `password_hash`
- **Items:** `D13` (1)
- **Depends on:** T10.2
- **Blocks:** T10.6, T10.8
- **Surface:** `server/`
- **Files:** `server/models.py`, new migration, `server/app.py`, `server/auth.py`,
`server/manage_users.py`
**Problem:** With binds doing the work, every password code path is dead weight and a
liability. It is also the criterion that makes this item irreversible, so it lands on its own
commit.
**Do:** Remove `password_hash` from `models.User` and drop the column in a new Alembic
revision with `down_revision = 'a1b8c6d4e2f9'` (the current head — confirm with
`alembic heads` rather than trusting this line). Remove from `auth.py`: `hash_password`,
`verify_password`, `password_problem`, `MIN_PASSWORD_LEN`, `_COMMON_PASSWORDS`,
`create_reset_token`, `decode_reset_token`, `RESET_MINUTES`. Remove from `app.py`:
`/api/auth/forgot-password`, `/api/auth/reset-password`, `/api/auth/reset-available`,
`/api/auth/password`, `/api/auth/users/{user_id}/password`, and the `_reset_last` throttle
with `reset_body`. Remove `reset-password` from `manage_users.py` and the password prompt
from `create` / `create-admin`.
`token_version` **stays.** It is still the session-revocation mechanism — role changes and
deactivation should bump it even though password changes no longer exist.
**Do not** remove the account-management endpoints themselves. `/api/auth/users/{id}/role`
is criterion 4 and must keep working.
**Done when:**
- [x] `grep -rn "password_hash\|hash_password\|verify_password\|password_problem" server/` returns nothing outside the migration — only the migration and one docstring naming the dropped column
- [x] `alembic upgrade head` then `downgrade -1` round-trips on SQLite and on Postgres (16.15, the compose image — Aug 24; needed a pre-existing T8.6 migration bug fixed first, see `495d87d`)
- [x] the migration's `downgrade()` recreates the column nullable, not `NOT NULL` — there are no hashes to put back — verified on SQLite and Postgres
- [x] `token_version` still invalidates an already-issued session — but **not** via a role change, which was the wrong premise. Nothing in `app.py` bumps it any more: `get_current_user` re-reads the account every request, so `role` and `is_active` changes take effect immediately without it. Its one remaining trigger is `manage_users` on disable. Tested by bumping it directly: the old cookie 401s and every other session is untouched. See `BL-030`.
- [x] `manage_users.py list`, `disable`, `enable` still work; `reset-password` is gone
- [x] no CLI command prompts for a password — superseded by `T10.9`, which removed `create-admin` and `create` outright; the first admin is now bootstrapped by signing in and then `promote`
---
### T10.4 — Just-in-time provisioning, without trampling existing accounts
- **Items:** `D13` (2, 4)
- **Depends on:** T10.2
- **Blocks:** T10.7
- **Surface:** `server/`
- **Files:** `server/app.py` (`login`), `server/auth.py`
**Problem:** Criteria 2 and 4 pull in opposite directions. Provisioning on first login must
create accounts that do not exist, and must not touch the role of accounts that do — an
existing `admin` signing in for the first time after this wave must still be an admin
afterwards.
**Do:** On a successful bind that passes the group check, look the account up with
`auth.find_user` (already case-insensitive across username **and** email). If it exists,
update only `last_login_at` and — if empty locally — `full_name` and `email` from the
directory. **Never write `role`.** If it does not exist, create it at
`ROLE_PROJECT_USER` with `full_name`/`email` from the directory.
**A JIT account gets NO project access, and that is deliberate.** An earlier draft of this
task said to honour the `auto_add_projects` machinery so a new account "lands in the right
projects" — that was wrong about how the flag works. `auto_add_projects` is evaluated when a
**project** is created (`app.py:245`), marking accounts that should join every *new* job; it
cannot retroactively add a new account to existing ones. There is no correct default, so
least privilege applies: the account exists, can sign in, and sees nothing until someone
grants access. That is a real UX cliff — a successful sign-in into an empty app — so it has
to be visible to admins rather than silent, which is what the `AuditLog` row is for.
Note the flush-order warning in the `models.py` docstring: `create_user` in `app.py` handles
account-then-membership correctly in one flush — follow it if you add rows.
Write an `AuditLog` row for each JIT creation. An account appearing without an administrator
creating it is exactly the kind of event that record exists for.
**Done when:**
- [x] an unknown username with a valid bind and group membership gets a `users` row at `project_user`
- [x] `full_name` and `email` are populated from the directory on creation
- [x] an existing `admin` signing in is still `admin` afterwards — asserted, not assumed — asserted in `ldap_auth_check` against a real server
- [x] an existing account with a locally-set `full_name` does not have it overwritten
- [x] a JIT account has NO `ProjectMember` rows and sees no projects
- [x] the new account appears in the Admin console user list so access can be granted — asserted through `GET /api/auth/users`, the request the console makes
- [x] each JIT creation writes an `AuditLog` row
- [x] a failed bind creates **no** row
- [x] a bind that succeeds but fails the group check creates **no** row
---
### T10.5 — CLOSED, NOT BUILT (Aug 24 2026): the required group stays an env var
- **Items:** `D13` (3)
- **Status:** **won't build.** `LDAP_REQUIRED_GROUP` in the environment is the answer.
**Do not build this later by reading the original task and assuming it was skipped.**
It was proposed, examined and rejected on purpose, and the reasoning is below.
**What it was going to be:** the required group moved out of the environment into an
Admin console setting, with a validate-on-save guard that resolved the group in the
directory and confirmed the saving admin was a member.
**Why it is not being built:**
1. **The console requirement was invented here, not asked for.** D13 criterion 3 says
*"An AD group is configured"* — not "configurable from the console". The env var
satisfies the criterion as written.
2. **The validate-on-save guard existed only to defend against a risk the console
itself introduced.** A feature whose complexity exists to defend against itself is
usually the wrong feature.
3. **The lockout scenario it defended against is already handled.** A group that does
not resolve raises `LookupError` in `member_of`, which `verify()` maps to
`GROUP_NOT_FOUND`, which `is_config_problem` classifies as ours — so `login()`
answers **503**, not 401, and the log says *"required group 'X' does not resolve in
DC=prime,DC=local — refusing the sign-in. This is a configuration fault, not a bad
password."* A genuine non-member still gets 401. The two are already distinguishable
in both the log and the response.
4. **A redeploy is deliberate and reviewable; a text box is not.** The group is set
once and effectively never changes — it is not SMTP configuration.
5. **The console version creates a circular failure.** Fixing a lockout would require
the console the lockout prevents you from reaching. Editing the env var does not.
**What is genuinely lost, and accepted:**
- **Discoverability.** An app admin cannot see which group is required without
Portainer or shell access. A read-only line in the Admin console diagnostics would
give the useful half without the dangerous half; it was offered and declined on
Aug 24 as not needed.
- **Deploy-time validation.** Nothing confirms the group resolves until the first
sign-in attempt. This is not fixable: resolving a group needs an authenticated
search, anonymous bind is disabled on this estate, and there is no service account
by design. The first-attempt 503 is the earliest possible detection.
### T10.6 — Strip the password UI
- **Items:** `D13` (1)
- **Depends on:** T10.3
- **Blocks:** nothing
- **Surface:** `html/`
- **Files:** `html/login.html`, `html/login.js`, `html/admin.js`, `html/users.js`,
`html/auth-guard.js`
**Problem:** `login.html` has three views — sign-in, forgot-password, set-new-password — and
two of them now point at endpoints that no longer exist. `auth-guard.js` has a
change-password dialog, and the user-admin UI has a "reset password" action per row.
**Decided Aug 21: "Forgot password?" is KEPT and repointed at
`https://primecontrols.okta.com/`.** An earlier draft of this task removed the link, and a
plain sentence saying "contact IT" was proposed instead. Okta is the better answer — it is a
real self-service path, and with no app password and no break-glass it is the only recovery
route that exists. Note this is the first sign of an Okta tenancy on this estate; see the
new backlog entry.
Mechanics that matter: it is a plain `<a href>`, not a form post, so the `form-action 'self'`
in the CSP does not apply and no `navigate-to` directive is set — off-origin link navigation
is allowed as-is. `target="_blank"` needs `rel="noopener noreferrer"`, and there must be NO
click handler on `#forgot-link`: the old one called `preventDefault()` to swap views, and
leaving it would silently swallow the navigation.
**Do:** Remove the `#view-forgot` and `#view-reset` sections and the reset-token handling in
`login.js`, and repoint `#forgot-link` as above. Remove the change-password dialog from `auth-guard.js`
(`backlog.md:180` refers to it) and the per-row password reset from the users UI. Keep the
sign-in form, and relabel the password field's hint to say it is the Windows/domain password
— people need to know which password to type.
Keep the role-granting controls exactly as they are. That is criterion 4.
**Done when:**
- [x] `grep -rn "forgot\|reset-password\|new-password" html/` returns nothing but prose
- [x] "Forgot password?" opens `https://primecontrols.okta.com/` in a new tab — by inspection of the markup: `target="_blank"` with `rel="noopener noreferrer"`
- [x] `#forgot-link` has NO click handler (a `preventDefault()` would swallow the navigation)
- [x] a 503 from the login endpoint says sign-in is unavailable, not that the password is wrong
- [x] the sign-in form still submits, and a failure still announces through `role="alert"` (`login.html` already does this correctly — do not regress it) — `url_state_check` drives the real form end to end; the `role="alert"` region is untouched
- [x] the password field says which password to enter — see the screenshots
- [x] no dead `<a href="#">` or handler remains for a removed view
- [x] granting admin to an existing user still works from the console — asserted through `POST /api/auth/users/{id}/role`, the request `users.js` sends, and the role really changes
- [x] exercised at 390px and at 1440px, screenshots in the PR — `docs/reference/baseline/before-wave10/` and `after-wave10/`, captured Aug 24 with `tests/baseline_shots.py`. "Before" comes from a detached worktree at `main` (`a8e28bf`) so each half was shot against its own server.
- [x] no raw hex added to any stylesheet (the token rule) — no CSS was added at all — the hint reuses the `.hint` class the page already had
---
### T10.7 — Make the suite testable without a domain controller
- **Items:** `D13` (1, 2, 3)
- **Depends on:** T10.1, T10.4
- **Blocks:** T10.8
- **Surface:** `tests/` + `server/`
- **Files:** `server/ldap_auth.py`, `tests/browser_check.py`, `tests/launcher_check.py`,
`tests/console_dialogs_check.py`, `tests/url_state_check.py`, `server/smoketest.py`,
`server/seed_demo.py`, new `tests/ldap_auth_check.py`
**Problem:** This is the task most likely to be underestimated. Four existing checks build
users with `password_hash=auth.hash_password(PW)` and sign in over HTTP;
`console_dialogs_check.py` drives the password-reset prompt specifically. `smoketest.py` and
`seed_demo.py` both sign in. None of them can reach a DC, and CI has no domain.
**Do:** Make the bind injectable — a module-level seam in `ldap_auth.py` that a test can
substitute (a fake directory: usernames, passwords, groups, full names), selected by an env
var that is refused when a real `DATABASE_URL` is configured, mirroring how
`auth._load_secret` refuses an ephemeral key in production. Port the four checks onto it.
Delete the password-reset half of `console_dialogs_check.py` — the flow it covers no longer
exists — and say so in the PR rather than leaving a skipped test.
Add `tests/ldap_auth_check.py` covering the non-negotiables from the decision doc: empty
password, empty username, `CERT_REQUIRED` asserted by inspecting the constructed `Tls`,
nested group membership, group-check refusal creating no user, and existing-admin role
preservation.
**Done when:**
- [x] the stub backend cannot be selected when a non-SQLite `DATABASE_URL` is set — asserted by a test
- [x] `tests/ldap_auth_check.py` covers the cases above — 20/20
- [x] the empty-password case is proved by nulling `Connection`, so any call to `bind()` would raise — it asserts the guard returns *before* the transport, not merely that the result is a failure
- [x] `smoketest.py` and `seed_demo.py` document which credentials they now need (T10.8)
- [x] the removed password-reset checks are called out, not silently dropped — `console_dialogs_check.py`'s docstring records the coverage loss and where the prompt kit is still covered
- [x] the full `tests/` suite passes with no DC reachable — **1284/1288 checks across 40 files**, Aug 24. Two files fail, both proven pre-existing and unrelated (`BL-028` `assets_check`, `BL-029` `generalinfo_check`): `git diff main...HEAD` shows this branch touches neither file. `token_check.py` is not a member of the suite — it is a capture/diff tool that requires `--out` or `--compare`, and a sweep script that runs it bare gets a usage message and exit 2.
**Scope note.** This task was estimated as far larger than it turned out to be. The
premise was that four checks sign in and would all need the seam; in fact `seed()` mints
a token with `auth.create_token()` and sets the cookie directly, so **no** browser check
signs in except `url_state_check`'s deep-link case. `browser_check` and `launcher_check`
needed one kwarg deleted each — and since 39 files import `seed`/`start_server` from
`browser_check`, that single line unblocked nearly the whole suite.
---
### T10.8 — Documentation and the deploy runbook
- **Items:** `D13`
- **Depends on:** T10.3, T10.5, T10.7
- **Blocks:** nothing
- **Surface:** docs
- **Files:** `DEPLOYMENT.md`, `server/README.md`, `DEPLOY-login-portal.md`, `CLAUDE.md`,
`IMPLEMENTATION.md`
**Problem:** `DEPLOY-login-portal.md` documents creating the first admin with a password and
is the page an admin will reach for. `DEPLOYMENT.md` describes `AUTH_SECRET_KEY` and SMTP but
knows nothing about a directory. `CLAUDE.md`'s verification section tells anyone touching the
frontend to run the smoke test, which changes here.
**Do:** Document the LDAP variables, how to produce the CA bundle from the two thumbprints
in the decision doc, and the `prime.local`-not-an-IP rule with the reason. Add the
`openssl s_client -CAfile` check as the first-line diagnostic. Rewrite the
`DEPLOY-login-portal.md` bootstrap step: the first admin is now an existing directory account
promoted with `manage_users.py`, not an account created with a password. State plainly what
happens when the DC is unreachable, whatever `T10.2` decides.
**Done when:**
- [x] every new env var is documented in `server/.env.example` and `DEPLOYMENT.md`
- [x] the CA bundle procedure is reproducible by an admin who has not read this thread — thumbprints and a `Get-ChildItem` one-liner in `DEPLOYMENT.md`
- [x] `DEPLOY-login-portal.md` no longer instructs anyone to set a password — rewritten, with a note saying what it replaced so an admin holding the old copy is not misled
- [x] the DC-unreachable behaviour is stated explicitly, with the diagnostic commands
- [x] `IMPLEMENTATION.md` section 4 lists wave 10
- [x] no doc still claims passwords are stored as bcrypt hashes (swept; remaining matches all say the opposite)
- [x] `CLAUDE.md` carries the four load-bearing auth rules, next to the token rule
---
### T10.9 — D14: the CLI authenticates, and stops creating accounts
- **Items:** `D14`
- **Depends on:** T10.3
- **Blocks:** T10.8
- **Surface:** `server/`
- **Files:** `server/manage_users.py`
**Problem:** `manage_users.py` writes to the `users` table with no authentication at all.
It also still offers `create-admin` / `create`, which are redundant now that accounts
provision themselves — and worse than redundant, because a hand-typed username can end up
matching no directory identity.
**Do:** As stated in `D14`. Remove the two create commands, add `promote` / `demote`, gate
every state-changing command on a prompted domain bind, and write an `AuditLog` row naming
the operator. Write the audit row by hand rather than importing `log_event` from `app.py`
that would pull FastAPI and the whole application into a CLI startup for one INSERT.
**Done when:**
- [x] `create-admin`, `create` and `reset-password` are rejected as invalid choices
- [x] `list` works with no credential and with no LDAP configured
- [x] a state-changing command with LDAP misconfigured refuses instead of proceeding
- [x] there is no `--password` flag on any command
- [x] `promote` raises a role; `demote` returns an account to `project_user`
- [x] promoting YOURSELF is allowed and recorded with `self: true`
- [x] the last active admin cannot be demoted
- [x] an unknown account gives an error that says accounts are made on first sign-in
- [x] every change writes an `AuditLog` row naming the operator
- [x] verified against a real domain bind — `promote` and `demote` confirmed working Aug 24 2026