Files
Project-SDE-WP-Suite/DEPLOY-runbook-2026-09-03.md
Matt Mabrey d6eae0d846 T10.9: rollback-aware deploy runbook for the Okta cutover
New DEPLOY-runbook-2026-09-03.md, separate from the 2026-08-04 runbook.
Names the five new OKTA_* env vars, treats the pre-deploy backup as the
only way back once 1d60a608bb51 (drop_local_password) commits since its
downgrade() restores the column but not the data, and splits Rollback
into the fixable case (Okta app integration misconfigured, fix and
redeploy api, no data at risk) versus the severe case (abandoning Okta
for local-password code, which only a destructive backup restore can
reach). States D16's no-break-glass posture plainly.

D17 records the decision and why: staged-deploy-sequence docs and
backlog-only were both considered and declined in favor of the runbook.

Logged to backlog.md rather than fixed here: okta_auth.describe() has
no caller (BL-028), users.failed_attempts/locked_until are vestigial
(BL-029), DEPLOY-login-portal.md is fully stale (BL-030).
2026-09-03 14:25:18 -07:00

356 lines
14 KiB
Markdown

# Deploy runbook: WP Suite Okta cutover (wave 10)
**For:** IT / whoever administers the Docker host and Portainer
**From:** m.mabrey@prime-controls.com
**Revised:** 2026-09-03. First version of this runbook.
**Expected duration:** 20-30 minutes, including the backup and the live sign-in check
**Expected downtime:** under a minute, while containers are recreated
This is a separate runbook from `DEPLOY-runbook-2026-08-04.md`, not a revision of it.
That one still applies to its own deploy. Read this one in full before starting.
It has a step that the earlier one does not: this deploy removes the local password
system entirely, and one part of that removal cannot be undone by the usual means.
See "What this deploy changes" below.
---
## Fill these in before handing this over
| Thing | Value |
|---|---|
| Docker host (SSH target) | `________________` |
| Stack name in Portainer | `________________` |
| Site URL | `https://________________` |
| Stack directory on the host (holds `docker-compose.yml` / `backups/`) | `________________` |
| A real Okta account assigned to the app integration, for the live check in Step 4 | `________________` |
Container names are fixed by the compose file and are the same on every host:
`nginx_webserver`, `wp_api`, `wp_db`, `wp_db_backup`.
---
## What this deploy changes
Sign-in changes from a local username/password to Okta OIDC, completely. Not a
toggle, not a fallback. Three things make this more than a routine deploy:
1. **Five new environment variables are required**, and without them nobody can
sign in: `OKTA_ISSUER`, `OKTA_CLIENT_ID`, `OKTA_CLIENT_SECRET`,
`OKTA_REDIRECT_URI`, and optionally `OKTA_IDENTITY_CLAIM`. This is a genuine
change from how deploys usually go here. Do not skip Step 1.5.
2. **A migration drops the `password_hash` column** (`1d60a608bb51`), and it is
**one-way in practice.** Its `downgrade()` re-adds the column, but empty.
The real password hashes are gone the moment this commits, and no `alembic
downgrade` brings them back. If this deploy needs to be undone after that
point, going back to the old local-password code does not work on its own.
See the Rollback section, Case C. This is why Step 1's backup is not
optional the way it sometimes reads in other runbooks.
3. **There is no break-glass path**, by design (recorded decision D16). If Okta
is unreachable or misconfigured after this deploy, the app is unreachable for
everyone, admins included, until Okta is fixed. That is expected behavior,
not a bug to roll back from. See Rollback, Case A/B, before assuming
something is broken.
The visible change for people using the app: the login page becomes "Sign in
with Okta" instead of a username/password form. Nothing else in the app's
day-to-day behavior changes.
---
## Step 0: Record the current state (needed for rollback)
SSH to the Docker host and run:
```bash
docker exec wp_api alembic -c server/alembic.ini current
docker inspect nginx_webserver --format 'nginx image: {{.Image}}'
docker inspect wp_api --format 'api image: {{.Image}}'
```
**Copy the output into your ticket.** Also note the Git commit the Portainer
stack is currently on (Portainer, the stack, the Git reference / last-updated
commit). Without these, rollback is guesswork.
Expected output of the first command before this deploy: `a1b8c6d4e2f9 (head)`.
If it shows anything else, stop and check with me before continuing. This
runbook assumes that starting point.
---
## Step 1: Back up the database
On the Docker host:
```bash
docker exec wp_db_backup /scripts/db-backup.sh
```
Expected output ends with a line like:
```
[db-backup] wrote 1.4M /backups/wpsuite-20260903-141233Z.sql.gz.enc
```
Confirm the file is on the host (substitute the stack directory):
```bash
ls -lt <stack-dir>/backups | head -3
```
**Record that filename.** Do not continue until you have seen the `wrote ...`
line and the file in that listing.
This backup matters more than usual for this deploy. Once the migration in
Step 3 commits, this file becomes the *only* way to get local password hashes
back, for any reason. Treat it as the point you would restore to, not routine
housekeeping.
- A `.sql.gz.enc` extension means backups are encrypted. Expected and correct.
- A `.sql.gz` extension plus a `WARNING: BACKUP_ENC_PASSPHRASE not set` line
means backups are unencrypted. Not a blocker for this deploy; report it back.
- **No SSH access?** Portainer, **Containers**, `wp_db_backup`, **Console**,
connect with `/bin/sh`, then run `/scripts/db-backup.sh`. Same result: the
dump lands on the host, because `/backups` is a bind mount.
---
## Step 1.5: Confirm the Okta app integration is actually ready
Do this before redeploying, not after. Everything here is checked in Okta's own
admin console and in the values that will go into the stack's environment
variables. Nothing touches the WP Suite host yet.
1. The Okta app integration exists (Sign-in method: OIDC, Authorization Code,
Application type: Web Application), and the account listed in the fill-in
table above is assigned to it.
2. `OKTA_REDIRECT_URI` matches a "Sign-in redirect URI" registered on that app
integration **exactly**: scheme, host, and path, including whether it ends
in `/api/auth/okta/callback`.
3. `OKTA_ISSUER`, `OKTA_CLIENT_ID`, and `OKTA_CLIENT_SECRET` are the values from
that same app integration, not a different one.
4. If your Okta configuration puts the directory identity somewhere other than
the `preferred_username` claim, `OKTA_IDENTITY_CLAIM` is set to the right
claim name. If unsure, leave it unset; `preferred_username` is the default.
Add all five to the stack's **Environment variables** in Portainer now, before
Step 2. `OKTA_CLIENT_SECRET` should be handled the same way `AUTH_SECRET_KEY`
already is: not typed anywhere it will be logged.
If any of items 1-3 above are not yet confirmed, stop here and get them
confirmed first. A wrong redirect URI or an unassigned account will not corrupt
anything, but it does mean nobody signs in after this deploy until it is fixed.
See Rollback, Case B, which is the ordinary way that gets fixed and does not
involve the database at all.
---
## Step 2: Redeploy the stack in Portainer
1. Portainer, **Stacks**, select the stack.
2. **Pull and redeploy**, with re-pull / re-build **enabled**.
3. Wait for it to report success.
A plain "restart" or "stop/start" will not pick up new code, and will not pick
up the environment variables added in Step 1.5 either.
---
## Step 3: Confirm the containers came up
```bash
docker ps --filter name=nginx_webserver --filter name=wp_api --filter name=wp_db
```
All three must be `Up`, and `wp_db` should show `(healthy)`. Then check the API
applied its migration cleanly:
```bash
docker logs wp_api --tail 40
```
You are looking for an Alembic `Running upgrade a1b8c6d4e2f9 -> 1d60a608bb51`
line followed by gunicorn starting up, and no traceback. The API refuses to
start if a migration fails, so a restarting `wp_api` container means it failed.
Go to Rollback, Case B, and read the "did the migration commit" note there
before doing anything to the database.
Confirm the database landed on the new revision:
```bash
docker exec wp_api alembic -c server/alembic.ini current
```
Expected: `1d60a608bb51 (head)`. **Once you see this, you have passed the
point of no return described above.** The backup from Step 1 is now the only
way back to a working local-password system, if that is ever needed.
Then verify nginx's own view of its config:
```bash
docker exec nginx_webserver nginx -t
```
Expected: `syntax is ok` / `test is successful`.
---
## Step 4: Confirm Okta sign-in actually works, live
This is the step that matters most for this deploy. A clean container start
does not by itself prove sign-in works, and there is currently no startup log
line that confirms Okta config is good (logged separately as a follow-up, not
fixed as part of this runbook). The only real proof is a live sign-in.
1. Open the site's normal URL in a private/incognito window. It should land on
`login.html` with a "Sign in with Okta" button, not a username/password
form.
2. Click it. You should be redirected to your organization's actual Okta
sign-in page (the real Okta domain from `OKTA_ISSUER`, not this app's own
domain).
3. Sign in with the account from the fill-in table. You should land back on
the WP Suite site, signed in.
4. If this is the account's first-ever sign-in, it is now JIT-provisioned as a
regular user (`project_user`). To make it an admin, on the Docker host:
```bash
docker exec -it wp_api python -m server.manage_users promote <username> --role admin
```
This only works on an account that has already signed in once through Okta.
It promotes an existing row; it does not create one. That is deliberate
(recorded decision D16): there is no other admin-bootstrap path.
If step 2 or 3 fails (redirected to an Okta error page, redirected back to
`login.html` with an error, or nothing happens), this is almost always a
configuration problem from Step 1.5, not a code or database problem. Go to
Rollback, Case B, before considering anything more drastic.
Also confirm the API is reachable through the proxy and the redirect itself is
wired up:
```bash
curl -s https://<site-url>/api/health # -> {"ok": true}
curl -sI https://<site-url>/api/auth/okta/login | grep -i ^location # -> your Okta authorize URL
```
---
## Step 5: Hard-reload once in a browser, then sanity-check
Press **Ctrl+Shift+R** (Cmd+Shift+R on macOS) once. The app uses a service
worker; a normal reload can serve the previous version.
1. Signed in as the account from Step 4, the home page offers to select or
create a project, same as before.
2. Open **Admin Console** as the promoted admin account. The user table shows
the account you just signed in with. There is no password column, no
"reset password" action anywhere in the UI.
3. Open a project and confirm a work package can be opened and edited
normally. Sign-in is the only thing this deploy changes, so the rest of
the app should look untouched.
**Deploy complete.** Please report back: the Step 0 output, the backup
filename from Step 1, and confirmation that Step 4's live sign-in worked.
---
## Rollback
Read this before assuming a rollback is needed. Cases A and B below do **not**
touch the database and are the far more likely outcome of something going
wrong here. Okta configuration is fiddly and easy to get slightly wrong. Case
C is the severe, destructive one, and should be a last resort, not a first
reaction.
### Case A: nginx won't start, or containers won't come up at all
Same as any other deploy: the database is untouched by container start-up
failures. In Portainer, redeploy the stack pinned to the **previous Git
commit** recorded in Step 0, then re-run Step 3.
```bash
docker logs nginx_webserver --tail 100
docker logs wp_api --tail 100
```
Send me whichever of those is relevant.
### Case B: containers are up, but Okta sign-in doesn't work
This is a configuration problem, not a data problem, and does **not** need a
code rollback or a database restore. Check, in order:
1. Is `wp_api`'s log showing anything at all when a sign-in is attempted?
`docker logs wp_api --tail 100`.
2. Do the five `OKTA_*` values in the stack's environment variables actually
match the Okta app integration (Step 1.5)? A copy-paste error in
`OKTA_CLIENT_SECRET` or a redirect URI that's off by a trailing slash are
the two most common causes.
3. Is the account assigned to the Okta app integration? An unassigned account
gets denied by Okta itself, before it ever reaches this app.
4. If a specific person can't sign in but others can, check
`OKTA_IDENTITY_CLAIM`. The claim it reads may not carry that person's
directory identity in the format expected. Confirm with security which
claim Okta is actually issuing.
Fix the environment variable(s) in Portainer, then redeploy (Pull and redeploy
is fine; the migration already applied and does not run again). No backup
restore, no code rollback.
If it's still not working after checking all four, send me the `wp_api` log
from item 1 along with which of items 2-4 you already ruled out.
### Case C: the decision is made to abandon Okta and restore local-password sign-in
This is the case the point-of-no-return warning in "What this deploy changes"
is about. Only reach for this if Case A and B do not apply. That is, Okta
itself is working correctly but a decision has been made to go back to the old
system entirely.
**This cannot be done with a code rollback alone.** The old code expects a
real `password_hash` on every user row. After Step 3 commits, that column is
either gone or (if `alembic downgrade` is run) present but empty. Either way,
nobody's stored password survived, including admins'. The only way to get a
working local-password system back is to restore the full database from the
Step 1 backup, which also rolls back every other change made since that
backup: new work packages, comments, uploaded files, everything.
**Do not do this without confirming with me first.** If it is confirmed, for
an encrypted dump, on the Docker host, in the `backups` directory:
```bash
export BACKUP_ENC_PASSPHRASE='<the passphrase, from the stack env vars>'
openssl enc -d -aes-256-cbc -pbkdf2 -pass env:BACKUP_ENC_PASSPHRASE \
-in wpsuite-<timestamp>.sql.gz.enc \
| gunzip \
| docker exec -i wp_db psql -U wpsuite -d wpsuite
unset BACKUP_ENC_PASSPHRASE
```
For an unencrypted dump, drop the `openssl` stage and pipe `gunzip` straight
into `psql`. Substitute the real values if `POSTGRES_USER` / `POSTGRES_DB` are
not `wpsuite`. After restoring, redeploy pinned to the Git commit recorded in
Step 0, since the restored database matches the old schema, not this one.
Reach me at m.mabrey@prime-controls.com.
---
## Notes
- Do not run `docker compose down -v`. The `-v` flag deletes the `pgdata`
volume and with it the entire database.
- `docker exec <container-name>` is used throughout rather than
`docker compose ...`, because a Portainer-managed Git stack's compose project
lives under Portainer's own data directory and usually isn't reachable from
an ad hoc SSH session the same way.
- There is currently no startup log line confirming Okta config is valid
(`okta_auth.describe()` exists but nothing calls it yet, logged as a
follow-up). Step 4's live sign-in is the real verification until that's
wired in.
- Full background documentation: `DEPLOYMENT.md` and `server/README.md` in the
repository. `docs/waves/decisions-2026-09-03.md` (D16, D17) records why
there is no break-glass path and why this runbook exists as its own document.