Files
Project-SDE-WP-Suite/DEPLOY-runbook-2026-08-04.md
n.siegfried 4ace2afb1c Move user administration to its own page; add Project Super User
User accounts lived in the Admin Console, which is admins-only. Project admins
need to create the accounts on their own jobs without an app admin on the phone,
so accounts move to a new User Directory page and a new role carries the right.

server/auth.py, server/app.py
  New permissions role `project_super_user`, between admin and project_admin:
  everything a project admin may do, plus user administration SCOPED to the
  projects they hold the role on. Four limits make it safe to hand out, all
  enforced server-side:

    * Scope comes from projects, not the job title. It resolves per membership
      (managed_project_ids), so an ordinary account can hold it on one job via
      ProjectMember.role, and a super user demoted on one job administers
      nobody there. No projects, no authority.
    * Account-level changes (password, disable, rename, permissions, delete)
      require EXCLUSIVE scope: refused when the target is also on a project the
      caller does not administer, because those changes are global. The
      directory renders such rows read-only with the reason.
    * No admin or super-user targets, and neither role can be granted by a
      super user -- that is the line that stops it becoming app-wide control.
    * PUT .../projects rebuilds only the caller's own slice; memberships on
      projects they do not administer are left untouched. A payload that simply
      omits them must not cut someone off a job the caller cannot see.

  Creating requires naming at least one of your own projects: an account with
  none would be one the creator instantly cannot manage.

  /api/auth/users is now scoped rather than admin-only, and carries a per-row
  `manageable` verdict plus the reason. Non-managers get a contact card only --
  a project user has no business reading colleagues' login history. New
  /api/auth/user-scope tells the page what it may offer. Administrative
  password resets are now audited; they were the one account change that left
  no trace. Settings, feature flags and the auto-add rule stay admin-only.

  While here: one definition of "is a user manager", derived from the managed
  set. An account-role-only version disagreed with the scoped one and locked
  per-project super users out of routes they were entitled to.

html/users.html, html/users.js
  The directory: three renderings from one page -- admin (everything), super
  user (controls per row, read-only where scope is shared), everyone else (a
  read-only directory of the people on their own projects).

html/console.css, html/console-util.js
  Extracted from admin.html/admin.js so both console pages share them. A
  divergent jsq() is an XSS and a divergent role list offers permissions the
  server refuses, so neither may exist twice.

html/wp-sidenav.{js,css}
  Global nav drawer, role-gated, carrying ?project= across links. Mounted on
  the field view (which had no way to anywhere) plus both console pages.

No migration: users.role is already String(20) and the new value fits.

Verified: 93 scope/gate tests, 29 live HTTP tests through the real dependency
stack, 33 static JS checks. Not verified in a browser -- no JS engine on this
machine -- so users.html and field.html want one manual load.

server/smoketest.py still fails with 401s. Pre-existing: it has no login code,
so auth_gate refuses it. Confirmed unchanged by stashing this work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 17:36:14 -07:00

11 KiB
Raw Blame History

Deploy runbook — WP Suite

For: IT / whoever administers the Docker host and Portainer From: n.siegfried@prime-controls.com Revised: 2026-08-05 — this replaces the 2026-08-04 version. Same procedure, but the deploy now carries a second database migration and a new admin screen. If you already have the earlier copy, work from this one instead. Expected duration: 1015 minutes, including the backup Expected downtime: under a minute, while containers are recreated


Fill these in before handing this over

Thing Value
Docker host (SSH target) ________________
Stack name in Portainer ________________
Site URL https://________________
Stack directory on the host (holds docker-compose.yml / backups/) ________________

Container names are fixed by the compose file and are the same on every host: nginx_webserver, wp_api, wp_db, wp_db_backup.


What this deploy changes

Front-end and nginx changes, a new admin screen, plus pending database migrations that run automatically. Three things make it more than a routine restart:

  1. The nginx config and the entire html/ directory are baked into the container image at build time. A plain restart deploys nothing — the stack must be re-pulled and re-built.
  2. A pending migration rewrites existing rows in the users.role column (b41c7ae90d52, values userproject_user). That is why step 1 is a backup and not optional. If a previous deploy already applied it, it will not run again — step 0 tells you which of these you are actually about to run.
  3. A second migration adds new columns (a7c31f9e5b02: an archive timestamp on projects, and two default-membership fields on users). This one is additive and has database defaults for existing rows, so it does not rewrite anything.

For the people using the app, the visible changes are: projects can now be archived from the Admin Console (they disappear from the pickers and go read-only, and can be brought back), certain users can be set to join every new project automatically, and the Admin Console has been rebuilt so the user table fits on screen.

Migrations run themselves when the wp_api container starts. There is nothing to type and no new environment variables — do not change the stack's environment variables.


Step 0 — Record the current state (needed for rollback)

SSH to the Docker host and run:

docker exec wp_api alembic -c server/alembic.ini current
docker inspect nginx_webserver --format 'nginx image: {{.Image}}'
docker inspect wp_api          --format 'api image:   {{.Image}}'

Copy the output into your ticket. Also note the Git commit the Portainer stack is currently on (Portainer → the stack → the Git reference / last-updated commit). Without these, rollback is guesswork.

The first command prints the migration the database is currently on. Use it to see which migrations this deploy will actually run:

alembic current shows What will run What that means
c93f2b1d7e04 or earlier both migrations The users.role rewrite is included — the backup in step 1 matters most in this case.
d15b8c4ef207 only a7c31f9e5b02 The users.role rewrite already happened on an earlier deploy. This one is additive only.
a7c31f9e5b02 nothing The database is already up to date; this is a code-only deploy.

Take the backup either way.


Step 1 — Back up the database

On the Docker host:

docker exec wp_db_backup /scripts/db-backup.sh

This triggers the stack's existing backup sidecar once, on demand. Expected output ends with a line like:

[db-backup] wrote 1.4M /backups/wpsuite-20260804-141233Z.sql.gz.enc

Confirm the file is on the host (substitute the stack directory):

ls -lt <stack-dir>/backups | head -3

Record that filename. Do not continue until you have seen the wrote … line and the file in that listing.

  • A .sql.gz.enc extension means backups are encrypted — expected and correct.
  • A .sql.gz extension plus a WARNING: BACKUP_ENC_PASSPHRASE not set line means backups are unencrypted. Not a blocker for this deploy; report it back.
  • No SSH access? Portainer → Containerswp_db_backupConsole → connect with /bin/sh, then run /scripts/db-backup.sh. Same result: the dump lands on the host, because /backups is a bind mount.

Step 2 — Redeploy the stack in Portainer

  1. Portainer → Stacks → select the stack.
  2. Pull and redeploy — with re-pull / re-build enabled.
  3. Wait for it to report success.

A plain "restart" or "stop/start" will not deploy this change. See "What this deploy changes" above.


Step 3 — Confirm the containers came up

docker ps --filter name=nginx_webserver --filter name=wp_api --filter name=wp_db

All three must be Up, and wp_db should show (healthy). Then check the API applied its migrations cleanly:

docker logs wp_api --tail 40

You are looking for Alembic Running upgrade … lines followed by gunicorn starting up, and no traceback. The last one should end at a7c31f9e5b02. The API deliberately refuses to start if a migration fails, so a restarting wp_api container means the migration failed — go to Rollback.

Confirm the database landed on the new revision:

docker exec wp_api alembic -c server/alembic.ini current

Expected: a7c31f9e5b02 (head).

Then verify nginx's own view of its config:

docker exec nginx_webserver nginx -t

Expected: syntax is ok / test is successful.


Step 4 — Confirm the response headers

curl -sI https://<site-url>/work-package-suite.html | grep -Ei 'cache-control|content-security-policy'

Add -k if the site uses an internal or self-signed certificate.

Both lines must come back. Expected, approximately:

cache-control: no-cache, must-revalidate
content-security-policy: default-src 'self'; script-src 'self' 'unsafe-inline'; ...

If the content-security-policy line is missing while cache-control is present, the deploy is bad — go to Rollback and send me the nginx log. (This is the specific regression this deploy fixes; the two headers must coexist.)

Also confirm the API is reachable through the proxy:

curl -s https://<site-url>/api/health      # → {"ok": true}

Step 5 — Hard-reload once in a browser

Open the site and press Ctrl+Shift+R (Cmd+Shift+R on macOS) once. The app uses a service worker; a normal reload can serve the previous version and make a good deploy look broken.

Sanity checks — all four should take under a minute:

  1. Log in. The home page offers to select or create a project.

  2. Open User Directory (the Users link in the top-right menu, or the tile on the home page). The table should read as one line per user — if rows are three lines tall and the table spills outside its white card, you are still on the old cached files: hard-reload again.

    Changed since this runbook was written: user accounts moved out of the Admin Console into users.html when the Project Super User role was added, so that a project admin can create accounts on their own job. If you are deploying a build from before that change, read this step as "Admin Console → the user table".

  3. Open Admin Console (admin account required). Two new cards are present and load: Projects, and Default members on new projects. Both should list rows, not an error.

  4. In the Projects card, click Archive on a project you don't mind hiding (a DEMO- one if there is one), confirm the prompt, then tick Show archived — it should reappear marked archived. Click Unarchive to put it back. That round trip proves the new migration and the new endpoint are both live.

Deploy complete. Please report back: the step 0 output (including which migrations ran), the backup filename, and the two header lines from step 4.


Rollback

Pick the case that matches.

Case A — nginx won't start, or the CSP header is missing

The database is untouched by this, so this is a code-only rollback. In Portainer, redeploy the stack pinned to the previous Git commit recorded in step 0 (Portainer → the stack → change the Git reference to that commit → Pull and redeploy). Then re-run step 3 and step 4.

Before you do: grab the log, because it is what I need to fix this.

docker logs nginx_webserver --tail 100

Send me that output. If the container is in a restart loop the log still works.

Case B — wp_api is restarting / a migration failed

docker logs wp_api --tail 100

Send me that output. Do not restore the database and do not roll the API back without contacting me first. Which migration got as far as committing decides what is safe, and they are not the same:

  • a7c31f9e5b02 (the new columns) is additive. If only this one ran, rolling the API back to the previous image is safe on its own — the old code simply ignores the extra columns. Nothing needs converting.
  • b41c7ae90d52 (the users.role rewrite) is not. If that one committed, rolling the API back without converting those values back will break logins. That conversion is a one-line command, but it has to match what actually ran.

The alembic current output from step 0, plus the Running upgrade … lines in the log above, are exactly what tells us which case you are in — please include both.

Reach me at n.siegfried@prime-controls.com.

Case C — restoring the backup (only if I ask for it)

Destructive: this drops and recreates the current schema and data. For an encrypted dump, on the Docker host, in the backups directory:

export BACKUP_ENC_PASSPHRASE='<the passphrase — from the stack env vars>'
openssl enc -d -aes-256-cbc -pbkdf2 -pass env:BACKUP_ENC_PASSPHRASE \
  -in wpsuite-<timestamp>.sql.gz.enc \
  | gunzip \
  | docker exec -i wp_db psql -U wpsuite -d wpsuite
unset BACKUP_ENC_PASSPHRASE

For an unencrypted dump, drop the openssl stage and pipe gunzip straight into psql. Substitute the real values if POSTGRES_USER / POSTGRES_DB are not wpsuite.


Notes

  • Do not add or change environment variables for this deploy.
  • Do not run docker compose down -v — the -v flag deletes the pgdata volume and with it the entire database.
  • docker compose … commands are avoided throughout this runbook on purpose: for a Portainer-managed Git stack the compose project lives under Portainer's own data directory, so docker compose from an SSH session usually can't find it. The docker exec <container-name> form used here works from any directory.
  • Full background documentation: DEPLOYMENT.md in the repository.