Files
Project-SDE-WP-Suite/DEPLOY-runbook-2026-08-04.md
n.siegfried 4ace2afb1c Move user administration to its own page; add Project Super User
User accounts lived in the Admin Console, which is admins-only. Project admins
need to create the accounts on their own jobs without an app admin on the phone,
so accounts move to a new User Directory page and a new role carries the right.

server/auth.py, server/app.py
  New permissions role `project_super_user`, between admin and project_admin:
  everything a project admin may do, plus user administration SCOPED to the
  projects they hold the role on. Four limits make it safe to hand out, all
  enforced server-side:

    * Scope comes from projects, not the job title. It resolves per membership
      (managed_project_ids), so an ordinary account can hold it on one job via
      ProjectMember.role, and a super user demoted on one job administers
      nobody there. No projects, no authority.
    * Account-level changes (password, disable, rename, permissions, delete)
      require EXCLUSIVE scope: refused when the target is also on a project the
      caller does not administer, because those changes are global. The
      directory renders such rows read-only with the reason.
    * No admin or super-user targets, and neither role can be granted by a
      super user -- that is the line that stops it becoming app-wide control.
    * PUT .../projects rebuilds only the caller's own slice; memberships on
      projects they do not administer are left untouched. A payload that simply
      omits them must not cut someone off a job the caller cannot see.

  Creating requires naming at least one of your own projects: an account with
  none would be one the creator instantly cannot manage.

  /api/auth/users is now scoped rather than admin-only, and carries a per-row
  `manageable` verdict plus the reason. Non-managers get a contact card only --
  a project user has no business reading colleagues' login history. New
  /api/auth/user-scope tells the page what it may offer. Administrative
  password resets are now audited; they were the one account change that left
  no trace. Settings, feature flags and the auto-add rule stay admin-only.

  While here: one definition of "is a user manager", derived from the managed
  set. An account-role-only version disagreed with the scoped one and locked
  per-project super users out of routes they were entitled to.

html/users.html, html/users.js
  The directory: three renderings from one page -- admin (everything), super
  user (controls per row, read-only where scope is shared), everyone else (a
  read-only directory of the people on their own projects).

html/console.css, html/console-util.js
  Extracted from admin.html/admin.js so both console pages share them. A
  divergent jsq() is an XSS and a divergent role list offers permissions the
  server refuses, so neither may exist twice.

html/wp-sidenav.{js,css}
  Global nav drawer, role-gated, carrying ?project= across links. Mounted on
  the field view (which had no way to anywhere) plus both console pages.

No migration: users.role is already String(20) and the new value fits.

Verified: 93 scope/gate tests, 29 live HTTP tests through the real dependency
stack, 33 static JS checks. Not verified in a browser -- no JS engine on this
machine -- so users.html and field.html want one manual load.

server/smoketest.py still fails with 401s. Pre-existing: it has no login code,
so auth_gate refuses it. Confirmed unchanged by stashing this work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 17:36:14 -07:00

291 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Deploy runbook — WP Suite
**For:** IT / whoever administers the Docker host and Portainer
**From:** n.siegfried@prime-controls.com
**Revised:** 2026-08-05 — **this replaces the 2026-08-04 version.** Same procedure,
but the deploy now carries a second database migration and a new admin screen. If
you already have the earlier copy, work from this one instead.
**Expected duration:** 1015 minutes, including the backup
**Expected downtime:** under a minute, while containers are recreated
---
## Fill these in before handing this over
| Thing | Value |
|---|---|
| Docker host (SSH target) | `________________` |
| Stack name in Portainer | `________________` |
| Site URL | `https://________________` |
| Stack directory on the host (holds `docker-compose.yml` / `backups/`) | `________________` |
Container names are fixed by the compose file and are the same on every host:
`nginx_webserver`, `wp_api`, `wp_db`, `wp_db_backup`.
---
## What this deploy changes
Front-end and nginx changes, a new admin screen, plus pending database migrations
that run automatically. Three things make it more than a routine restart:
1. **The nginx config and the entire `html/` directory are baked into the
container image at build time.** A plain restart deploys nothing — the stack
must be re-pulled and re-built.
2. **A pending migration rewrites existing rows** in the `users.role` column
(`b41c7ae90d52`, values `user``project_user`). That is why step 1 is a backup
and not optional. If a previous deploy already applied it, it will not run again —
step 0 tells you which of these you are actually about to run.
3. **A second migration adds new columns** (`a7c31f9e5b02`: an archive timestamp on
projects, and two default-membership fields on users). This one is additive and
has database defaults for existing rows, so it does not rewrite anything.
For the people using the app, the visible changes are: projects can now be
**archived** from the Admin Console (they disappear from the pickers and go
read-only, and can be brought back), certain users can be set to join **every new
project automatically**, and the Admin Console has been rebuilt so the user table
fits on screen.
Migrations run themselves when the `wp_api` container starts. There is nothing
to type and **no new environment variables** — do not change the stack's
environment variables.
---
## Step 0 — Record the current state (needed for rollback)
SSH to the Docker host and run:
```bash
docker exec wp_api alembic -c server/alembic.ini current
docker inspect nginx_webserver --format 'nginx image: {{.Image}}'
docker inspect wp_api --format 'api image: {{.Image}}'
```
**Copy the output into your ticket.** Also note the Git commit the Portainer
stack is currently on (Portainer → the stack → the Git reference / last-updated
commit). Without these, rollback is guesswork.
The first command prints the migration the database is currently on. Use it to see
which migrations this deploy will actually run:
| `alembic current` shows | What will run | What that means |
|---|---|---|
| `c93f2b1d7e04` or earlier | both migrations | The `users.role` rewrite is included — the backup in step 1 matters most in this case. |
| `d15b8c4ef207` | only `a7c31f9e5b02` | The `users.role` rewrite already happened on an earlier deploy. This one is additive only. |
| `a7c31f9e5b02` | nothing | The database is already up to date; this is a code-only deploy. |
Take the backup either way.
---
## Step 1 — Back up the database
On the Docker host:
```bash
docker exec wp_db_backup /scripts/db-backup.sh
```
This triggers the stack's existing backup sidecar once, on demand. Expected
output ends with a line like:
```
[db-backup] wrote 1.4M /backups/wpsuite-20260804-141233Z.sql.gz.enc
```
Confirm the file is on the host (substitute the stack directory):
```bash
ls -lt <stack-dir>/backups | head -3
```
**Record that filename.** Do not continue until you have seen the `wrote …`
line and the file in that listing.
- A `.sql.gz.enc` extension means backups are encrypted — expected and correct.
- A `.sql.gz` extension plus a `WARNING: BACKUP_ENC_PASSPHRASE not set` line
means backups are unencrypted. Not a blocker for this deploy; report it back.
- **No SSH access?** Portainer → **Containers**`wp_db_backup`**Console**
connect with `/bin/sh`, then run `/scripts/db-backup.sh`. Same result: the
dump lands on the host, because `/backups` is a bind mount.
---
## Step 2 — Redeploy the stack in Portainer
1. Portainer → **Stacks** → select the stack.
2. **Pull and redeploy** — with re-pull / re-build **enabled**.
3. Wait for it to report success.
A plain "restart" or "stop/start" will **not** deploy this change. See "What
this deploy changes" above.
---
## Step 3 — Confirm the containers came up
```bash
docker ps --filter name=nginx_webserver --filter name=wp_api --filter name=wp_db
```
All three must be `Up`, and `wp_db` should show `(healthy)`. Then check the API
applied its migrations cleanly:
```bash
docker logs wp_api --tail 40
```
You are looking for Alembic `Running upgrade …` lines followed by gunicorn
starting up, and **no** traceback. The last one should end at `a7c31f9e5b02`. The
API deliberately refuses to start if a migration fails, so a restarting `wp_api`
container means the migration failed — go to Rollback.
Confirm the database landed on the new revision:
```bash
docker exec wp_api alembic -c server/alembic.ini current
```
Expected: `a7c31f9e5b02 (head)`.
Then verify nginx's own view of its config:
```bash
docker exec nginx_webserver nginx -t
```
Expected: `syntax is ok` / `test is successful`.
---
## Step 4 — Confirm the response headers
```bash
curl -sI https://<site-url>/work-package-suite.html | grep -Ei 'cache-control|content-security-policy'
```
Add `-k` if the site uses an internal or self-signed certificate.
**Both lines must come back.** Expected, approximately:
```
cache-control: no-cache, must-revalidate
content-security-policy: default-src 'self'; script-src 'self' 'unsafe-inline'; ...
```
If the `content-security-policy` line is **missing** while `cache-control` is
present, the deploy is bad — go to Rollback and send me the nginx log. (This is
the specific regression this deploy fixes; the two headers must coexist.)
Also confirm the API is reachable through the proxy:
```bash
curl -s https://<site-url>/api/health # → {"ok": true}
```
---
## Step 5 — Hard-reload once in a browser
Open the site and press **Ctrl+Shift+R** (Cmd+Shift+R on macOS) once. The app
uses a service worker; a normal reload can serve the previous version and make a
good deploy look broken.
Sanity checks — all four should take under a minute:
1. Log in. The home page offers to select or create a project.
2. Open **User Directory** (the `Users` link in the top-right menu, or the tile on the
home page). The table should read as **one line per user** — if rows are three lines
tall and the table spills outside its white card, you are still on the old cached
files: hard-reload again.
> Changed since this runbook was written: user accounts moved out of the Admin
> Console into `users.html` when the **Project Super User** role was added, so that
> a project admin can create accounts on their own job. If you are deploying a build
> from before that change, read this step as "Admin Console → the user table".
3. Open **Admin Console** (admin account required). Two new cards are present and
load: **Projects**, and **Default members on new projects**. Both should list rows,
not an error.
4. In the **Projects** card, click **Archive** on a project you don't mind hiding
(a `DEMO-` one if there is one), confirm the prompt, then tick **Show archived**
it should reappear marked `archived`. Click **Unarchive** to put it back. That
round trip proves the new migration and the new endpoint are both live.
**Deploy complete.** Please report back: the step 0 output (including which
migrations ran), the backup filename, and the two header lines from step 4.
---
## Rollback
Pick the case that matches.
### Case A — nginx won't start, or the CSP header is missing
The database is untouched by this, so this is a code-only rollback. In Portainer,
redeploy the stack pinned to the **previous Git commit** recorded in step 0
(Portainer → the stack → change the Git reference to that commit → Pull and
redeploy). Then re-run step 3 and step 4.
**Before you do:** grab the log, because it is what I need to fix this.
```bash
docker logs nginx_webserver --tail 100
```
Send me that output. If the container is in a restart loop the log still works.
### Case B — `wp_api` is restarting / a migration failed
```bash
docker logs wp_api --tail 100
```
Send me that output. **Do not restore the database and do not roll the API back
without contacting me first.** Which migration got as far as committing decides what
is safe, and they are not the same:
- **`a7c31f9e5b02`** (the new columns) is additive. If only this one ran, rolling
the API back to the previous image is safe on its own — the old code simply
ignores the extra columns. Nothing needs converting.
- **`b41c7ae90d52`** (the `users.role` rewrite) is not. If that one committed,
rolling the API back without converting those values back **will break logins**.
That conversion is a one-line command, but it has to match what actually ran.
The `alembic current` output from step 0, plus the `Running upgrade …` lines in the
log above, are exactly what tells us which case you are in — please include both.
Reach me at n.siegfried@prime-controls.com.
### Case C — restoring the backup (only if I ask for it)
Destructive: this drops and recreates the current schema and data. For an
encrypted dump, on the Docker host, in the `backups` directory:
```bash
export BACKUP_ENC_PASSPHRASE='<the passphrase — from the stack env vars>'
openssl enc -d -aes-256-cbc -pbkdf2 -pass env:BACKUP_ENC_PASSPHRASE \
-in wpsuite-<timestamp>.sql.gz.enc \
| gunzip \
| docker exec -i wp_db psql -U wpsuite -d wpsuite
unset BACKUP_ENC_PASSPHRASE
```
For an unencrypted dump, drop the `openssl` stage and pipe `gunzip` straight
into `psql`. Substitute the real values if `POSTGRES_USER` / `POSTGRES_DB` are
not `wpsuite`.
---
## Notes
- Do not add or change environment variables for this deploy.
- Do not run `docker compose down -v` — the `-v` flag deletes the `pgdata`
volume and with it the entire database.
- `docker compose …` commands are avoided throughout this runbook on purpose:
for a Portainer-managed Git stack the compose project lives under Portainer's
own data directory, so `docker compose` from an SSH session usually can't find
it. The `docker exec <container-name>` form used here works from any directory.
- Full background documentation: `DEPLOYMENT.md` in the repository.