diff --git a/.github/workflows/backup.yml b/.github/workflows/backup.yml index 3ba5def72..bb2ceb99a 100644 --- a/.github/workflows/backup.yml +++ b/.github/workflows/backup.yml @@ -57,6 +57,16 @@ jobs: cd $PROJECT_DIR ./infra/scripts/backup-verify.sh $BACKUP_DIR/db-$BACKUP_DATE.sql.gz + # The nightly full SSG rebuild lives here, not in every deploy: the box is idle at this + # hour and the backup has finished reading the disk. Async — ssg-worker renders it + # (~25 min); a failed or stale rebuild shows up in /health/ready, which health-check.yml + # watches every 5 min. Runs even if the backup failed: the two are unrelated. + - name: Queue nightly SSG rebuild + if: always() + run: | + curl -sf -X POST http://localhost:8080/internal/ssg/rebuild-all && echo "SSG rebuild queued" \ + || { echo "::error::could not queue the nightly SSG rebuild"; exit 1; } + # Runs even when the backup failed — a full disk is the usual reason it does. # Failing here fires the same failure notification as a broken backup. - name: Disk space alarm (>= 85%) diff --git a/.github/workflows/deploy.yml b/.github/workflows/deploy.yml index 24e702431..ccd37234c 100644 --- a/.github/workflows/deploy.yml +++ b/.github/workflows/deploy.yml @@ -8,6 +8,10 @@ on: rollback_commit: description: 'Commit SHA to rollback to (optional)' required: false + rebuild_ssg: + description: 'Full SSG rebuild now (otherwise it runs nightly after the backup)' + type: boolean + default: false concurrency: group: deploy-production @@ -309,7 +313,13 @@ jobs: UNHEALTHY=$(cd $PROJECT_DIR && docker compose --profile mcp -f docker-compose.yml -f docker-compose.gpu.yml ps --format '{{.Service}} {{.Health}}' | awk '$2 == "unhealthy" {print $1}') [ -z "$UNHEALTHY" ] || { echo "::error::unhealthy services: $UNHEALTHY"; exit 1; } + # A full rebuild is ~25 min of Puppeteer on the prod box and held the only runner for it. + # Not needed per deploy: the snapshot step above keeps the old SSG tree AND the assets it + # references, so crawlers keep complete pages; publishing a book rebuilds that book. The + # full rebuild runs nightly after the backup (backup.yml). Tick `rebuild_ssg` on a manual + # run when a release changes how SEO pages render. - name: Queue SSG rebuild + if: inputs.rebuild_ssg run: | # Stamp the moment we asked, so the wait below can require a swap that happened # *after* it rather than guessing from a file's age. @@ -366,6 +376,7 @@ jobs: echo "Bot map found in nginx config" - name: Wait for SSG rebuild to complete + if: inputs.rebuild_ssg run: | # Queue SSG rebuild step is async. Poll until the job finishes. # ssg-worker bind mount: ./apps/web/dist:/app/dist, so files atomic-swapped diff --git a/.github/workflows/health-check.yml b/.github/workflows/health-check.yml index 77c73eb57..27c71f593 100644 --- a/.github/workflows/health-check.yml +++ b/.github/workflows/health-check.yml @@ -60,11 +60,9 @@ jobs: # Two ways SSG dies, and they need different patience. # # A failing rebuild is a malfunction whatever the cadence — the next one - # will not help either — so it alarms at once. Staleness alone needs a - # generous window: rebuilds here are deploy-driven, and the production - # history shows a healthy 19-hour gap between them on a quiet day. 72 - # hours is well past anything normal and still catches the failure that - # actually happened, when nothing rebuilt for five weeks. + # will not help either — so it alarms at once. Staleness: the full rebuild + # runs nightly after the backup (backup.yml), so a healthy tree is <25h + # old. 36 hours means one night was missed — worth an alarm, not a page. - name: SSG freshness run: | set -euo pipefail @@ -92,8 +90,8 @@ jobs: exit 0 fi # bash has no floats; compare whole hours. - if [ "${age%.*}" -ge 72 ]; then - echo "::error::Newest SSG rebuild is ${age}h old. Nothing has regenerated in three days." + if [ "${age%.*}" -ge 36 ]; then + echo "::error::Newest SSG rebuild is ${age}h old. The nightly rebuild did not run." exit 1 fi echo "SSG last rebuilt ${age}h ago" diff --git a/CHANGELOG.md b/CHANGELOG.md index 377809771..1f07fa8f0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -23,6 +23,7 @@ the archive; if it broke production, it belongs in `docs/incidents/`. See ## [Unreleased] +- **Ops** — deploys no longer run a full SSG rebuild (~25 min of CPU, held the only runner): it runs nightly after the backup; `rebuild_ssg` input on a manual deploy for releases that change SEO rendering; health check alarms at 36h stale (was 72h) — infra - **Security** — hardening from the architecture review: stored files are served with a sandbox CSP + nosniff (nginx and the API), `/internal/*` is refused at nginx and its network check is one tested helper, two copyrighted test PDFs removed (image-only test now uses a generated PDF) — backend, infra - **Discover** — the "Learn a language by reading real books" card no longer fills a small screen: it scrolls with the page under the search box, and a × hides it for good (`onboarding.startReadingCard.dismissed`) — closed-test report, Unihertz Titan 2 — mobile - **Docs** — docs checked against the code before the architecture review: STATUS, CLAUDE.md, architecture + system docs, ADR status lines, feature/ops/dev docs; `[Unreleased]` cut into deploy-date headings; stale plans marked historical — docs diff --git a/CLAUDE.md b/CLAUDE.md index 6c25c9fe6..f8311cc4c 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -12,7 +12,7 @@ Free book library w/ Kindle-like reader. Upload EPUB/PDF → parse → SEO pages **Prerequisites**: Docker, .NET 10 SDK, Node.js 18+, pnpm -**CI/CD**: Push to `main` → auto-deploy. SSG rebuild: admin panel or `make rebuild-ssg`. +**CI/CD**: Push to `main` → auto-deploy (no SSG rebuild). Full SSG rebuild: nightly after the backup (`backup.yml`); on demand via the admin panel, `make rebuild-ssg`, or a manual deploy with `rebuild_ssg`. ## Where to write things down @@ -490,7 +490,7 @@ That single command builds the AAB and pushes it to Internal Testing. Service ac **GitHub Actions workflows** (`.github/workflows/`): - **ci.yml** — runs on PR + push to main. Jobs: backend (build, lint, migrations, search tests), frontend (web + admin build), docker (integration tests), e2e (Playwright) -- **deploy.yml** — self-hosted runner on server. Pre-deploy backup → git pull → frontend build → docker compose up → health checks → SSG rebuild queue → image cleanup +- **deploy.yml** — self-hosted runner on server. Pre-deploy backup → git pull → frontend build → docker compose up → health checks → SSG content check → image cleanup. Full SSG rebuild only with the `rebuild_ssg` input; otherwise nightly in backup.yml - **backup.yml** — daily at 3 AM UTC. DB dump + storage tar.gz, keeps 5 newest of each - **health-check.yml** — every 5 min. Checks API + both frontends