Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .github/workflows/backup.yml
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,16 @@ jobs:
cd $PROJECT_DIR
./infra/scripts/backup-verify.sh $BACKUP_DIR/db-$BACKUP_DATE.sql.gz

# The nightly full SSG rebuild lives here, not in every deploy: the box is idle at this
# hour and the backup has finished reading the disk. Async — ssg-worker renders it
# (~25 min); a failed or stale rebuild shows up in /health/ready, which health-check.yml
# watches every 5 min. Runs even if the backup failed: the two are unrelated.
- name: Queue nightly SSG rebuild
if: always()
run: |
curl -sf -X POST http://localhost:8080/internal/ssg/rebuild-all && echo "SSG rebuild queued" \
|| { echo "::error::could not queue the nightly SSG rebuild"; exit 1; }

# Runs even when the backup failed — a full disk is the usual reason it does.
# Failing here fires the same failure notification as a broken backup.
- name: Disk space alarm (>= 85%)
Expand Down
11 changes: 11 additions & 0 deletions .github/workflows/deploy.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,10 @@ on:
rollback_commit:
description: 'Commit SHA to rollback to (optional)'
required: false
rebuild_ssg:
description: 'Full SSG rebuild now (otherwise it runs nightly after the backup)'
type: boolean
default: false

concurrency:
group: deploy-production
Expand Down Expand Up @@ -309,7 +313,13 @@ jobs:
UNHEALTHY=$(cd $PROJECT_DIR && docker compose --profile mcp -f docker-compose.yml -f docker-compose.gpu.yml ps --format '{{.Service}} {{.Health}}' | awk '$2 == "unhealthy" {print $1}')
[ -z "$UNHEALTHY" ] || { echo "::error::unhealthy services: $UNHEALTHY"; exit 1; }

# A full rebuild is ~25 min of Puppeteer on the prod box and held the only runner for it.
# Not needed per deploy: the snapshot step above keeps the old SSG tree AND the assets it
# references, so crawlers keep complete pages; publishing a book rebuilds that book. The
# full rebuild runs nightly after the backup (backup.yml). Tick `rebuild_ssg` on a manual
# run when a release changes how SEO pages render.
- name: Queue SSG rebuild
if: inputs.rebuild_ssg
run: |
# Stamp the moment we asked, so the wait below can require a swap that happened
# *after* it rather than guessing from a file's age.
Expand Down Expand Up @@ -366,6 +376,7 @@ jobs:
echo "Bot map found in nginx config"

- name: Wait for SSG rebuild to complete
if: inputs.rebuild_ssg
run: |
# Queue SSG rebuild step is async. Poll until the job finishes.
# ssg-worker bind mount: ./apps/web/dist:/app/dist, so files atomic-swapped
Expand Down
12 changes: 5 additions & 7 deletions .github/workflows/health-check.yml
Original file line number Diff line number Diff line change
Expand Up @@ -60,11 +60,9 @@ jobs:
# Two ways SSG dies, and they need different patience.
#
# A failing rebuild is a malfunction whatever the cadence — the next one
# will not help either — so it alarms at once. Staleness alone needs a
# generous window: rebuilds here are deploy-driven, and the production
# history shows a healthy 19-hour gap between them on a quiet day. 72
# hours is well past anything normal and still catches the failure that
# actually happened, when nothing rebuilt for five weeks.
# will not help either — so it alarms at once. Staleness: the full rebuild
# runs nightly after the backup (backup.yml), so a healthy tree is <25h
# old. 36 hours means one night was missed — worth an alarm, not a page.
- name: SSG freshness
run: |
set -euo pipefail
Expand Down Expand Up @@ -92,8 +90,8 @@ jobs:
exit 0
fi
# bash has no floats; compare whole hours.
if [ "${age%.*}" -ge 72 ]; then
echo "::error::Newest SSG rebuild is ${age}h old. Nothing has regenerated in three days."
if [ "${age%.*}" -ge 36 ]; then
echo "::error::Newest SSG rebuild is ${age}h old. The nightly rebuild did not run."
exit 1
fi
echo "SSG last rebuilt ${age}h ago"
Expand Down
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@ the archive; if it broke production, it belongs in `docs/incidents/`. See

## [Unreleased]

- **Ops** — deploys no longer run a full SSG rebuild (~25 min of CPU, held the only runner): it runs nightly after the backup; `rebuild_ssg` input on a manual deploy for releases that change SEO rendering; health check alarms at 36h stale (was 72h) — infra
- **Security** — hardening from the architecture review: stored files are served with a sandbox CSP + nosniff (nginx and the API), `/internal/*` is refused at nginx and its network check is one tested helper, two copyrighted test PDFs removed (image-only test now uses a generated PDF) — backend, infra
- **Discover** — the "Learn a language by reading real books" card no longer fills a small screen: it scrolls with the page under the search box, and a × hides it for good (`onboarding.startReadingCard.dismissed`) — closed-test report, Unihertz Titan 2 — mobile
- **Docs** — docs checked against the code before the architecture review: STATUS, CLAUDE.md, architecture + system docs, ADR status lines, feature/ops/dev docs; `[Unreleased]` cut into deploy-date headings; stale plans marked historical — docs
Expand Down
4 changes: 2 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ Free book library w/ Kindle-like reader. Upload EPUB/PDF → parse → SEO pages

**Prerequisites**: Docker, .NET 10 SDK, Node.js 18+, pnpm

**CI/CD**: Push to `main` → auto-deploy. SSG rebuild: admin panel or `make rebuild-ssg`.
**CI/CD**: Push to `main` → auto-deploy (no SSG rebuild). Full SSG rebuild: nightly after the backup (`backup.yml`); on demand via the admin panel, `make rebuild-ssg`, or a manual deploy with `rebuild_ssg`.

## Where to write things down

Expand Down Expand Up @@ -490,7 +490,7 @@ That single command builds the AAB and pushes it to Internal Testing. Service ac

**GitHub Actions workflows** (`.github/workflows/`):
- **ci.yml** — runs on PR + push to main. Jobs: backend (build, lint, migrations, search tests), frontend (web + admin build), docker (integration tests), e2e (Playwright)
- **deploy.yml** — self-hosted runner on server. Pre-deploy backup → git pull → frontend build → docker compose up → health checks → SSG rebuild queue → image cleanup
- **deploy.yml** — self-hosted runner on server. Pre-deploy backup → git pull → frontend build → docker compose up → health checks → SSG content check → image cleanup. Full SSG rebuild only with the `rebuild_ssg` input; otherwise nightly in backup.yml
- **backup.yml** — daily at 3 AM UTC. DB dump + storage tar.gz, keeps 5 newest of each
- **health-check.yml** — every 5 min. Checks API + both frontends

Expand Down
Loading