diff --git a/README.md b/README.md index 304764c48..e328116b6 100644 --- a/README.md +++ b/README.md @@ -8,6 +8,7 @@ video library. [![Journal integrity](https://github.com/ActiveInferenceInstitute/ActiveInferenceJournal/actions/workflows/journal-integrity.yml/badge.svg)](https://github.com/ActiveInferenceInstitute/ActiveInferenceJournal/actions/workflows/journal-integrity.yml) Learn more: https://activeinference.institute/learning/ · +Site: https://activeinferenceinstitute.github.io/ActiveInferenceJournal/ · Tooling: https://github.com/ActiveInferenceInstitute/Journal_Utilities ## Layout diff --git a/TO-DO.md b/TO-DO.md index 6c270c15c..07bcec15c 100644 --- a/TO-DO.md +++ b/TO-DO.md @@ -107,3 +107,30 @@ overhaul, cross-cutting refactors. - **Curated assets elsewhere** — after the image-repair pass, no `.md` under `assets/` has broken relative refs and no machine paths remain anywhere in the repo (verified by grep). + +## M4-pages wave (feat/m4-pages, 2026-09-23) + +- ✓ **I4 verified on the real tree** — builder reads only lowercase + `translations/` (Journal-Utilities `src/journal_utilities/site/builder.py:121`); + tree holds 13 items with lowercase `translations/` (~336 SRTs) vs 131 items + with capital-T `Translations/` (~2,880 SRTs); live `manifest.json` shows + languages for 13/573 items (129 language entries). Fix + spec in + [`docs/m4-site-spec.md`](docs/m4-site-spec.md) §0; permanent fix = M2 + translation migration. +- ✓ **Textbook Cohort 2 Meeting 20 SRT misfiled in + `ModelStream_011/captions/`** — `git mv` to canonical + `TextbookGroup/ParrPezzuloFriston2022/Cohort_2/Meeting_020/captions/` + (video `QjGcN1l6NXg` via INDEX.json). 57 further byte-identical copies + remain scattered (GuestStream/MathStream/ModelStream/MorphStream/Courses/ + symposium/Cohort_4) — owned by the J3/I10 captions-naming validator, see + [`docs/misfiling-findings.md`](docs/misfiling-findings.md). +- ✓ **2022 Robotics translations inside 2021 Symposium `Translations/`** — + verified near-duplicate second copy (22 files, diffs are re-translated + lines); documented rather than moved (a move would clobber or double-count); + to be resolved during the M2 per-series migration — + [`docs/misfiling-findings.md`](docs/misfiling-findings.md) §2. +- ✓ **README.md** — Pages site link added. +- ✓ **`docs/m4-site-spec.md`** — M4 builder spec (static + `/item///index.html`, schema.org VideoObject JSON-LD, + canonical/og, sitemap.xml + robots.txt), prepared for the Journal_Utilities + builder agent; not yet implemented. diff --git a/data/video/activeinferenceinstitute/ModelStream/ModelStream_011/captions/ActInf Textbook Group ~ Cohort 2 ~ Meeting 20 (Chapter 9, part 1).en.srt b/data/video/activeinferenceinstitute/TextbookGroup/ParrPezzuloFriston2022/Cohort_2/Meeting_020/captions/ActInf Textbook Group ~ Cohort 2 ~ Meeting 20 (Chapter 9, part 1).en.srt similarity index 100% rename from data/video/activeinferenceinstitute/ModelStream/ModelStream_011/captions/ActInf Textbook Group ~ Cohort 2 ~ Meeting 20 (Chapter 9, part 1).en.srt rename to data/video/activeinferenceinstitute/TextbookGroup/ParrPezzuloFriston2022/Cohort_2/Meeting_020/captions/ActInf Textbook Group ~ Cohort 2 ~ Meeting 20 (Chapter 9, part 1).en.srt diff --git a/docs/README.md b/docs/README.md index 8599e5064..72f3d4b14 100644 --- a/docs/README.md +++ b/docs/README.md @@ -33,3 +33,10 @@ materials from the Active Inference Institute video library. - `main` — content without audio (lightweight). - `audio` — `main` + `/audio/.64k.m4a` (64 kbps media), same layout. + +See also: [`m4-site-spec.md`](m4-site-spec.md) — spec for the M4 static +per-item pages + sitemap/robots (for the Journal_Utilities builder agent) — +and [`misfiling-findings.md`](misfiling-findings.md) — captions/translations +misfiling dispositions (Textbook Cohort 2 Meeting 20 SRT; 2021 Symposium +`Translations/` robotics duplicates). + diff --git a/docs/m4-site-spec.md b/docs/m4-site-spec.md new file mode 100644 index 000000000..7c754331c --- /dev/null +++ b/docs/m4-site-spec.md @@ -0,0 +1,135 @@ +# M4 site spec — static per-item pages, sitemap, robots.txt + +Status: **specification for implementation, not yet implemented.** The current +deployed site (hash-routed SPA fed by `manifest.json` + `data/*.json`, built by +Journal-Utilities `src/journal_utilities/site/builder.py` via +`scripts/build_pages_site.py`, deployed by +[`.github/workflows/deploy-pages.yml`](../.github/workflows/deploy-pages.yml)) +exposes exactly one crawlable URL. This file states what the M4 wave must emit +so the Journal_Utilities builder agent (or a later wave) can implement it. +Handoff context: `/tmp/aii_journal_handoff.md` milestone M4. + +## 0. Prerequisite fix bundled with M4 (handoff I4) + +`builder.py` reads only lowercase `translations/` (`tr_dir = item_dir / +"translations"`, builder.py:121) while the journal tree holds **131 items with +capital-T `Translations/`** (~2,880 SRTs) vs **13 items with lowercase +`translations/`** (~336 SRTs). Verified 2026-09-23 on this tree and against the +live deployed `manifest.json` (13/573 items with `languages`, 129 language +entries; the site ships 3,552 fewer translation files than the journal holds: +6,432 total SRTs vs 3,552 readable). + +Fix (permanent fix remains the M2 `Translations/ → translations/` migration): + +```python +tr_dir = next( + (d for d in (item_dir / "translations", item_dir / "Translations") if d.is_dir()), + None, +) +``` + +Also strip legacy ISO-639-2 parenthesized tags when deriving `lang` +(builder.py:124 derives it from `name.split(".")[-2]`, so +`...chi(translated)` yields `chi(translated)` today). Map `chi→zh-Hans` +or `zh-Hant`, `dut→nl`, `fre→fr`, `ger→de`, `jpn→ja`, `kor→ko`, `rus→ru`, +`por→pt`, `spa→es`, `ita→it` per the M2 table; `translate_subtitles_openrouter.py:381` +must not re-translate files that already exist under the normalized name. + +## 1. URL scheme + +Base URL (canonical, used everywhere below): +`https://activeinferenceinstitute.github.io/ActiveInferenceJournal/` + +- SPA stays at `/` (hash routing preserved for backward compatibility). +- New per-item page per INDEX item: + `/item///index.html` → canonical URL + `.../item///` (trailing slash, no `index.html` in canonicals). +- `` and `` are the exact on-disk directory names from + `INDEX.json items[].path` (`data/video/activeinferenceinstitute//`) — + including spaces (e.g. `Applied Active Inference Symposium`) — percent-encoded + in emitted URLs. Directory names ARE the slugs; M4 must not invent new slugs + (folder slugification is DAF decision §6/J11, out of scope). +- `data/_.json` payload files keep their existing location; item + pages reference them for the interactive transcript view. + +## 2. Per-item page contract + +Emit `output/item///index.html` for **every** item in +`INDEX.json items[]` (573 items at the time of writing). Requirements: + +1. **Server-rendered (baked) content** — no JS required to read the core + metadata: title, series, date (`parts[].upload_date`, ISO 8601), guests, + summary/abstract from `metadata.json`, chapter list + (`parts[].chapters`, present only with provenance), full transcript text + (from `transcript.txt`, `parts`-tagged when multi-part). +2. **Embedded video** — YouTube iframe per part using `parts[].video_id` + (`https://www.youtube-nocookie.com/embed/`), part titles as + headings. +3. **Structured data** — one `application/ld+json` block per part: + schema.org `VideoObject` with + `name`, `description` (abstract from metadata), `uploadDate` (ISO 8601), + `contentUrl`/`embedUrl` (the YouTube URL from `parts[].url`), + `thumbnailUrl` (`https://i.ytimg.com/vi//hqdefault.jpg`), + `inLanguage` (source language, `en`), `transcript` reference, and + `hasPart` → `Clip` entries from gated chapters (`name`, `startOffsetTime`, + `endOffsetTime`, `url` with `?t=` fragment). Top level: an `Dataset` node + for the item (name, `license` = CC-BY-4.0, `isPartOf` → the journal Dataset, + `citation` → Zenodo DOI `10.5281/zenodo.7299755`). +4. **Head tags** — ``, + Open Graph (`og:type=video`, `og:title`, `og:description`, `og:url`, + `og:image` = thumbnail), Twitter `summary_large_image`. +5. **Language list** — translations actually present for the item (post-I4-fix + set), linking to the interactive view; do not fabricate languages. +6. **Navigation** — breadcrumb (journal → series → item), links to + prev/next item in series, link back to the SPA for the interactive player. +7. **Non-goals**: no search, no client-side filtering, no new JS frameworks; + static HTML + the existing CSS. + +## 3. sitemap.xml + +Emit at `/sitemap.xml`: + +- `` with **one + `` per item** (573): `` = canonical item URL, + `` = max `parts[].upload_date` for the item (ISO 8601, omit if + unknown — never fabricate), **no** `changefreq`/`priority` (noise). +- Plus the root `` = `/`. +- Single file is fine (well under the 50k-URL / 50 MB limits); if it ever + exceeds them, switch to a sitemap index `/sitemap-index.xml`. +- Post-build assertion: `sitemap` URL count == processed item count (builder + already returns `items_processed`); fail the build on mismatch. + +## 4. robots.txt + +Emit at `/robots.txt`: + +``` +User-agent: * +Allow: / +Sitemap: https://activeinferenceinstitute.github.io/ActiveInferenceJournal/sitemap.xml +``` + +Both files are generated by the builder (not committed by hand) and must be +included in the Pages artifact alongside the existing `.nojekyll`. + +## 5. Deployment notes (ties to handoff E5/J7, owned by the workflow wave) + +- Builder changes land in Journal-Utilities; pin the workflow's utilities + checkout to a tag/SHA, sparse-checkout, plain `python` (no torch/CUDA), and + keep a single Pages deployment (drop the dual actions+gh-pages deploy or + make branch-mode the fallback behind a flag). +- Post-deploy smoke test (CI): fetch `/sitemap.xml`, assert URL count + equals `manifest.json` `total_items`, spot-check Rich Results Test on one + sample item page, assert `/robots.txt` returns 200. + +## 6. Acceptance criteria + +- Every INDEX item has a static URL returning 200 with server-rendered title, + guests, date, summary, chapters, transcript, and a valid VideoObject + JSON-LD block per part. +- `sitemap.xml` lists the root + all item URLs; `robots.txt` references it. +- Rich Results Test validates a sampled item page (no VideoObject errors). +- Site language counts reflect the union of both `translations/` spellings + (post-fix: 144 items with ≥1 translation at current tree state). +- Canonical URLs resolve against `INDEX.json` `items[].path` (series/item + exactly); a test asserts emitted `` ↔ item-path bijection. diff --git a/docs/misfiling-findings.md b/docs/misfiling-findings.md new file mode 100644 index 000000000..08445caf3 --- /dev/null +++ b/docs/misfiling-findings.md @@ -0,0 +1,53 @@ +# Misfiling findings — captions and translations triage + +Recorded: 2026-09-23 (M4-pages wave, fleet handoff `/tmp/aii_journal_handoff.md`). +Scope: the two known misfilings called out in the handoff baseline, verified +against the real tree. Generator-owned rewrites should consult this file. + +## 1. Textbook Cohort 2 Meeting 20 SRT in `ModelStream_011/captions/` — FIXED + +- Symptom: `ModelStream/ModelStream_011/captions/` held + `ActInf Textbook Group ~ Cohort 2 ~ Meeting 20 (Chapter 9, part 1).en.srt` + alongside the item's own `youtube_captions.txt` for video `Y9hP79tBXHo` + (ModelStream #011.1 ~ Poisson Variational Autoencoder). +- Correct home identified: `TextbookGroup/ParrPezzuloFriston2022/Cohort_2/Meeting_020` + (canonical video `QjGcN1l6NXg`, resolved via `INDEX.json`). +- Disposition: `git mv` into that item's `captions/` (non-destructive rename, + committed on `feat/m4-pages`). +- Context: this file was one of **58 byte-identical copies** of the same + YouTube-derived SRT scattered across unrelated items (GuestStream 051–070, + MathStream 006–012, ModelStream 008–013, MorphStream 001–005, Courses, + symposium items, TextbookGroup Cohort_4). Commit `0bbc97d1` (2026-08-11) + already deleted one copy (Insights_005) and repurposed 4 more as translation + sources; the other 57 copies remain and are tracked as the broader J3/I10 + captions-naming work item in Journal-Utilities, not re-fixed here. The copy + in `Meeting_020` was chosen because that item lacked a `.en.srt` variant + (only `.eng(transcribed).srt`), so the move fills the canonical item's + YouTube-captions slot without overwriting anything. + +## 2. 2022 Robotics translations inside `2021 Symposium .../Translations/` — DOCUMENTED, NOT MOVED + +- Symptom: + `Applied Active Inference Symposium/2021 Symposium with Karl Friston/Translations/` + (capital-T, item videos `INRaCBikpso`, `X2GwqUVLlcs`, `hW9IiOujS1E` — the + three Prof. Karl Friston symposium parts) contains 22 files named + `2nd Applied Active Inference Symposium on Robotics ~ {1st,2nd} session..srt`. + These belong to `2022 Symposium on Robotics` (videos `zm2d9o5n0PU`, + `dTVHHenms_Y`), which already has its own lowercase `translations/` with the + same 22 files. +- Verified: every one of the 22 pairs is near-identical (cue counts equal, + e.g. 5638 cues for the 1st session de pair); diffs are a handful of + re-translated lines. The 2021 item's directory additionally holds the 33 + genuine Friston-symposium translations (3 parts × 11 languages). +- Why not moved: the robotics files are **not an orphaned misfiling but a + near-duplicate second copy**; `git mv` would either clobber the 2022 item's + (newer-translation) files or create a `migrate/` duplicate that the builder + would count twice — both regressions. The M2 translation migration + (`git mv Translations/ → translations/`, normalize `..srt`, + record `previous_paths` in `metadata.json`) is the right vehicle: during that + per-series PR the duplicates should be diffed once more and the loser deleted, + keeping the 2021 item's directory for the 33 genuine Friston files only. +- Interim state: left untouched on `feat/m4-pages`; case-sensitivity already + hides the capital-T directory from the site builder (see + [`m4-site-spec.md`](m4-site-spec.md) and handoff item I4), so the duplicates + are not user-visible on the site.