Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ video library.
[![Journal integrity](https://github.com/ActiveInferenceInstitute/ActiveInferenceJournal/actions/workflows/journal-integrity.yml/badge.svg)](https://github.com/ActiveInferenceInstitute/ActiveInferenceJournal/actions/workflows/journal-integrity.yml)

Learn more: https://activeinference.institute/learning/ ·
Site: https://activeinferenceinstitute.github.io/ActiveInferenceJournal/ ·
Tooling: https://github.com/ActiveInferenceInstitute/Journal_Utilities

## Layout
Expand Down
27 changes: 27 additions & 0 deletions TO-DO.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,3 +107,30 @@ overhaul, cross-cutting refactors.
- **Curated assets elsewhere** — after the image-repair pass, no `.md` under
`assets/` has broken relative refs and no machine paths remain anywhere in
the repo (verified by grep).

## M4-pages wave (feat/m4-pages, 2026-09-23)

- ✓ **I4 verified on the real tree** — builder reads only lowercase
`translations/` (Journal-Utilities `src/journal_utilities/site/builder.py:121`);
tree holds 13 items with lowercase `translations/` (~336 SRTs) vs 131 items
with capital-T `Translations/` (~2,880 SRTs); live `manifest.json` shows
languages for 13/573 items (129 language entries). Fix + spec in
[`docs/m4-site-spec.md`](docs/m4-site-spec.md) §0; permanent fix = M2
translation migration.
- ✓ **Textbook Cohort 2 Meeting 20 SRT misfiled in
`ModelStream_011/captions/`** — `git mv` to canonical
`TextbookGroup/ParrPezzuloFriston2022/Cohort_2/Meeting_020/captions/`
(video `QjGcN1l6NXg` via INDEX.json). 57 further byte-identical copies
remain scattered (GuestStream/MathStream/ModelStream/MorphStream/Courses/
symposium/Cohort_4) — owned by the J3/I10 captions-naming validator, see
[`docs/misfiling-findings.md`](docs/misfiling-findings.md).
- ✓ **2022 Robotics translations inside 2021 Symposium `Translations/`** —
verified near-duplicate second copy (22 files, diffs are re-translated
lines); documented rather than moved (a move would clobber or double-count);
to be resolved during the M2 per-series migration —
[`docs/misfiling-findings.md`](docs/misfiling-findings.md) §2.
- ✓ **README.md** — Pages site link added.
- ✓ **`docs/m4-site-spec.md`** — M4 builder spec (static
`/item/<series>/<item>/index.html`, schema.org VideoObject JSON-LD,
canonical/og, sitemap.xml + robots.txt), prepared for the Journal_Utilities
builder agent; not yet implemented.
7 changes: 7 additions & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,3 +33,10 @@ materials from the Active Inference Institute video library.

- `main` — content without audio (lightweight).
- `audio` — `main` + `<item>/audio/<name>.64k.m4a` (64 kbps media), same layout.

See also: [`m4-site-spec.md`](m4-site-spec.md) — spec for the M4 static
per-item pages + sitemap/robots (for the Journal_Utilities builder agent) —
and [`misfiling-findings.md`](misfiling-findings.md) — captions/translations
misfiling dispositions (Textbook Cohort 2 Meeting 20 SRT; 2021 Symposium
`Translations/` robotics duplicates).

135 changes: 135 additions & 0 deletions docs/m4-site-spec.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,135 @@
# M4 site spec — static per-item pages, sitemap, robots.txt

Status: **specification for implementation, not yet implemented.** The current
deployed site (hash-routed SPA fed by `manifest.json` + `data/*.json`, built by
Journal-Utilities `src/journal_utilities/site/builder.py` via
`scripts/build_pages_site.py`, deployed by
[`.github/workflows/deploy-pages.yml`](../.github/workflows/deploy-pages.yml))
exposes exactly one crawlable URL. This file states what the M4 wave must emit
so the Journal_Utilities builder agent (or a later wave) can implement it.
Handoff context: `/tmp/aii_journal_handoff.md` milestone M4.

## 0. Prerequisite fix bundled with M4 (handoff I4)

`builder.py` reads only lowercase `translations/` (`tr_dir = item_dir /
"translations"`, builder.py:121) while the journal tree holds **131 items with
capital-T `Translations/`** (~2,880 SRTs) vs **13 items with lowercase
`translations/`** (~336 SRTs). Verified 2026-09-23 on this tree and against the
live deployed `manifest.json` (13/573 items with `languages`, 129 language
entries; the site ships 3,552 fewer translation files than the journal holds:
6,432 total SRTs vs 3,552 readable).

Fix (permanent fix remains the M2 `Translations/ → translations/` migration):

```python
tr_dir = next(
(d for d in (item_dir / "translations", item_dir / "Translations") if d.is_dir()),
None,
)
```

Also strip legacy ISO-639-2 parenthesized tags when deriving `lang`
(builder.py:124 derives it from `name.split(".")[-2]`, so
`...chi(translated)` yields `chi(translated)` today). Map `chi→zh-Hans`
or `zh-Hant`, `dut→nl`, `fre→fr`, `ger→de`, `jpn→ja`, `kor→ko`, `rus→ru`,
`por→pt`, `spa→es`, `ita→it` per the M2 table; `translate_subtitles_openrouter.py:381`
must not re-translate files that already exist under the normalized name.

## 1. URL scheme

Base URL (canonical, used everywhere below):
`https://activeinferenceinstitute.github.io/ActiveInferenceJournal/`

- SPA stays at `/` (hash routing preserved for backward compatibility).
- New per-item page per INDEX item:
`/item/<series>/<item>/index.html` → canonical URL
`.../item/<series>/<item>/` (trailing slash, no `index.html` in canonicals).
- `<series>` and `<item>` are the exact on-disk directory names from
`INDEX.json items[].path` (`data/video/activeinferenceinstitute/<series>/<item>`) —
including spaces (e.g. `Applied Active Inference Symposium`) — percent-encoded
in emitted URLs. Directory names ARE the slugs; M4 must not invent new slugs
(folder slugification is DAF decision §6/J11, out of scope).
- `data/<series>_<item>.json` payload files keep their existing location; item
pages reference them for the interactive transcript view.

## 2. Per-item page contract

Emit `output/item/<series>/<item>/index.html` for **every** item in
`INDEX.json items[]` (573 items at the time of writing). Requirements:

1. **Server-rendered (baked) content** — no JS required to read the core
metadata: title, series, date (`parts[].upload_date`, ISO 8601), guests,
summary/abstract from `metadata.json`, chapter list
(`parts[].chapters`, present only with provenance), full transcript text
(from `transcript.txt`, `parts`-tagged when multi-part).
2. **Embedded video** — YouTube iframe per part using `parts[].video_id`
(`https://www.youtube-nocookie.com/embed/<video_id>`), part titles as
headings.
3. **Structured data** — one `application/ld+json` block per part:
schema.org `VideoObject` with
`name`, `description` (abstract from metadata), `uploadDate` (ISO 8601),
`contentUrl`/`embedUrl` (the YouTube URL from `parts[].url`),
`thumbnailUrl` (`https://i.ytimg.com/vi/<id>/hqdefault.jpg`),
`inLanguage` (source language, `en`), `transcript` reference, and
`hasPart` → `Clip` entries from gated chapters (`name`, `startOffsetTime`,
`endOffsetTime`, `url` with `?t=` fragment). Top level: an `Dataset` node
for the item (name, `license` = CC-BY-4.0, `isPartOf` → the journal Dataset,
`citation` → Zenodo DOI `10.5281/zenodo.7299755`).
4. **Head tags** — `<link rel="canonical" href="<base>item/<series>/<item>/">`,
Open Graph (`og:type=video`, `og:title`, `og:description`, `og:url`,
`og:image` = thumbnail), Twitter `summary_large_image`.
5. **Language list** — translations actually present for the item (post-I4-fix
set), linking to the interactive view; do not fabricate languages.
6. **Navigation** — breadcrumb (journal → series → item), links to
prev/next item in series, link back to the SPA for the interactive player.
7. **Non-goals**: no search, no client-side filtering, no new JS frameworks;
static HTML + the existing CSS.

## 3. sitemap.xml

Emit at `<base>/sitemap.xml`:

- `<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">` with **one
`<url>` per item** (573): `<loc>` = canonical item URL,
`<lastmod>` = max `parts[].upload_date` for the item (ISO 8601, omit if
unknown — never fabricate), **no** `changefreq`/`priority` (noise).
- Plus the root `<loc>` = `<base>/`.
- Single file is fine (well under the 50k-URL / 50 MB limits); if it ever
exceeds them, switch to a sitemap index `<base>/sitemap-index.xml`.
- Post-build assertion: `sitemap` URL count == processed item count (builder
already returns `items_processed`); fail the build on mismatch.

## 4. robots.txt

Emit at `<base>/robots.txt`:

```
User-agent: *
Allow: /
Sitemap: https://activeinferenceinstitute.github.io/ActiveInferenceJournal/sitemap.xml
```

Both files are generated by the builder (not committed by hand) and must be
included in the Pages artifact alongside the existing `.nojekyll`.

## 5. Deployment notes (ties to handoff E5/J7, owned by the workflow wave)

- Builder changes land in Journal-Utilities; pin the workflow's utilities
checkout to a tag/SHA, sparse-checkout, plain `python` (no torch/CUDA), and
keep a single Pages deployment (drop the dual actions+gh-pages deploy or
make branch-mode the fallback behind a flag).
- Post-deploy smoke test (CI): fetch `<base>/sitemap.xml`, assert URL count
equals `manifest.json` `total_items`, spot-check Rich Results Test on one
sample item page, assert `<base>/robots.txt` returns 200.

## 6. Acceptance criteria

- Every INDEX item has a static URL returning 200 with server-rendered title,
guests, date, summary, chapters, transcript, and a valid VideoObject
JSON-LD block per part.
- `sitemap.xml` lists the root + all item URLs; `robots.txt` references it.
- Rich Results Test validates a sampled item page (no VideoObject errors).
- Site language counts reflect the union of both `translations/` spellings
(post-fix: 144 items with ≥1 translation at current tree state).
- Canonical URLs resolve against `INDEX.json` `items[].path` (series/item
exactly); a test asserts emitted `<loc>` ↔ item-path bijection.
53 changes: 53 additions & 0 deletions docs/misfiling-findings.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# Misfiling findings — captions and translations triage

Recorded: 2026-09-23 (M4-pages wave, fleet handoff `/tmp/aii_journal_handoff.md`).
Scope: the two known misfilings called out in the handoff baseline, verified
against the real tree. Generator-owned rewrites should consult this file.

## 1. Textbook Cohort 2 Meeting 20 SRT in `ModelStream_011/captions/` — FIXED

- Symptom: `ModelStream/ModelStream_011/captions/` held
`ActInf Textbook Group ~ Cohort 2 ~ Meeting 20 (Chapter 9, part 1).en.srt`
alongside the item's own `youtube_captions.txt` for video `Y9hP79tBXHo`
(ModelStream #011.1 ~ Poisson Variational Autoencoder).
- Correct home identified: `TextbookGroup/ParrPezzuloFriston2022/Cohort_2/Meeting_020`
(canonical video `QjGcN1l6NXg`, resolved via `INDEX.json`).
- Disposition: `git mv` into that item's `captions/` (non-destructive rename,
committed on `feat/m4-pages`).
- Context: this file was one of **58 byte-identical copies** of the same
YouTube-derived SRT scattered across unrelated items (GuestStream 051–070,
MathStream 006–012, ModelStream 008–013, MorphStream 001–005, Courses,
symposium items, TextbookGroup Cohort_4). Commit `0bbc97d1` (2026-08-11)
already deleted one copy (Insights_005) and repurposed 4 more as translation
sources; the other 57 copies remain and are tracked as the broader J3/I10
captions-naming work item in Journal-Utilities, not re-fixed here. The copy
in `Meeting_020` was chosen because that item lacked a `.en.srt` variant
(only `.eng(transcribed).srt`), so the move fills the canonical item's
YouTube-captions slot without overwriting anything.

## 2. 2022 Robotics translations inside `2021 Symposium .../Translations/` — DOCUMENTED, NOT MOVED

- Symptom:
`Applied Active Inference Symposium/2021 Symposium with Karl Friston/Translations/`
(capital-T, item videos `INRaCBikpso`, `X2GwqUVLlcs`, `hW9IiOujS1E` — the
three Prof. Karl Friston symposium parts) contains 22 files named
`2nd Applied Active Inference Symposium on Robotics ~ {1st,2nd} session.<lang>.srt`.
These belong to `2022 Symposium on Robotics` (videos `zm2d9o5n0PU`,
`dTVHHenms_Y`), which already has its own lowercase `translations/` with the
same 22 files.
- Verified: every one of the 22 pairs is near-identical (cue counts equal,
e.g. 5638 cues for the 1st session de pair); diffs are a handful of
re-translated lines. The 2021 item's directory additionally holds the 33
genuine Friston-symposium translations (3 parts × 11 languages).
- Why not moved: the robotics files are **not an orphaned misfiling but a
near-duplicate second copy**; `git mv` would either clobber the 2022 item's
(newer-translation) files or create a `migrate/` duplicate that the builder
would count twice — both regressions. The M2 translation migration
(`git mv Translations/ → translations/`, normalize `<video_id>.<bcp47>.srt`,
record `previous_paths` in `metadata.json`) is the right vehicle: during that
per-series PR the duplicates should be diffed once more and the loser deleted,
keeping the 2021 item's directory for the 33 genuine Friston files only.
- Interim state: left untouched on `feat/m4-pages`; case-sensitivity already
hides the capital-T directory from the site builder (see
[`m4-site-spec.md`](m4-site-spec.md) and handoff item I4), so the duplicates
are not user-visible on the site.
Loading