Skip to content

v1.1.0: felt cues, downbeat hit, levelled swooshes, footnotes on the charts - #1

Merged
MaxGhenis merged 10 commits into
mainfrom
felt-cues
Sep 27, 2026
Merged

MaxGhenis merged 10 commits into
mainfrom
felt-cues

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

The v1.1.0 cut: every change comes from Max's review of v1.0.0. v1.0.0 and its release assets stay as they are.

What changes

  • Cues ("the chimes are a bit cheesy"). bell() (FM) becomes felt(), a soft felted mallet: mostly fundamental, low-passed at 2.4 kHz, an octave lower. The logo's bell cascade is gone. The family curve now brightens a held chord in place of a pitch glide.

  • Downbeat ("the burst on beat 4 around 0:11 makes it sound like a 3/4 measure"). The riser runs over beats 3 and 4 (11.0–11.95 s). The "+$1,600" line and its hit land on the 12.0 s downbeat.

  • Swooshes ("the swoosh sounds … moving the 2200 are a bit loud; they're all throughout the video"). level_sweeps() mixes the rest of the score first. It then sets each noise sweep (three whooshes, and the noise in the 11 s riser and the 25 s swell) SWEEP_UNDER = 6 LU below everything else at the sweep's loudest 400 ms, measured K-weighted (BS.1770). Before, the fixed-gain whooshes peaked 8.3, 3.0 and 1.5 LU over the music. The whoosh band also tops out at 5 kHz instead of 7.

  • Footnotes ("move those footnotes on the last page into smaller things on the relevant graphs … nothing below policyengine.org on the close"). Each chart now carries its own small print in its bottom-right corner:

    • curve: policyengine.py 6.1.1 · static
    • map and stats: policyengine.py 6.1.1 · populace_us_2024 · static / poverty: Supplemental Poverty Measure
    • deciles: policyengine.py 6.1.1 · populace_us_2024 · static

    The close keeps only the wordmark, the tagline and the URL. data/video.json swaps provenance for sources, built from the versions each run recorded.

Invariants (all tested)

  • Sweep level. Before mastering, every sweep's loudest window sits exactly 6 LU under the rest of the score. This is measured independently with pyloudnorm's filters on the returned audio.
    • After mastering, measured through the master's own gain envelope, every sweep sits within half an LU of that. The exception is the swell (≥ 4.3), whose window shares a limited clap.
  • Window bounds. The window never reaches the cue a sweep leads into. Short sweeps are measured over their own length, and what plays after a sweep cannot change its level (property). Sweeps clipped at either end of the score are levelled correctly.
  • Scale and bed. The levelled result ignores a sweep's own gain and scales one-for-one with the bed (property).
  • Conservation. The mastered bed plus the mastered sweeps equals the score, to 1e-9.
  • K-weighting. Our BS.1770 filter bank agrees with pyloudnorm's, to within the 0.04 dB offset between the two filter designs (differential, Hypothesis).
  • Unchanged elsewhere. Before mastering, everything outside the sweeps and their reverb tails is unchanged from v1.0.0 to 3e-13. Normalising to −14 LUFS plays it 0.5 to 1 dB louder.
  • Beat grid. The hit, the drop and the logo land on bar lines. The riser ends just before the hit, and the swell ends on the logo.
  • Mastering and determinism. −14 LUFS and true peak ≤ −2 dBTP in the WAV. Each frame depends only on t (both layouts).
  • Source tags. Each tag names the version and dataset its run recorded, and no provenance remains on the close. Rebuilding video.json is a no-op.

Mutation-checked: sweeps halved or muted, a short-sweep window running past its end, the last 50 ms of windows skipped, a missing empty-span guard, an old event log (hit at 11.5 s), and the limiter at −1.5 dB each fail a test.

Reviews

Three independent review rounds:

  • Codex on the cues and downbeat: changes requested, fixed.
  • Opus plus Astra on the sweep levelling: changes requested, fixed, then approved.
  • Opus on the source tags.

Renders

out/ renders pass tools/check_outputs.py: 30.000 s, 1,800 frames, −14.2 LUFS, true peak −1.5 dBTP after AAC, 0 single-frame glitches. After merge, they ship as release v1.1.0.

🤖 Generated with Claude Code

MaxGhenis and others added 2 commits September 25, 2026 17:29
The FM bells on the lock, citation, stats, deciles and household light-ups
read as cheesy. They are now a felted mallet (mostly fundamental, fast-decaying
second harmonic, muffled contact noise, low-passed at 2.4 kHz) an octave lower;
the bell cascade under the wordmark is gone; the family curve brightens a held
chord (350 Hz to 2.6 kHz low-pass following the computed gain) instead of a
pitch glide. 2-8 kHz energy falls about 5 dB in the cue-heavy sections;
integrated loudness stays at -14.0 LUFS. Picture unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The hit sat on beat 4 of bar 6 (11.5 s) right after the drums and bass drop
out under the riser, so the bar read as 3/4 until the groove returned. The
second line and its hit now land on the 12.0 s downbeat, the riser covers
beats 3-4, and the note follows at 12.15 s.

Also make the clean-edges test state the real property: the fades start and
end at exactly zero and stay under the limiter ceiling times the ramp (the old
1e-3 bound on the 10th sample was arbitrary and failed on a 2% louder intro).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis MaxGhenis changed the title Replace the chimes with a soft felted mallet Soften the cues and land the family-scene hit on the downbeat Sep 25, 2026
MaxGhenis and others added 6 commits September 25, 2026 18:32
- audio/events.json: the regenerated cue list (riser 11.00-11.95 s, hit 12.0 s).
  Without it a fresh checkout's score still hit on beat 4.
- test_heavy_cues_land_on_downbeats: hit, drop and logo sit on even seconds (bar
  lines at 120 BPM), the riser ends just before the hit, the swell ends on the logo.
- The true-peak check now enforces the -2 dBTP ceiling (was -1).
- Curve cue: validate the samples and clamp the normalised gain to [0, 1], so a
  gain above gmax cannot push the filter cutoff past Nyquist. Current data is
  already in range; the rebuilt score.wav is byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The whooshes (the dive to $2,200, its flight into the code, the move to the
map) had a fixed gain. Measured K-weighted, they peaked 8.3, 3.0 and 1.5 LU
over whatever else was playing. The noise in the 11 s riser and the 25 s
swell sat level with the bed.

Now level_sweeps() mixes the rest of the score first, then scales each noise
sweep so its loudest 400 ms window sits SWEEP_UNDER (6) LU under the rest of
the mix in that same window. A quiet intro and a busy bridge get the same
relationship. The windows stay inside the sweep, so the riser is never
compared with the 12 s hit it leads into; the limiter squashes that hit.

The whoosh band now tops out at 5 kHz instead of 7, so it hisses less. The
rng draws are unchanged, so every other sound in the score is sample-identical.

Tests (tests/test_soundtrack.py):
- Every sweep is exactly 6 LU under the bed, and its window never reaches the
  next cue.
- The same property holds on the mastered audio: the score minus a sweep-free
  score, measured with pyloudnorm's own K-weighting, is at least 4 LU under.
  The limiter trims the swell's margin to 4.6 because a clap shares its
  window.
- Differential: our BS.1770 filter bank agrees with pyloudnorm's to within
  0.1 LU on random band-limited noise (Hypothesis). The two filter designs
  differ by a constant 0.04 dB.
- Property: the levelled result ignores the sweep's own gain and scales
  one-for-one with the bed.
- Mutation-checked: SWEEP_UNDER = 0, a window running 0.3 s past the sweep,
  and sweeps mixed 6 dB hot all fail.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review of 975572b found four problems.

- A sweep shorter than 400 ms had its window run past its end. Audio playing
  after the sweep could then set its level (+48 dB in a probe). Now a short
  sweep is measured over its own length.
- A sweep cut off by the end of the score raised IndexError, and one placed
  wholly outside it crashed on an empty span. Now the first is levelled over
  what remains and the second is dropped.
- Windows sat on a 10 ms grid that skipped the last one. The riser and swell
  peak at their ends, so they were levelled 5.70 and 5.48 LU under, not 6.
  Every per-sample window is now a candidate.
- The tests trusted the report's own numbers, so halved or muted sweeps
  passed.

Tests now:
- build(stems=...) exposes the bed and the levelled sweeps before and after
  the master. The master is one filter and one gain envelope, shared, so the
  mastered stems sum to the score (asserted to 1e-9).
- The 6 LU margin is measured independently with pyloudnorm's filters on the
  returned audio, both before and after the master. The mastered check is
  bounded on both sides: 4 LU at least (the limiter trims the swell to about
  4.6), SWEEP_UNDER + 0.3 at most.
- An exhaustive scan finds no window inside a sweep's span louder than the one
  it was levelled at.
- Property: raising the bed after a sweep ends leaves the sweep's level
  unchanged, for sweeps from 50 ms to 1.5 s.
- Sweeps at 29.2 to 30.5 s and before 0 s render without error.
- Mutations: sweeps halved, sweeps muted, a short-sweep window running past its
  end, the last 50 ms of windows skipped, and no empty-span guard. Each one
  fails a test.

A note on 975572b's "every other sound is sample-identical": that holds before
the master. The master normalises the whole score to -14 LUFS, so with quieter
sweeps it plays the rest 0.5 to 1 dB louder.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Max asked for the sign-off's three lines of small print to become small tags on
the relevant graphs, with nothing below policyengine.org on the close.

Each chart now carries its own source line in its bottom-right corner, in the
mono "machine voice", and it fades with the chart:
- The earnings curve: "policyengine.py 6.1.1 · static". In landscape it sits on
  the axis-title line; in portrait it drops a line below the chart, where it
  would otherwise run into "Earnings".
- The map and national stats: "policyengine.py 6.1.1 · populace_us_2024 ·
  static", then "poverty: Supplemental Poverty Measure". The poverty measure
  stays on screen because the child poverty figure depends on it.
- The decile bars: "policyengine.py 6.1.1 · populace_us_2024 · static".

The close keeps the wordmark, the tagline and "Free and open source ·
policyengine.org". The count of distinct households behind the 12,000 dots
(6,976) moves to the README.

data/video.json replaces "provenance" with "sources". The build takes the tags
from the versions and dataset each run recorded. The curve's tag comes from the
sweep's metadata and the national tags from national.json's, and a test asserts
both, plus that no provenance remains. Every other key is byte-identical, and
rebuilding the data is a no-op.

tests/conftest.py drops Hypothesis's per-example deadline for the whole suite.
These are invariant tests, and a test that allocates a 23 MB score buffer per
example can overrun 200 ms on a loaded machine; one run failed intermittently
and could not be reproduced.

Determinism: identical frames at every checked timestamp, both layouts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follow-ups from the re-review of 0408fa9 (verdict APPROVE):

- A sweep shorter than 400 ms is still measured over its own length, but the
  bed is now measured over the 400 ms that end where the sweep's window ends.
  A single drum hit inside a 50 ms slice can no longer swing its gain by about
  10 dB, and nothing after the sweep can set its level. Every sweep in the
  score is at least 0.75 s long, so score.wav is byte-identical.
- momentary() raises ValueError on windows that don't fit. It used a bare
  assert, which python -O strips.
- The mastered-audio test has per-sweep floors: the swell (limited by the
  clap at 25.5 s) must stay at least 4.3 LU under and every other sweep at
  least 5.5. The old single floor of 4.0 would have let a whoosh drift 2 LU
  louder.
- The edge test now checks that sweeps clipped at 29.2-29.999 s, and one
  starting at -0.5 s, sit exactly SWEEP_UNDER below the bed, not just that
  they render.
- The window-bounds check allows one sample of rounding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each tag line is now its own block with text-wrap: balance. A statement that
wraps (the UK map's ONS licence lines in portrait) splits evenly instead of
leaving an orphan. In portrait the map tag may wrap within 540 px, the free
right column beside the third stat.

The US posters at 12.9, 18.9, 23.6 and 28.8 s are pixel-identical to the
renders from ca03434 in both layouts, so the US videos need no re-render.
Frames match at every checked timestamp.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis MaxGhenis changed the title Soften the cues and land the family-scene hit on the downbeat v1.1.0: felt cues, downbeat hit, levelled swooshes, footnotes on the charts Sep 26, 2026
MaxGhenis and others added 2 commits September 27, 2026 17:35
Follow-ups from the independent review of the source-tag commits (APPROVE):

- The curve tag says "policyengine.py 6.1.1", but the sweep computes the curve
  with policyengine_us directly, and policyengine.py spot-checked only 7
  points. New data/compute/curve_policyengine_py.py runs the full $250 grid,
  601 points, through policyengine.py's pe.us.calculate_household with an
  earnings axis. It matches the plotted change in net income at every point to
  $0.004. The build refuses to write the tag unless that check covers every
  point and ran on the tagged version, and a test asserts the same.
- "static" and "Supplemental Poverty Measure" are now checked, not just typed.
  The build asserts that no gov.simulation behavioural parameter is in either
  reform, that the national run recorded no behavioural responses, and that
  the on-screen child poverty stat is the SPM figure.
- A missing input no longer yields a literal "MOCK" tag with no banner. The
  build adds "sources" to mock_parts, so the page paints the banner.
- Tests now pin the close ("Free and open source · policyengine.org" and
  nothing else) and assert the page draws all three tags and no provenance.
- The level_sweeps docstring says the window lies within the dry sweep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The true-peak limiter took its 3 ms lookahead as np.minimum.reduce over 144
np.roll copies of the full-length gain curve. That is about 1.7 GB of
temporaries per mastering pass, and the test suite peaked at 4.3 GB of
resident memory, enough for the OS to kill a local run.

lookahead_min() computes the same wraparound window minimum with
scipy.ndimage.minimum_filter1d.
- A Hypothesis differential test checks it against the original np.roll
  formulation for any length and window: equal, not approximately equal.
- The rebuilt audio/score.wav is byte-identical.
- The suite now peaks at 1.0 GB and runs faster.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis merged commit c592471 into main Sep 27, 2026
2 checks passed
@MaxGhenis
MaxGhenis deleted the felt-cues branch September 30, 2026 10:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant