Skip to content

perf(tm): run long Turing machine runs through a table-driven fast lane - #104

Merged
thethinkmachine merged 2 commits into
mainfrom
perf/fast-tm-lane
Oct 1, 2026
Merged

thethinkmachine merged 2 commits into
mainfrom
perf/fast-tm-lane

Conversation

@thethinkmachine

@thethinkmachine thethinkmachine commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Running BB(5) to its halt took about 14 seconds headless (~17 s in the browser with ⏭). Almost none of that was the machine. Each step went through a generator, a Map tape of strings, a loop fingerprint, three calls into the tape log and step columns, a transition lookup by name, and a reactive Set asked whether the state accepts. Studying Busy Beaver Blaze showed the same steps as table lookups over a byte array take a few nanoseconds. This brings that approach into the player without changing what a run records.

Measured

BB(5) to its halt (47,176,870 steps), Node 24, Ryzen 7 7435HS, main and this branch alternating, median of 3:

before after
Run to the end (traceMachine, what ⏭ does minus painting) 14.3 s 1.43 s
Same, in ⏭'s 500-step slices 13.6 s 1.46 s
Run + space-time diagram built 15.5 s 4.3 s
One step at a time (playback at Max) 0.29 µs/step 0.25 µs/step
Memory ~6 B/step ~6 B/step

The diagram row improves less because building the diagram itself still takes ~3 s. npm run bench -- player shows the TM cases 84–90% faster. Every other machine type is within its noise, and I re-checked Moore, NFA and DFA directly.

Re-measured after merging main (77ac4c6)

main moved 63 commits past this PR's base, so the numbers were taken again on the merged head (423105c) against current main. This was a different machine (4-thread Xeon @ 2.8 GHz, Node 22, a shared cloud container), so the absolute times are slower than the table above. Compare the ratios, not the times. BB(5) to its halt, alternating trees, median of 3:

main 77ac4c6 this PR 423105c
Run to the end (traceMachine) 25.6 s 2.35 s 10.9×
Same, in ⏭'s 500-step slices 24.2 s 2.20 s 11.0×
Run + space-time diagram (index + overview) 29.6 s 8.02 s 3.7×
One step at a time, 3M pulls (playback at Max) 0.55 µs/step 0.41 µs/step −26%
Peak RSS, run to the end 397 MB 375 MB

Every run ended the same way on both trees: 47,176,871 steps, accept, 4,098 ones.

npm run bench -- --baseline <main run> (80 cases, 3 rounds each, main recorded on the same container just before): 0 worse, 1 better, 0 changed checks. The TM player cases moved: BB(5), first 1M steps −91% (the one case outside its noise, ▼ faster), and TM sweeper, 100k steps −85%. Everything else, including decide, search, layout, space-time and memory per step, is within its run-to-run spread. The horizon cases (BB(5) run to halt, with and without the diagram) hold at 285 MB.

How it works

  • Hand-over. Once the loop detector stops looking (after 5,000 configurations, or from step 0 when loop detection is off), streamTM hands the rest of the run to the new js/machines/fast-tm.js. Short runs, where loops are caught, never leave the old loop. It covers TM and ITM.
  • Same record. States and symbols become small integers and δ becomes flat typed arrays, filled the first time each (state, symbol) pair is met. The tape is the tape log's own copy in codes. Steps are written into the same columns begin/noteWrite/step fill, including checkpoints and the undo record for explicit blanks. The player, scrubber, trace log, space-time diagram and StateMate are untouched.
  • Batches. The run cursor now tells a generator how many steps it wants (next(n)), and a producer can answer with batch(count, last, at). ⏭ still asks in 500-step slices, so block boundaries are found at the same granularity. playEagerly and for…of still get one step per next().
  • Single steps skip the table. Playback asks for one step at a time, so that path asks the lookup directly, as the old loop did. An early version re-checked the table per step and made playback at Max 2.8× slower. That's fixed, and now slightly faster than main.
  • Edits to a paused run. The old loop picks up edits to a transition's target, write and direction, and to accepting marks, on the next step. The fast lane re-reads those before every batch, so it behaves the same.

Test plan

  • npm test: 2,353 tests, 0 failures (1 skipped, the C-compiler test, as on main)
  • After merging main: npm test 2,634 tests, 0 failures (same 1 skipped), and CI green on ubuntu/macOS/windows × Node 20/24
  • New tests/fast-tm.test.js: every step compared against the old loop (setFastLane(false)) on state, transition, note, verdict, journal row, checkpoints, and the tape read forwards, backwards and scattered. It covers BB(5), 160 random machines (wildcards, missing transitions, stay-put moves, both tape models, loop detection on and off, explicit blanks), pull sizes from 1 to a whole run, a run edited while paused, simTM, and a trailing explicit blank stepped back over.
  • Deliberately breaking the fast lane three ways (dropping the explicit-blank undo record, skipping the re-check, using a stale rule on resume) each fails the new tests.
  • BB(5) to its halt: 47,176,870 steps, 4,098 ones, accept, identical before and after.
  • Not yet checked in the browser or Electron. The numbers above are headless Node.

Not in this PR

  • The TM decider (decideMachine) still peaks at ~4.7 GB on BB(5) to its halt, from its per-configuration fingerprint table.
  • LBA and MTM could use the same fast lane.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XKgcDmTSkNcY5B2bfTtt6b

A TM step cost about 300ns, and almost none of it was the machine: the
generator handing over a step object, the cursor pulling it, a Map tape of
strings, the loop fingerprint, three calls into the tape log and step
columns, a transition lookup by name, and App.accepts asked through the
reactive Set. The same steps as table lookups over a byte array take a few
nanoseconds.

Once the loop detector stops looking (after 5,000 configurations, or from
step 0 when loop detection is off), streamTM hands the rest of the run to
js/machines/fast-tm.js: states and symbols as small integers, delta as flat
typed arrays filled the first time each pair is met, the tape log's own
copy of the tape as the tape, and steps recorded straight into the same
columns begin/noteWrite/step would have filled. Every reader of a run is
untouched.

The run cursor now tells a generator how many steps it wants (next(n)),
and a producer can answer with batch(count, last, at). The drain still
asks in 500-step slices, so block boundaries are found at the same
granularity; playEagerly and for...of ask for nothing and get one step per
next(). A single step skips the table and asks the lookup directly, so
playback at Max does not pay for a re-check. Before each multi-step batch
the table re-reads the transition fields and accepting marks the slow loop
reads every step, so an edit to a paused run is obeyed the same way.

BB(5) to its halt, Node 24: traceMachine 14.3s -> 1.43s; the drain's
500-step slices 13.6s -> 1.46s; with the space-time diagram built 15.5s ->
4.3s. One-step pulls 0.29 -> 0.25us. Memory unchanged at ~6 bytes a step.

tests/fast-tm.test.js compares every step against the slow loop
(setFastLane(false)) across BB(5), 160 random machines, pull sizes from 1
to a whole run, a run edited while paused, simTM, and an explicit blank
stepped back over. Removing the undo exception, the re-check, or the live
read on resume each fails it.
Copilot AI lite review requested due to automatic review settings September 27, 2026 19:39
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, you can upgrade your account or add credits to your account and enable them for code reviews in your settings.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Both conflicts were the fast lane and tm-behaviour.js landing in the same
two lists — the js/machines/ listing in CLAUDE.md and the harness's module
imports and namespaces. Kept both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKgcDmTSkNcY5B2bfTtt6b
@thethinkmachine
thethinkmachine merged commit 5333542 into main Oct 1, 2026
18 checks passed
@thethinkmachine
thethinkmachine deleted the perf/fast-tm-lane branch October 1, 2026 20:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants