Skip to content

A swap's redial dials until logd admits one, bounded by the swap's window alone - #566

Merged
Japabu merged 16 commits into
mainfrom
wt/toyos-redial
Sep 28, 2026
Merged

Japabu merged 16 commits into
mainfrom
wt/toyos-redial

Conversation

@Japabu

@Japabu Japabu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

A swap of netd reddened lan_swap, swap_netd and swap_crash_rolls_back with "the stream's redial was turned away 64 time(s), its ceiling of 64, and gave up", on boots where the guest finished every swap: the re-claimed netd leased, and init said in service or restored. The red came from the host's own dial count. A count is neither the event nor a timeout.

What changed, per decision

  • The redial waits on the event, bounded by the swap's window. The event is logd admitting a dial. It can only be seen by dialing, and each dial already waits on the machine's answer. Stream::redial(by) dials until a connection carries a line or by passes, and a redial whose bound passes says so in Stream::unopened. metalswap::swap bounds the redial by what is left of the swap's window. turned_since_dial and past_ceiling are deleted: the first dial's count is turned_away itself, since nothing is dialled before it.
  • On a forward, a failed connect ends the redial at once. Peer::At's only user is QEMU's slirp hostfwd, which takes every connect for as long as QEMU lives. A refused or reset connect there means QEMU has exited, and without a dial ceiling nothing else stopped that redial spinning a host core until the window passed. count_or_give_up ends it on the first failed connect and says why in unopened. A forward's first dial keeps its ceiling.
  • A redial's bound is read in one place, and it names the redial's latest dial. serve used to read the clock after a connection closed before a line, and open read it again before its first dial, starting from "never asked". A bound passing between the two reads ended a redial that had dialled, with its close counted, as … by the bound: never asked. serve no longer reads the bound. It hands the close's reason to open as its starting last, so open's one check ends every redial and names what its latest dial got: … by the bound: the latest connection ended before a line: <how>. serve's second reason format is deleted. read returns how a connection it counted as turned away ended, in place of a bool.
  • TURNED_AWAY_CEILING bounds first dials only, and it is private to metaltalk. Every caller passed it, so the ceiling parameter of Stream::connect and State::ceiling are gone, and open is told whether it is redialling.
  • swap prints the milliseconds to the admission. It does not judge them. The judge's said line already carries how many dials were turned away. When the redial admits nothing, the judge reds on connections (1, 1) and the missing word, with no Err of swap's own.
  • The judge's ceiling finding is deleted. turned_away stays in the report as a number and is never a verdict.
  • The three rows come off src/redlist.rs, and issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md is deleted with them. Its exit condition was "Stream::redial gives up on its time bound alone, never on a dial count, and a swap whose refusal window is staged long is green". Hold-green meets it (below). The slug has no other citation in the tree (git grep).
  • issues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.md records that the T14's redial may ask the link faster than RFC 6762 §5.2's floor. That rate is unmeasured on metal, and the issue's exit removes it.
  • Tests.
    • A redial asks past 1000 refusals and resets on a name until a dial is taken. It is staged through Reach on Peer::Named, because refusals are real there (the T14).
    • A redial asks past 200 closes before a line until one carries a line.
    • A redial ends at its bound alone and says so, both on a refusing name and on closes before a line. The closes arm does not assume its reader dials inside the 50 ms bound (see "The two load-flaky tests" below). A redial that dialled must name its close, never "never asked".
    • Every test close of an accepted connection is a shutdown(Shutdown::Both) before the drop (close). On macOS std sets close-on-exec only after accept returns, and a child spawned in that window holds a copy, so a drop alone does not give the reader its EOF. close_every's closer ends on its wake-up connection rather than closing it.
    • A redial on a forward that refuses ends at the refusal. The refusal is staged through Reach (TurnedAway), because a loopback listener this test process drops still takes connects while any child that another test is spawning holds its fd. The test asserts that the first dial counts 0 and the redial counts 1.
    • A first dial refused is asked again up to its ceiling of 64. Its refusals are staged through Reach (Refusing) for the same reason.
    • The judge's test requires 10 000 dials turned away to be no verdict.

The forward test's flake

a_redial_on_a_forward_that_refuses_ends_at_the_refusal went red under host load with turned_away() == 2.

  • The landing round's theory is refuted. It held that the first dial retries a transient refusal from a live listener. But read_first's listener is bound before the dial and never has more than one pending connection, so a loopback SYN to it is taken, never refused. In all 9 reds of 33cfc348's binary, its new pre-redial assertion (turned_away() == 0) held, and the count of 2 was the redial's.
  • Root cause: the test dialled a listener it had dropped, and a dropped listener is not provably closed. A child that another test is spawning holds a copy of every fd until its exec closes the close-on-exec ones. The redial's connect was taken by the still-open listener and reset once the child closed it, which read counts. It was then dialled again and refused, which count_or_give_up counts. That makes two counts.
  • Measured with a scratch probe (fdprobe/): bind 127.0.0.1:0, drop, connect, 20000 times per run.
    • No concurrent spawn: 0 taken, twice.
    • A thread spawning children beside it: 34 taken and 4 reset, then 26 taken and 7 reset.
  • A second, smaller window: accepted's helper thread drops its try_clone only after its send. A 300 ms sleep after the send reproduced left: 2, right: 1 every time. Closing the clone first (33cfc34) still went red 9 times in 41 full suites, the same count as the unfixed binary in the same interleaved runs. 8cc19ef reverts that change, since no test depends on a dropped listener closing any more.
  • The product code is unchanged. It counts that sequence correctly. The test now stages the forward's refusal through Reach, as the other redial tests stage theirs, and issues/build/a-forwards-redial-negative-control-counts-a-live-listeners-transient-refusal.md is deleted. The slug had no other citation.

Proof: 200 iterations beside a looping cargo test --workspace --exclude toyos-build (11 runs, all EXIT=0), at one-minute load averages up to 43. Each iteration ran three things.

run this test green
the test alone, at 8cc19ef1 200/200
inside the full --lib suite, at 8cc19ef1 200/200
inside the full --lib suite, 941f0bda's binary, interleaved 157/200 (43 red)

A first loop, with the test alone, went 200/200 on both the unfixed and the fixed binary at load average 3.5, so it had no power to detect the race.

The two load-flaky tests (e47d1f0)

Both tests the loaded suite found are fixed here, and their issue files are deleted.

  • a_refused_first_dial_is_asked_again_up_to_its_ceiling dialled a dropped listener. That is the forward test's mechanism: a taken dial closed before a line ends a first dial with unopened unset. Its refusals are now staged through Reach by Refusing, which refuses every dial and treats a resolve or an ask as unreachable!, since an address is only dialled.
  • a_redial_ends_at_its_bound_alone_and_says_so's closes arm assumed its reader dials inside a 50 ms bound. No test can make that true. redial computes the bound on the test's thread and the reader reads it on its own, and staging a clock would be a production change. So the arm accepts either history, and its reason must name what happened:
    • no dial counted: … by the bound: never asked;
    • one dial counted: by the bound: the latest connection ended before a line;
    • more than one: the test fails.
      close_every now closes a connection only once the bound has passed since it took it. A redial that dialled therefore ends on its first close.
  • Deterministic reproduction of the filed red: probe-late-reader.patch sleeps 100 ms at the top of a redial's open. With it, the test as 54ce11ea has it goes red with the filed message, At(127.0.0.1:57899) was not serving its log by the bound: never asked (EXIT=101). The test as e47d1f06 has it stays green (EXIT=0). Both builds exit 0, and the patch is applied checked and restored (mutate.sh, mutations.txt).

Loop at e47d1f06: 200 runs of the full --lib test binary beside a looping cargo test --workspace --exclude toyos-build (6 runs, all EXIT=0), at one-minute load averages from 4.70 to 53.49 (redial-r4/loop.sh, loop-results.txt, load-exits.txt).

test green
a_refused_first_dial_is_asked_again_up_to_its_ceiling 200/200
a_redial_ends_at_its_bound_alone_and_says_so 200/200
a_redial_on_a_forward_that_refuses_ends_at_the_refusal 200/200
the full suite 200/200 EXIT=0

At 1 red in 200, a loop of 200 has little power to show a race gone. The staged refusals and the late-reader probe are the evidence that it is gone. The loop only shows nothing new appeared.

Loop at 76cbf2d7: the same method, 200 runs of the full --lib test binary built from a clean tree at 76cbf2d7 (bin-provenance.txt), beside a looping cargo test --workspace --exclude toyos-build (5 runs, all EXIT=0), at one-minute load averages from 6.36 to 71.69 (redial-r5/loop.sh, loop-results.txt, load-exits.txt). Each run checked all 15 metaltalk::tests.

run green
all 15 metaltalk tests, every run 200/200
the full suite 197/200 EXIT=0

The three suite reds are off this PR: buildlock::tests::a_key_being_built_is_waited_for_and_another_key_is_not at src/buildlock.rs:945 in runs 92 and 135, and sysroot::tests::a_sweep_removes_what_no_worktree_names_and_nobody_uses at src/sysroot.rs:942 in run 134, where the sweep after drop(using) removed nothing (full-92.log, full-134.log, full-135.log).

Mutations of the product behaviour each test guards. Each is applied as a checked patch, built (EXIT=0), run with --exact and restored. The tree was compared with the fix after each one (mutations.txt).

mutation test exit red on
m1a-ceiling-off-by-one: turned_away > TURNED_AWAY_CEILING ceiling 101 left: 65, right: 64
m1b-first-dial-not-asked-again: a first dial gives up at its first refusal ceiling 101 left: 1, right: 64
m2b-named-redial-ends-at-a-refusal: a named redial ends at a refusal the way a forward's does bound 101 … a forward fails a connect only once QEMU has exited

The redial's-reason fix's controls, at 76cbf2d7. Each is a checked patch against the fix, built (EXIT=0), run with --exact on the bound test, and restored. The tree was compared with the fix after each (redial-r5/mutate.sh, mutations.txt), and the metaltalk tests were green after the last restore (EXIT=0).

control exit red on
nc1-unfixed-production: every production hunk of the fix reverted onto its base, the new test kept 101 At(127.0.0.1:61545) admitted no connection by the redial's bound: the latest ended before a line, the deleted format
nc2-m2a-on-unfixed: nc1 plus round 4's m2a, which holds open the window between serve's clock read and open's 101 At(127.0.0.1:61551) was not serving its log by the bound: never asked
nc3-close-not-carried: the fix, with serve seeding last with "never asked" instead of the close 101 At(127.0.0.1:61561) was not serving its log by the bound: never asked

On the unfixed code the window is two adjacent clock reads with no seam between them, so no host test reaches it without a mutation. nc2 is that window held open, and nc3 is the fix with the defect put back.

Gates at 76cbf2d7

Logs are in the job scratchpad, redial-r5/, from a clean tree.

gate exit
cargo test --workspace --exclude toyos-build 0 (gate-workspace.log)
cargo test -p toyos-build --lib 0; 392 passed, 2 ignored (gate-lib.log)
cargo run -- --clippy 0 (gate-clippy.log)

Negative controls

Host-level:

  • a_redial_on_a_forward_that_refuses_ends_at_the_refusal was added onto 74f7d717. That tree built (exit 0), and the test was red on "a refusing forward ends the redial at once" (exit 101, after 5.02 s; redial-r2/b1-red.sh, b1-red-run.log).
  • At 8cc19ef1, deleting only the forward arm of count_or_give_up (forward-arm.patch, applied checked and restored) built. The staged test went red on "a refusing forward ends the redial at once" (exit 101, after 5.02 s; mutation-forward-arm-staged.log).

Guest-level, run by the orchestrator at 7e06a657.

  • cargo test --test toyos-build -- swap_netd: EXIT=0 (566r2-swap_netd.log).
  • cargo test --test toyos-build -- lan_swap: EXIT=0 (566r2-lan_swap.log).
  • cargo test --test toyos-build -- swap_crash_rolls_back: EXIT=0 (566r2-swap_crash_rolls_back.log).
  • hold-green (hold.patch on the branch; cargo test --test toyos-build -- swap_netd): EXIT=0. The redial was turned away 8125 times, and logd admitted the stream again 7019 ms after the redial (566r2-hold-green.log).
  • hold-red (hold.patch plus fix.patch reverted, which is the whole code change reverted onto its base; cargo test --test toyos-build -- swap_netd): EXIT=1 on "the stream's redial was turned away 64 time(s), its ceiling of 64, and gave up" (566r2-hold-red.log).
  • cargo test --test toyos-build (the Fast tier, no filter): 399 passed, 399 total, EXIT=0 (566r2-fast.log).

fix.patch at that head is git diff af817e51 7e06a657 over src/metal.rs, src/metalswap.rs, src/metaltalk.rs and tests/common/logstream.rs. It leaves out the issue files, which are prose, and src/redlist.rs, whose rows would disable the test being run. Against a scratch index of 7e06a657, git apply --cached --check hold.patch exits 0, git apply --cached -R --check fix.patch exits 0, and so does fix.patch reverted on top of hold.patch. Both arms build at 7e06a657: with each applied, cargo run -- --build-only and cargo test --test toyos-build --no-run exit 0 (arms-build.sh, arms.status).

Guest-level, run by the orchestrator at 76cbf2d7:

  • cargo test --test toyos-build -- swap_netd: EXIT=0 (566r5-swap_netd.log).
  • cargo test --test toyos-build -- lan_swap: EXIT=0 (566r5-lan_swap.log).
  • cargo test --test toyos-build -- swap_crash_rolls_back: EXIT=0 (566r5-swap_crash_rolls_back.log).
  • cargo test --test toyos-build (the Fast tier, no filter): 398 passed, 398 total, EXIT=0 (566r5-fast.log).

Independent oracle: hold-green's guest console carries logd's own word that it admitted the redialled reader, in its second serving line: {7.721 tid=1 logd} logd: serving this boot's log to 10.0.2.2:59913. The first was {0.435 tid=1 logd} … 10.0.2.2:51775. Nothing this host counts produces that line.

Unsure

  • The bound test cannot show, without a clock seam, which arm a given run took. It asserts the reason that matches what happened, and a reader that never dialled is a green never asked with no dial counted. The loops did not record which arm each run took.

🤖 Generated with Claude Code

https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j

…ndow alone

The netd-swap tests went red with "the stream's redial was turned away 64
time(s), its ceiling of 64, and gave up" on boots whose guest finished every
swap: the re-claimed netd leased and init said "in service" or "restored". The
red was the host's own dial count, which is neither the event nor a timeout.

The event, "logd admits a dial", is seen only by dialing, and each dial already
waits on the machine's answer. So `Stream::redial` now dials until a connection
carries a line or its bound passes, and loses its ceiling parameter:

- the reader's ceiling is the first dial's only (`State::ceiling` is `None`
  once a redial begins), so `turned_since_dial` and `past_ceiling` go;
- a redial whose bound passes with every connection closed before a line now
  says so (`Stream::unopened`), where it used to go quiet;
- `metalswap::swap` bounds the redial by what is left of the swap's window,
  waits on the admitted connection, prints how long the refusal window was and
  how many dials it turned away, and returns `Err` naming both when the window
  passes with none admitted;
- the judge's ceiling finding is deleted; `turned_away` stays a reported number;
- `TURNED_AWAY_CEILING` now bounds first dials only, so it moves beside the
  reader in `metaltalk`.

The two ceiling tests become tests that a redial asks past 1000 refusals and
resets and past 200 closes before a line, and that it ends at its bound alone
and says so on both paths. The judge's test now holds 10 000 dials turned away
to be no verdict.

The diagnostics issue on the redial's spin keeps its exit; what bounds the spin
is now the window alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review r1 at 74f7d717: CI has not run at this head. Run 36388162347's host job concluded skipped, because ci.yml gates it on draft == false and the PR is still a draft. So cargo run -- --ci host and gate-stage have no exit code here, and the PR body's own cargo test --test toyos-build -- --list row is 101.

NOT READY FOR REVIEW

@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review r1 of #566 at 74f7d717

Gate. CI host concluded success at 74f7d717 (run 36410033729: the cargo run -- --ci host and gate-stage steps both succeeded; host runs cargo test --lib, src/ci.rs:437). Guest runs by the orchestrator: swap_netd, lan_swap and swap_crash_rolls_back EXIT=0; hold-green EXIT=0; hold-red EXIT=1.

Checked, no finding

  • The window is a legitimate hang ceiling, not a hidden timing verdict. The redial's by is what is left of the rig's CEILING (tests/common/swap.rs:33, "A liveness guard on a guest that stopped talking, never a verdict"), and on metal it is the operator's wait_secs. When it expires, the only verdict is "nothing admitted", which is a hang. The "admitted again N ms" line is printed and never judged. The deleted 64-dial count was the hidden clock: 64 × the forward's per-dial latency. The one exception is BLOCKER 1, where the bound replaces an event the harness already has.
  • The negative control is the whole change, reverted. redial-r1/fix.patch is git diff 16d1fbea 74f7d717 minus the issue-file hunk, which is prose with no behaviour. So hold-red is the whole code change reverted onto 16d1fbea. It is red on the ceiling line itself. Even without the hold, swap_netd at this head was turned away 85 times (more than 64), so the base reds on an ordinary boot too.
  • The oracle is independent. Hold-green's guest console carries {7.616 … logd} logd: serving this boot's log to 10.0.2.2:52608, the second such line. That is the guest's own word; no host count produces it.
  • The Fast-tier red is not this branch's. quiesce_stops_the_machine runs through tests/common/power.rs, which uses none of metaltalk, metalswap or logstream. Main already records the same red in issues/kernel/quiesce-stops-the-machine-stayed-up-beside-other-guests.md. In the same run, every carrier swap was admitted within about 6 s (34 to 55 dials), so there was no host spin.
  • Deleted code. turned_since_dial equals turned_away until the first redial, because no dial comes before the first dial. past_ceiling only formatted a message. The judge's ceiling finding is replaced by the time bound. One thing only the redial ceiling guarded is now lost: see BLOCKER 1.
  • Net lines. Overall +152 −115. Production +64 −58, net +6: metaltalk +45 −40, metalswap +18 −17, metal.rs +1 −1. Tests +83 −50, net +33. Issue prose +5 −7.
  • Merging origin/main has no textual conflict. git merge-tree --write-tree HEAD origin/main exits 0 at tree 78f4efa3. The semantic conflict is BLOCKER 2. Main also deletes issues/build/lan-swap-redial-spent-its-ceiling-on-a-nightly-shard.md and issues/build/swap-crash-rolls-back-reds-when-its-redial-spends-its-ceiling-under-load.md, which at this head still name metalswap::TURNED_AWAY_CEILING. Take main's deletions.

BLOCKER

  1. src/metaltalk.rs:427, :456-465 — A redial against a forward that refuses spins a host core with no wait until the window ends (120 s).
    • Why it spins: Peer::At's only user is QEMU's slirp hostfwd (tests/common/logstream.rs:112, tests/common/qemu.rs:4671). That listener accepts for as long as QEMU lives, so a refused or reset connect on it means QEMU has exited, for example on a guest reset under -no-reboot.
    • What changed: the deleted redial ceiling was the only thing that stopped this loop. The refusal is already the event, and it is final.
    • Why it matters: main's issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md records sibling guests going red under spinning host threads. A spin on a dead guest is defensive code that hides the failure instead of failing fast.
    • Fix: on Peer::At, a refused or reset connect ends the redial at once, with unopened saying so.
    • Test to add, red at 74f7d717:
      #[test]
      fn a_redial_on_a_forward_that_refuses_ends_at_the_refusal() {
          let dir = toyos_tmpdir::TempDir::new("metaltalk-forward-gone");
          let server = TcpListener::bind("127.0.0.1:0").unwrap();
          let (stream, first) = read_first(&server, Arc::new(Net), &dir);
          drop((server, first));
          stream.redial(Duration::from_secs(60));
          assert_eq!(stream.wait_for_connection(1, Duration::from_secs(5)), None);
          let why = stream.unopened().expect("a refusing forward ends the redial at once");
          assert!(why.contains("refused"), "{why}");
          assert_eq!(stream.turned_away(), 1);
      }
    • Knock-on: a_redial_asks_past_every_refusal_and_reset_until_a_dial_is_taken (:1541) stages 1000 refusals on Peer::At, which no Peer::At peer produces. It moves to Peer::Named through Reach::ask, where refusals are real (the T14).
  2. src/redlist.rs (after merging origin/main) — The three rows and main's issue file must come off in this PR. Merge origin/main, then:

NOTE

  • src/metalswap.rs:208-216 — Deleting the wait_for_connection Err block passes every test, and no guest run reached it. The claim "fails loudly with what it waited for" rests on no measurement. Delete the block, since the judge already reds with connections (1, 1) and the missing word, or show it red once.
  • src/metaltalk.rs:170, :174 — ceiling is a parameter with one production value. Both callers pass TURNED_AWAY_CEILING. Drop the parameter and make the const private. That also deletes the two repoints (src/metal.rs:1794, tests/common/logstream.rs:112).
  • src/metalswap.rs:217-221 — the printed turned-away count repeats the judge's said line (:329-334). Print the milliseconds only.
  • src/metaltalk.rs:1584 — close_then_replay(&server, usize::MAX) leaves a thread blocked in accept for the rest of the test process.
  • src/metaltalk.rs:428 — Uncapped T14 redial. On the T14 (Peer::Named), each refused redial dial asks the link again as soon as the old netd answers, with no cap for up to wait_secs. That is below ASK_WAIT's cited RFC 6762 §5.2 floor, and the T14 is unmeasured, as the diagnostics issue itself says.
  • PR body:
    • The --list row must carry the orchestrator's measurement: the command, EXIT=0 and the log.
    • The guest controls must carry measured results, not "owed" ones: hold-green EXIT=0, 97 turned away, and both oracle lines; hold-red EXIT=1 on "turned away 64 time(s), its ceiling of 64"; the three tests EXIT=0.

REMOVE

  • issues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.md:19-20 — the "What bounds it: the swap's window alone…" sentence was rewritten instead of deleted, and Stream::redial's doc already states the bound.
  • src/metaltalk.rs:424-426 — the clause ", so the first dial's ceiling or a redial's bound, never a wait, bounds a forward that refuses" was corrected instead of deleted, and it is false once BLOCKER 1 is fixed.
  • PR body:
    • The --list row's .git/modules/rust/config explanation is false at this head.
    • The "I did not turn on the reader's echo…" paragraph is not load-bearing for main's record.
    • "src/redlist.rs is untouched. The three tests' rows … come off once both have landed" is false once BLOCKER 2 is done.
    • The "## Unsure" section: its second bullet is false (the orchestrator built and ran the control), and its first restates the issue.
    • "(120 s in the rig)" restates a count.

SEND BACK

Japabu and others added 2 commits September 28, 2026 12:59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…tests come off the redlist

QEMU's slirp forward takes every connect for as long as QEMU lives, so a
refused or reset connect on a `Peer::At` redial means QEMU has exited.
With the dial ceiling gone, nothing else stopped that redial spinning a
host core until the swap's window passed. `count_or_give_up` now ends a
forward's redial at the first failed connect and says why in `unopened`.
`a_redial_on_a_forward_that_refuses_ends_at_the_refusal` drops the
listener and requires the end within 5 s of a 60 s redial with one dial
turned away. It is red at 74f7d71 ("a refusing forward ends the redial
at once", exit 101) and red again with only the forward arm deleted.

The refusal and reset tests move to `Peer::Named` through a `Reach` that
answers the name at once, since that is where refusals are real (the
T14).

`TURNED_AWAY_CEILING` bounds first dials only, and every caller passed
it, so the `ceiling` parameter and `State::ceiling` are gone. The const
is private, and `open` is told whether it is redialling.

`metalswap::swap` no longer returns an `Err` when the redial admits
nothing. No guest run reached it, and the judge already reds on
`connections (1, 1)` and the missing word. It prints only the
milliseconds to the admission, because the judge's `said` line already
carries the count. The bound test's closer thread now ends instead of
staying blocked in `accept`.

`lan_swap`, `swap_netd` and `swap_crash_rolls_back` come off the
redlist, and
`issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md`
is deleted. Its exit condition was "`Stream::redial` gives up on its
time bound alone, never on a dial count, and a swap whose refusal window
is staged long is green". The orchestrator's hold-green run at 74f7d71
met it: init held the old netd for 1 s, the swap was turned away 97
times, it was green, and the guest console carried logd's second
`serving this boot's log to 10.0.2.2:52608` line. The same hold with the
change reverted was red on "turned away 64 time(s), its ceiling of 64".

The diagnostics issue records that the T14's redial may ask the link
faster than RFC 6762 §5.2's floor between queries. That rate is
unmeasured on metal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Japabu added a commit that referenced this pull request Sep 28, 2026
…ows, and file the harness's misreport

quiesce_stops_the_machine was disabled behind a finding whose title and exit
rested on the harness's message that the guest asked for a reboot and stayed
up. PR #566's capture at 74f7d71 refutes that: writer 5's first pass ran
from 2.170 s to 7.434 s, the job printed "5 of 6 writers reached their loop
in 5s" at 6.874 s and exited 1, and it never printed "asking for the reset".
No stop began. PR #524's capture at 235c5a5 shows the same with "3 of 6",
and nightly run 36351950439 on PR #555 at d265676 with "4 of 6".

- The slow pass is the defect quiesce_dump_holds_the_stopped's issue already
  tracks, whose exit names a first write-and-fsync pass over 5 s. That issue
  is renamed to what both tests show, and gains these sightings. Both rows
  point at it.
- The stops finding is folded into it and deleted. It carried no durable line
  for a module header; its three sightings move with their evidence. Its
  a58abf5 sighting also had no stop: record, and whether that job printed
  its give-up line was not recorded.
- stopped_boot waits its whole QMP budget and then calls
  returned_to_firmware before it reads the console, so a job that never asked
  is reported as a guest that asked. In the #566 capture every scheduler
  heartbeat from 10.750 s to 253.244 s was idle, and the test went red after
  266 s. Filed as tooling, held by the orchestrator.
- The park issue is renamed: its records show 0 block operations open, so
  both threads were running, and one was the held thread, which
  last::hold keeps spinning while a sweep counts 2. It now names each
  sighting's PR and head, adds the 98e803c stop that gave up, labels the
  dispose_yield suspect as a hypothesis, and records that
  woken_by_its_threads has no enabled caller. Its exit asks for an
  instrument that names each thread still running, and for that coverage
  back.
- Deleted: the scratch log names and paths, "after a stop that reported
  every thread stopped", "on a loaded host", and the build/ park issue's
  "--known-red answers NO".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review r2 of #566 at 7e06a657

Gate. CI host at 7e06a657 (run 36414146625): success. The --ci host and gate-stage steps both succeeded. The new tests are green in redial-r2/gate-lib.log. The orchestrator's guest runs at this head:

  • swap_netd, lan_swap, swap_crash_rolls_back: EXIT=0 each, with the rows off.
  • hold-green: EXIT=0.
  • hold-red: EXIT=1 on the stream's redial was turned away 64 time(s), its ceiling of 64, and gave up (566r2-hold-red.log:36).
  • Fast: 399/399.

redial-r2/fix.patch is byte-identical to git diff af817e51 7e06a657 over the four code files, so hold-red is the whole code change reverted onto its merge base.

Earlier BLOCKERs

  1. CLOSED. A redial against a forward that refuses no longer spins.
    • src/metaltalk.rs:383 ends a redial at a forward's first failed connect.
    • At 74f7d717, a_redial_on_a_forward_that_refuses_ends_at_the_refusal was red: exit 101, on "a refusing forward ends the redial at once", after 5.02 s (b1-red.out, b1-red-run.log).
    • At head, with only the forward arm deleted, it is red on the same line: exit 101 (m1.out, m1-run.log).
    • At head it is green (CI, gate-lib.log:340).
    • Neighbouring mutations are also caught. Dropping if again reds a_refused_first_dial_is_asked_again_up_to_its_ceiling (it expects 64). Dropping !again && reds the 1000-refusal test, now on Peer::Named.
  2. CLOSED. The three rows and the issue came off with the merge.
    • The three rows are gone from src/redlist.rs, and the issue file is deleted.
    • git merge-tree --write-tree HEAD origin/main exits 0 at tree 35ac6ad3, with main at fb34346e.
    • In that tree, a git grep for the slug, for metalswap::TURNED_AWAY_CEILING and for both issue files main deleted finds nothing. The quoted row names do not appear in src/redlist.rs.
    • The three tests are EXIT=0 at head (566r2-{swap_netd,lan_swap,swap_crash_rolls_back}.log).

Earlier NOTEs and REMOVEs

  • CLOSED:
    • the Err block;
    • the ceiling parameter, now a private const at metaltalk.rs:91;
    • the repeated count, now milliseconds only (metalswap.rs:209);
    • the leaked closer thread (close_every joins its thread, metaltalk.rs:1529);
    • the --list row;
    • every r1 REMOVE.
  • OPEN, as filed: the T14 rate (below).
  • PARTLY: the guest controls in the body are measured at 74f7d717. The head's runs are listed as "owed" (below).

The 8125 dials in hold-green

This is not a BLOCKER.

  • Not a CPU spin. Each dial blocks in connect and read on the guest's answer, so the loop runs at the guest's refusal rate.

  • No other lawful option. A pace between dials would be a timer, not the event, which is the flat wait the rules forbid. The event itself is the filed exit of issues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.md.

  • Its size. Unstaged swaps at this head turned away 15, 21 and 19 dials (the three tests), and 6, 32 and 18 in Fast.

  • Its cost, which the record does not carry. The guest pays for the storm, not the host:

    • hold-green: 16231 userdev interrupts, 85.3% of the guest's 19034, all on cpu0;
    • unstaged swap_netd: 181.

    So under a long gap the harness is most of the load on the machine it is judging. Main's deleted TURNED_AWAY_CEILING doc called this loop "spinning". That is the NOTE below.

BLOCKER

None open.

NOTE

  • issues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.md:19 — the record carries no evidence of the forward's cost.
    • This branch removed the compromise's only stop. The issue still has no measurement for QEMU and no owner.
    • Add the hold-green line: 8125 dials turned away in 7019 ms, and 16231 of the guest's 19034 interrupts. A compromise is recorded with owner, evidence and exit.
  • src/metaltalk.rs:419, :452 — the T14 still re-asks mDNS after every refusal (r1 NOTE, still open).
    • Without a dial ceiling, a gap where the old or new netd sends RSTs sends multicast queries at LAN round-trip rate for up to wait_secs. That is below the RFC 6762 §5.2 floor the code cites at ASK_WAIT.
    • A timer-free fix: set on_the_link only on a host-absent failure (EHOSTDOWN, EHOSTUNREACH, ENETUNREACH, or a failure after WAITED). A refusal or reset is the machine answering at that address, so it is redialled without a new ask.
  • src/metalswap.rs:214 with src/metaltalk.rs:260 — after a redial has ended, wait_until for init's final word still waits out the window.
    • wait_until does not wake when the stream can no longer carry a line (unopened set, no current, no redial). So a forward that failed, now named at once, still leaves the swap waiting up to 120 s.
    • hold-red took 124 s this way (566r2-hold-red.log:507). This predates the branch.
    • Fix: give wait_until the dialing exit that wait_for_connection already has (metaltalk.rs wait_for_connection).
  • A guest gap for the orchestrator to file.
    • In hold-red at this head, logd re-bound 41337 at 1.758 s. Init said in service at 6.739 s (566r2-hold-red.log:330, :332).
    • The guest (identical in hold-green, since fix.patch is host-only) admitted the redial only at 7.721 s (hold-green oracle). The T14 echo answered 7336 ms after the ask.
    • So connects are turned away for about 6 s after logd listens again. That, not logd's closed listener, is what makes the redial's window long.
  • PR body: the measurements at the head that lands are missing.
    • The guest-level section is at 74f7d717 (97 dials, 6866 ms), which understates the cost.
    • Replace it with the orchestrator's 7e06a657 runs: the three tests EXIT=0; hold-green EXIT=0 with 8125 turned away, 7019 ms, and oracle lines {0.435 …:51775} and {7.721 …:59913}; hold-red EXIT=1 on the ceiling line; Fast 399/399. Each with its command, exit code and log.

REMOVE

  • PR body: "## Owed at the merged head, run by the orchestrator". False now: it has been run.
  • PR body: "The closer thread is ended when the test finishes." Narrates a review fix.
  • PR body, worktree-hygiene narration that is not main's record:
    • "The patch was removed and the tree left clean";
    • "The patch was reverted, and the file compared byte for byte";
    • "Each was then reverted and git status --porcelain came back empty".

Net lines

git diff --shortstat origin/main...HEAD: 7 files, +207 −217.

  • Production (metaltalk outside its tests, metalswap outside its tests, metal.rs, tests/common/logstream.rs): +63 −78, net −15.
  • src/redlist.rs: −12.
  • Tests (the metaltalk and metalswap test modules): +137 −61, net +76. That is the forward test, close_every, a working Reach for TurnedAway, and named, each requested in r1.
  • Issues: +7 −66.

LAND AFTER NAMED CHANGES

Japabu and others added 8 commits September 28, 2026 15:48
…d an owner, and three NOTEs become their own issues

The redial's forward cost was a NOTE with no evidence: add hold-green's
8125 dials turned away in 7019 ms, 16231 of the guest's 19034 interrupts on
cpu0, and hold the diagnostics issue for the orchestrator alongside it.

Three more NOTEs from the round-2 review are filed rather than fixed here:
the T14's redial still re-asking mDNS after every refusal below RFC 6762
§5.2's floor, `wait_until` for init's final word never waking once the
stream is dead, and a ~6 s guest-side gap after `logd` re-binds that a swap's
redial actually pays for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…und's brief

Running the round's own cargo test --lib gate under this machine's current
concurrent load reproduced a red on a-redial_on_a_forward_that_refuses_ends_at_the_refusal
twice in six runs, always green alone or on a quiet host. Filed rather than
fixed: it is not named in this round's ruling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…edials

a_redial_on_a_forward_that_refuses_ends_at_the_refusal went red under host
load with turned_away() == 2. The landing round blamed the first dial
retrying a transient refusal from a live listener. That cannot happen here:
the listener is bound before the dial and never holds more than one pending
connection, so a loopback SYN to it is taken, never refused.

The real race was in the test helper `accepted`. Its thread accepted on a
try_clone of the listener and dropped that clone only after sending the
connection back. So the test's drop(server) did not close the listener when
the helper thread had not yet run past its send. The redial's connect was
then taken into the still-live accept queue, reset once the clone closed
(read counts it: a connection ended before a line), dialled again, and
refused (count_or_give_up counts it and ends the redial). That makes two
counts. The product code counted that sequence correctly; the stimulus was
wrong.

Holding the window open with a 300 ms sleep after the send reproduced it
every time (left: 2, right: 1). With the fix, the same sleep stays green.

The helper now drops its clone before it sends, so the listener is closed
before `accepted` returns. The test also asserts that the first dial counts
nothing before the redial, so the first dial's count and the redial's are
measured apart.

Deletes issues/build/a-forwards-redial-negative-control-counts-a-live-listeners-transient-refusal.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…istener

This corrects 33cfc34's root cause, which was incomplete. Closing
`accepted`'s clone before its send closed one window (a 300 ms sleep after
the send reproduced the red every time). But in the full `--lib` suite
under load, that binary still went red on this test 9 times in 41 runs.
The pre-fix binary went red 9 times in the same 41, interleaved with it.

The listener has a second holder: any child process this test process is
spawning. A spawned child holds a copy of every fd until its exec closes
the close-on-exec ones, and many lib tests spawn processes (git, cargo, the
test binary itself). A scratch probe measured it: bind 127.0.0.1:0, drop,
connect, 20000 times.
- No concurrent spawn: 0 taken, 0 reset, 20000 refused (twice).
- A thread spawning children beside it: 34 taken and 4 reset, then 26 taken
  and 7 reset.

So no test can make a dropped loopback listener provably gone. The redial's
connect was taken by the still-open listener and reset once the child's
exec closed it, which `read` counts. It was then dialled again and refused,
which `count_or_give_up` counts. That makes two counts. The product code
counts that sequence correctly and is unchanged.

The test now stages the forward through `TurnedAway`, as the other redial
tests stage their refusals. The first dial is real and taken, and every
dial after it is refused or reset by the test. It asserts that the first
dial counts nothing and that the redial counts its one refused connect. The
`accepted` change is reverted, since no test depends on a dropped listener
closing any more.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
… forward redial's fix

The proof loop for 8cc19ef ran 200 full `--lib` suites of its test binary
and of 941f0bd's, interleaved, beside `cargo test --workspace --exclude
toyos-build` at one-minute load averages up to 43. Three other tests went
red. None is on this round's brief, so all three are filed and not fixed.

- a_refused_first_dial_is_asked_again_up_to_its_ceiling: once for 8cc19ef
  and twice for 941f0bd. It dials a dropped listener, which is the measured
  spawn-held-fd mechanism.
- a_redial_ends_at_its_bound_alone_and_says_so: once for 941f0bd. Its reader
  woke past a 50 ms bound and never dialled.
- buildlock's a_key_being_built_is_waited_for_and_another_key_is_not: once
  for 8cc19ef. A dropped lock was still held; the cause is unmeasured.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…e scheduler

Two of this PR's lib tests reddened under load, each 1 in 200 full `--lib`
suite runs beside `cargo test --workspace --exclude toyos-build`.

`a_refused_first_dial_is_asked_again_up_to_its_ceiling` dialled a dropped
listener. A child that another test in the same process is spawning holds a
copy of that fd until its exec, and while it does the listener takes
connects (8cc19ef measured this). Its refusals are now staged through
`Reach` by `Refusing`, as the forward test's are. Every dial is refused, and
the stage refuses a resolve or an ask loudly, since an address is only
dialled.

`a_redial_ends_at_its_bound_alone_and_says_so` assumed its reader dials
inside a 50 ms bound. The bound is computed in `redial` on the test's thread
and read on the reader's, so no test can make that dial happen without a
clock seam, and adding one is a production change. The test now accepts
either history and checks the reason names what happened. A redial that
never dialled ends "by the bound: never asked". A redial that dialled ends
on its one close, which `close_every` makes only once the bound has passed
since it took the connection, so `serve`'s check ends the redial and names
the close. More than one counted dial fails the test. Injecting a 100 ms
delay before a redial's first dial reproduces the filed red on the old test
(EXIT=101, "never asked") and passes the new one (EXIT=0).

Along the way I found a window this test no longer reaches: `serve` continues
inside the bound, `open` then re-reads the clock, and a redial that dialled
can end saying "never asked". It is filed as
issues/diagnostics/a-redial-that-dialled-can-end-saying-never-asked.md.

Both issue files are deleted. The keyed-lock issue's citation of the first
one goes with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 4, at efa141b

Gate. CI host COMPLETED SUCCESS at efa141b0. The orchestrator's guest runs at efa141b0: swap_netd, lan_swap and swap_crash_rolls_back EXIT=0 each, and Fast EXIT=0. Between f8283ca0 and e47d1f06 the branch changed only src/metaltalk.rs's test module (hunks from :1309 on) and issues/. efa141b0 is a main merge that touches none of the four code files.

Earlier BLOCKERs. r2 left none open.

Checked, no finding

  • "Accepts either history" can still fail.
    • m2a (serve reads until + FOREVER after a close) is red at :1637 with turned_away 1: redial-r4/m2a-close-ignores-the-bound-run.log, EXIT=101.
    • Arm 1 is deterministic. redial() sets until before the reader's dial, and close_every holds each taken connection BOUND past its accept, so the close always lands past until. That leaves exactly one count and serve's message.
    • A mutation that ends a redial at its first close whatever the bound passes this test but is red in a_redial_asks_past_every_close_before_a_line_until_one_carries_a_line. A redial that never dials passes arm 0 but is red in that test and in a_redial_asks_past_a_refusal_and_keeps_each_line_once.
    • m1a/m1b are red on the ceiling test (left: 65 / left: 1), and m2b is red on the refused arm (:1626).
  • Neither test depends on a dropped listener closing. The ceiling test dials through Refusing and opens no socket at all. In the bound test's closes arm, the listener stays alive until the test ends. What arm 1 does depend on is a dropped accepted socket's close reaching the reader (see NOTE 1).
  • The 200/200 is recorded.
    • The PR body's "Loop at e47d1f06" points at redial-r4/. loop-results.txt has 200 rows, and all 200 have full_exit=0 with the three tests ok.
    • load-exits.txt has 6 load runs, each EXIT=0.
    • bin-provenance.txt records that the binary was built at e47d1f06 from a clean tree.
    • The loop did not record which arm the bound test took (the body's "Unsure"). So under load it shows only that the test stayed green, not that arm 1 was exercised.
  • The filed defect is true of the code. serve :348 continues while now < until. open then resets last to "never asked" at :407, fails its while at :411, and returns :461 … by the bound: never asked, with turned_away ≥ 1. m2a's own red message is exactly that string at turned_away 1.
  • Net lines (git diff --shortstat origin/main...HEAD): +386 −222 across 12 files.
    • Production +62 −78.
    • Tests +181 −65.
    • Issues +142 −67.
    • src/redlist.rs −12.
    • Rounds 3 and 4 added no production lines.

BLOCKER

None.

NOTE

  1. src/metaltalk.rs:1563 — arm 1 of the bound test depends on drop(taken) reaching the reader as a close.
    • On macOS, std's accept sets CLOEXEC only after the accept returns (rust/library/std/src/sys/net/connection/socket/unix.rs:267-269). A spawn in that window hands a child a copy that is not close-on-exec, and the reader's EOF then waits for that child to exit. If that takes more than 5 s, the test goes red on "the redial ended within 5 s and said why".
    • On Linux a child's copy only delays the close, which is still past the bound.
    • Fix: call shutdown(Shutdown::Both) on the taken stream before dropping it. That sends the FIN whatever holds a copy.
    • The same applies to close_then_replay (:1536) and to drop(accepted(..)) (:1465).
  2. issues/diagnostics/a-redial-that-dialled-can-end-saying-never-asked.md — the branch rewrote the block that opens this window (serve :345-356), so closing it here deletes code rather than adding an issue.
    • Pass the latest close's reason from serve into open as its starting last, in place of "never asked".
    • serve's second reason format (:351-353) then goes, and one path names every redial's end.
  3. issues/build/a-dropped-keyed-lock-is-still-held-in-the-loaded-lib-suite.md — a flaky test in the required host check is filed and left enabled. src/redlist.rs cannot disable a host test, so "a flaky test is disabled at once" has nothing to disable it with. The orchestrator decides; it is off this PR.

REMOVE

  • issues/diagnostics/a-redial-that-dialled-can-end-saying-never-asked.md:15 — "The metal loop records that reason as the redial's end." — false.
    • After a redial, no production code reads unopened(): metalswap.rs:206-210 reads only wait_for_connection.
    • Its only readers are converse (metaltalk.rs:935) and logstream::reader (tests/common/logstream.rs:115), both on the first dial.
  • PR body: "The keyed-lock issue, which this PR files, cited the deleted ceiling issue by path. The citation is deleted, and the issue is otherwise untouched." — it narrates an in-PR edit to a file main never had.
  • PR body: "The orchestrator ruled that a PR does not land flaky tests of its own into the required host check." — review narration, not main's record.

LAND AFTER NAMED CHANGES

Japabu and others added 3 commits September 28, 2026 18:14
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…sts' closes reach the reader whatever a child holds

`serve` read the clock after a connection closed before a line, and
`open` read it again before its first dial, starting from "never asked".
A bound passing between the two reads ended a redial that had dialled,
and counted the close, with "... by the bound: never asked". `serve` no
longer reads the bound: it hands the close's reason to `open` as its
`last`, so `open`'s one bound check names every redial's end, and
`serve`'s second reason format is gone. `read` returns how a connection
it counted as turned away ended, in place of a bool.

`a_redial_ends_at_its_bound_alone_and_says_so` now requires a redial
that dialled once to end with "by the bound: the latest connection ended
before a line". Against the unfixed production code it is red on the
deleted format; with round 4's m2a on that code, and with the fix's
reason not carried, it is red on "by the bound: never asked".

The tests' three closes of an accepted connection shut it down before
dropping it: on macOS std sets close-on-exec only after `accept`
returns, and a child spawned in that window holds a copy that keeps the
reader from seeing EOF. `close_every`'s closer ends on the wake-up
connection rather than closing it.

Deletes issues/diagnostics/a-redial-that-dialled-can-end-saying-never-asked.md,
the defect being fixed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
src/redlist.rs conflict: main's #564 deleted so_cache_refusals and
usb_disk_index_stable rows (their tests are gone); this branch's redial
fix deleted swap_crash_rolls_back, swap_netd and lan_swap rows (their
tests now pass). Both sets of deletions kept; no row this branch changed
was reverted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 5, at e0edce3

Gate. CI host COMPLETED SUCCESS at e0edce3d. The orchestrator's guest runs at 76cbf2d7: swap_netd, lan_swap and swap_crash_rolls_back each test result: ok, with the redial turned away 16, 55 and 32 times before logd admitted it. Those runs go through the rewritten close-before-a-line continue. Fast: 398 passed, 398 total. git merge-tree --write-tree HEAD origin/main at bde7b569 (#572) EXIT=0, and the merged tree differs from main by this branch's diff alone.

Earlier BLOCKERs. None were open after round 4.

  • r4 NOTE 1 (a drop is no close while a child holds a copy): CLOSED. close() shuts the socket down at all three sites.
  • r4 NOTE 2 ("never asked"): CLOSED, by deletion, measured by the controls below.
  • r4 REMOVEs: CLOSED. The issue file is deleted, and both PR-body lines are gone.

Checked, no finding

  • Every path that sets or reads last names the redial's latest dial.
    • last is moved into open. Each way back to open reassigns it first: the continue at :349 sets the close's reason, and the redial take at :362 sets "never asked". The compiler rejects any path that skips the reassignment.
    • open overwrites last with each later refusal, reset, silent ask or failed resolve. So its bound error at :464, its stopped error at :416 and count_or_give_up's forward message all name the latest event.
    • A redial or stop that lands while a turned-away connection is being read fails the guard at :347. It then falls to the wait loop, which resets last or returns. No old close is carried into a new redial.
    • read returns None only when a line was carried. There unopened was already cleared on admission, so no reason is owed.
    • No reader outside the tests parses the deleted format (git grep).
  • The controls revert what they claim.
    • fix.patch is byte-identical to git diff cd82f651 76cbf2d7 -- src/metaltalk.rs (cmp 0).
    • nc1 reverts every production hunk in serve, open and read, and keeps the test. It is red at EXIT=101.
    • nc2 is the defect's window held open on the unfixed code. It is red on never asked, which can only be arm 1, since arm 0 accepts that suffix.
    • nc3 is the fix without the carried reason, red the same way.
    • After the controls the tree was restored to the fix, and all 15 tests were green (EXIT=0).
  • The 200-run loop.
    • The binary was built from a clean tree at 76cbf2d7. loop-results.txt has 200 rows, and 197 have full_exit=0 with no metaltalk test missing its ok.
    • The 3 reds are buildlock at :945 in runs 92 and 135, and sysroot at :942 in run 134. None is in metaltalk.
    • The check grep -x makes errs strict: an interleaved line counts as not ok.
    • The 5 load runs were each EXIT=0.
  • The e0edce3d resolution. The remerge-diff drops exactly the five conflicted rows. The resulting src/redlist.rs differs from 807f4561 only by this branch's three rows.
  • Net lines: +387 −233 across 11 files.
    • src/metaltalk.rs production code: +26 −22 in the landing round. That includes the 7-line reflow of open's signature and the deleted second format.
    • Tests: +17 −7.

BLOCKER

None.

NOTE

  1. src/metaltalk.rs:630 — the string read's early return gives is never read.
    • That path holds stop || redial.is_some(), and nothing between it and serve's guard at :347 clears either.
    • So the guard fails, and the wait loop resets last or returns.
    • Any other string, or None, passes every test by construction.
  2. sysroot::tests::a_sweep_removes_what_no_worktree_names_and_nobody_uses — its red in run 134 (src/sysroot.rs:942, full-134.log) is filed nowhere, in this branch or in main. It is off this PR; filing it is the orchestrator's.
  3. PR body, hold-green and hold-red — the PR's oracle and its whole-change control were measured at 7e06a657.
    • 76cbf2d7 rewrote the close-before-a-line path those runs exercised, and they were not re-run.
    • The r5 guest runs cover that path green, but not as the held-window pair.
    • The metal harness is not high-risk, so this is a NOTE.

REMOVE

  • issues/build/a-dropped-keyed-lock-is-still-held-in-the-loaded-lib-suite.md — a second record of the defect main already files as issues/build/a-key-being-built-is-waited-for-and-another-key-is-not-reds-under-host-load.md (same test, same assertion at buildlock.rs:945), and it names no owner. The PR body's "Unsure" bullet citing it goes with it.
  • PR body: the m2a row and "These four were run at e47d1f06. m2a's site is gone at 76cbf2d7, since serve no longer reads the bound." — a mutation of code that main never receives.
  • PR body: "and each is a keyed lock read as held right after its guard dropped" — unmeasured for the sysroot red.
  • PR body: the ec0a91ad/cd82f651 merge paragraph — it narrates an in-PR merge and a file that main never had.
  • PR body: the 807f4561/e0edce3d merge paragraph — resolution narration, which the merge commit's message already carries.
  • PR body: "76cbf2d7 changes production code again, in metaltalk's serve, open and read, so these runs predate it." and ", after the redial's-reason fix that moved metaltalk's serve, open and read again" — in-PR chronology.

LAND AFTER NAMED CHANGES

Japabu and others added 2 commits September 28, 2026 19:40
… the second buildlock red's duplicate issue file is deleted.

serve's guard at :347 always fails on the path that leads to read's early
exit, so its string was never read; it now returns None, which serve treats
the same as a line having been carried.

issues/build/a-dropped-keyed-lock-is-still-held-in-the-loaded-lib-suite.md
duplicated the record main already keeps under
a-key-being-built-is-waited-for-and-another-key-is-not-reds-under-host-load.md
for the same test and assertion, and named no owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Conflict in src/redlist.rs: main still disables swap_crash_rolls_back and
swap_netd behind issues/build/a-swaps-redial-races-a-hard-dial-ceiling-
against-an-unbounded-guest-gap.md, the issue this branch's fix resolves and
already deleted at 7e06a65; those two rows are dropped, keeping main's
unrelated new syscall_window_nmi row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu
Japabu added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit bc9ccad Sep 28, 2026
1 check passed
@Japabu
Japabu deleted the wt/toyos-redial branch September 28, 2026 18:23
Japabu added a commit that referenced this pull request Sep 28, 2026
Brings in #572 (host QEMU's edk2), #580 and #579 (disabled reds), #549
(a kill never waits on its victim) and #566 (metaltalk redial).

- src/redlist.rs: one row each for user_copy_races_munmap and
  quiesce_leaves_the_volume_whole, which both sides added. main's new rows
  stay (netd_refused_accept, quiesce_wakes_on_the_last_teardown,
  root_chunk_refused_on_a_usb_stick, syscall_window_nmi). The rows for
  tests or issues this branch deleted go (hda_tone, doom_sound_flood,
  latency_wake, sched_check_build), and so does lan_swap, whose issue
  main deleted with swap_netd's and swap_crash_rolls_back's rows.
- The two issue files both sides added take main's text.
- tests/common/power.rs: main's woken_by_the_held_thread, shared by the
  new quiesce_wakes_on_the_last_teardown, without the two clock verdicts
  this branch took off QEMU (stopped_the_machine in stopped_boot, and
  woken_by_its_threads).
- tests/common/qemu.rs: qemu_command takes main's firmware_vars and has
  no audio_wav, so profile_argv passes six paths. The
  too_many_arguments allow goes, because seven parameters do not
  trigger it.
- kill_while_blocked.rs: main's text. After #549 a kill does not park
  in retire_task, so this branch's doc for arm 4 was false. main's arm
  also has no clock.
- tests/toyos.rs check_rust_result: this branch's single-print form,
  which already carries the stdout main added.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Japabu added a commit that referenced this pull request Sep 29, 2026
Main's rust pin has not moved since the last merge, so the fork is unchanged.
Every conflict, and how it was resolved:

Modify/delete, main deleted:
- issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md:
  #566 fixed the defect and deleted the issue. This branch had added one sighting to
  it, and a sighting of a fixed defect has no home, so the file stays deleted.
- src/heartbeat.rs: #562 deleted `kernel_heartbeat`'s CPU-mask and gap verdicts with
  the file. This branch had given its done-line table blockd and fsd rows. The table
  goes with the verdict it served.
- tests/doomcase/system.toml: #562 moved the doom audio tests to metal and deleted
  their QEMU config. This branch had added blockd and fsd rows to it. Nothing boots
  it now.

Modify/delete, this branch deleted:
- tests/toyos-rust-tests/src/bin/ftruncate_flush_race.rs,
  tests/toyos-rust-tests/src/bin/quiesce_fsync.rs,
  issues/build/ftruncate-flush-race-reds-intermittently-and-nothing-says-why.md,
  issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md
  and issues/kernel/a-root-metadata-read-refused-on-budget-is-not-retried.md: main's
  hunks remove timing from them or note its own runs. They are about the kernel FAT
  flush, the stop's kernel sync and the kernel's metadata read, which this branch
  deletes, so they stay deleted.

Content:
- kernel/src/actuator.rs: main's `quiesce_last_teardown` (#549) is kept. The kernel
  FAT actuators `fat_flush_meta_refuse`, `resize_evict_window` and
  `resize_fault_refuse` stay deleted. `process_reopen_selftest` stays where this
  branch has it, with main's doc (#549 also opens every kernel thread's pid).
- src/redlist.rs: both conflicted rows go. `doom_sound_flood` left QEMU with #562,
  and this branch deletes `ftruncate_flush_race`.
- tests/common/gpt.rs: this branch's `device_saying` and decoy `boot` are kept. Main
  drops the `drain_serial` window, so its `qemu` binding is no longer `mut`.
- tests/common/inspect.rs: main's "nothing plays audio" (#562 deleted
  `inspect_plays`) is taken, with this branch's clause on the boot stick.
- tests/common/iommu.rs: main's `panic-reboot-fast` and its wait for the fatal path's
  reset are kept. This branch's `iommu_empty_domain` reads the xHCI's DCBAAP over
  QMP, and QEMU has exited by the time that reset is seen. So `fault_boot` now takes
  a `holding` read, which it runs after the fault line and before it waits for the
  reset, while the fatal path holds its panel. `iommu_context_absent` reads nothing
  there.
- tests/common/origin.rs: main's judgement of `log_ring_keeps_the_owners_slots` is
  taken whole: init says it waited a flush out, or its stop line is missing. That
  drops the millisecond inference between two records, whose record this branch had
  changed from `Syncing filesystems...` to the stop record (#562: no QEMU test
  measures time).
- tests/common/volumes.rs: main's timing edit to `ftruncate_flush_race` goes with
  the test.
- tests/logstallcase/system.toml: main drops `power` and the `shutdown` symlink, since
  the metal row reads `/log` without a stop. This branch's blockd and fsd rows are
  kept, because fsd holds `/log`.
- tests/toyos-rust-tests/src/bin/blockd_io.rs: main's `claim_when_free`, now generic
  and with no deadline, is taken inside this branch's `if let Some(syscap)`. `bench`
  is this branch's blockd-only arm with main's timing removed: no MiB/s, and the
  line says only how many Flushes each run took. The module doc's "timed" goes.
- tests/toyos-rust-tests/src/roster.rs (add/add): both sides wrote one roster
  decoder. Main's is taken whole, because five binaries read it and it has no
  deadline (#562). This branch's copy had a 5 s give-up.
- tests/toyos-rust-tests/src/bin/process_lifecycle.rs: main's is taken whole. This
  branch's only change to it was the move onto its own roster.rs.
- tests/toyos-rust-tests/src/bin/process_stats.rs: main's `refused_calls_are_counted`
  and its roster wait for the held child are kept, and so are this branch's two
  connection arms. The system capability is taken once in `main` and passed to the
  three arms that read the roster, since a second take of the label finds nothing.
  The connection arms now wait on main's `threads_of` for the child's main thread
  to be blocked, with no deadline.
- tests/toyos-rust-tests/src/bin/quiesce_twice.rs: main's `Duration`-only import.
  This branch deletes the owed file, so `File` and `Write` go.
- tests/toyos.rs:
  - RUST_SKIP: main's audio rows are taken. `audio_tone_load` goes, since main
    deleted it. `log_volume_reread` goes, since this branch deletes it.
  - MACHINE_TESTS: `quiesce_leaves_the_volume_whole` stays deleted.
    `quiesce_wakes_on_the_last_teardown` comes from main with main's comment.
    `blockd_serves_nothing` is kept. `hda_tone` and `hda_client_stall` went to metal
    with #562, and `hda_two_live_refused` takes main's comment.
  - CARRIES and dispatch: the same.
  - `nvme_wide_sector`: this branch's blockd arm, which already had no drain window.
- toyos-quiesce/src/lib.rs: this branch's `FILES_MS`, `FLUSH_MS` and `SYNC_MS` are
  kept, with main's `LAST_THREAD` doc, which names both quiesce-last actuators.
- userland/logd/src/policy.rs: this branch deletes the module doc and the
  `LOG_WRITE_BUDGET` paragraphs main edited one line of, so they stay deleted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant