diff --git a/issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md b/issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md new file mode 100644 index 00000000000..0df0869b2aa --- /dev/null +++ b/issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md @@ -0,0 +1,59 @@ +--- +status: expected-red +kind: defect +opened: 2026-09-28 +--- + +# A swap's redial races a hard dial ceiling against an unbounded guest gap + +One mechanism, six sightings across `lan_swap`, `swap_netd` and +`swap_crash_rolls_back`: + +- `lan_swap`, Fast tier. +- `swap_netd` and `swap_crash_rolls_back`, Fast tier. +- `lan_swap`, PR #535's nightly (run 36314576406, `guest (1)`, `a4f68c5a`, KVM, QEMU 11.1.0). +- `swap_crash_rolls_back`, on origin/main's netd and on a branch's, under ten spinning host threads beside the run. +- `swap_crash_rolls_back`, main's nightly (`1ce71831`, run 36290616312), not seen on the nightly before #527 (run 36285169430). +- `swap_netd`, the rust-lld branch's fast tier (`e4317d3f`, PR #532), the one red of 400 besides `lan_mdns_answer`'s `SUN_LEN`. + +The guest side finished on every sighting that shows a console: today's three +boots reached DHCP lease, `logd` back on port 41337, and init's own +`restored`/`in service` line; `lan_swap`'s nightly guest reached `logd: serving +this boot's log on port 41337` at 1.176 s and `init: swap netd: in service` at +6.141 s, with no second `serving this boot's log to 10.0.2.2:…` line; both +`swap_crash_rolls_back` sightings' consoles show the rollback completing +(`restored`, then `logd` serving again). Only the host's redial gave up first, +every time. Every sighting's `cargo run -- --known-red` answered NO (not +quarantined), and an alone re-run is reliably green: 4 dials turned away on +`lan_swap`'s nightly, `swap_netd` green in 10 s on the rust-lld branch, +`swap_crash_rolls_back` green twice on main's nightly and once with netd +reverted to origin/main. + +## What the code shows + +`metalswap::swap` arms `Stream::redial` after logd's `CARRIER_LEAVING` and the +`go` (`src/metalswap.rs`), and `Stream::redial` opens a plain TCP dial to +`logd`'s log-stream port (`toyos_logstream::CARRIER`). `serve`'s loop +(`src/metaltalk.rs`) counts every dial that is refused, reset, or closes +before a line, and redials **at once** — there is no wait between attempts, +the comment names the refusal itself as the event — until a line arrives or +`metalswap::TURNED_AWAY_CEILING` (64) is reached. So the redial spends a +**fixed count** of dials against a gap whose length the *guest* sets: the old +netd's exit, the new one's spawn and DHCP lease, and init's whole 5000 ms "in +service" probe before it falls back to the one it replaced. This is +`issues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.md`'s +compromise. + +## What the measurement shows + +In `swap_netd`, 64 dials were refused in 326 ms between logd's `netd is being +replaced` (3.061 s) and its re-listen (3.387 s), each taking 5 ms or less. +Green swaps in the same runs were refused 3, 6 and 22 times. The dials got +faster on the red runs, not slower, which points at the refusal window logd +holds open while the old netd is still up. + +## Exit condition + +`Stream::redial` gives up on its time bound alone, never on a dial count, and a +swap whose refusal window is staged long is green; the three rows come off with +it. Owner: `src/metaltalk.rs`'s `Stream::redial`; held by the orchestrator. diff --git a/issues/build/lan-swap-redial-spent-its-ceiling-on-a-nightly-shard.md b/issues/build/lan-swap-redial-spent-its-ceiling-on-a-nightly-shard.md deleted file mode 100644 index e76e81fa8f1..00000000000 --- a/issues/build/lan-swap-redial-spent-its-ceiling-on-a-nightly-shard.md +++ /dev/null @@ -1,44 +0,0 @@ ---- -status: open -kind: defect -opened: 2026-09-27 ---- - -# `lan_swap`'s redial spent its ceiling on a nightly shard - -PR #535's nightly (run 36314576406, `guest (1)`, a4f68c5a, KVM, QEMU 11.1.0): - -``` -FAIL lan_swap: 2 finding(s): - init's words on netd were ["accepted"] ending in None, where InService is owed (the stream had 1 connection(s) before the ask and 1 after) - the stream's redial was turned away 64 time(s), its ceiling of 64, and gave up -``` - -The guest completed the swap. Its console has `init: swap netd: in service` -at 6.141 s, and `logd: serving this boot's log on port 41337` at 1.176 s -through the new netd. After that `logd` admitted no reader: no second -`serving this boot's log to 10.0.2.2:…` line. So none of the host's 64 dials -reached `logd` once it listened again. That is consistent with all of them -ending inside the guest's gap, which runs from `logd`'s `netd is being -replaced` (1.000 s), through init stopping the old netd (1.080 s), to `logd` -listening again (1.176 s). The alone re-run was green with 4 dials turned away. - -This is `issues/build/swap-crash-rolls-back-reds-when-its-redial-spends-its-ceiling-under-load.md` -on the 82574 bench: the compromise -`issues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.md` -records, reached. A redial asks again at once, and the ceiling counts dials, -not time. `lan_swap`'s path (`Ssh::swap`, `metalswap::swap`, `Stream::redial`, -`logd`, netd, init) has no change on #535. Main's nightly at 16d2e645 (run -36306830048) was green. Main's nightly at 1ce71831 (run 36290616312) had the -same two findings on `swap_crash_rolls_back`. - -Dev host, QEMU 11.1.1, TCG, one named run each: `nightly-green2` at 877b8c95, -`EXIT=0`, 12 dials turned away; `main` at 16d2e645, `EXIT=0`, 11. Not measured: -how fast a KVM guest turns a dial away, and so how many dials fit inside the -gap there. - -`cargo run -- --known-red lan_swap` answers NO. - -**Exit**: the redial waits on a guest-side event (the diagnostics issue's exit). -Until then `lan_swap` reds on the nightly at a rate. It should go on #542's -disabled list when that lands, citing this file. diff --git a/issues/build/parallel-tests-red-under-other-suites.md b/issues/build/parallel-tests-red-under-other-suites.md index df067417384..23580b7a723 100644 --- a/issues/build/parallel-tests-red-under-other-suites.md +++ b/issues/build/parallel-tests-red-under-other-suites.md @@ -505,11 +505,3 @@ mechanism for it. parallel run — a loaded full fast tier in which `i8042_undecoded_bytes`' first mute line names nothing and its second names the sequence, or the retirement's clause narrowed to the conditions under which it holds. - -- **`swap_netd`, again on its recorded signature.** The rust-lld branch's fast - tier at `e4317d3f` (PR #532), on a dev host whose load averages read 39.8, - 31.1 and 37.5 as it started, from other worktrees: `the stream's redial was - turned away 64 time(s), its ceiling of 64, and gave up`, with init's words on - netd `["accepted"]` ending in `None`; the one red of 400 besides - `lan_mdns_answer`'s `SUN_LEN`, and `ALONE swap_netd: GREEN` in 10 s. The - branch touches neither netd, swap nor init. Not investigated here. diff --git a/issues/build/swap-crash-rolls-back-redial-turned-away-once-on-mains-nightly.md b/issues/build/swap-crash-rolls-back-redial-turned-away-once-on-mains-nightly.md deleted file mode 100644 index f9a3abde491..00000000000 --- a/issues/build/swap-crash-rolls-back-redial-turned-away-once-on-mains-nightly.md +++ /dev/null @@ -1,24 +0,0 @@ ---- -status: open -kind: finding -opened: 2026-09-27 ---- - -# `swap_crash_rolls_back`'s redial was turned away to its ceiling once on main's nightly - -Main's nightly at 1ce71831 (run 36290616312), one guest shard, wide: - -``` -FAIL swap_crash_rolls_back: 2 finding(s): - init's words on netd were ["accepted"] ending in None, where Restored is owed (the stream had 1 connection(s) before the ask and 1 after) - the stream's redial was turned away 64 time(s), its ceiling of 64, and gave up -``` - -`ALONE swap_crash_rolls_back: GREEN` twice. It was not red on the nightly -before #527 (run 36285169430). -`cargo run -- --known-red swap_crash_rolls_back` answers NO. - -Not shown: what turned the redial away 64 times, and why. - -**Exit**: a cause for a redial turned away to its ceiling on a swap that -rolled back, or a rate with enough runs to call it gone. diff --git a/issues/build/swap-crash-rolls-back-reds-when-its-redial-spends-its-ceiling-under-load.md b/issues/build/swap-crash-rolls-back-reds-when-its-redial-spends-its-ceiling-under-load.md deleted file mode 100644 index b5b806e18d2..00000000000 --- a/issues/build/swap-crash-rolls-back-reds-when-its-redial-spends-its-ceiling-under-load.md +++ /dev/null @@ -1,30 +0,0 @@ ---- -status: open -kind: defect -opened: 2026-09-26 ---- - -# swap_crash_rolls_back reds when its redial spends its ceiling under host load - -`swap_crash_rolls_back` is red with the same two findings on origin/main's netd -and on a branch's: the host stream's redial after the swap was turned away -`metalswap::TURNED_AWAY_CEILING` (64) times and gave up, so init's `restored` -never reached the host, although the guest's own console shows the rollback -completing (`init: swap netd: restored`, then `logd: serving this boot's log on -port 41337`). It is the compromise -`issues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.md` -records, reached: the redial asks again at once, and on a loaded host the -ceiling runs out before `logd` listens again. - -Seen with ten spinning host threads beside the run, on the review branch of -the netd receive-pipe fix: once in a full `-- swap` run of five, red again -alone in that same run; and once in three `-- swap_crash_rolls_back` runs -with netd reverted to origin/main (4b235d27), where the harness's alone re-run -was green and called the `Sched::Parallel` classification wrong. The other -five of those six runs, three on the branch's netd and two on main's, were -green. `cargo run -- --known-red -swap_crash_rolls_back` answers that it is not quarantined. - -Exit condition: the redial waits on a guest-side event (the linked issue's -exit condition), or the test is shown green over a stated number of loaded -runs. diff --git a/issues/kernel/a-held-disk-waits-for-a-pass-no-cpu-takes-when-every-cpu-is-in-a-call-on-it.md b/issues/kernel/a-held-disk-waits-for-a-pass-no-cpu-takes-when-every-cpu-is-in-a-call-on-it.md index d9bfc536b7b..df18b1914ba 100644 --- a/issues/kernel/a-held-disk-waits-for-a-pass-no-cpu-takes-when-every-cpu-is-in-a-call-on-it.md +++ b/issues/kernel/a-held-disk-waits-for-a-pass-no-cpu-takes-when-every-cpu-is-in-a-call-on-it.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-09-22 --- @@ -43,3 +43,8 @@ the held disk; the interleaving comes at a rate. **Exit**: a disk call that waits for a device does not hold a CPU with `IF` clear — the wait parks — or a held call's CPU may bind within its bound without making the bind the caller's cost. Either is the owner's ruling to revisit. + +Related records — `usb_transport_break`'s three other open red modes, not this +one: `issues/kernel/a-shutdown-on-a-held-usb-disk-left-a-cpu-deaf-to-a-tlb-shootdown.md`, +`issues/build/usb-transport-break-flushedstick-can-break-after-the-reboot.md`, and +`issues/boot-media/a-disk-whose-port-went-away-panics-the-boot-at-roots-hold.md`. diff --git a/src/redlist.rs b/src/redlist.rs index 9adcffd4c98..301519afafe 100644 --- a/src/redlist.rs +++ b/src/redlist.rs @@ -46,6 +46,10 @@ pub const DISABLED: &[Disabled] = &[ issue: "issues/hardware/i8042-mouse-ends-four-packets-short-with-a-clean-exit.md", }, Disabled { test: "kill_while_blocked", issue: "issues/kernel/deferred-release-outlives-its-syscall.md" }, + Disabled { + test: "lan_swap", + issue: "issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md", + }, Disabled { test: "latency_wake", issue: "issues/build/latency-wake-reds-on-the-dev-host-at-a-rate.md" }, Disabled { test: "partition_claim_departure", @@ -75,10 +79,22 @@ pub const DISABLED: &[Disabled] = &[ test: "so_cache_refusals", issue: "issues/kernel/so-cache-refusals-saw-the-kernel-refuse-nothing-once.md", }, + Disabled { + test: "swap_crash_rolls_back", + issue: "issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md", + }, + Disabled { + test: "swap_netd", + issue: "issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md", + }, Disabled { test: "usb_disk_index_stable", issue: "issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md", }, + Disabled { + test: "usb_transport_break", + issue: "issues/kernel/a-held-disk-waits-for-a-pass-no-cpu-takes-when-every-cpu-is-in-a-call-on-it.md", + }, Disabled { test: "xhci_flap", issue: "issues/hardware/a-collapsed-replug-is-enumerated-only-when-another-port-event-arrives.md",