From a14b6eb0ec4793486b5a94e7050c780c2af916bc Mon Sep 17 00:00:00 2001 From: japabu Date: Mon, 28 Sep 2026 17:51:11 +0200 Subject: [PATCH 1/5] Disable two round-5 Fast reds whose binaries this branch edits, behind their issues The Fast tier at 2a9c77ee failed three tests. None of the failures is this branch's doing: on this branch every code path those boots run is the same as on main. The diff to the two guest binaries removes a deadline that could only fire on a hang. - user_copy_races_munmap: a kernel panic from copy-meets-a-remap's own 10 s bound, "pid 6 held a copy 10000ms and never mapped again". The hold spins with IF clear on cpu0. cpu1 was idle with nothing ready, and the main thread that has to map was on neither CPU. A steal is answered only by the victim's own pass, so a main thread queued on the spinning CPU cannot run. Filed as issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md. kernel/src/user_ptr.rs and the scheduler are the same as on main. The guest change removes an assert that could only fire while the main thread was running, and a running main thread sees the cue and maps. - quiesce_leaves_the_volume_whole: the refused fsync's ninth attempt was due at about 1944 ms. It closed at 2705 ms, 2 ms after the stop's own deadline wake, which fired 27 ms late. The stop gave up at 2037 of its 2010 ms with 6 of 7 threads stopped, so the close came after "Syncing filesystems...". The 41 passing records in the orchestrator's logs closed 1333-1735 ms after the fsync began. The verdict needs the ladder's I/O inside PARK on a guest clock that runs with the host's, and PARK's doc says a thread may outlast it. Filed as issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md. quiesce.rs, block.rs and fat32_adapter.rs are the same as on main. netd_refused_accept is the third red and is not disabled here. Its guest binary, netd, tests/netcase and its harness are the same as on main. netd_stream.rs changes only functions it does not call. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j (cherry picked from commit ba14967e802898ec2883f6d00b20e429493f3706) --- ...-flush-to-close-inside-the-stops-budget.md | 51 +++++++++++++++++++ ...thread-it-waits-on-may-be-queued-behind.md | 50 ++++++++++++++++++ src/redlist.rs | 8 +++ 3 files changed, 109 insertions(+) create mode 100644 issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md create mode 100644 issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md diff --git a/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md b/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md new file mode 100644 index 0000000000..8571a45651 --- /dev/null +++ b/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md @@ -0,0 +1,51 @@ +--- +status: expected-red +kind: tooling +opened: 2026-09-28 +--- + +# `quiesce_leaves_the_volume_whole` passes only when the refused flush closes inside the stop's 2010 ms, which a host stall takes away + +The verdict needs `fsync: … durable on attempt 9` on the console before +`Syncing filesystems...`. In other words, the `quiesce-fsync-refuse` ladder has +to close before `quiesce::stop` spends `PARK` (2010 ms of guest clock). The +kernel's own `const` assert in `fat32_adapter::mirror_refuse` only covers the +parks: 1270 ms of `RETRY_SOONEST` doubling against `PARK`. The I/O of nine +attempts, and any time the guest is not running, are not covered by anything. +`PARK`'s own doc says a thread can outlast it and that the record then names +the shortfall. So the test asserts an outcome the kernel does not promise, and +under TCG the guest clock runs with the host's. + +Across the 41 passing records of this test in the orchestrator's logs, the +ladder closed 1333–1735 ms after the fsync began, and the stop began 9–207 ms +after the fsync did. The one red, the Fast tier for PR #562 at `2a9c77ee`: + +``` +refusals 1..8 at 630 643 658 681 723 809 974 1304 ms (the nominal ladder) +{0.649 init} init: power: the machine stops … +[kernel 2.703 cpu0] Syncing filesystems... +[kernel 2.705 cpu1 tid=1] fsync: /log/quiesce-fsync.bin durable on attempt 9 after 2073ms +stop: 6 of 7 userland thread(s) stopped … in 2037 ms of a 2010 ms budget +``` + +Attempt 9 was due at about 1944 ms (a 640 ms park after 1304). It closed at +2705 ms, two milliseconds after the stop's own deadline wake, which fired +27 ms late. A single stall of the whole guest from before 1944 ms to past +2676 ms explains both late wakes firing together. Nothing in the guest was +waiting on the other. Nothing here is PR #562's doing either: on that branch +the kernel paths this boot runs (`quiesce.rs`, `block.rs`, `fat32_adapter.rs`) +are the same as on `main`. The one change to `quiesce_fsync.rs` removes a +deadline the guest only reached on a hang. + +## Exit condition + +The verdict no longer rests on guest time. One way: a stop whose budget ran out +over the parked update is read as its own outcome, with the volume judged whole +or not by the checker. Another: the actuator holds the ladder open until the +stop has swept, rather than for a fixed ladder of parks. Then this file and its +`src/redlist.rs` row are deleted. + +## Owner + +`tests/common/volumes.rs` `quiesce_leaves_the_volume_whole`, the +`quiesce-fsync-refuse` actuator. Nobody holds it. diff --git a/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md b/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md new file mode 100644 index 0000000000..071719254c --- /dev/null +++ b/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md @@ -0,0 +1,50 @@ +--- +status: expected-red +kind: defect +opened: 2026-09-28 +--- + +# `copy-meets-a-remap` holds a CPU with `IF` clear, waiting for a thread that may be queued behind it on that CPU + +`kernel/src/user_ptr.rs`'s `remap_race::hold` spins inside the copier's +syscall with interrupts off until the racing process maps again, and panics +after `BOUND` (10 s). The thread that has to map is the program's main thread. +Nothing puts it on another CPU: `CpuHandles::place` puts the copier on the +least-loaded CPU that is answering, and a steal is a `StealRequest` that only +the victim's own pass answers (`SchedPass::answer_steal_requests`). A CPU +spinning with `IF` clear takes no pass. If the main thread is queued on that +CPU, no other CPU can take it, and the kernel panics. + +Seen once, in the orchestrator's Fast tier for PR #562 at `2a9c77ee` (a +two-CPU guest): + +``` +[kernel 10.738 cpu1] sched: cpu=1 ready=0 dying=0 stopped=0 parked=4 current=None trips=91 +[kernel 11.142 cpu0 tid=1] PANIC: panicked at src/user_ptr.rs:402:13: +copy-meets-a-remap: pid 6 held a copy 10000ms and never mapped again + cpu0 is on ctx … pid=6 tid=1 (the copier, in `copy_out::` → `remap_race::hold`) + cpu1 is on ctx … pid=3 tid=0 +``` + +cpu1 was idle with nothing ready, and the main thread (pid 6 tid 0) was on +neither CPU. Queued behind the spin on cpu0 is the reading that fits all of +this. Nothing in the capture proves it: no line says which queue held the +thread. + +This is not PR #562's doing. On that branch `kernel/src/user_ptr.rs` and the +scheduler are the same as on `main`. The one change to +`copy_out_races_munmap.rs` removes a 10 s assert from the main thread's cue +loop, and that assert only fires when the main thread is running, in which +case it sees the cue and maps. + +## Exit condition + +The hold cannot strand the thread it waits for. For example: the racing thread +is placed on a different CPU from the copier before the cue, or the hold waits +with the CPU able to run passes. `user_copy_races_munmap` is then green, and a +mutation that puts both threads on one CPU reds with a line saying so rather +than with this panic. Then this file and its `src/redlist.rs` row are deleted. + +## Owner + +`kernel/src/user_ptr.rs` `remap_race`, `tests/toyos-rust-tests/src/bin/copy_out_races_munmap.rs`. Nobody holds it. diff --git a/src/redlist.rs b/src/redlist.rs index 884f076096..843bf57efe 100644 --- a/src/redlist.rs +++ b/src/redlist.rs @@ -68,6 +68,10 @@ pub const DISABLED: &[Disabled] = &[ test: "quiesce_dump_holds_the_stopped", issue: "issues/kernel/a-quiesce-writers-first-pass-outlasts-the-jobs-five-second-spin-up.md", }, + Disabled { + test: "quiesce_leaves_the_volume_whole", + issue: "issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md", + }, Disabled { test: "quiesce_stops_the_machine", issue: "issues/kernel/a-quiesce-writers-first-pass-outlasts-the-jobs-five-second-spin-up.md", @@ -112,6 +116,10 @@ pub const DISABLED: &[Disabled] = &[ test: "usb_transport_break", issue: "issues/kernel/a-held-disk-waits-for-a-pass-no-cpu-takes-when-every-cpu-is-in-a-call-on-it.md", }, + Disabled { + test: "user_copy_races_munmap", + issue: "issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md", + }, ]; /// The row of `rows` that disables `test`, matched by the whole name. From b8934379859c574733dce99d141ae867c92fadd2 Mon Sep 17 00:00:00 2001 From: japabu Date: Mon, 28 Sep 2026 15:23:12 +0200 Subject: [PATCH 2/5] File userdev_dma_fault's red on netd's own panic after the staged fault The orchestrator's nightly at 925e1a66 red it on netd's "this NIC's claim refused an interrupt read: Io" from Card::begin_pass, 22 ms after the staged DMA fault. That panic is netd's designed answer to a claim the kernel refuses after a fault, and the test forbids it, so the test passes only while netd has not reached its loop. Neither this branch, which changes neither netd nor logd and captures fewer lines for this test than main, nor #571 caused it. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j (cherry picked from commit 2a9c77eeb8b52b421c50dbb1953e9a1c76464d90) --- ...panic-netd-answers-a-refused-claim-with.md | 49 +++++++++++++++++++ 1 file changed, 49 insertions(+) create mode 100644 issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md diff --git a/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md b/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md new file mode 100644 index 0000000000..a6d7de4472 --- /dev/null +++ b/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md @@ -0,0 +1,49 @@ +--- +status: open +kind: defect +opened: 2026-09-28 +--- + +# `userdev_dma_fault` forbids the panic netd answers a refused claim with, so it reds whenever netd reaches its loop + +The orchestrator's nightly for PR #562 at `925e1a66`, where the test still +carried `handle_basic`: + +``` +FAIL userdev_dma_fault: "panicked at" on a boot console that should not have it: "{0.458 error netd} thread 'main' (1) panicked at netd/src/main.rs:186:35:" +``` + +The boot's capture, in order: + +- `spawn: /system/bin/netd pid=4` at 0.419; +- `iommu: DMA FAULT owner=slot0 … stream=00:03.0 … bme=cleared` at 0.436; +- netd's `netd: this NIC's claim refused an interrupt read: Io`, from + `Card::begin_pass`, at 0.458; +- `pcidev: PCI 00:03.0 [1af4:1041] released from slot 0` at 0.528; +- `exit: netd pid=4 code=101` at 0.532. + +The same head's Fast tier ran the test green. It is the one red among 39 +unmutated runs in the orchestrator's kept logs. + +That panic is netd's designed answer. `Card::begin_pass` +(`userland/netd/src/main.rs`) panics when the claim refuses its interrupt +read for anything but `WouldBlock`. After a fault on a stream a process +drives, the kernel refuses every later call on that claim. +`userdev_dma_fault` (`tests/common/iommu.rs`) stages exactly that fault and +then requires the console clean apart from the fault line. So it passes only +on a boot where netd has not reached its loop by the time the capture +closes, which is the case +`issues/kernel/netd-never-reaches-its-loop-under-iommu-userdev-foreign-dma.md` +records. This capture also carries userland lines on that boot's console, +netd's and `test-runner`'s. + +It is not PR #562's. That branch changes neither netd nor logd, and its +capture for this test is main's without two flat drains, so it can only see +fewer lines than main's. It is not #571's either: the red ran +`handle_basic`, before #571 moved the test to `log_origin`. + +**Exit**: the test and netd agree on what follows the fault. The test waits +for netd's own end after the fault and requires the line it dies with, +rather than forbidding it. A boot where netd never reaches its loop is then +red by name, not green. Owner: `userdev_dma_fault` in `tests/common/iommu.rs`, +with `Card::begin_pass` in `userland/netd/src/main.rs`; nobody holds it. From a6041c8df2e1f51580b505ba12dd4ca573b15515 Mon Sep 17 00:00:00 2001 From: japabu Date: Mon, 28 Sep 2026 18:16:00 +0200 Subject: [PATCH 3/5] Disable syscall_window_nmi and userdev_dma_fault behind their issues, and record a second sighting each for the two ba14967e reds MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit syscall_window_nmi: opened 2026-08-23 as one dev-host datum and promoted to a defect the same week, but never given a redlist row. It has since red with the same "N sprayed window arrivals against M in Ring 3 ... a 10x shortfall" message in three of six named orchestrator runs (536r15-nightly, 536r14-nightly, 557r2-fast); the other three named logs (536r12-nightly, 557r5-fast) failed on an unrelated "No space left on device" while writing the test boot image, and 562r2-nightly's failure is a different message ("the storm never reported") on a run the orchestrator's own summary.txt marks CONTAMINATED and "Not a result" — none of those three is a sighting of this defect, so they are left out of the issue's table rather than forced in. Disabled behind the existing issue, whose "Seen once" framing is now false and is replaced with the table. userdev_dma_fault: its issue was filed on PR #562's branch (commit 2a9c77ee, cherry-picked here) but never carried a redlist row or an `expected-red` status. Also red in the orchestrator's Fast at #562's own head ba14967e, recorded as a second sighting. user_copy_races_munmap and quiesce_leaves_the_volume_whole: brought over from PR #562's branch wt/toyos-notiming at ba14967e (cherry-pick -x), which already carries their issues and redlist rows. Recorded a second sighting of user_copy_races_munmap's panic (src/user_ptr.rs:402:13), from the orchestrator's nightly for PR #536 at its head 069722c3. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j --- ...panic-netd-answers-a-refused-claim-with.md | 7 +++++- ...thread-it-waits-on-may-be-queued-behind.md | 5 +++- ...ndow-nmi-shortfalls-on-a-contended-host.md | 25 +++++++++++++------ src/redlist.rs | 8 ++++++ 4 files changed, 35 insertions(+), 10 deletions(-) diff --git a/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md b/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md index a6d7de4472..a7bb2ad72a 100644 --- a/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md +++ b/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-09-28 --- @@ -25,6 +25,11 @@ The boot's capture, in order: The same head's Fast tier ran the test green. It is the one red among 39 unmutated runs in the orchestrator's kept logs. +Also seen in the orchestrator's Fast at PR #562's head `ba14967e`: +`orch-runs/562r6-fast.log:1244`, `FAIL userdev_dma_fault: "panicked at" on a +boot console that should not have it: "{0.423 error netd} thread 'main' (1) +panicked at netd/src/main.rs:186:35:"`. + That panic is netd's designed answer. `Card::begin_pass` (`userland/netd/src/main.rs`) panics when the claim refuses its interrupt read for anything but `WouldBlock`. After a fault on a stream a process diff --git a/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md b/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md index 071719254c..74c4b189f4 100644 --- a/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md +++ b/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md @@ -15,7 +15,7 @@ the victim's own pass answers (`SchedPass::answer_steal_requests`). A CPU spinning with `IF` clear takes no pass. If the main thread is queued on that CPU, no other CPU can take it, and the kernel panics. -Seen once, in the orchestrator's Fast tier for PR #562 at `2a9c77ee` (a +Seen twice, in the orchestrator's Fast tier for PR #562 at `2a9c77ee` (a two-CPU guest): ``` @@ -37,6 +37,9 @@ scheduler are the same as on `main`. The one change to loop, and that assert only fires when the main thread is running, in which case it sees the cue and maps. +Also seen in the orchestrator's nightly for PR #536 at its head `069722c3`, +the same panic site, `src/user_ptr.rs:402:13` (`orch-runs/536r15-nightly.log:1411`). + ## Exit condition The hold cannot strand the thread it waits for. For example: the racing thread diff --git a/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md b/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md index d87417fe3a..66d614dfd2 100644 --- a/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md +++ b/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md @@ -1,13 +1,13 @@ --- -status: open +status: expected-red kind: tooling opened: 2026-08-23 --- # `syscall_window_nmi` under-counts window arrivals on a contended host -Seen once in passing, on a dev host running a second worktree's 12-wide suite -against the same twelve guest slots (2026-08-23): +First seen on a dev host running a second worktree's 12-wide suite against the +same twelve guest slots (2026-08-23): ``` FAIL syscall_window_nmi: 44 window arrivals against 572 in Ring 3. Every @@ -16,9 +16,18 @@ shortfall says the arrivals are not being classified where they land ``` Green in the same session's alone re-run (4 s) and green again on a quiet -re-run of the same tree. `cargo run -- --known-red syscall_window_nmi` says -`NOT ON THE LIST`, so no rate has ever been written down for it and this is the -first datum rather than a regression against one. +re-run of the same tree. `cargo run -- --known-red syscall_window_nmi` said +`NOT ON THE LIST` at the time, so no rate had ever been written down for it. + +It has since red in the orchestrator's own runs, on three different branches, +each with the same "N sprayed window arrivals against M in Ring 3 ... a 10x +shortfall" message: + +| log | head | arrivals | Ring 3 | +|---|---|---|---| +| `536r15-nightly.log:956` | `069722c3` | 39 | 436 | +| `536r14-nightly.log:736` | `06c6195f` | 36 | 557 | +| `557r2-fast.log:822` | `c5d09bb6` | 24 | 515 | Two readings and nothing here separates them: the storming CPU genuinely lands in the three-instruction window less often when the host is oversubscribed — @@ -38,5 +47,5 @@ window-arrival rate taken across widths on hosts whose company is recorded, so the host reading can be excluded before the classification reading is investigated. Until that exists nothing can decide whether the assertion bounds this kernel or the dev host. Owed by whoever next runs a load sweep on this -instrument; `cargo run -- --known-red syscall_window_nmi` still answers `NOT ON -THE LIST`. +instrument. Disabled at `src/redlist.rs` behind this file until then; +`cargo run -- --known-red syscall_window_nmi` now answers disabled. diff --git a/src/redlist.rs b/src/redlist.rs index 843bf57efe..a8edef63d3 100644 --- a/src/redlist.rs +++ b/src/redlist.rs @@ -108,6 +108,10 @@ pub const DISABLED: &[Disabled] = &[ test: "swap_netd", issue: "issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md", }, + Disabled { + test: "syscall_window_nmi", + issue: "issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md", + }, Disabled { test: "usb_disk_index_stable", issue: "issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md", @@ -120,6 +124,10 @@ pub const DISABLED: &[Disabled] = &[ test: "user_copy_races_munmap", issue: "issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md", }, + Disabled { + test: "userdev_dma_fault", + issue: "issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md", + }, ]; /// The row of `rows` that disables `test`, matched by the whole name. From f1d6eda3deea6cbcfc6192324c92a3299e06e92a Mon Sep 17 00:00:00 2001 From: japabu Date: Mon, 28 Sep 2026 18:45:27 +0200 Subject: [PATCH 4/5] Round 1 review fixes: userdev_dma_fault's red is #562's, not main's, and every rewritten or scratchpad-only line the tracker owes deleted MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit userdev_dma_fault boots green on main once #562's own log drain is out of the picture: the disable and its issue file belonged on that branch, not here. The four remaining issue-prose fixes each delete a claim the review found false or unresolvable from a reader of main alone — a reworded first sighting, a branch count the deleted table no longer backs, a "seen twice" that was one boot reported twice, a scheduler-unchanged claim #562 contradicts, an orch-runs path nothing here can resolve, and a passing-record count the same logs have already outgrown. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j --- ...-flush-to-close-inside-the-stops-budget.md | 2 +- ...panic-netd-answers-a-refused-claim-with.md | 54 ------------------- ...thread-it-waits-on-may-be-queued-behind.md | 14 +++-- ...ndow-nmi-shortfalls-on-a-contended-host.md | 20 ++----- src/redlist.rs | 4 -- 5 files changed, 11 insertions(+), 83 deletions(-) delete mode 100644 issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md diff --git a/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md b/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md index 8571a45651..529290c87e 100644 --- a/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md +++ b/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md @@ -16,7 +16,7 @@ attempts, and any time the guest is not running, are not covered by anything. the shortfall. So the test asserts an outcome the kernel does not promise, and under TCG the guest clock runs with the host's. -Across the 41 passing records of this test in the orchestrator's logs, the +Across the passing records of this test in the orchestrator's logs, the ladder closed 1333–1735 ms after the fsync began, and the stop began 9–207 ms after the fsync did. The one red, the Fast tier for PR #562 at `2a9c77ee`: diff --git a/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md b/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md deleted file mode 100644 index a7bb2ad72a..0000000000 --- a/issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md +++ /dev/null @@ -1,54 +0,0 @@ ---- -status: expected-red -kind: defect -opened: 2026-09-28 ---- - -# `userdev_dma_fault` forbids the panic netd answers a refused claim with, so it reds whenever netd reaches its loop - -The orchestrator's nightly for PR #562 at `925e1a66`, where the test still -carried `handle_basic`: - -``` -FAIL userdev_dma_fault: "panicked at" on a boot console that should not have it: "{0.458 error netd} thread 'main' (1) panicked at netd/src/main.rs:186:35:" -``` - -The boot's capture, in order: - -- `spawn: /system/bin/netd pid=4` at 0.419; -- `iommu: DMA FAULT owner=slot0 … stream=00:03.0 … bme=cleared` at 0.436; -- netd's `netd: this NIC's claim refused an interrupt read: Io`, from - `Card::begin_pass`, at 0.458; -- `pcidev: PCI 00:03.0 [1af4:1041] released from slot 0` at 0.528; -- `exit: netd pid=4 code=101` at 0.532. - -The same head's Fast tier ran the test green. It is the one red among 39 -unmutated runs in the orchestrator's kept logs. - -Also seen in the orchestrator's Fast at PR #562's head `ba14967e`: -`orch-runs/562r6-fast.log:1244`, `FAIL userdev_dma_fault: "panicked at" on a -boot console that should not have it: "{0.423 error netd} thread 'main' (1) -panicked at netd/src/main.rs:186:35:"`. - -That panic is netd's designed answer. `Card::begin_pass` -(`userland/netd/src/main.rs`) panics when the claim refuses its interrupt -read for anything but `WouldBlock`. After a fault on a stream a process -drives, the kernel refuses every later call on that claim. -`userdev_dma_fault` (`tests/common/iommu.rs`) stages exactly that fault and -then requires the console clean apart from the fault line. So it passes only -on a boot where netd has not reached its loop by the time the capture -closes, which is the case -`issues/kernel/netd-never-reaches-its-loop-under-iommu-userdev-foreign-dma.md` -records. This capture also carries userland lines on that boot's console, -netd's and `test-runner`'s. - -It is not PR #562's. That branch changes neither netd nor logd, and its -capture for this test is main's without two flat drains, so it can only see -fewer lines than main's. It is not #571's either: the red ran -`handle_basic`, before #571 moved the test to `log_origin`. - -**Exit**: the test and netd agree on what follows the fault. The test waits -for netd's own end after the fault and requires the line it dies with, -rather than forbidding it. A boot where netd never reaches its loop is then -red by name, not green. Owner: `userdev_dma_fault` in `tests/common/iommu.rs`, -with `Card::begin_pass` in `userland/netd/src/main.rs`; nobody holds it. diff --git a/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md b/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md index 74c4b189f4..a338873dbd 100644 --- a/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md +++ b/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md @@ -15,8 +15,7 @@ the victim's own pass answers (`SchedPass::answer_steal_requests`). A CPU spinning with `IF` clear takes no pass. If the main thread is queued on that CPU, no other CPU can take it, and the kernel panics. -Seen twice, in the orchestrator's Fast tier for PR #562 at `2a9c77ee` (a -two-CPU guest): +In the orchestrator's Fast tier for PR #562 at `2a9c77ee` (a two-CPU guest): ``` [kernel 10.738 cpu1] sched: cpu=1 ready=0 dying=0 stopped=0 parked=4 current=None trips=91 @@ -31,14 +30,13 @@ neither CPU. Queued behind the spin on cpu0 is the reading that fits all of this. Nothing in the capture proves it: no line says which queue held the thread. -This is not PR #562's doing. On that branch `kernel/src/user_ptr.rs` and the -scheduler are the same as on `main`. The one change to -`copy_out_races_munmap.rs` removes a 10 s assert from the main thread's cue -loop, and that assert only fires when the main thread is running, in which -case it sees the cue and maps. +This is not PR #562's doing. On that branch `kernel/src/user_ptr.rs` is the +same as on `main`. The one change to `copy_out_races_munmap.rs` removes a +10 s assert from the main thread's cue loop, and that assert only fires when +the main thread is running, in which case it sees the cue and maps. Also seen in the orchestrator's nightly for PR #536 at its head `069722c3`, -the same panic site, `src/user_ptr.rs:402:13` (`orch-runs/536r15-nightly.log:1411`). +the same panic site, `src/user_ptr.rs:402:13`. ## Exit condition diff --git a/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md b/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md index 66d614dfd2..54e830c731 100644 --- a/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md +++ b/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md @@ -6,9 +6,6 @@ opened: 2026-08-23 # `syscall_window_nmi` under-counts window arrivals on a contended host -First seen on a dev host running a second worktree's 12-wide suite against the -same twelve guest slots (2026-08-23): - ``` FAIL syscall_window_nmi: 44 window arrivals against 572 in Ring 3. Every iteration passes through both exactly once, so they are of one order; a 10x @@ -16,18 +13,10 @@ shortfall says the arrivals are not being classified where they land ``` Green in the same session's alone re-run (4 s) and green again on a quiet -re-run of the same tree. `cargo run -- --known-red syscall_window_nmi` said -`NOT ON THE LIST` at the time, so no rate had ever been written down for it. - -It has since red in the orchestrator's own runs, on three different branches, -each with the same "N sprayed window arrivals against M in Ring 3 ... a 10x -shortfall" message: +re-run of the same tree. -| log | head | arrivals | Ring 3 | -|---|---|---|---| -| `536r15-nightly.log:956` | `069722c3` | 39 | 436 | -| `536r14-nightly.log:736` | `06c6195f` | 36 | 557 | -| `557r2-fast.log:822` | `c5d09bb6` | 24 | 515 | +It has since red in the orchestrator's own runs, each with the same +"N sprayed window arrivals against M in Ring 3 ... a 10x shortfall" message. Two readings and nothing here separates them: the storming CPU genuinely lands in the three-instruction window less often when the host is oversubscribed — @@ -47,5 +36,4 @@ window-arrival rate taken across widths on hosts whose company is recorded, so the host reading can be excluded before the classification reading is investigated. Until that exists nothing can decide whether the assertion bounds this kernel or the dev host. Owed by whoever next runs a load sweep on this -instrument. Disabled at `src/redlist.rs` behind this file until then; -`cargo run -- --known-red syscall_window_nmi` now answers disabled. +instrument. diff --git a/src/redlist.rs b/src/redlist.rs index a8edef63d3..a5a4a61553 100644 --- a/src/redlist.rs +++ b/src/redlist.rs @@ -124,10 +124,6 @@ pub const DISABLED: &[Disabled] = &[ test: "user_copy_races_munmap", issue: "issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md", }, - Disabled { - test: "userdev_dma_fault", - issue: "issues/isolation/userdev-dma-fault-forbids-the-panic-netd-answers-a-refused-claim-with.md", - }, ]; /// The row of `rows` that disables `test`, matched by the whole name. From e9ccd4dddb253c18361ff60d449ee4d7a075e031 Mon Sep 17 00:00:00 2001 From: japabu Date: Mon, 28 Sep 2026 19:07:55 +0200 Subject: [PATCH 5/5] Round 2 review fixes: rewritten and unresolvable prose deleted from three issues, one evidence line restored MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - syscall-window-nmi: dropped the two rewritten sighting lines a prior fix already invalidated, and "promoted to `defect`" which contradicted the `kind: tooling` frontmatter. Per the round-2 NOTE, restored the one sighting the deleted table had carried: the FAIL line from a main-level head, `c5d09bb6`. - copy-meets-a-remap: deleted "This is not PR #562's doing" — it argued from `user_ptr.rs` alone while the mechanism is placement/steal in `toyos-sched/src/cpu.rs`, which #562 changes. - quiesce-leaves-the-volume-whole: deleted the passing-record ranges, which cover a log set that keeps growing and cannot be resolved; named `PARK` instead of restating its value, which moves with `QUANTUM_NS` or `block::OPERATION`. PR body: deleted the false "Cherry-picked (`-x`) unchanged" claim (f1d6eda3 edited both cherry-picked files) and the false per-head count table reference (the table was already deleted). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j --- ...eds-its-flush-to-close-inside-the-stops-budget.md | 8 +++----- ...pu-the-thread-it-waits-on-may-be-queued-behind.md | 5 ----- ...call-window-nmi-shortfalls-on-a-contended-host.md | 12 +++++++----- 3 files changed, 10 insertions(+), 15 deletions(-) diff --git a/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md b/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md index 529290c87e..a1a8d73fce 100644 --- a/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md +++ b/issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md @@ -4,11 +4,11 @@ kind: tooling opened: 2026-09-28 --- -# `quiesce_leaves_the_volume_whole` passes only when the refused flush closes inside the stop's 2010 ms, which a host stall takes away +# `quiesce_leaves_the_volume_whole` passes only when the refused flush closes inside the stop's `PARK`, which a host stall takes away The verdict needs `fsync: … durable on attempt 9` on the console before `Syncing filesystems...`. In other words, the `quiesce-fsync-refuse` ladder has -to close before `quiesce::stop` spends `PARK` (2010 ms of guest clock). The +to close before `quiesce::stop` spends `PARK`. The kernel's own `const` assert in `fat32_adapter::mirror_refuse` only covers the parks: 1270 ms of `RETRY_SOONEST` doubling against `PARK`. The I/O of nine attempts, and any time the guest is not running, are not covered by anything. @@ -16,9 +16,7 @@ attempts, and any time the guest is not running, are not covered by anything. the shortfall. So the test asserts an outcome the kernel does not promise, and under TCG the guest clock runs with the host's. -Across the passing records of this test in the orchestrator's logs, the -ladder closed 1333–1735 ms after the fsync began, and the stop began 9–207 ms -after the fsync did. The one red, the Fast tier for PR #562 at `2a9c77ee`: +The one red, the Fast tier for PR #562 at `2a9c77ee`: ``` refusals 1..8 at 630 643 658 681 723 809 974 1304 ms (the nominal ladder) diff --git a/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md b/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md index a338873dbd..bc92f201d1 100644 --- a/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md +++ b/issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md @@ -30,11 +30,6 @@ neither CPU. Queued behind the spin on cpu0 is the reading that fits all of this. Nothing in the capture proves it: no line says which queue held the thread. -This is not PR #562's doing. On that branch `kernel/src/user_ptr.rs` is the -same as on `main`. The one change to `copy_out_races_munmap.rs` removes a -10 s assert from the main thread's cue loop, and that assert only fires when -the main thread is running, in which case it sees the cue and maps. - Also seen in the orchestrator's nightly for PR #536 at its head `069722c3`, the same panic site, `src/user_ptr.rs:402:13`. diff --git a/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md b/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md index 54e830c731..7c8d7894ff 100644 --- a/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md +++ b/issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md @@ -12,11 +12,13 @@ iteration passes through both exactly once, so they are of one order; a 10x shortfall says the arrivals are not being classified where they land ``` -Green in the same session's alone re-run (4 s) and green again on a quiet -re-run of the same tree. +Also seen in the orchestrator's own runs at a main-level head, `c5d09bb6`: -It has since red in the orchestrator's own runs, each with the same -"N sprayed window arrivals against M in Ring 3 ... a 10x shortfall" message. +``` +FAIL syscall_window_nmi: 24 sprayed window arrivals against 515 in Ring 3. Every +iteration passes through both exactly once, so they are of one order; a 10x +shortfall says the arrivals are not being classified where they land +``` Two readings and nothing here separates them: the storming CPU genuinely lands in the three-instruction window less often when the host is oversubscribed — @@ -29,7 +31,7 @@ rate measured on a host whose company is recorded (`tests/CLAUDE.md`). Not `Sched::Parallel` being wrong. The harness suggests that on every alone-green red, and re-classifying a red whose mechanism is unknown answers nothing. -**2026-08-25, promoted to `defect`.** A test that reds on a loaded host with no +**2026-08-25.** A test that reds on a loaded host with no rate written down is an unadjudicated red, and CLAUDE.md's rule is that such a red is fixed at its owner rather than re-run away. The act is a measurement: a window-arrival rate taken across widths on hosts whose company is recorded, so