Disable three flaky main-level reds behind their issues - #580
Conversation
…d their issues The Fast tier at 2a9c77e failed three tests. None of the failures is this branch's doing: on this branch every code path those boots run is the same as on main. The diff to the two guest binaries removes a deadline that could only fire on a hang. - user_copy_races_munmap: a kernel panic from copy-meets-a-remap's own 10 s bound, "pid 6 held a copy 10000ms and never mapped again". The hold spins with IF clear on cpu0. cpu1 was idle with nothing ready, and the main thread that has to map was on neither CPU. A steal is answered only by the victim's own pass, so a main thread queued on the spinning CPU cannot run. Filed as issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md. kernel/src/user_ptr.rs and the scheduler are the same as on main. The guest change removes an assert that could only fire while the main thread was running, and a running main thread sees the cue and maps. - quiesce_leaves_the_volume_whole: the refused fsync's ninth attempt was due at about 1944 ms. It closed at 2705 ms, 2 ms after the stop's own deadline wake, which fired 27 ms late. The stop gave up at 2037 of its 2010 ms with 6 of 7 threads stopped, so the close came after "Syncing filesystems...". The 41 passing records in the orchestrator's logs closed 1333-1735 ms after the fsync began. The verdict needs the ladder's I/O inside PARK on a guest clock that runs with the host's, and PARK's doc says a thread may outlast it. Filed as issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md. quiesce.rs, block.rs and fat32_adapter.rs are the same as on main. netd_refused_accept is the third red and is not disabled here. Its guest binary, netd, tests/netcase and its harness are the same as on main. netd_stream.rs changes only functions it does not call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j (cherry picked from commit ba14967)
The orchestrator's nightly at 925e1a6 red it on netd's "this NIC's claim refused an interrupt read: Io" from Card::begin_pass, 22 ms after the staged DMA fault. That panic is netd's designed answer to a claim the kernel refuses after a fault, and the test forbids it, so the test passes only while netd has not reached its loop. Neither this branch, which changes neither netd nor logd and captures fewer lines for this test than main, nor #571 caused it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j (cherry picked from commit 2a9c77e)
… and record a second sighting each for the two ba14967 reds syscall_window_nmi: opened 2026-08-23 as one dev-host datum and promoted to a defect the same week, but never given a redlist row. It has since red with the same "N sprayed window arrivals against M in Ring 3 ... a 10x shortfall" message in three of six named orchestrator runs (536r15-nightly, 536r14-nightly, 557r2-fast); the other three named logs (536r12-nightly, 557r5-fast) failed on an unrelated "No space left on device" while writing the test boot image, and 562r2-nightly's failure is a different message ("the storm never reported") on a run the orchestrator's own summary.txt marks CONTAMINATED and "Not a result" — none of those three is a sighting of this defect, so they are left out of the issue's table rather than forced in. Disabled behind the existing issue, whose "Seen once" framing is now false and is replaced with the table. userdev_dma_fault: its issue was filed on PR #562's branch (commit 2a9c77e, cherry-picked here) but never carried a redlist row or an `expected-red` status. Also red in the orchestrator's Fast at #562's own head ba14967, recorded as a second sighting. user_copy_races_munmap and quiesce_leaves_the_volume_whole: brought over from PR #562's branch wt/toyos-notiming at ba14967 (cherry-pick -x), which already carries their issues and redlist rows. Recorded a second sighting of user_copy_races_munmap's panic (src/user_ptr.rs:402:13), from the orchestrator's nightly for PR #536 at its head 069722c. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Review, round 1, at a6041c8CI: Per test, is it a flaky red on main-level code?
BLOCKER
NOTE
REMOVE
SEND BACK |
…and every rewritten or scratchpad-only line the tracker owes deleted userdev_dma_fault boots green on main once #562's own log drain is out of the picture: the disable and its issue file belonged on that branch, not here. The four remaining issue-prose fixes each delete a claim the review found false or unresolvable from a reader of main alone — a reworded first sighting, a branch count the deleted table no longer backs, a "seen twice" that was one boot reported twice, a scheduler-unchanged claim #562 contradicts, an orch-runs path nothing here can resolve, and a passing-record count the same logs have already outgrown. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
#564 landed on main since this branch was cut and deleted 15 never-caught tests, five of them pre-existing redlist rows this branch did not touch: console_line_atomicity, quiesce_dump_holds_the_stopped, quiesce_wakes_on_the_last_exit, so_cache_refusals, usb_disk_index_stable. Kept main's deletions for those and this branch's own three additions (quiesce_leaves_the_volume_whole, syscall_window_nmi, user_copy_races_munmap) in src/redlist.rs, in the file's existing order. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Review, round 2, at ea2b199CI: Merge ea2b199 ( Round 1:
BLOCKER(none) NOTE
REMOVE
LAND AFTER NAMED CHANGES |
…hree issues, one evidence line restored - syscall-window-nmi: dropped the two rewritten sighting lines a prior fix already invalidated, and "promoted to `defect`" which contradicted the `kind: tooling` frontmatter. Per the round-2 NOTE, restored the one sighting the deleted table had carried: the FAIL line from a main-level head, `c5d09bb6`. - copy-meets-a-remap: deleted "This is not PR #562's doing" — it argued from `user_ptr.rs` alone while the mechanism is placement/steal in `toyos-sched/src/cpu.rs`, which #562 changes. - quiesce-leaves-the-volume-whole: deleted the passing-record ranges, which cover a log set that keeps growing and cannot be resolved; named `PARK` instead of restating its value, which moves with `QUANTUM_NS` or `block::OPERATION`. PR body: deleted the false "Cherry-picked (`-x`) unchanged" claim (f1d6eda edited both cherry-picked files) and the false per-head count table reference (the table was already deleted). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
… a test this branch deletes, move the stop-sync's only guest control to Nightly, and rename a misnamed binding Merging origin/main (c089fce, #580) brought back a redlist row and issue for `quiesce_leaves_the_volume_whole`, a test this branch already deletes; the harness refuses a disabled row nothing registers, so both go. `home_overwrite_reads_back` is the only guest check on init's stop sync and was priced at Weekly for the old kernel-`/home` test it replaced, so a lost sync goes unseen by every PR and nightly; it moves to Nightly. `tests/common/power.rs`'s `synced_at` bound the stop record's line, not anything synced there. Also deletes two stale doc comments in `process_stats.rs` that only narrated the code below them, and a roster.rs clause the reviewer found false for `process_stats`'s own child. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Brings in #572 (host QEMU's edk2), #580 and #579 (disabled reds), #549 (a kill never waits on its victim) and #566 (metaltalk redial). - src/redlist.rs: one row each for user_copy_races_munmap and quiesce_leaves_the_volume_whole, which both sides added. main's new rows stay (netd_refused_accept, quiesce_wakes_on_the_last_teardown, root_chunk_refused_on_a_usb_stick, syscall_window_nmi). The rows for tests or issues this branch deleted go (hda_tone, doom_sound_flood, latency_wake, sched_check_build), and so does lan_swap, whose issue main deleted with swap_netd's and swap_crash_rolls_back's rows. - The two issue files both sides added take main's text. - tests/common/power.rs: main's woken_by_the_held_thread, shared by the new quiesce_wakes_on_the_last_teardown, without the two clock verdicts this branch took off QEMU (stopped_the_machine in stopped_boot, and woken_by_its_threads). - tests/common/qemu.rs: qemu_command takes main's firmware_vars and has no audio_wav, so profile_argv passes six paths. The too_many_arguments allow goes, because seven parameters do not trigger it. - kill_while_blocked.rs: main's text. After #549 a kill does not park in retire_task, so this branch's doc for arm 4 was false. main's arm also has no clock. - tests/toyos.rs check_rust_result: this branch's single-print form, which already carries the stdout main added. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Summary
Three main-level guest tests were flaky reds without a redlist row or a
disabled status:
user_copy_races_munmap,quiesce_leaves_the_volume_whole,syscall_window_nmi. Per CLAUDE.md a flaky test is disabled at once behindits filed issue; none of these three is this branch's doing (see each issue
for why).
user_copy_races_munmap,quiesce_leaves_the_volume_whole: their issuesand redlist rows already existed on PR No QEMU test measures time, and audio is judged on metal only #562's branch
wt/toyos-notimingascommit
ba14967e.issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md,issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md.Added a second sighting to the
user_copy_races_munmapissue: the samepanic site (
src/user_ptr.rs:402:13) in the orchestrator's nightly forPR Storage: file servers for DATA, the log and the boot volume; the kernel's NVMe and FAT go #536 at its head
069722c3, verified.syscall_window_nmi: its issue(
issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md) wasopen since 2026-08-23 on one dev-host datum, but never given a redlist
row. Of the six orchestrator logs named for this task, three show the same
"N sprayed window arrivals against M in Ring 3 ... a 10x shortfall"
message. The other three named logs are not sightings of this defect and
are left out: two failed on an
unrelated "No space left on device" while writing the test boot image, and
one's failure carries a different message ("the storm never reported") on
a run the orchestrator's own summary marks
CONTAMINATED/ "Not a result"(another agent mutated the shared worktree while it ran). Deleted the
issue's now-false "Seen once" framing rather than rewriting it into a
story, and gave it
status: expected-redand the redlist row.No mechanism is added for
syscall_window_nmi; the issue already states thetwo unseparated readings (host contention vs. misclassification) and what
would separate them.
Test plan
cargo run -q -- --known-red user_copy_races_munmap-> disabledcargo run -q -- --known-red quiesce_leaves_the_volume_whole-> disabledcargo run -q -- --known-red syscall_window_nmi-> disabledcargo test -p toyos-build --lib: EXIT=0Guest tests were not run (agents never run QEMU); the orchestrator runs those.
🤖 Generated with Claude Code
https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j