diff --git a/issues/build/a-stopped-boot-whose-job-never-asked-waits-out-the-reset-budget-and-says-it-asked.md b/issues/build/a-stopped-boot-whose-job-never-asked-waits-out-the-reset-budget-and-says-it-asked.md new file mode 100644 index 00000000000..ef355adc122 --- /dev/null +++ b/issues/build/a-stopped-boot-whose-job-never-asked-waits-out-the-reset-budget-and-says-it-asked.md @@ -0,0 +1,23 @@ +--- +status: assigned +kind: tooling +opened: 2026-09-28 +--- + +# A stopped boot whose job never asked waits out the reset budget and says it asked + +`stopped_boot` (`tests/common/power.rs`) waits `qemu.budget(WAIT)` for QEMU's +`SHUTDOWN` event and then calls `returned_to_firmware`. That function turns +every `None` into `QEMU never reported stopping: the guest asked for a reboot +and stayed up`. So a boot whose job exited without asking for the reset waits +out the whole budget and then reports that the guest asked. + +In PR #566's fast tier at `74f7d717`, `quiesce_stops_the_machine`'s job exited 1 +at 10.303 s without asking. The guest's scheduler then reported both CPUs idle +(`current=None`, `parked=2`) from 10.750 s through its last heartbeat at +253.244 s. The test went red after 266 s with that message. + +**Exit**: `stopped_boot` stops waiting when the guest's job ends without +asking, and its red names that job's exit and last error line. A job that exits +before asking is then red in seconds, for its own reason. Owner: +`tests/common/power.rs`; held by the orchestrator. diff --git a/issues/build/quiesce-wakes-on-the-last-park-lost-its-serial-ready-beside-other-guests.md b/issues/build/quiesce-wakes-on-the-last-park-lost-its-serial-ready-beside-other-guests.md index 8af3bd54ce3..54d8ef7f738 100644 --- a/issues/build/quiesce-wakes-on-the-last-park-lost-its-serial-ready-beside-other-guests.md +++ b/issues/build/quiesce-wakes-on-the-last-park-lost-its-serial-ready-beside-other-guests.md @@ -13,7 +13,6 @@ until the stop waits on it alone`, `stop: 4 of 7 userland thread(s) stopped ... in 2010 ms of a 2010 ms budget`, `usb-quiesce: disk 0 SYNCHRONIZE CACHE ok` and `Rebooting.`, then `shutdown: /log did not answer in 2000ms`; the uart captured `nothing at all`. The harness's re-run alone was green in 2 s. -`cargo run -- --known-red` answers NO. The same shape as `issues/build/quiesce-wakes-on-the-last-exit-lost-its-serial-ready-beside-other-guests.md`, diff --git a/issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md b/issues/kernel/a-quiesce-writers-first-pass-outlasts-the-jobs-five-second-spin-up.md similarity index 54% rename from issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md rename to issues/kernel/a-quiesce-writers-first-pass-outlasts-the-jobs-five-second-spin-up.md index 25c97e68686..af7028ec9a1 100644 --- a/issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md +++ b/issues/kernel/a-quiesce-writers-first-pass-outlasts-the-jobs-five-second-spin-up.md @@ -4,11 +4,20 @@ kind: defect opened: 2026-09-25 --- -# quiesce_dump_holds_the_stopped reds wide, with two USB transport breaks, and is green alone +# A `quiesce_writers` writer's first write-and-fsync pass outlasts the job's 5 s spin-up -Seen once, in the fast tier on the logd branch (PR #492), dev host: red wide -with `QEMU never reported stopping: the guest asked for a reboot and stayed -up`, then `ALONE ... GREEN`. +`quiesce_dump_holds_the_stopped` and `quiesce_stops_the_machine` both boot +`quiesce_writers`. It asks for the reset only once each of its six writers has +finished one pass: a create, 64 KiB of writes and an fsync. If a writer is +still in its first pass after 5 s, the job prints `quiesce_writers: of 6 +writers reached their loop in 5s` and exits 1 without asking, so no stop +begins. Every sighting below then reads `QEMU never reported stopping: the +guest asked for a reboot and stayed up`; that misreport is +`issues/build/a-stopped-boot-whose-job-never-asked-waits-out-the-reset-budget-and-says-it-asked.md`. + +`quiesce_dump_holds_the_stopped`, in the fast tier on the logd branch +(PR #492), dev host: red wide with `QEMU never reported stopping: the guest +asked for a reboot and stayed up`, then `ALONE ... GREEN`. The boot's console shows the stick's transport breaking twice on `SCSI 0x2a` (`no answer in the status phase in 2000 ms`, recovered each time), then @@ -40,6 +49,24 @@ after their ` 0` lines. Writer 1 printed pass 33 at 6.905 s, its passes 93 to 306 ms apart. The branch changes no kernel or guest code. A red on a one-guest lane is not the other suites' load. +**`quiesce_stops_the_machine`, PR #566's fast tier at `74f7d717`.** Writer 5 +began its first pass at 2.170 s (`quiesce-writer: 5 0`) and ended it at +7.434 s (`5 1`), 5.264 s later. The other five had ended theirs by 3.738 s. At +6.874 s the job printed `quiesce_writers: 5 of 6 writers reached their loop in +5s`, and at 10.303 s the runner printed `===TEST_END test_rs_quiesce_writers +exit=1===`. The job never printed `6 writers are running; asking for the +reset`, and the capture has no `stop:` record. Red after 266 s. + +The same test, with the same harness message and no `stop:` record: + +- PR #524's fast tier at `235c5a5b`, load average 20 to 28: after 266 s, with + `quiesce_writers: 3 of 6 writers reached their loop in 5s`. +- PR #555's nightly at `d2656765` (run 36351950439, `guest (3)`): after 47 s, + with `4 of 6`. +- PR #511's merged head `a58abf50`: after 306 s, with a `usb-storage` transport + break on `SCSI 0x2a` that recovered. Whether its job printed the give-up + line was not recorded. + **Exit condition.** Re-enabled when a reproduction names what holds a writer's first write-and-fsync pass for over 5 s while another writer passes in under a third of a second, and the fix is shown against it. Owner: the `/log` write and diff --git a/issues/kernel/quiesce-stops-the-machine-stayed-up-beside-other-guests.md b/issues/kernel/quiesce-stops-the-machine-stayed-up-beside-other-guests.md deleted file mode 100644 index 436dffd5fe4..00000000000 --- a/issues/kernel/quiesce-stops-the-machine-stayed-up-beside-other-guests.md +++ /dev/null @@ -1,28 +0,0 @@ ---- -status: open -kind: finding -opened: 2026-09-26 ---- - -# `quiesce_stops_the_machine` stayed up after asking for a reboot, beside other guests - -Fast tier at `a58abf50` (PR #511's merged head; another worktree's FAT suite -ran on the host at the same time): `QEMU never reported stopping: the guest -asked for a reboot and stayed up`, after 306 s. Its capture holds the -writers' progress lines and a `usb-storage ... transport broke on SCSI 0x2a: -no answer in the status phase in 2000 ms` that recovered after one break. The -harness's re-run alone was green in 3 s (`stop: 9 of 9 userland thread(s) -stopped across 2 cpu(s) in 29 ms of a 2010 ms budget over 2 sweep(s)`), and so -was `cargo test --test toyos-build -- quiesce_stops_the_machine` alone -afterwards (EXIT=0). `cargo run -- --known-red` answers NO. - -Not shown: where the loaded run's stop went, since its capture has no `stop:` -record. - -**Exit**: the stop's record, or its absence, explained on a loaded run. - -Again in the fast tier on PR #524's branch at `235c5a5b`: the same `QEMU -never reported stopping: the guest asked for a reboot and stayed up`, after -266 s, and the harness's re-run alone was green. The host was loaded -throughout by another worktree's spinner at 397% CPU, with the load average -between 20 and 28. diff --git a/issues/kernel/quiesce-wakes-on-the-last-park-gave-up-on-one-thread-beside-the-held-one.md b/issues/kernel/quiesce-wakes-on-the-last-park-gave-up-on-one-thread-beside-the-held-one.md new file mode 100644 index 00000000000..5a9246a9d9b --- /dev/null +++ b/issues/kernel/quiesce-wakes-on-the-last-park-gave-up-on-one-thread-beside-the-held-one.md @@ -0,0 +1,66 @@ +--- +status: expected-red +kind: defect +opened: 2026-09-28 +--- + +# `quiesce_wakes_on_the_last_park`'s stop gave up on one thread beside the held one + +Three Fast tiers carry the identical failure: + +``` +FAIL quiesce_wakes_on_the_last_park: the stop gave up on 2 thread(s) that never reached a safe point: + stop: 3 of 5 userland thread(s) stopped across 2 cpu(s) in 2010 ms of a 2010 ms budget over 2 sweep(s), 0 of N userland block operation(s) still open; this reset lands wherever the other 2 are +``` + +- PR #539 at `2d6d228f`. +- PR #555 at `d2656765`. +- PR #559 at `ac948e6a`. + +The earliest is PR #510 at `98e803cb`, recorded in +`issues/build/quiesce-wakes-on-the-last-park-lost-its-serial-ready-beside-other-guests.md`: +`stop: 4 of 7 userland thread(s) stopped ... in 2010 ms of a 2010 ms budget`. +That boot then lost its READY, so the harness reported the READY and not the +stop. + +One of the two threads is the held thread by construction: `quiesce::last::hold` +yields until the latest sweep counts 1 running, so a sweep that counts 2 keeps +it spinning. The defect is the other thread. No sighting can name it, because +the `stop:` record carries only counts. + +**Hypothesis A, untested.** The hold's yield loop keeps its CPU busy. A Ready +thread queued on that CPU then runs only if `dispose_yield` +(`toyos-sched/src/cpu.rs`) re-inserts the spinner behind it. + +**Hypothesis B, untested.** `tests/quiescelastcase/system.toml` starts `logd`, +which fsyncs `/log`. `begin_update`'s only caller is `SYS_FSYNC` +(`kernel/src/object/ops.rs:626`), and that update spans every retry, including +the park in `between_attempts`. A thread parked there is `Blocked` with +`MID_UPDATE`; `stop_if_blocked` refuses it (`toyos-sched/src/task.rs:450`), the +sweep counts it running, and it adds 0 to `in_flight` — so `logd` parked +between refused fsync attempts past the 2010 ms budget is a candidate the +records cannot exclude, and it sits on the same `/log` fsync path the writers +issue owns. + +**What no enabled guest test checks while this is disabled.** +`woken_by_its_threads` (`tests/common/power.rs`) has no enabled caller. So no +enabled test checks any of these: + +- that a band, a park or an exit wakes the stop, rather than its deadline; +- `in_flight == 0` with `begun > 0`; +- the thread census; +- the `console-queue-at-the-stop` drain. + +`quiesce_refuses_a_second_shutdown` stays green over a lost post. It judges the +stop only by `stopped_the_machine`, so a stop that spends its budget and then +finds everything stopped passes it. + +**Exit**: + +- an instrument that names each thread still running when the stop gives up, + with its name, tid, cpu and scheduler state; +- the mechanism it names fixed; +- this test green in a Fast tier beside other guests; +- an enabled guest test checking each claim listed above. + +Owner: the stop path, `kernel/src/quiesce.rs`; held by the orchestrator. diff --git a/src/redlist.rs b/src/redlist.rs index 8a148a67b2e..ec48d2a75f5 100644 --- a/src/redlist.rs +++ b/src/redlist.rs @@ -62,12 +62,20 @@ pub const DISABLED: &[Disabled] = &[ }, Disabled { test: "quiesce_dump_holds_the_stopped", - issue: "issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md", + issue: "issues/kernel/a-quiesce-writers-first-pass-outlasts-the-jobs-five-second-spin-up.md", + }, + Disabled { + test: "quiesce_stops_the_machine", + issue: "issues/kernel/a-quiesce-writers-first-pass-outlasts-the-jobs-five-second-spin-up.md", }, Disabled { test: "quiesce_wakes_on_the_last_exit", issue: "issues/build/quiesce-wakes-on-the-last-exit-lost-its-serial-ready-beside-other-guests.md", }, + Disabled { + test: "quiesce_wakes_on_the_last_park", + issue: "issues/kernel/quiesce-wakes-on-the-last-park-gave-up-on-one-thread-beside-the-held-one.md", + }, Disabled { test: "sched_check_build", issue: "issues/build/the-pass-cost-gates-ci-sample-is-eight-days-stale-twice.md",