diff --git a/.github/pull_request_template.md b/.github/pull_request_template.md index 6b3972cceb5..8e360993f59 100644 --- a/.github/pull_request_template.md +++ b/.github/pull_request_template.md @@ -26,6 +26,6 @@ command. An estimate says it is one. --> ## Anything a reader of `main` must not miss diff --git a/CLAUDE.md b/CLAUDE.md index 3b300578402..577449ced1b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -66,7 +66,7 @@ The bar is not yet the tree. The standing failures are declared rather than remo ## Build & test -The testing rules live where they are enforced: instruments and known reds in `src/redlist.rs`, tiers in `src/tiers.rs`, the PR gate and the nightly in `.github/workflows/`. Operationally: +The testing rules live where they are enforced: known reds in `src/redlist.rs`, tiers in `src/tiers.rs`, the PR gate and the nightly in `.github/workflows/`. Operationally: - `cargo run` builds everything (toolchain, kernel, bootloader, userland, image) and launches QEMU; `--build-only` skips the launch. `cargo test` runs the QEMU harness; `cargo test --workspace --exclude toyos-build` runs every host-crate suite. - **Agents verify through `cargo test`, never `cargo run`** — the run path opens a QEMU window on the owner's desktop by design; the harness runs headless. @@ -119,7 +119,7 @@ system.toml What to build and boot - **Every written number comes from a command that was run.** An estimate or datasheet bound says so. Write commit messages with `git commit -F `, never `-m` — a double-quoted `-m` substitutes backticks and the shell runs them. - **Commit freely on your branch; land through a pull request.** `main` moves only through a merged PR, and `cargo run -- --pr` is the whole local half. `gh pr create --draft` at the first push — CI runs on PRs and nothing else; `gh pr ready` plus a written `--title`/`--body-file` when finished (never `--fill`); `gh pr merge --auto --merge` enqueues on `main`'s required merge queue, which builds each merge's exact composition and runs the required checks on it before `main` moves; `cargo run -- --sync` after it lands. Never merge into `main` by hand; `gate-stage` reads the protection back. The PR's title and body become the merge commit's: write them as main's record. A modify/delete conflict is resolved by accounting for every hunk of the modified side, never by checking its headings survived. A merge that deletes a document also deletes every citation to it in the same merge, checked by searching the bare name as well as the path. An ABI change lands with the work that needs it: every worktree builds the toolchain its own sources name, so branches that change the ABI run side by side and none waits on another. Every merge leaves `main`'s tip compiling. A branch lands after a review against `.claude/agents/reviewer.md`, spawned by the orchestrator with its brief and judged by it. - **Never rewrite history, and never touch `main`.** No `--amend`, no `rebase`, no `--force` — on your own branch as much as anywhere: a pushed hash may already be cited. `main` is protected — PR required, no force-push, no deletion, no bypass. -- **A red is known only if `cargo run -- --known-red ` says so** (`src/redlist.rs`). A PR red not about the author's diff is adjudicated there and fixed at its owner, never re-run away. +- **A red test is a defect unless `src/redlist.rs` disables it with its issue (`cargo run -- --known-red `); a flaky test is disabled at once, never re-run.** - **A high-risk change names its two checks.** Security boundaries, the scheduler, the ABI, filesystems, devices, memory management, concurrency primitives: the PR names the negative control or mutation that fails if the implementation is wrong, and one epistemically independent oracle — an external specification, a differential implementation, real hardware, a third-party checker, a formal model, or a recorded real failure. A second agent is not independence: five artifacts from one wrong model still agree. A mutation is a negative control only if it reverts the *whole* change onto the base the green arm was measured on — a one-line revert of a change that moved two things measures neither. - **Host load is not an excuse.** A load-coincident audio failure is investigated as a real defect, never re-run away as noise; evidence against that assumption goes to the owner, not into quiet workarounds. - **Subagents wait in the foreground** — background notifications do not reliably re-wake them: explicit `timeout`s, and for longer work background once and block with a few long foreground waits, polling before each sleep. diff --git a/issues/README.md b/issues/README.md index 3954ced5c9b..646c3154d8f 100644 --- a/issues/README.md +++ b/issues/README.md @@ -27,7 +27,7 @@ Four fields, three required, no defaults. |---|---|---| | `status` | `open` | it is work, and nobody is holding it | | | `assigned` | it is work, and somebody is — the body says who or which task | -| | `expected-red` | a test fails on this today and `src/redlist.rs` quarantines it | +| | `expected-red` | a test fails on this today and `src/redlist.rs` disables it | | | `owner` | it is the owner's to decide, and nobody else may | | | `none` | nothing is owed | | `kind` | `defect` | real, reproducible, someone should fix it | @@ -138,7 +138,7 @@ checkout (`git grep `): `rg` skips dotfile directories without `--hidden`, and `.github/` holds citations too. Then read where the hits are. One in a comment under `toyos-abi/src`, `toyos/src` or a published crate changes no identity (`src/identity.rs`), so it owes no version and builds no sysroot. One in -`src/redlist.rs` is a quarantine row's `issue`: the row goes with the file. +`src/redlist.rs` is a disabled test's `issue`: the row goes with the file. ## Two area notes, carried over from the file this replaced diff --git a/issues/audio/doom-sound-flood-played-full-scale-once.md b/issues/audio/doom-sound-flood-played-full-scale-once.md index 13dbd938551..5349b6284c4 100644 --- a/issues/audio/doom-sound-flood-played-full-scale-once.md +++ b/issues/audio/doom-sound-flood-played-full-scale-once.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-09-04 --- @@ -36,7 +36,6 @@ whether it is the mixer's sum overflowing or the analysis reading a wrapped value, and whether a listener would hear it. Nothing in the captured line distinguishes those, and the WAV that would is not kept by a CI job. -**Exit condition.** A run that reproduces the peak with the capture retained, -and one sentence naming which of the three it is. `src/redlist.rs` carries the -rate (1 of the 3 nightlies of that week); the two rows for this name that were -about `timed out after 88s` are retired on the change of shape. +**Exit condition.** The full-scale sample's cause is fixed, and +`doom_sound_flood` green on the KVM `guest` shards where it went red. Owner: +orchestrator. diff --git a/issues/audio/hda-tone-phase-check.md b/issues/audio/hda-tone-phase-check.md index 7e6a8a4995b..3daec584bb9 100644 --- a/issues/audio/hda-tone-phase-check.md +++ b/issues/audio/hda-tone-phase-check.md @@ -7,13 +7,10 @@ task: 88 # HDA: the captured tone is not one sine -**This is the write-up `tests/toyos.rs`'s `EXPECTED_FAILURES` names for `hda_tone`.** Do not delete it without moving that pointer. - `hda_tone` plays the same 3.0 s 440 Hz tone the virtio arm plays, out of an `intel-hda` controller soundd drives itself, and the capture comes back with -**8 to 16 phase discontinuities** where the virtio arm has none. Declared in -`EXPECTED_FAILURES` against the message "the captured tone is not one sine"; -every other assertion that test makes still reds the run. +**8 to 16 phase discontinuities** where the virtio arm has none. Every other +assertion that test makes still reds the run. What is *not* wrong, measured on this host (QEMU 11.0.3, 2026-08-07): the tone is present at full amplitude, there is **no mid-tone silence at all** (`gaps @@ -103,3 +100,7 @@ fresh one holds 139,253-142,325 of the 144,256 submitted frames, so the adjacent-frame *pairs* with |period| in the hundreds, not the 118-frame clusters. And the load dependence is sharp where it used to be a correlation: 0 of 8 alone against 3 of 11 beside other guests, same tree, same hour. + +**Exit condition.** The adjacent-frame-pair breaks' cause is fixed, and +`hda_tone` reads 0 phase breaks beside other guests on the dev host and on +CI's KVM shards. Owner: orchestrator. diff --git a/issues/audio/hda-tone-red-beyond-its-exemption.md b/issues/audio/hda-tone-red-beyond-its-exemption.md index 2ed6399a795..5cd53a280d9 100644 --- a/issues/audio/hda-tone-red-beyond-its-exemption.md +++ b/issues/audio/hda-tone-red-beyond-its-exemption.md @@ -15,23 +15,11 @@ FAIL hda_tone: 1 mid-tone silences in the capture: total 1 [1p×1] the entry covers ["the captured tone is not one sine"] ``` -The `EXPECTED_FAILURES` entry does what it is supposed to: it pins the assertion -rather than the test, so a *second* defect in the same test still reds the run -and says which. What is red is the mid-tone-silence assertion — a gap in the -capture, which is gate A's harm verdict — and not #88's spectral one. The entry -still says so at the site: it names only `"the captured tone is not one sine"`, -beside a comment listing "no mid-tone silence" among the things that red the run -"because each of those is the milestone rather than the open question". +**What has changed since, and what has not.** `hda_tone` is `Tier::Nightly` for +`Why::TimerAnchored` (`src/tiers.rs`), so a plain `cargo test` no longer runs +it and a landing whose gate is `cargo test` no longer meets this red at all. -**What has changed since, and what has not.** Two things narrow the harm this -was filed for. `hda_tone` is `Tier::Nightly` for `Why::TimerAnchored` -(`src/tiers.rs`), so a plain `cargo test` no longer runs it and a landing whose -gate is `cargo test` no longer meets this red at all; and `src/redlist.rs` -carries rows for the name — the dev-host-alone one sourced here retired on -3 of 3 green alone on 2026-09-04, and one at 4 of 5 on CI — so an agent who -does meet it is told whose it is rather than reading it as theirs. - -Neither touches the verdict, and the verdict was owed a fresh sample: every +That doesn't touch the verdict, and the verdict was owed a fresh sample: every capture behind it had gone through QEMU's 48000→44100 resampler, since removed. **Re-judged 2026-08-29, and it stands.** On the current instrument (QEMU diff --git a/issues/audio/idle-suspend-reds-on-a-loaded-host-and-on-main.md b/issues/audio/idle-suspend-reds-on-a-loaded-host-and-on-main.md index 4c71c1d9009..aa6a51815fd 100644 --- a/issues/audio/idle-suspend-reds-on-a-loaded-host-and-on-main.md +++ b/issues/audio/idle-suspend-reds-on-a-loaded-host-and-on-main.md @@ -53,11 +53,6 @@ nothing could see an idle wake when this was filed. The 2026-08-29 section below is that instrument existing; a red taken since then decides the split by itself. -Not on `src/redlist.rs` — `cargo run -- --known-red audio_idle_suspend` answers -`NOT ON THE LIST`. Adjudicating it there needs the rate above and a decision -about whether the row records a soundd defect or a `Sched::Parallel` -misclassification, and those are not the same row. - Gate A is unaffected and green throughout: `audio_tone` and `audio_tone_load` at smp=1 and smp=8 all pass on `cf72c3dc` with 440.0 Hz, phase-breaks 0, gaps none, 0 underruns and 0 drains. diff --git a/issues/boot-media/kernel-log-file-reds-beside-other-guests-and-is-green-alone.md b/issues/boot-media/kernel-log-file-reds-beside-other-guests-and-is-green-alone.md index 894af150e25..9c2557cf80c 100644 --- a/issues/boot-media/kernel-log-file-reds-beside-other-guests-and-is-green-alone.md +++ b/issues/boot-media/kernel-log-file-reds-beside-other-guests-and-is-green-alone.md @@ -24,11 +24,7 @@ paid at 1.55x width ``` `cargo run -- --known-red kernel_log_file` answered `NOT ON THE LIST` when this -was filed and answers `KNOWN-RED` now: `src/redlist.rs` carries a row for the -name, sourced here, `Finding::Seen` with `no rate`. So the next reader of the -name gets the sighting instead of silence — but the row records the same single -observation this file does, and a `Seen` with no rate is what it says it is. -What is owed is unchanged and is the rate. +was filed. What is owed is unchanged and is the rate. **The company is recorded, because the runner is the instrument.** A second agent's `cargo test --workspace` was running in `toyos-banner` on the same @@ -50,7 +46,5 @@ deletion of an unreachable `if rflags & TF != 0` branch in ## Promoted 2026-08-25 -A known-red test with a `Seen`-and-no-rate row is real, owed work: the -`src/redlist.rs` row still records the single observation this file does. Owed to whoever next runs a session free to measure the rate against an unchanged tree. diff --git a/issues/boot-media/log-flush-retry-two-older-failure-modes-with-no-home.md b/issues/boot-media/log-flush-retry-two-older-failure-modes-with-no-home.md index cef96988ee4..be1ba2f83d2 100644 --- a/issues/boot-media/log-flush-retry-two-older-failure-modes-with-no-home.md +++ b/issues/boot-media/log-flush-retry-two-older-failure-modes-with-no-home.md @@ -35,7 +35,7 @@ is known fixed by that work. Re-measure `log_flush_retry` (wide and alone) enough times, on a tree that carries ROOT-in-memory, to say whether either mode still occurs; if seen -again, a `src/redlist.rs` row carries it with a measured rate, or it is fixed -at its owner (`kernel/src/drivers/xhci` for the transport line, the boot +again, it is disabled with a `src/redlist.rs` row and this file, or it is +fixed at its owner (`kernel/src/drivers/xhci` for the transport line, the boot harness for the timeout). If not seen in that many runs, this file closes with the count that supports it. diff --git a/issues/boot-media/screen-fatal-halt-reds-on-ci-with-a-usb-storage-transport-break-during-boot.md b/issues/boot-media/screen-fatal-halt-reds-on-ci-with-a-usb-storage-transport-break-during-boot.md index cbfc0bc7503..78faa7317de 100644 --- a/issues/boot-media/screen-fatal-halt-reds-on-ci-with-a-usb-storage-transport-break-during-boot.md +++ b/issues/boot-media/screen-fatal-halt-reds-on-ci-with-a-usb-storage-transport-break-during-boot.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-09-05 --- @@ -24,8 +24,12 @@ Two facts the log holds. test-runner was spawned at 0.635 s and its first act is to print `===READY===`, which never reached the console in 31 s. The SYNCHRONIZE CACHE (0x35) the boot issued to the USB disk got no status-phase answer in 2000 ms and the transport was declared broken at 2.718 s. Whether the -second stalls the first is the question; the redlist row -(`src/redlist.rs`, `screen_fatal_halt`, `Instrument::Ci`) records the one -observation, and this file is its owner's: the usb-storage wait path -(`kernel/src/drivers/xhci/wait/msc.rs`) and whatever the runner's stdout -blocked on. +second stalls the first is the question, and this file is its owner's: the +usb-storage wait path (`kernel/src/drivers/xhci/wait/msc.rs`) and whatever the +runner's stdout blocked on. + +**Exit condition.** Re-enabled when a reproduction pins whether the boot +runner's stdout was blocked behind the usb-storage transport break or the two +are independent, and the wait path stops holding the runner's own output +hostage. Owner: the usb-storage wait path, +`kernel/src/drivers/xhci/wait/msc.rs`; held by the orchestrator. diff --git a/issues/build/a-defect-only-contention-exposes-is-classified-as-a-wrong-sched.md b/issues/build/a-defect-only-contention-exposes-is-classified-as-a-wrong-sched.md deleted file mode 100644 index 3a63445eca4..00000000000 --- a/issues/build/a-defect-only-contention-exposes-is-classified-as-a-wrong-sched.md +++ /dev/null @@ -1,36 +0,0 @@ ---- -status: open -kind: tooling -opened: 2026-09-26 ---- - -# A defect only contention exposes is classified as a wrong `Sched::Parallel` - -When a test fails in the wide run and passes in the lone re-run, and the wide -run shared the host, `alone_line` (`tests/toyos.rs`) prints one verdict: -`GREEN — it fails only beside other guests, so its Sched::Parallel is wrong`. -That names a cause. What the two runs establish is only that the failure needs -the timing a loaded host produces, and a kernel race that only a slow or -preempted vCPU opens produces exactly that pair. - -Recorded case, PR #524's fast tier at `f67863c2`: 270 passed, 130 failed, and -254 of the failures were one kernel panic, `vconsole: no tx slot`. Each of the -130 was green alone, so each got the scheduling verdict, 130 times. The cause -was one kernel defect: `serial::BackendGuard::try_lock` built, and so dropped -and unlocked, a guard for the CPU that lost the exchange (fixed in -`1e5ef5c1`). Changing any test's `Sched` would have hidden it. - -What reads it wrong: - -- The per-test verdict names a cause the evidence does not decide. The - `tests/CLAUDE.md` caveat already says it is a hypothesis; the line itself - says it is a finding. -- Nothing groups the failures. Tests that fail wide on one shared headline (the - same panic message, `toyos_build::alone::same_failure`) are one finding, and - the report prints them as N separate classifications. - -**Exit condition**: the shared-host green arm states the observation (fails -only beside other guests) without naming `Sched` as the cause, and the run's -summary groups wide failures that share a headline, printing the group's count -before any per-test classification. The staged pair for the grouping is the -`f67863c2` shape: many tests, one panic headline. diff --git a/issues/build/console-line-atomicity-loses-five-of-a-thousand-lines-on-ci.md b/issues/build/console-line-atomicity-loses-five-of-a-thousand-lines-on-ci.md new file mode 100644 index 00000000000..432fda64247 --- /dev/null +++ b/issues/build/console-line-atomicity-loses-five-of-a-thousand-lines-on-ci.md @@ -0,0 +1,25 @@ +--- +status: expected-red +kind: defect +opened: 2026-08-20 +--- + +# `console_line_atomicity` loses five of a thousand lines on CI, one guest per machine + +First sighting on the CI instrument: PR #166 run `32364721784`, `guest (10)`, +`writer A declared 1000 whole lines and the capture carries 995`, `ALONE: +GREEN` in the same job. The diff it rode on is an issues-and-prose audit, so +the tree is not a suspect. + +CI runs one guest per machine with `--jobs 1`, so whatever loses five of a +writer's thousand lines there is not host contention between suites — which +sharpens the question the test was built to ask rather than settling it: the +loss is inside one guest's own console path. + +Split out of `issues/build/parallel-tests-red-under-other-suites.md`, whose +rate table was never this test's — CI's single-guest-per-machine shards rule +out the contention shape that file is about. + +**Exit condition.** The lost lines' cause is fixed, and +`console_line_atomicity` green on CI's `guest` shards, one guest per machine. +Owner: orchestrator. diff --git a/issues/build/latency-wake-reds-on-the-dev-host-at-a-rate.md b/issues/build/latency-wake-reds-on-the-dev-host-at-a-rate.md index b59e4c1a63e..299b2482cba 100644 --- a/issues/build/latency-wake-reds-on-the-dev-host-at-a-rate.md +++ b/issues/build/latency-wake-reds-on-the-dev-host-at-a-rate.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: finding opened: 2026-09-07 --- @@ -28,9 +28,11 @@ A seventh sighting, in the whole-branch review's twelve-wide `cargo test` on 7294acef: `258 past the 4096us histogram`, and the harness's own isolated re-run green. -So this is a rate on this host, not a classification and not a regression. The -rate is on the list — two `src/redlist.rs` rows under this name, `FIRES 1 of 6` -alone and one `SEEN` under load, both citing this file. What is still owed is -the other reading of the same evidence: whether cyclictest's 4,096-bucket -histogram is simply too low for a TCG guest on a loaded laptop, which is a -change to the instrument and not to the kernel. +What is still owed is the other reading of the same evidence: whether +cyclictest's 4,096-bucket histogram is simply too low for a TCG guest on a +loaded laptop, which is a change to the instrument and not to the kernel. + +**Exit condition.** Re-enabled when the base's p99 is shown to genuinely +exceed 4096 us and that is fixed at the timer-interrupt entry the deadline arm +already touches. Owner: the boot-deadline work (`kernel/src/sched`'s +timer-interrupt entry); held by the orchestrator. diff --git a/issues/build/parallel-tests-red-under-other-suites.md b/issues/build/parallel-tests-red-under-other-suites.md index 3d120889eed..04047234b5c 100644 --- a/issues/build/parallel-tests-red-under-other-suites.md +++ b/issues/build/parallel-tests-red-under-other-suites.md @@ -68,31 +68,28 @@ changes. the profile used to seat a second desktop beside, `desktop_window_child`, is `Tier::Nightly` and so never in a pull request's parallel phase. - **`desktop_locale_detect`** — retired 2026-09-04, green 5 of 5 beside a full - fast tier; `src/redlist.rs` carries the runs. + fast tier. - **`netd_connection_caps`** — retired 2026-09-04, green 5 of 5 beside a full - fast tier; `src/redlist.rs` carries the runs. + fast tier. - **`metal_sim_pointer_churn`** — observed once, on a host carrying three other suites *and* a `toyos-sched-sim` run. Not investigated. Still `Sched::Parallel`. - **`dump_nmi_probe`** — retired 2026-09-04, green 3 of 3 beside a full fast - tier; `src/redlist.rs` carries the runs. `4ad8875` made it `Sched::Serial`, + tier. `4ad8875` made it `Sched::Serial`, which shows what serialising buys and what it does not: within one run the phase is quiet, across runs nothing but `buildlock::guest_slot` spans worktrees and twelve slots is not one guest. - **`blocked_dump`** — retired 2026-09-04, green 3 of 3 beside a full fast - tier; `src/redlist.rs` carries the runs. + tier. - **`screen_console_scroll`** — retired 2026-09-04, green 3 of 3 beside a full - fast tier; `src/redlist.rs` carries the runs. + fast tier. - **`hda_tone`** — added 2026-08-07, hours after the test itself landed. In a full run on a host carrying another worktree's suite: `2 mid-tone silences in the capture: total 2 [3p×1 4p×1]`, `dither 3.3%`, `phase-breaks 92`. Alone on the same tree eight minutes later: `gaps none`, `phase-breaks 16` — the declared #88 failure and nothing else. It is `Sched::Serial`, so like `dump_nmi_probe` the harness never re-runs it alone and the run simply reds. - Its `EXPECTED_FAILURES` entry covers the phase-break message alone, which is - why a *dropout* under load reaches the verdict, and that is correct: **do not - widen it.** A silence and a phase break are two different defects and an entry - that covered both would stop saying anything. The tree it was seen on differed + The tree it was seen on differed from main only in `src/`, so the guest image was byte-identical to main's. **Three times the same day**, all three in landing gates of that one build-system branch and all three confirmed alone within ten minutes: `2 @@ -103,7 +100,7 @@ changes. the branch last merged, which reads as the branch's own work and is not. - **`xhci_hid_break`** — retired 2026-09-04, green 3 of 3 beside a full fast - tier; `src/redlist.rs` carries the runs. It is one of the three longest jobs + tier. It is one of the three longest jobs in the suite by `longest_first`'s own profile, so it is dispatched early and runs beside everything. @@ -147,9 +144,7 @@ changes. immediately afterwards: `PASS handle_kill_policy (615ms)`. A third sighting of the same census, on a third unrelated branch, is what the mechanism above predicts — and the three together are why it is no longer only this file's - record: `src/redlist.rs` carries an `Instrument::Ci` row for - `handle_kill_policy` as of the CI sighting, so `cargo run -- --known-red - handle_kill_policy` now answers it. + record. - **`wall_clock_file`** — added 2026-08-17, same session, **1 of 6**, `ALONE … GREEN`, green on all twelve shards of the same tree. Not @@ -220,8 +215,7 @@ changes. record of which children died and which calls their parent made. - **`console_line_atomicity`** — added 2026-08-20, the name's first sighting - on the CI instrument (its standing rows are the loaded dev host's, 1 of 3 - there): PR #166 run 32364721784, `guest (10)`, `writer A declared 1000 + on the CI instrument: PR #166 run 32364721784, `guest (10)`, `writer A declared 1000 whole lines and the capture carries 995`, `ALONE: GREEN` in the same job. CI runs one guest per machine, so whatever loses five of a writer's thousand lines there is not host contention — which sharpens this file's @@ -275,7 +269,7 @@ changes. yet. `ALONE … GREEN` in **5 s** in the same session, reporting the storm in full: `3000 sent, 3000 taken, 43 in the window, 140 in Ring 3, 663 syscalls made under the storm`. `cargo run -- --known-red syscall_window_nmi` answered - `NOT ON THE LIST` when it was filed; `src/redlist.rs` carries a row now. + `NOT ON THE LIST` when it was filed. **Not the branch it was found on**: that branch changed the syscall entry's displacement *spelling* — `const` operands for the same immediates, @@ -415,8 +409,7 @@ by construction, since a gate's builds are these builds. What it does **not** bound is anything that never enters `src/build.rs` — a `toyos-sched-sim measure`, a hand-run `cargo build` in a fork clone, the primary's `./x.py`. -**What to do about a red on any of these names:** read the `ALONE` line under it -before anything else. `GREEN` there means the host, not the kernel. What none of +What none of them should get is a widened bound — a gate that tolerates one lost byte tolerates the defect it was written for. The two fixes above are the two shapes that are legitimate: make the verdict independent of the rate, or scale a @@ -426,8 +419,8 @@ admits twelve guests across every worktree, so the four-suite regime these were observed in cannot recur. A looser assertion is still not the answer. -**But `ALONE … red again — the defect is real` is not evidence, and the protocol -above leans on it.** The re-run happens inside the same process, moments after +**But `ALONE … red again — the defect is real` is not evidence.** +The re-run happens inside the same process, moments after twelve guests have been torn down and while another worktree's suite may still own the host — so it is alone in the suite's bookkeeping and not on the machine. Measured 2026-08-06 on the xHCI port-machine branch, whose kernel delta is @@ -452,12 +445,6 @@ re-run included. A verdict that flips between "GREEN, it is the host" and "red again, the defect is real" for one test on one tree twenty minutes apart is measuring the host in both directions. -Consequence for the protocol: `ALONE: GREEN` still means what it says, because a -green cannot be produced by load. `ALONE: red again` means nothing on its own -and must be confirmed against `main` in the same session before it is believed — -which is the A/B the audio rules already require and which this line currently -invites an agent to skip. - **2026-08-23 — the host-speed correction was blind to wide-SMP oversubscription, and now is not.** Each CI `guest` shard is its own four-core `ubuntu-24.04` runner (AMD EPYC, nested KVM) running one guest at `--jobs 1`, so there is no @@ -517,7 +504,7 @@ describes. `ALONE: GREEN` both times, and green again when re-run alone by hand (3 s, 2 s). So the name flakes for two different reasons and only one of them is the ceiling this section corrects; what a contended host does to `/system/bin/init`'s handle accounting on a *refused* launch is not explained here, and nobody has a -mechanism for it. `src/redlist.rs` carries the sighting. +mechanism for it. - **`i8042_undecoded_bytes`** — added 2026-09-07 on the metal branch's pre-pull-request fast tier over the merged tip `8d895a15`, one sighting: diff --git a/issues/build/quiesce-wakes-on-the-last-exit-lost-its-serial-ready-beside-other-guests.md b/issues/build/quiesce-wakes-on-the-last-exit-lost-its-serial-ready-beside-other-guests.md index 2c18c525d26..03940d2f0fe 100644 --- a/issues/build/quiesce-wakes-on-the-last-exit-lost-its-serial-ready-beside-other-guests.md +++ b/issues/build/quiesce-wakes-on-the-last-exit-lost-its-serial-ready-beside-other-guests.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: finding opened: 2026-09-26 --- @@ -14,18 +14,21 @@ sweep(s)`, `usb-quiesce: disk 0 SYNCHRONIZE CACHE ok`, `Rebooting.` But the uart captured `nothing at all`. Before it, `usb-storage: 00:02.0 slot 1 transport broke on SCSI 0x28: no answer in the data phase in 2000 ms`, and the test-runner's spawn reported `layout=2069ms`. The harness's re-run alone was -green in 3 s, 2 sweeps. `cargo run -- --known-red` answers NO. +green in 3 s, 2 sweeps. Not shown: why the uart saw none of the test-runner's output in a boot whose console carried the whole stop. The READY wait reports the missing marker, not its cause. -**Exit**: a cause for the empty uart on a boot that rebooted as designed, or -the marker waited for where the boot's reboot cannot race it. - Again in the fast tier at `8846c021` (PR #524's branch, alone on the host): the same `QEMU died before ===READY=== (status: exit 0)` with `uart: nothing at all`, after `stop: 5 of 5 userland thread(s) stopped`, `usb-quiesce: disk 0 SYNCHRONIZE CACHE ok` and `Rebooting.`; this time no usb-storage transport -break before it. The re-run alone was green. `cargo run -- --known-red` -answers NO. +break before it. The re-run alone was green. + +On main's lineage, with `the stopped-boot drain carried no kernel output at all +(38 bytes)`: f231c43e (run 36280285913, `guest (3)`), 67a430c8 (run +36285169430) and a55d62c6 (run 36287592139). + +**Exit**: the marker waited for where the boot's reboot cannot race it, and the +test green in a fast tier beside other guests. Owner: orchestrator. diff --git a/issues/build/the-alone-classifier-cannot-see-a-reading-with-no-unit.md b/issues/build/the-alone-classifier-cannot-see-a-reading-with-no-unit.md deleted file mode 100644 index 057925c7c21..00000000000 --- a/issues/build/the-alone-classifier-cannot-see-a-reading-with-no-unit.md +++ /dev/null @@ -1,50 +0,0 @@ ---- -status: open -kind: tooling -opened: 2026-09-18 ---- - -# Four assertions render a reading with no unit beside it, so the `ALONE:` line reports one of them twice as two different failures - -`src/alone.rs` decides whether the harness's isolated re-run found the failure -the wide run found. It takes a number out of a sentence only when a unit follows -it, because a number with no unit is an identity — `slot 1` against `slot 2`, an -opcode, a CPU, a count — and two identities are two observations, which is the -larger of the two findings. That rule errs toward "different" on purpose, and -this is the price: an assertion that prints its reading *without* a unit has -that reading read as an identity, so one defect reproduced at two readings is -reported as two failures. - -Owned by `src/alone.rs` and the `ALONE:` line in `tests/toyos.rs` that calls it; -nobody is holding it. - -## The four sites in this tree - -Each writes a headline, so each reaches the classifier. - -| site | what it renders | two runs read as | -|---|---|---| -| `tests/common/audio.rs`, `check_physical`'s wake-lateness fault | the same reading twice — `{}us` and `({:.1} pipeline depths)` | two failures: only the `us` copy is masked | -| `tests/toyos.rs`, the dither floor | `only {:.1}% of silent samples are non-zero` | two failures | -| `tests/toyos.rs`, the tone peak | `tone too quiet: peak {}` | two failures | -| `tests/toyos.rs`, the log-drain verdict | `stops at {} bytes` — `bytes` spelled out is not `B` | two failures | - -`src/alone.rs`'s `a_reading_rendered_twice_keeps_the_copy_with_no_unit`, -`two_percentages_are_two_failures`, `two_bare_counts_are_two_failures` and -`bytes_spelled_out_is_not_a_unit` assert exactly this, so the limit is pinned -rather than latent: changing it reds those tests and lands as a decision. - -## Exit - -Either of these closes it, and the four tests above are rewritten or deleted by -whichever does: - -- each of those four assertions renders its reading with a unit the scan already - knows — `bytes` becomes `B`, the pipeline-depth rendering carries one or goes, - and the percentage and the peak print the quantity they measure; or -- the classifier stops taking an assertion's identity from its text and takes it - from the assertion's source location instead, which makes the units list and - this whole class unnecessary. - -The second is the real answer and the first is what one landing can do. Neither -has an owner yet. diff --git a/issues/build/the-console-input-path-can-stop-after-a-ps2-overflow.md b/issues/build/the-console-input-path-can-stop-after-a-ps2-overflow.md index e45e0abe74f..bcb37ef76a9 100644 --- a/issues/build/the-console-input-path-can-stop-after-a-ps2-overflow.md +++ b/issues/build/the-console-input-path-can-stop-after-a-ps2-overflow.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-09-01 --- @@ -58,8 +58,7 @@ controller" from "the console stopped reading the kernel's queue", and it needs denominator**: the same STALL sentence, **891 s**, `ALONE: GREEN — it fails only beside other guests`, in **1 of 6** full `cargo test` fast tiers run in one worktree that day, each with single-test runs of the same suite beside - it. Green in the three nightly `ci` runs of the same week. `src/redlist.rs` - carries it as this name's first `DevHostLoaded` row. + it. Green in the three nightly `ci` runs of the same week. The two do not reconcile at a common rate: at the 2-of-5 arm's own p = 0.4, P(0 of 20) is 3.66e-05. **The tree is not the difference** — the branch's @@ -72,9 +71,9 @@ available on a wedged boot, and the reported `drained 0` *while bytes were still being injected* is the ISR side having stopped taking them off the controller — not the console having stopped reading the kernel's queue. -Exit: reproduce a wedge with the counter visible and read whether `RX_BYTES` is -still rising while the panel is frozen. One number decides it, and it is -already within reach of the armed arm rather than blocked on a new instrument. +**Exit condition.** The input path is fixed so that no PS/2 overflow stops +it, the staging above takes every line after the first on twenty boots, and +`console_locale_detect` is green beside other guests. Owner: orchestrator. Not a reason to leave `shell_type_once` unpaced: pacing keeps the harness from provoking this, and the tracker keeps the defect. diff --git a/issues/build/the-pass-cost-gates-ci-sample-is-eight-days-stale-twice.md b/issues/build/the-pass-cost-gates-ci-sample-is-eight-days-stale-twice.md index 8e268f16a72..b322bbd081d 100644 --- a/issues/build/the-pass-cost-gates-ci-sample-is-eight-days-stale-twice.md +++ b/issues/build/the-pass-cost-gates-ci-sample-is-eight-days-stale-twice.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: tooling opened: 2026-09-13 --- @@ -34,11 +34,9 @@ further crossing costs a landing an hour and answers nothing. 2026-09-01 — every `guest` shard that ran this name, with the runner's reported CPU model beside each reading — and the line re-derived from it with the same rule; if the readings split by runner model, the gate names -the model it judges and refuses to judge the rest. Until then, a CI red -under this name is this file and the redlist row, and the landing it -dequeues is re-queued once, not re-run until green. +the model it judges and refuses to judge the rest. Owner: orchestrator. -A second signature, which this file's quarantine row does not cover. The fast +A second signature: the fast tier on PR #524's branch at `235c5a5b` reds `sched_check_build` on `cpu0: 85 passes … a 90th percentile needs at least 100 samples behind it and this has 85`. The host was loaded then by another worktree's spinner at 397% CPU. The diff --git a/issues/build/the-shard-split-prices-a-boot-and-not-the-image-behind-it.md b/issues/build/the-shard-split-prices-a-boot-and-not-the-image-behind-it.md index 0087d0ce121..032f92fe7c5 100644 --- a/issues/build/the-shard-split-prices-a-boot-and-not-the-image-behind-it.md +++ b/issues/build/the-shard-split-prices-a-boot-and-not-the-image-behind-it.md @@ -120,13 +120,6 @@ runs carry main's `d44b4978`, so they are not the same tree as the first three which is why the claim here is about the instrument's spread and not about any tree.) -**The mechanism, measured by bundle 16 on 2026-09-03**: `sysret_ss_reload`'s -probe line (`sysret-ss: reloaded`) lands BEFORE `===READY===`, so `boot_log()` -already holds it and `drain_until`'s predicate can only time out — the ceiling -(10 s scaled by width and host speed) is spent on every run: 334, 452, 388, -352, 433 s across five tiers. The fix is the test's (probe after READY, or -check `boot_log` first); this record is tooling. - So a taker needs, in order: the build clock keyed by config rather than by thread; a committed per-config profile that only `Shard::keep` reads, merged the way `cargo run -- --merge-durations` merges the test profile; and **two diff --git a/issues/hardware/a-collapsed-replug-is-enumerated-only-when-another-port-event-arrives.md b/issues/hardware/a-collapsed-replug-is-enumerated-only-when-another-port-event-arrives.md index c1ae44693af..4fbcd4a15d7 100644 --- a/issues/hardware/a-collapsed-replug-is-enumerated-only-when-another-port-event-arrives.md +++ b/issues/hardware/a-collapsed-replug-is-enumerated-only-when-another-port-event-arrives.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-09-27 --- @@ -50,6 +50,11 @@ both events are drained before the completion. With the one-line wake added it is green, `EXIT=0`, `3 seen as such`. - One named run of `xhci_flap` as committed is green on `nightly-green2` (`EXIT=0`) and on `main` at 16d2e645 (`EXIT=0`). +- PR #542's nightly at 059c5de7 (run 36328646395, `guest (2)`), whose diff + touches no driver or guest code, was red with the same sentence. The + collapse torn down at 1.523 s was enumerated at 2.225 s, when the next + cycle's edges arrived; the one torn down at 3.328 s had nothing after it + before the guest's input ended at 4.234 s. So the gate as committed passes by parity wherever every collapse loses its wake. On the dev host it cannot go red. It reds on CI's KVM shards only when @@ -64,5 +69,6 @@ A port torn down with its device still in it is looked at again without waiting for another event, shown by a gate that goes red on the lost wake on every host. `xhci_flap` at an odd cycle count is one such gate. An assertion that every collapsed teardown is enumerated before the next cycle's edges is -another. Until then `xhci_flap` reds on main's nightly at a rate. It should go -on #542's disabled list when that lands, citing this file. +another. `xhci_flap` is disabled in `src/redlist.rs` until then, and the change +that meets this deletes its row. Owner: the xHCI driver's port stepping +(`kernel/src/drivers/xhci/mod.rs`); held by the orchestrator. diff --git a/issues/hardware/eleven-names-red-on-ci.md b/issues/hardware/eleven-names-red-on-ci.md index 967bc1f916e..77b5456abb6 100644 --- a/issues/hardware/eleven-names-red-on-ci.md +++ b/issues/hardware/eleven-names-red-on-ci.md @@ -6,13 +6,6 @@ opened: 2026-08-08 # Eleven names are red on CI, at a rate that is now measured -**Do not answer "is this test known-red" out of this file.** Every table and -every list below is transcribed into `src/redlist.rs`, one row per measurement, -and `cargo run -- --known-red ` is what answers. A `rg` here hits the -twelve names that came *off* the list exactly as readily as the eleven that are -on it, and that has been read the wrong way round. What is here is the reasoning -and the evidence; what is there is the verdict. - Supersedes *a runner reds a rotating handful every run, and the rate is unmeasured*, whose whole ask was this run. `probe-rate.yml`, run `31258202923`, tree `f8f73e1`: **five reps of the exact twelve-shard configuration `ci.yml` @@ -53,17 +46,13 @@ is #172's signature away from the T14: two clients connect, both tones say **The top five reproduce, so they are defects and not a rate.** The bottom six fire one or two runs in five, which is 20–40% and is not "noise" either: the bar this was measured against tolerates one in fifty *with the failure named*, and -none of these six has been looked at. **No entry here is a candidate for -`EXPECTED_FAILURES`** — an exemption names a defect and a write-up, and "fires -40% of the time for reasons nobody has looked at" is neither. +none of these six has been looked at. **`metal_sim_null_audio` and `hda_two_live_refused` are the first two off this table**, closed when soundd stopped racing to present its null sink. **Re-taken 2026-09-04 against the last three nightly `ci` runs on `main`** (`33485669019`, `33603832656`, `33728852421`): of the names above, only -`usb_disk_index_stable` reds in any of them, and `src/redlist.rs` carries that -row at 1 of 3. Every other row this file sources is retired there on those -runs, never deleted. +`usb_disk_index_stable` reds in any of them. **Six of the eleven are `Sched::Serial`, and until 2026-08-08 the harness re-ran none of them**: the retry loop was written for the parallel phase and branched on diff --git a/issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md b/issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md new file mode 100644 index 00000000000..8ad470d8c29 --- /dev/null +++ b/issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md @@ -0,0 +1,20 @@ +--- +status: expected-red +kind: defect +opened: 2026-08-08 +--- + +# `usb_disk_index_stable` reds 1 of 5 on CI: nothing enumerated on the first controller + +From the twelve-shard CI probe (`probe-rate.yml` run `31258202923`, tree +`f8f73e1`, five reps of the exact `ci.yml` configuration): `usb_disk_index_stable` +red 1 of 5, shard 2, `Sched::Parallel`, `nothing enumerated on the first +controller`. Re-taken 2026-09-04 against the last three nightly `ci` runs on +`main` (`33485669019`, `33603832656`, `33728852421`): of the eleven names that +probe found, only this one still reds. + +Split out of `issues/hardware/eleven-names-red-on-ci.md`, which covers eleven +names and has no exit condition for this one in particular. + +**Exit condition.** The cause of the empty first controller is fixed, and +`usb_disk_index_stable` green on CI's `guest` shards. Owner: orchestrator. diff --git a/issues/kernel/deferred-release-outlives-its-syscall.md b/issues/kernel/deferred-release-outlives-its-syscall.md index 3ed48138262..343e88157ec 100644 --- a/issues/kernel/deferred-release-outlives-its-syscall.md +++ b/issues/kernel/deferred-release-outlives-its-syscall.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-08-19 --- @@ -78,11 +78,6 @@ tail. The census is immune to *another binary's churn*, which is what the free-memory verdicts are not, and it reds anyway. So the shared boot was never the common factor between these three names; the release latency is. -`handle_kill_policy` is on the redlist already (`src/redlist.rs`, dev host -loaded, 1 of 3, 2026-08-18) with a contention reading, and it is **not touched -here** — this entry records the mechanism, and re-adjudicating that row is its -owner's to do with a measurement rather than with this argument. - **A fourth witness, hosted CI, 2026-08-25.** `handle_transfer` red on run 32876917304 `guest (3)` — the census found one extra live `PipeRead` (2 → 3) after its deferred-release scenarios, red again in the shard's own alone @@ -140,14 +135,6 @@ no behaviour change anywhere; `ZERO_QUEUE` and `ZERO_PENDING` arrived with `6c39b1b4` and this test with `8f74272d`, so the shape has been reachable since the queue existed and no landing is a suspect. -**`kill_while_blocked`'s `ALONE … GREEN` must not be read as the harness reads -it.** That line says the name's `Sched::Parallel` is wrong, which is the wrong -conclusion for this defect and is already ruled out for its sibling: -`src/redlist.rs`'s `handle_lifetime` row records that `Sched::Serial` would have -retired nothing. What is owed at the name is that its rows stay, so a landing -gate that hits it has a rate to check the red against and nobody re-runs it away -or re-classifies its `Sched`. - So the sentence below is no longer the whole of it: this is a *semantic* event riding a release the caller cannot wait for, which is what `kernel/src/object/mod.rs`'s own header says must never happen — *"every diff --git a/issues/kernel/desktop-window-child-freeze.md b/issues/kernel/desktop-window-child-freeze.md index 05b8ce253df..690905b6a05 100644 --- a/issues/kernel/desktop-window-child-freeze.md +++ b/issues/kernel/desktop-window-child-freeze.md @@ -72,28 +72,12 @@ beside it. The teardown is not a regression from the deadline fix (`add6aeb`, exits, and is not a descendant of it. **What this does *not* settle is the freeze**, and the entry stays. It stays -`Sched::Parallel` for the same reason as before, `EXPECTED_FAILURES` keeps its -declaration to its review date, and a green run still proves nothing — the -signature at the top of this entry is a guest that goes *silent*, and none of -the eleven boots in this session produced one. What has changed is that the -test can now reach the snake rounds where the freeze was seen, which it could -not before. Judge the next occurrence by the signature, never by a run. - -**Landing while it is red** needs nothing special: `desktop_window_child` is -declared in `EXPECTED_FAILURES` (`tests/toyos.rs`) and the gate is the ordinary -one. The declaration reports it by name on every run, is red -if the test *passes* where the entry says a pass is proof, and is red on -`2026-09-06` regardless — this entry is intermittent, so its own expiry is a -date rather than a green run. The `--skip` flag that used to be the answer is -deleted: an exclusion nobody reviews cannot expire, and this one has to. - -**What the declaration will and will not absorb.** Its `says` list covers the -six of this test's messages whose failure is *the desktop ceasing to answer -after a window closed*. The other five red the run — the client binary missing, -the desktop never coming up, a window never being created, and the client -leaving on its own deadline. That pins which assertion failed and not why, so -the log-tail discriminator above is still a human's to apply; the run prints the -pointer to this section beside every `XFAIL` line for exactly that reason. +`Sched::Parallel` for the same reason as before, and a green run still proves +nothing — the signature at the top of this entry is a guest that goes +*silent*, and none of the eleven boots in this session produced one. What has +changed is that the test can now reach the snake rounds where the freeze was +seen, which it could not before. Judge the next occurrence by the signature, +never by a run. **One thing #156's capture leaned on is closed, and it is not this.** The deadline was stored twice — `ParkedEntry.deadline` and `DeadlineHeap` — and @@ -168,8 +152,7 @@ above is the measured one and not the whole of it. The test stopped at its **first** probe: the windowed child asked for a window, was answered `NotEndowed`, and printed `WINDOW-CHILD-REFUSED this program was -given no compositor` — while `EXPECTED_FAILURES`'s `the windowed child never -reported leaving` absorbed it, so no run said so. The client is a harness +given no compositor`. The client is a harness binary, no `[programs]` row can name one, and `/system/bin/init` endows a name the manifest does not carry with nothing. @@ -186,3 +169,12 @@ left both ways and the shell kept its prompt — which is the outcome this entry calls "#156 did not fire this run, which proves nothing". What changed is that a red now means the desktop stopped answering, which is what the declaration was written about. + +**Exit condition and owner.** Re-enabled when #156 is fixed and a `sched::dump` +NMI probe taken on a reproduction confirms no CPU stopped taking scheduler +passes during the freeze — nothing short of that instrument distinguishes this +signature from a green run, which this entry has already shown proves nothing +either way. Owner: `toyos-sched`, the placement track that closed the +CPU-selection half of this family (`CpuHandle::answering`, +`toyos-sched/src/cpu.rs`) and is nearest the remaining half; held by the +orchestrator. diff --git a/issues/kernel/handle-kill-policy-census-grew-one-sharedmem-on-two-nightlies.md b/issues/kernel/handle-kill-policy-census-grew-one-sharedmem-on-two-nightlies.md index 5d2870a54ab..ff31fdeb266 100644 --- a/issues/kernel/handle-kill-policy-census-grew-one-sharedmem-on-two-nightlies.md +++ b/issues/kernel/handle-kill-policy-census-grew-one-sharedmem-on-two-nightlies.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-09-27 --- @@ -34,9 +34,12 @@ second census. The red boot's kernel reports a TLB shootdown wait of up to Not shown: which process held the tenth `SharedMem`, or that its release was the one in flight. -`cargo run -- --known-red handle_kill_policy` answers NO. +PR #542's nightly at 059c5de7 (run 36328646395, `guest (9)`), KVM, QEMU +11.1.0, red with the same sentence, numbers included; that boot's kernel +reports `tlb: shootdowns=89 wait=354978us max=66850us`. The branch changes no +kernel or guest code. -**Exit**: the census names the owner of a grown kind, and the red is -attributed or the release is shown to finish before `wait` returns. Until then -`handle_kill_policy` reds on main's nightly at a rate. It should go on #542's -disabled list when that lands, citing this file. +**Exit condition.** A killed process's deferred releases finish before `wait` +returns, the fix `issues/kernel/deferred-release-outlives-its-syscall.md` +names under "What to do", and `handle_kill_policy` green on the KVM `guest` +shards where it went red. Owner: orchestrator. diff --git a/issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md b/issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md index 4faff8fa22d..25c97e68686 100644 --- a/issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md +++ b/issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md @@ -1,6 +1,6 @@ --- -status: open -kind: finding +status: expected-red +kind: defect opened: 2026-09-25 --- @@ -8,8 +8,7 @@ opened: 2026-09-25 Seen once, in the fast tier on the logd branch (PR #492), dev host: red wide with `QEMU never reported stopping: the guest asked for a reboot and stayed -up`, then `ALONE ... GREEN`. `cargo run -- --known-red -quiesce_dump_holds_the_stopped` says it is not quarantined. +up`, then `ALONE ... GREEN`. The boot's console shows the stick's transport breaking twice on `SCSI 0x2a` (`no answer in the status phase in 2000 ms`, recovered each time), then @@ -17,8 +16,6 @@ The boot's console shows the stick's transport breaking twice on `SCSI 0x2a` `test_rs_quiesce_writers exit=1`. The branch touches logd, netd and the console, not the USB stack, quiesce or its config. -Exit: a cause, or a recurrence that shows it is not load-bound. - **Recurrence, PR #524's fast tier at `396f5b4d`, dev host.** Red wide after 325 s with the same `QEMU never reported stopping: the guest asked for a reboot and stayed up`, then `ALONE ... GREEN` in 4 s. The capture has the same @@ -31,3 +28,20 @@ The branch reorders xHCI ring writes (`dma_wmb`) but touches no quiesce code. A same-session A/B then gave 5 of 5 green on each arm, main at `e48604c0` and the branch: the red did not come back alone or beside the other arm, so it is still not shown to be anything but load-bound. + +**Recurrence, PR #542's nightly at 059c5de7 (run 36328646395, `guest (7)`), +KVM, one guest on the runner.** The same `QEMU never reported stopping: the +guest asked for a reboot and stayed up`, the same `quiesce_writers: 4 of 6 +writers reached their loop in 5s` and `test_rs_quiesce_writers exit=1`, and no +transport break in the capture. A writer's first `quiesce-writer: 1` line +is its first pass done: writer 1 at 0.681 s, 0 at 1.061 s, 3 at 1.844 s, 5 at +2.014 s, and writers 2 and 4 only at 6.610 s and 6.778 s, 5.8 s and 5.9 s +after their ` 0` lines. Writer 1 printed pass 33 at 6.905 s, its passes 93 +to 306 ms apart. The branch changes no kernel or guest code. A red on a +one-guest lane is not the other suites' load. + +**Exit condition.** Re-enabled when a reproduction names what holds a writer's +first write-and-fsync pass for over 5 s while another writer passes in under a +third of a second, and the fix is shown against it. Owner: the `/log` write and +sync path `tests/toyos-rust-tests/src/bin/quiesce_writers.rs` drives; held by +the orchestrator. diff --git a/issues/kernel/short-sleep-livelock-stalls-on-ci-with-one-sleeper-never-returning.md b/issues/kernel/short-sleep-livelock-stalls-on-ci-with-one-sleeper-never-returning.md index 76d51a9d567..aefa6afcd7e 100644 --- a/issues/kernel/short-sleep-livelock-stalls-on-ci-with-one-sleeper-never-returning.md +++ b/issues/kernel/short-sleep-livelock-stalls-on-ci-with-one-sleeper-never-returning.md @@ -1,5 +1,5 @@ --- -status: open +status: expected-red kind: defect opened: 2026-09-06 --- @@ -26,7 +26,9 @@ sleeps of 100000 ns returned Four of the program's sleepers finished their 100000 ns round and exited on cpu1; the fifth never printed its return and the guest said nothing more for 63 s. That is the shape the test exists to catch, seen once on the CI -instrument; the redlist row (`src/redlist.rs`, `short_sleep_livelock`, -`Instrument::Ci`) records it, and the owner is the sleep path in -`kernel/src/sched` that the test's write-up (`tests/toyos.rs`, -`short_sleep_livelock`) names. +instrument, and the owner is the sleep path in `kernel/src/sched` that the +test's write-up (`tests/toyos.rs`, `short_sleep_livelock`) names. + +**Exit condition.** The fifth sleeper's stall is fixed in the sleep path, and +`short_sleep_livelock` green on CI's KVM `guest` shards. +Owner: the sleep path, `kernel/src/sched`; held by the orchestrator. diff --git a/issues/kernel/so-cache-refusals-saw-the-kernel-refuse-nothing-once.md b/issues/kernel/so-cache-refusals-saw-the-kernel-refuse-nothing-once.md index d48c2fba442..f21bdbf3331 100644 --- a/issues/kernel/so-cache-refusals-saw-the-kernel-refuse-nothing-once.md +++ b/issues/kernel/so-cache-refusals-saw-the-kernel-refuse-nothing-once.md @@ -13,5 +13,7 @@ alone both times` in the same job. The verdict: no "byte budget; refused" line — the kernel refused nothing: twelve 2 MiB images entered a cache whose test budget refuses at the second. -Owed: a mechanism. Nobody has one. `src/redlist.rs` quarantines the name until -then. +Owed: a mechanism. Nobody has one. + +**Exit condition.** The cause of the missing refusal is fixed, and +`so_cache_refusals` green on CI's KVM `guest` shards. Owner: orchestrator. diff --git a/issues/kernel/toyos-runs-on-arm64.md b/issues/kernel/toyos-runs-on-arm64.md index abf5a9b7218..78e2baceb0c 100644 --- a/issues/kernel/toyos-runs-on-arm64.md +++ b/issues/kernel/toyos-runs-on-arm64.md @@ -329,7 +329,7 @@ Each stage names its exit; "measured" means a number from a run. 8. **An aarch64 tier in the harness.** `tests/common/qemu.rs` takes an `Arch`: `virt`, edk2-aarch64, HVF on Apple hosts (TCG otherwise). - `src/tiers.rs` gains the arch axis. `src/redlist.rs` quarantines per arch. + `src/tiers.rs` gains the arch axis. The CI runner is picked here, after measuring hosted `macos-latest` (whether HVF is usable inside the runner VM) and hosted `ubuntu-24.04-arm` (whether `/dev/kvm` exists there). **Exit**: the fast diff --git a/issues/panic-path/a-double-panic-at-boots-edge-says-nothing-but-its-name.md b/issues/panic-path/a-double-panic-at-boots-edge-says-nothing-but-its-name.md index f93200e085c..0628f0b5195 100644 --- a/issues/panic-path/a-double-panic-at-boots-edge-says-nothing-but-its-name.md +++ b/issues/panic-path/a-double-panic-at-boots-edge-says-nothing-but-its-name.md @@ -8,7 +8,6 @@ opened: 2026-08-19 Two findings in one sighting, dev host under load (two worktrees' full suites interleaved over the shared twelve guest slots), 2026-08-19 22:21 UTC. -`src/redlist.rs` carries the row. **1. The kernel double-panicked under load.** `log_poll_outlives_a_close`, parallel phase, 25 s: the guest went quiet with every CPU halted, and the @@ -48,21 +47,10 @@ could not: the first crash's identity and site, the second panic's site, and the state the CPU was in when it arrived. What this sighting still does not establish is what that first crash was. The capture is `scratchpad/hkpfix-harness.log` in the 2026-08-20 orchestrator session; the -durable evidence is quoted here and in the redlist row. +durable evidence is quoted here. ## What closes it is a count, and the count is owed -There is a leading explanation and it is not this entry's to argue: the -`log_poll_outlives_a_close` row in `src/redlist.rs` records it as the -already-fixed missing-`cld` class, with the before/after boot counts that make -the case, and reasons that a machine-wide death at boot's edge under two suites -on a branch carrying no kernel byte is that class's shape. It is not shown here, -so the row stands and so does this. - -**Nothing about it is a decision.** The row names its own retirement condition — -three loaded suites of the fixed tree with no red under this name — and that is -an instrument run: the result read rather than argued. - **2026-09-04, run: six, and the row is retired.** Six full `cargo test` fast tiers in one worktree, each with single-test runs of the same suite beside it, four of the six contending with a second worktree (`toyos-rootfs3`) for guest diff --git a/src/CLAUDE.md b/src/CLAUDE.md index 7d61612a283..82b6c66c417 100644 --- a/src/CLAUDE.md +++ b/src/CLAUDE.md @@ -36,4 +36,4 @@ Loads when you read a file under `src/` — the root cargo project, package name - **Documentation carries no gates** — `src/redlist.rs` resolves doc paths only because it gates a Rust table, not a corpus. - **Every CI lane is GitHub-hosted and no workflow may name a self-hosted label** — a `runs-on:` naming one queues until it times out rather than failing, so `src/ci.rs`'s `workflows_run_against_main_on_hosted_runners` refuses it; a measurement owed on hardware goes to the metal loop, not to a runner. - **A workflow job that runs in a container adds `safe.directory` itself** — `actions/checkout` sets it into a temporary global config it discards when its step ends, so the first git command a container step runs after checkout dies on a dubiously-owned repository. -- **A red build may be the build system — re-run in isolation before believing any single red.** A `stage1-std//dist/deps` temp-dir error means a concurrent build, never a broken checkout; never repair or force-rebuild the toolchain. +- **A `stage1-std//dist/deps` temp-dir error means a concurrent build**, never a broken checkout; never repair or force-rebuild the toolchain. diff --git a/src/alone.rs b/src/alone.rs deleted file mode 100644 index 91ddc511b50..00000000000 --- a/src/alone.rs +++ /dev/null @@ -1,346 +0,0 @@ -//! Whether two failure sentences are the same failure — the one decision behind -//! the suite's `ALONE:` line. -//! -//! What differs between two readings of one assertion is a measurement; what -//! differs between two failures is an identity. So a sentence is compared with -//! only what is unmistakably a measurement taken out of it: a number carrying a -//! unit of time or size (`0.303 s`, `1007 ms`, `12 MiB`), and the time in a -//! kernel record's stamp (`[kernel 0.075 cpu0]`, which no two boots write the -//! same). Everything else stays, because it names *which* thing was observed — -//! `slot 1` against `slot 2`, a port, an APIC id, a CPU, an opcode, a register, -//! a count with no unit — and two sentences naming two of them are two -//! observations, which is the larger finding of the two. -//! -//! **The limit that rule carries**: a reading rendered without one of [`UNITS`] -//! beside it — an address, a percentage, a bare count, `bytes` spelled out — -//! reads as an identity, so one assertion printing two of them is reported as -//! two failures, which is the direction this rule errs in on purpose. A byte -//! count is the one unit this tree also uses to *name* a size class rather -//! than report one (`"blocks of 512 B"`, `"sectors of 4096 B"`): the word `of` -//! immediately before it is what tells "sized in units of N" from "N was -//! measured", so that shape stays an identity too. - -/// The units that make digits a reading, each one a spelling an assertion in -/// this tree prints; the test names the site per unit. -const UNITS: &[&str] = &["ns", "us", "ms", "s", "KiB", "MiB", "GiB", "MB", "GB", "B"]; - -/// Whether `one` and `other` are the same assertion firing, at possibly -/// different readings. -/// -/// Both arguments are headlines — one line, the sentence without the capture. -pub fn same_failure(one: &str, other: &str) -> bool { - skeleton(one) == skeleton(other) -} - -/// `text` with every measurement in it replaced by one `#`, which is what is -/// left of a sentence when its readings are taken out of it. -fn skeleton(text: &str) -> String { - let mut out = String::with_capacity(text.len()); - let mut i = 0; - while let Some(c) = text[i..].chars().next() { - let rest = &text[i..]; - if let Some((time, end)) = stamp(rest) { - out.push_str(&rest[..time]); - out.push('#'); - i += end; - } else if let Some(n) = numeral(rest) { - // The whole numeral moves at once, classified once: a byte count - // ruled out by `of` still has its own later digits, and re-asking - // from inside them would find no `of` right behind and take them - // for a reading `512` never was. - if is_reading(rest, n, &text[..i]) { - out.push('#'); - } else { - out.push_str(&rest[..n]); - } - i += n; - } else { - out.push(c); - i += c.len_utf8(); - } - } - out -} - -/// Where the boot time sits inside a kernel record's stamp at the head of -/// `text` — the offsets of `0.075` in `[kernel 0.075 cpu0]`. -/// -/// The `cpu` field is what tells a stamp from any other bracketed pair, and it -/// is an identity: which CPU wrote a record is an observation about the record. -fn stamp(text: &str) -> Option<(usize, usize)> { - let inside = text.strip_prefix('[')?; - let close = inside.find(']')?; - let mut fields = inside[..close].split(' '); - let (writer, time, cpu) = (fields.next()?, fields.next()?, fields.next()?); - let stamped = !writer.is_empty() - && !time.is_empty() - && time.bytes().all(|c| c.is_ascii_digit() || c == b'.') - && cpu.strip_prefix("cpu").is_some_and(|n| !n.is_empty() && n.bytes().all(|c| c.is_ascii_digit())) - && fields.next().is_none(); - let at = 1 + writer.len() + 1; - stamped.then_some((at, at + time.len())) -} - -/// The length of the numeral at the head of `text` — digits, and the fraction -/// behind a `.` with digits on both sides so `0.303` is one numeral and not -/// two — whether or not it turns out to be a reading. -fn numeral(text: &str) -> Option { - let bytes = text.as_bytes(); - if !bytes.first().is_some_and(u8::is_ascii_digit) { - return None; - } - let mut n = 0; - while n < bytes.len() && bytes[n].is_ascii_digit() { - n += 1; - } - while bytes.get(n) == Some(&b'.') && bytes.get(n + 1).is_some_and(u8::is_ascii_digit) { - n += 1; - while n < bytes.len() && bytes[n].is_ascii_digit() { - n += 1; - } - } - Some(n) -} - -/// Whether the numeral `text[..n]` is a reading: a unit follows it, and — -/// for a byte count — `before`, everything already read, does not put it -/// after the word `of`. -fn is_reading(text: &str, n: usize, before: &str) -> bool { - let after = &text[n..]; - let after = after.strip_prefix(' ').unwrap_or(after); - // A whole unit and not the head of a longer word: `2 sticks` counts nothing. - let Some(unit) = UNITS.iter().find(|unit| { - after - .strip_prefix(**unit) - .is_some_and(|tail| !tail.starts_with(|c: char| c.is_ascii_alphanumeric())) - }) else { - return false; - }; - *unit != "B" || !names_a_size_class(before) -} - -/// Whether `before` ends on the word `of`, which is how this tree writes "each -/// one sized N bytes" (`"blocks of 512 B"`, `"sectors of 4096 B"`) rather than -/// a byte count it measured — the two words this checks for a boundary around -/// so `"roof 512 B"` does not read as one. -fn names_a_size_class(before: &str) -> bool { - let head = before.trim_end_matches(' '); - head.ends_with("of") && head[..head.len() - 2].chars().next_back().is_none_or(|c| !c.is_ascii_alphanumeric()) -} - -#[cfg(test)] -mod tests { - use super::*; - - const WIDE: &str = "the controller started at 0.303 s, past the 0.3 s the ports are held \ - empty for, so nothing in this boot read a hidden port. The boot has \ - outgrown the injection window: widen SLOW_CONNECT_NS, not this gate"; - const ALONE: &str = "the controller started at 0.300 s, past the 0.3 s the ports are held \ - empty for, so nothing in this boot read a hidden port. The boot has \ - outgrown the injection window: widen SLOW_CONNECT_NS, not this gate"; - - #[test] - fn one_assertion_at_two_measurements_is_one_failure() { - assert_ne!(WIDE, ALONE, "the two runs did write different sentences"); - assert!(same_failure(WIDE, ALONE), "one assertion at two readings read as two"); - assert!(same_failure(WIDE, WIDE), "a sentence is itself"); - } - - /// Two assertions of the *same* test, which is the pair a looser rule would - /// collapse: `tests/common/usb.rs`'s floor and ceiling on when the first - /// port was named, each carrying two readings. - #[test] - fn two_assertions_stay_two_failures() { - let floor = "the first port was named at 0.395 s, before the 0.4 s the held-empty \ - window and the debounce behind it come to — the injection did not reach \ - the driver"; - let ceiling = "the first port was named at 0.980 s, 0.580 s after the connect became \ - visible — the settle did not end on the device appearing"; - assert!(!same_failure(floor, ceiling), "two assertions read as one reproduced defect"); - - let endpoints = "3 endpoint(s) were found Running after the break, want 2"; - let pointer = "input never came back: no pointer event moved by (2560, -1920)"; - assert!(!same_failure(endpoints, pointer)); - } - - /// A verdict that quotes a kernel record carries that record's stamp, and - /// two boots never stamp the same line the same. The sentence is - /// `tests/toyos.rs`'s `pci_cap_selftest`, which fires on a verdict that is - /// short of `14/14` and pastes the record it read. - #[test] - fn one_verdict_at_two_kernel_stamps_is_one_failure() { - let control = "not every crafted capability layout was answered: [kernel 0.075 cpu0] \ - virtio: pci cap selftest 13/14"; - let mutant = "not every crafted capability layout was answered: [kernel 0.082 cpu0] \ - virtio: pci cap selftest 13/14"; - assert_ne!(control, mutant); - assert!(same_failure(control, mutant), "one verdict at two kernel stamps read as two"); - - // The count is the finding, so the stamp rule does not reach it. - let fewer = control.replace("13/14", "12/14"); - assert!(!same_failure(control, &fewer)); - // Nor does it reach the CPU beside the time. - let elsewhere = control.replace("cpu0", "cpu3"); - assert!(!same_failure(control, &elsewhere)); - } - - /// An index is an identity however it is spelled — welded to its name or - /// separated from it — because two of them are two objects. - /// - /// Each pair is a whole line this tree writes, at two indices: the - /// transport break of `kernel/src/drivers/xhci/wait/msc.rs` (the one - /// `src/redlist.rs` quotes off CI), the durability check of - /// `tests/common/volumes.rs`, and the stall line of - /// `kernel/src/heartbeat.rs`. - #[test] - fn an_index_is_an_identity() { - let broke = |slot: u8, opcode: &str| { - format!( - "usb-storage: 00:02.0 slot {slot} transport broke on SCSI {opcode}: no answer \ - in the status phase in 2000 ms" - ) - }; - assert!(!same_failure(&broke(1, "0x35"), &broke(2, "0x35"))); - assert!(!same_failure( - "slot 3 holds 12 on the device — a write the guest confirmed durable is not in the \ - bytes the host reads", - "slot 4 holds 12 on the device — a write the guest confirmed durable is not in the \ - bytes the host reads" - )); - assert!(!same_failure( - "heartbeat: cpu5 last reached one 0.349s ago", - "heartbeat: cpu6 last reached one 0.349s ago" - )); - assert!(same_failure( - "heartbeat: cpu5 last reached one 0.349s ago", - "heartbeat: cpu5 last reached one 0.712s ago" - )); - // An opcode is one too, and so is a count that carries no unit. - assert!(!same_failure(&broke(1, "0x35"), &broke(1, "0x28"))); - assert!(!same_failure( - "3 endpoint(s) were found Running after the break, want 2", - "4 endpoint(s) were found Running after the break, want 2" - )); - } - - /// The limit the module header states: an address is digits with no unit - /// after them, so `tests/common/iommu.rs`'s translation assertions read as - /// two failures at two addresses rather than one at two readings. - #[test] - fn two_addresses_are_two_failures() { - assert!(!same_failure( - "the retired scanout at 0xfd000000 was given device address 0x40000000, which still \ - translates to 0x1000 while a holder maps the pages", - "the retired scanout at 0xfe000000 was given device address 0x40000000, which still \ - translates to 0x1000 while a holder maps the pages" - )); - } - - /// The same limit: `tests/common/audio.rs`'s wake-lateness fault renders one - /// reading twice and only the `us` copy carries a unit, so the depths beside - /// it hold the sentence apart. - #[test] - fn a_reading_rendered_twice_keeps_the_copy_with_no_unit() { - assert!(!same_failure( - "wake lateness 21000000us (904.4 pipeline depths) exceeds the whole 4.00s run it \ - was measured inside — the instrument is broken, not the scheduler", - "wake lateness 21400000us (921.7 pipeline depths) exceeds the whole 4.00s run it \ - was measured inside — the instrument is broken, not the scheduler" - )); - } - - /// The same limit at a percentage: `tests/toyos.rs`'s dither floor. - #[test] - fn two_percentages_are_two_failures() { - assert!(!same_failure( - "dither missing: only 8.3% of silent samples are non-zero (expected ~25%, floor \ - 10%) — soundd is not dithering, so the underrun detector is blind", - "dither missing: only 7.6% of silent samples are non-zero (expected ~25%, floor \ - 10%) — soundd is not dithering, so the underrun detector is blind" - )); - } - - /// The same limit at a bare count: `tests/toyos.rs`'s tone peak, which is a - /// sample value and carries no unit to be measured in. - #[test] - fn two_bare_counts_are_two_failures() { - assert!(!same_failure( - "tone too quiet: peak 3912 (expected >= 4000)", - "tone too quiet: peak 3874 (expected >= 4000)" - )); - } - - /// The same limit where the unit is spelled out: `tests/toyos.rs`'s - /// log-drain verdict writes `bytes`, which is not `B`. - #[test] - fn bytes_spelled_out_is_not_a_unit() { - let stops_at = |len: u32| { - format!( - "/log/2026-09-01-202502.log stops at {len} bytes and never carries \ - \"metal-panic-probe\" — this boot wrote no log at all, so the drain's own \ - verdict is not what is wrong here" - ) - }; - assert!(!same_failure(&stops_at(40960), &stops_at(45056))); - } - - /// `tests/common/usb.rs`'s `check_geometry` prints the sector size a driver - /// reported as `"blocks of {lba} B"`: 512 and 4096 are different driver - /// defects (the default-geometry path against the 4Kn one), not two - /// readings of one assertion, so they must stay two failures. - #[test] - fn a_size_class_is_not_a_reading() { - assert!(!same_failure( - "the driver did not report \"blocks of 512 B\"", - "the driver did not report \"blocks of 4096 B\"" - )); - } - - /// Every unit in [`UNITS`] is a spelling an assertion in this tree prints. A - /// unit with no such site only widens a merge that must err toward - /// "different", so the list and this table are one thing. - #[test] - fn every_unit_is_one_this_tree_prints() { - let printed_by = [ - ("ns", "tests/toyos-rust-tests/src/bin/tlb_shootdown_waits.rs's shootdown cost"), - ("us", "tests/common/audio.rs's wake-lateness limit"), - ("ms", "tests/toyos.rs's i8042 boot A/B"), - ("s", "tests/common/usb.rs's controller-start ceiling"), - ("KiB", "tests/toyos.rs's xhci pool size"), - ("MiB", "tests/toyos.rs's pmm accounting"), - ("GiB", "tests/toyos-rust-tests/src/bin/abuse_pipe_ring.rs's ring-header cases"), - ("MB", "kernel/src/process.rs's per-process peak, quoted into a headline"), - ("GB", "tests/toyos-rust-tests/src/bin/allocator_stress.rs's memory-total range"), - ("B", "tests/common/usb.rs's usbmon-capture-file-size check"), - ]; - assert_eq!(printed_by.len(), UNITS.len(), "a unit in the list that nothing here cites"); - for (unit, site) in printed_by { - assert!(UNITS.contains(&unit), "{site} prints {unit}, which left the list"); - assert_eq!(skeleton(&format!("took 7 {unit}")), format!("took # {unit}"), "{site}"); - assert_eq!(skeleton(&format!("took 7{unit}")), format!("took #{unit}"), "{site}"); - } - } - - #[test] - fn a_skeleton_keeps_everything_that_is_not_a_reading() { - assert_eq!(skeleton("started at 0.303 s"), "started at # s"); - assert_eq!(skeleton("read 1007 ms and 12 MiB"), "read # ms and # MiB"); - assert_eq!(skeleton("cpu6 0.349s"), "cpu6 #s"); - assert_eq!(skeleton("[kernel 0.075 cpu0] slot 1 at 0x35"), "[kernel # cpu0] slot 1 at 0x35"); - assert_eq!(skeleton("moved by (2560, -1920)"), "moved by (2560, -1920)"); - assert_eq!(skeleton("a sentence with no readings"), "a sentence with no readings"); - // A unit is a whole word or it is not one, and a bracketed pair that is - // not a record's stamp is not one either. - assert_eq!(skeleton("2 sticks in 4 seconds"), "2 sticks in 4 seconds"); - assert_eq!(skeleton("[budget 0.075 left]"), "[budget 0.075 left]"); - // "of" names a size class, not a reading — but only ahead of `B`: the - // driver did not report how many *seconds* something is sized in. - assert_eq!(skeleton("blocks of 512 B"), "blocks of 512 B"); - assert_eq!(skeleton("sectors of 4096 B, 12 MiB delivered"), "sectors of 4096 B, # MiB delivered"); - assert_eq!(skeleton("waited a multiple of 5 s"), "waited a multiple of # s"); - // "of" is a whole word: a byte count after "roof" or "thereof" is read - // exactly as it would be with no word before it at all. - assert_eq!(skeleton("roof 512 B"), "roof # B"); - assert_eq!(skeleton("thereof 512 B"), "thereof # B"); - } -} diff --git a/src/ci.rs b/src/ci.rs index 69f0c30ff37..5518b76c6b8 100644 --- a/src/ci.rs +++ b/src/ci.rs @@ -644,8 +644,7 @@ fn guest(root: &Path, suite: &[String]) -> Vec { } /// The suite's own count line and every line naming a verdict worth reading -/// without the log: a failure, whether it survived being run alone, and a -/// quarantined name that failed for something else. +/// without the log: a failure. fn verdicts(log: &str) -> String { let total = log .lines() @@ -655,10 +654,8 @@ fn verdicts(log: &str) -> String { .lines() .filter(|l| { l.starts_with("FAIL ") - || l.starts_with("XFAIL ") || (l.starts_with(' ') - && ["STALL ", "INVL ", "ALONE "].iter().any(|v| l.trim_start().starts_with(v))) - || l.contains("is quarantined for something else") + && ["STALL ", "INVL "].iter().any(|v| l.trim_start().starts_with(v))) }) .collect(); if named.is_empty() { @@ -969,11 +966,11 @@ mod tests { #[test] fn the_summary_keeps_the_count_and_the_verdicts() { let log = "test result: ok. 3 passed\n\ - FAIL rs::lan_talk: no exit code\n ALONE lan_talk: GREEN\nnoise\n\ + FAIL rs::lan_talk: no exit code\nnoise\n\ test result: FAILED. 40 passed; 1 failed, 41 total (300 s)\n"; let said = verdicts(log); assert!(said.starts_with("test result: FAILED. 40 passed; 1 failed, 41 total"), "{said}"); - assert!(said.contains("FAIL rs::lan_talk") && said.contains("ALONE lan_talk"), "{said}"); + assert!(said.contains("FAIL rs::lan_talk"), "{said}"); assert!(!said.contains("noise"), "{said}"); assert_eq!(verdicts(""), "no suite result line"); } diff --git a/src/lib.rs b/src/lib.rs index e44e0c0ced2..2ef6e4d30dc 100644 --- a/src/lib.rs +++ b/src/lib.rs @@ -1,9 +1,6 @@ /// The actuator-state coupling gate, read by nothing but its own tests. #[cfg(test)] pub mod actuatorstate; -/// What the suite's isolated re-run is allowed to call one failure; read by -/// `tests/toyos.rs` and by its own tests. -pub mod alone; pub mod arch; pub mod assets; pub mod bootlog; diff --git a/src/main.rs b/src/main.rs index d355c342f2e..e0d675c47a5 100644 --- a/src/main.rs +++ b/src/main.rs @@ -126,7 +126,7 @@ fn main() { return; } // Reads one table and prints. Here for the same reason again, and for one - // more: the question it answers — "is this red quarantined?" — is asked + // more: the question it answers — "is this test disabled?" — is asked // while a build is broken as often as while one works. if asked(&flags::KNOWN_RED) { toyos_build::redlist::dispatch(&args); diff --git a/src/redlist.rs b/src/redlist.rs index edc5ca2554a..f941c3eaccd 100644 --- a/src/redlist.rs +++ b/src/redlist.rs @@ -1,210 +1,255 @@ -//! The quarantine list: the tests known to fail on a defect somebody owns. +//! The disabled list: the tests that do not run, each with the issue that owns +//! its defect. //! -//! A row excuses one known failure, not the test. A quarantined test still runs -//! on every run; a failure whose text contains one of the row's -//! [`says`](Quarantined::says) fragments is reported by name and does not fail -//! the suite, and the same test failing any other way fails it like a test on no -//! list. `tests/toyos.rs` reads [`QUARANTINE`] for every verdict and refuses a -//! row whose name nothing registers. +//! A red is a red, so a failing test, flaky or not, is fixed or disabled here +//! with its issue. A disabled test runs nowhere: `tests/toyos.rs` skips it and +//! names it and its issue on every run, and refuses a row whose test nothing +//! registers or whose issue file does not exist. The change that fixes the test +//! deletes its row. //! -//! `cargo run -- --known-red ` answers yes or no with what the row -//! excuses; with no argument it prints the list. +//! `cargo run -- --known-red ` answers whether a test is disabled; with +//! no argument it prints the list. + +use std::path::Path; use crate::flags; -/// One quarantined failure of one test. +/// One disabled test. #[derive(PartialEq, Eq, Debug)] -pub struct Quarantined { +pub struct Disabled { /// The registered test name, exactly. pub test: &'static str, - /// The failure this row excuses, quoted from the test's own message: the - /// failure's text must contain one of these or the row does not apply. - /// Alternatives rather than conjuncts, because one defect can surface at - /// more than one of a test's assertions. A quotation, so that a reworded - /// assertion stops matching and the run reds asking about it — of the test's - /// own message and never of the harness's framing around it, which a gate - /// below refuses by name. - pub says: &'static [&'static str], /// The issue file that owns the defect. pub issue: &'static str, } -impl Quarantined { - /// Whether `failure` is the failure this row is about. - pub fn excuses(&self, failure: &str) -> bool { - self.says.iter().any(|fragment| failure.contains(fragment)) - } -} - -/// Every quarantined failure, by test name. -pub const QUARANTINE: &[Quarantined] = &[ - Quarantined { +/// Every disabled test. +pub const DISABLED: &[Disabled] = &[ + Disabled { test: "console_line_atomicity", - says: &["whole lines and the capture carries"], - issue: "issues/build/parallel-tests-red-under-other-suites.md", + issue: "issues/build/console-line-atomicity-loses-five-of-a-thousand-lines-on-ci.md", }, - Quarantined { + Disabled { test: "console_locale_detect", - says: &["waiting for the wizard to ask for a key under /system/bin/console — the console \ - did not lend it the keyboard — it never stopped talking and never got there"], issue: "issues/build/the-console-input-path-can-stop-after-a-ps2-overflow.md", }, - Quarantined { - test: "desktop_window_child", - says: &[ - "the windowed child never reported leaving", - "a windowed child exited by itself and the shell never answered again", - "GUI+Q never reached the compositor", - "the compositor closed the window and the client did not leave", - "snake did not leave when its window was closed in round", - "snake's window was closed, snake left, and the shell never answered again", - ], - issue: "issues/kernel/desktop-window-child-freeze.md", - }, - Quarantined { - test: "doom_sound_flood", - says: &["the device played a peak of 32768"], - issue: "issues/audio/doom-sound-flood-played-full-scale-once.md", + Disabled { test: "desktop_window_child", issue: "issues/kernel/desktop-window-child-freeze.md" }, + Disabled { test: "doom_sound_flood", issue: "issues/audio/doom-sound-flood-played-full-scale-once.md" }, + Disabled { + test: "handle_kill_policy", + issue: "issues/kernel/handle-kill-policy-census-grew-one-sharedmem-on-two-nightlies.md", }, - Quarantined { - test: "handle_transfer", - says: &["handle transfer left more live objects behind"], - issue: "issues/kernel/deferred-release-outlives-its-syscall.md", + Disabled { test: "handle_transfer", issue: "issues/kernel/deferred-release-outlives-its-syscall.md" }, + Disabled { test: "hda_tone", issue: "issues/audio/hda-tone-phase-check.md" }, + Disabled { test: "kill_while_blocked", issue: "issues/kernel/deferred-release-outlives-its-syscall.md" }, + Disabled { test: "latency_wake", issue: "issues/build/latency-wake-reds-on-the-dev-host-at-a-rate.md" }, + Disabled { + test: "quiesce_dump_holds_the_stopped", + issue: "issues/kernel/quiesce-dump-holds-the-stopped-reds-wide-with-usb-transport-breaks.md", }, - Quarantined { - test: "hda_tone", - says: &["the captured tone is not one sine"], - issue: "issues/audio/hda-tone-phase-check.md", + Disabled { + test: "quiesce_wakes_on_the_last_exit", + issue: "issues/build/quiesce-wakes-on-the-last-exit-lost-its-serial-ready-beside-other-guests.md", }, - Quarantined { - test: "kill_while_blocked", - says: &[ - "a pipe whose only reader was killed mid-read still took a write", - "a connection whose peer was killed mid-read still took a write", - ], - issue: "issues/kernel/deferred-release-outlives-its-syscall.md", - }, - Quarantined { - test: "latency_wake", - says: &["the p99 landed in the histogram's last bucket"], - issue: "issues/build/latency-wake-reds-on-the-dev-host-at-a-rate.md", - }, - Quarantined { + Disabled { test: "sched_check_build", - says: &["this distribution has mass the KVM, native x86-64 sample never showed"], issue: "issues/build/the-pass-cost-gates-ci-sample-is-eight-days-stale-twice.md", }, - Quarantined { + Disabled { test: "screen_fatal_halt", - says: &["transport broke on SCSI 0x35"], issue: "issues/boot-media/screen-fatal-halt-reds-on-ci-with-a-usb-storage-transport-break-during-boot.md", }, - Quarantined { + Disabled { test: "short_sleep_livelock", - says: &["of guard expired, and the guest had said nothing for the last"], issue: "issues/kernel/short-sleep-livelock-stalls-on-ci-with-one-sleeper-never-returning.md", }, - Quarantined { + Disabled { test: "so_cache_refusals", - says: &["no \"byte budget; refused\" line — the kernel refused nothing"], issue: "issues/kernel/so-cache-refusals-saw-the-kernel-refuse-nothing-once.md", }, - Quarantined { + Disabled { test: "usb_disk_index_stable", - says: &["nothing enumerated on the first controller; there is no renumbering to survive"], - issue: "issues/hardware/eleven-names-red-on-ci.md", + issue: "issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md", + }, + Disabled { + test: "xhci_flap", + issue: "issues/hardware/a-collapsed-replug-is-enumerated-only-when-another-port-event-arrives.md", }, ]; +/// The row of `rows` that disables `test`, matched by the whole name. +pub fn disabled<'a>(rows: &'a [Disabled], test: &str) -> Option<&'a Disabled> { + rows.iter().find(|row| row.test == test) +} + +/// Whether `issue` has the one shape a per-test issue is allowed: +/// `issues//.md` — an area directory and a file, nothing nested +/// deeper and nothing sitting directly in `issues/` itself (`issues/README.md` +/// is the tracker's own doc, not a test's issue). +fn is_per_test_issue_path(issue: &str) -> bool { + let Some(rest) = issue.strip_prefix("issues/") else { return false }; + let mut parts = rest.split('/'); + let (Some(_area), Some(slug)) = (parts.next(), parts.next()) else { return false }; + parts.next().is_none() && slug.ends_with(".md") +} + +/// Whether the file at `path` opens with a frontmatter block whose `status` +/// field is `expected-red` — the one status a disabled test's issue may carry +/// (`issues/README.md`'s own frontmatter table). +fn is_expected_red(path: &Path) -> bool { + let Ok(text) = std::fs::read_to_string(path) else { return false }; + let mut lines = text.lines(); + if lines.next() != Some("---") { + return false; + } + for line in lines { + if line == "---" { + return false; + } + if let Some(value) = line.strip_prefix("status:") { + return value.trim() == "expected-red"; + } + } + false +} + +/// Every row of `rows` against the tree under `root`: no test twice, each one +/// `registered`, and each issue an `issues//.md` file whose +/// frontmatter says `status: expected-red`. +pub fn check(rows: &[Disabled], registered: impl Fn(&str) -> bool, root: &Path) -> Result<(), String> { + for (at, row) in rows.iter().enumerate() { + if disabled(&rows[..at], row.test).is_some() { + return Err(format!("{} is disabled twice", row.test)); + } + if !registered(row.test) { + return Err(format!( + "{} is disabled and nothing registers it: a renamed or deleted test takes its row \ + with it", + row.test + )); + } + if !is_per_test_issue_path(row.issue) { + return Err(format!( + "{}: `{}` is not an `issues//.md` path", + row.test, row.issue + )); + } + let path = root.join(row.issue); + if !path.is_file() { + return Err(format!("{}: `{}` is not a file in this tree", row.test, row.issue)); + } + if !is_expected_red(&path) { + return Err(format!( + "{}: `{}` does not open with `status: expected-red`", + row.test, row.issue + )); + } + } + Ok(()) +} + /// `cargo run -- --known-red []`. pub fn dispatch(args: &[String]) { - print!("{}", answer(QUARANTINE, flags::CARGO_RUN.value(args, &flags::KNOWN_RED))); + print!("{}", answer(DISABLED, flags::CARGO_RUN.value(args, &flags::KNOWN_RED))); } -fn answer(rows: &[Quarantined], asked: Option<&str>) -> String { - let excused = |q: &Quarantined| { - q.says.iter().map(|fragment| format!("{fragment:?}")).collect::>().join(" or ") - }; +fn answer(rows: &[Disabled], asked: Option<&str>) -> String { let Some(test) = asked else { - return rows.iter().map(|q| format!("{} {} {}\n", q.test, q.issue, excused(q))).collect(); + return rows.iter().map(|row| format!("{} {}\n", row.test, row.issue)).collect(); }; - match rows.iter().find(|q| q.test == test) { - Some(q) => format!( - "{test}: YES, quarantined — a failure saying {} does not fail the suite, and any \ - other failure of it does.\n {}\n", - excused(q), - q.issue - ), - None => format!("{test}: NO, not quarantined — its failure fails the suite.\n"), + match disabled(rows, test) { + Some(row) => format!("{test}: YES, disabled — it does not run.\n {}\n", row.issue), + None => format!("{test}: NO, not disabled — it runs, and its red fails the suite.\n"), } } #[cfg(test)] mod tests { use super::*; - use std::collections::BTreeSet; - use std::path::Path; - /// How the harness frames a shared-boot failure the guest left no message - /// for: it heads every assertion of every such test, so a row quoting it — - /// or quoting any part of it — would excuse the test rather than a failure. - const FRAMING: &str = "exit code Some("; + fn root() -> &'static Path { + Path::new(env!("CARGO_MANIFEST_DIR")) + } + + /// A scratch `issues/build/.md` under `dir`, so a test can point a + /// row at a file it controls rather than one this tree tracks and might + /// close out from under it. + fn write_issue_fixture(dir: &Path, path: &str, status: &str) { + let full = dir.join(path); + std::fs::create_dir_all(full.parent().unwrap()).unwrap(); + std::fs::write(&full, format!("---\nstatus: {status}\nkind: defect\n---\n\n# fixture\n")).unwrap(); + } #[test] - fn every_row_names_one_test_a_failure_and_an_issue_file_that_exists() { - let root = Path::new(env!("CARGO_MANIFEST_DIR")); - let mut seen = BTreeSet::new(); - for q in QUARANTINE { - assert!(seen.insert(q.test), "{} is quarantined twice", q.test); - assert!( - !q.says.is_empty(), - "{} quotes no failure, so its row would excuse every failure or none", - q.test - ); - for fragment in q.says { - assert!( - !fragment.trim().is_empty(), - "{} carries an empty fragment, which every failure text contains", - q.test - ); - assert!( - !fragment.contains(FRAMING) && !FRAMING.contains(fragment), - "{}: {fragment:?} is the harness's framing {FRAMING:?} and not an assertion \ - of the test, so the row would excuse every failure of it", - q.test - ); - } - assert!( - q.issue.starts_with("issues/") && root.join(q.issue).is_file(), - "{}: `{}` is not an issue file in this tree", - q.test, - q.issue - ); - } + fn every_row_names_one_test_and_an_issue_file_that_exists() { + check(DISABLED, |_| true, root()).unwrap(); } #[test] - fn a_row_excuses_the_failure_it_quotes_and_no_other() { - let row = Quarantined { - test: "reds", - says: &["it said no", "it said nothing"], - issue: "issues/x.md", + fn a_row_is_refused_for_an_unregistered_test_a_missing_issue_or_a_second_row() { + let tmp = toyos_tmpdir::TempDir::new("redlist-check"); + const ISSUE: &str = "issues/build/a-fixture.md"; + write_issue_fixture(tmp.path(), ISSUE, "expected-red"); + let registered = |name: &str| name == "a_real_test"; + check(&[Disabled { test: "a_real_test", issue: ISSUE }], registered, tmp.path()).unwrap(); + check(&[], registered, tmp.path()).unwrap(); + let refused = |rows: &[Disabled], why: &str| { + let said = check(rows, registered, tmp.path()).unwrap_err(); + assert!(said.contains(why), "{said}"); }; - assert!(row.excuses("round 2: it said no:\n")); - assert!(row.excuses("it said nothing")); - assert!(!row.excuses("the client binary was not built")); - let quotes_nothing = Quarantined { test: "reds", says: &[], issue: "issues/x.md" }; - assert!(!quotes_nothing.excuses("it said no")); + refused(&[Disabled { test: "a_renamed_test", issue: ISSUE }], "nothing registers it"); + refused( + &[Disabled { test: "a_real_test", issue: "issues/build/no-such-issue.md" }], + "is not a file in this tree", + ); + refused( + &[Disabled { test: "a_real_test", issue: "issues/README.md" }], + "is not an `issues//.md` path", + ); + refused( + &[Disabled { test: "a_real_test", issue: "CLAUDE.md" }], + "is not an `issues//.md` path", + ); + refused( + &[Disabled { test: "a_real_test", issue: ISSUE }, Disabled { test: "a_real_test", issue: ISSUE }], + "disabled twice", + ); + } + + #[test] + fn a_row_is_refused_unless_its_issue_says_status_expected_red() { + let tmp = toyos_tmpdir::TempDir::new("redlist-check-status"); + let registered = |_: &str| true; + for (path, status) in [ + ("issues/build/a-fixture-open.md", "open"), + ("issues/build/a-fixture-assigned.md", "assigned"), + ("issues/build/a-fixture-owner.md", "owner"), + ("issues/build/a-fixture-none.md", "none"), + ] { + write_issue_fixture(tmp.path(), path, status); + let said = check(&[Disabled { test: "a_real_test", issue: path }], registered, tmp.path()).unwrap_err(); + assert!(said.contains("does not open with `status: expected-red`"), "{status}: {said}"); + } + const RED: &str = "issues/build/a-fixture-red.md"; + write_issue_fixture(tmp.path(), RED, "expected-red"); + check(&[Disabled { test: "a_real_test", issue: RED }], registered, tmp.path()).unwrap(); + } + + #[test] + fn a_row_disables_its_whole_name_and_nothing_that_extends_it() { + for row in DISABLED { + assert_eq!(disabled(DISABLED, row.test), Some(row)); + assert_eq!(disabled(DISABLED, &format!("{}_controls", row.test)), None); + } + assert_eq!(disabled(DISABLED, ""), None); } #[test] - fn the_answer_is_yes_or_no_with_what_is_excused() { - let rows = [Quarantined { test: "reds", says: &["it said no"], issue: "issues/x.md" }]; + fn the_answer_is_yes_or_no_with_the_issue() { + let rows = [Disabled { test: "reds", issue: "issues/x.md" }]; let yes = answer(&rows, Some("reds")); - assert!(yes.starts_with("reds: YES"), "{yes}"); - assert!(yes.contains("it said no") && yes.contains("issues/x.md"), "{yes}"); - let no = answer(&rows, Some("greens")); - assert!(no.starts_with("greens: NO"), "{no}"); - assert_eq!(answer(&rows, None).lines().count(), 1); + assert!(yes.starts_with("reds: YES") && yes.contains("issues/x.md"), "{yes}"); + assert!(answer(&rows, Some("greens")).starts_with("greens: NO")); + assert_eq!(answer(&rows, None), "reds issues/x.md\n"); } } diff --git a/tests/CLAUDE.md b/tests/CLAUDE.md index db0d9d8f609..6c90d844e88 100644 --- a/tests/CLAUDE.md +++ b/tests/CLAUDE.md @@ -1,15 +1,13 @@ # Tests -The mechanics live where the work is: profiles and shapes in `tests/common/`, registration and tiers in `tests/toyos.rs`, known reds in `src/redlist.rs`'s `QUARANTINE`, the fast tier's line in `src/tiers.rs` — read those, not this file, for how the harness works. +The mechanics live where the work is: profiles and shapes in `tests/common/`, registration and tiers in `tests/toyos.rs`, the fast tier's line in `src/tiers.rs` — read those, not this file, for how the harness works. ## Caveats that bite every agent - **The dev host is a laptop that sleeps mid-session, and the suite says so** — a run whose wall clock jumped against the monotonic one reports `INVL` per test and exits 2: re-run. A wild outlier *not* marked that way is a real finding. -- **A landing-gate red on a test that is green alone is not therefore the host** — `STALL` is red and named apart; `ALONE: red again` on a loaded host means nothing without a same-session A/B against `main`; none of this re-runs an audio harm verdict away. - **A machine-wide kernel panic reds whichever test was running** — that red's name is the workload, never the cause. `QEMU died before ===READY=== (status 0)` is the same thing said silently: a guest that reset itself is a kernel death, and its evidence is the boot log. - **A machine-wide death during boot is reproduced by boots, not by suites** — parallel `bootable.img` guests, each waited on its completion marker and never on a fixed timer; the baseline is measured in the same session as the arm; a death counts whether or not a marker printed, so run guests with `-action reboot=shutdown -action shutdown=pause` and read a silent one's registers over QMP; a T14 block that gained a CI container mid-run is discarded, never corrected; a defect whose rate is set by interrupts per unit of guest work is measured on the *slowest* instrument — TCG can be the stronger oracle. -- **`ALONE: GREEN — its Sched::Parallel is wrong` is a hypothesis, not a finding** — the harness suggests scheduling, the mechanism decides. -- **Gate A's thorough tier reds on the dev host, and the fast tier intermittently** — stash and re-run before believing a red is yours; a plain `cargo test` boots no audio config at all; read the nightly's verdict line, not its check status. +- **Gate A's thorough tier reds on the dev host, and the fast tier intermittently** — a plain `cargo test` boots no audio config at all; read the nightly's verdict line, not its check status. - **A C test's capture has other processes' lines removed before comparison**, on the boot config's list of who may speak (`tests/common/console.rs`); a line without a trailing newline is unjoined from the next writer's there too. - **`/system/bin/init` speaks in every program's name before that program runs** — a predicate keyed on a `: ` prefix is satisfied by the wrong speaker; wait for the whole line, in the constant the assertion also reads. - **A guest binary cannot ask what a handle it does not hold does** — the probe ends its caller with exit 139, so it runs in a child, one fault per child; `handle_kill_policy` is the pattern. @@ -21,7 +19,6 @@ The mechanics live where the work is: profiles and shapes in `tests/common/`, re - **A liveness ceiling scales by two host facts** — boot-derived host speed *and* the guest's own `vcpus/cores` oversubscription. Widen a *liveness* guard for this, never a correctness bound. - **A wedge verdict needs both the budget spent and the guest gone quiet** — a healthy idle guest can be silent for minutes, and a guest still talking past its budget is slow, not stuck; only a far backstop stands behind a guest that keeps talking. - **A measured bound is asserted against the derivation, never against the measurement** — a bound that has to be widened to pass is a finding. A test asserting a kernel `Budget` never expires asserts a bound the kernel does not promise; the red is only the outcome that is neither the answer nor the declared degradation. -- **A test named as evidence may itself be quarantined** — ask `cargo run -- --known-red ` before quoting it; a row's `says` matches a *message*, not a cause, so read the capture, never the row. - **A crafted-input test asserts the harm before the return value, and never Debug-prints a refused value** — an unrefused one is as large as the input asked for. - **A stimulus sent through a channel that can silently lose it is verified before its effect is asserted** — QEMU's PS/2 queue drops the seventeenth byte, so typed input paces against the guest's report (`shell_type_once`); a guest's console reaches the host as whole lines only, so a partial line exists on no channel. - **A harness field that can be silently inert is this suite's worst defect class** — where two options can describe the same guest they refuse each other by name, and an image is asked what it is armed with. diff --git a/tests/audio-baseline.toml b/tests/audio-baseline.toml index 00105094064..69132d3d6d6 100644 --- a/tests/audio-baseline.toml +++ b/tests/audio-baseline.toml @@ -51,9 +51,7 @@ # nor skip: they compare a fresh KVM sample against this one, as they already # did across the host and the accelerator # (`issues/audio/gate-a-has-no-runner-baseline.md`), and now across the QEMU -# version as well; a timing red from them is read against this paragraph. No row -# in `src/redlist.rs` can say so, because the thorough tier ends before the -# quarantine is read. +# version as well; a timing red from them is read against this paragraph. # # RE-RECORDED 2026-08-15 on tree 4c7d809, justified by an instrument change: # the dev host's QEMU moved 11.0.3 -> 11.1.0 (owner-approved upgrade; CI's diff --git a/tests/common/qemu.rs b/tests/common/qemu.rs index c6f140a0a1e..ed205a58079 100644 --- a/tests/common/qemu.rs +++ b/tests/common/qemu.rs @@ -425,10 +425,7 @@ pub fn host_scale_self_check() -> Result<(), String> { /// is the part of "how fast is the host today" the harness knows. It does not /// know the rest, and a retry loop bounded by elapsed time has that ceiling for /// a *verdict* the moment the rest moves: a guest that is merely late reports -/// exactly what a wedged one reports. `issues/design-debt/` is the bill — -/// `desktop_audio_client` 385 s wide against 13 s alone, a landing gate that is -/// a coin toss, and six reds in four suites every one of which was -/// `ALONE: GREEN`. +/// exactly what a wedged one reports. /// /// The two are distinguishable and the console is what distinguishes them: a /// guest still printing is a guest still working. So the ceiling here is time in @@ -536,13 +533,6 @@ pub fn guest_liveness() -> Liveness { } /// A kernel line without its `[kernel cpu] ` stamp. -/// -/// The stamp is the instance and the rest is the finding, and which of the two -/// a verdict quotes decides an adjudication. `alone_line` compares the wide -/// run's sentence against the lone re-run's, and two runs of one deterministic -/// panic differ in the stamp alone — quoted whole, a staged double fault read -/// `red again on a DIFFERENT failure`, which is the harness reporting two -/// defects where there is one. The stamp is still in the capture underneath. fn without_stamp(line: &str) -> &str { if !is_kernel_line(line) { return line; @@ -551,13 +541,6 @@ fn without_stamp(line: &str) -> &str { } /// The sentence a wait gives when what stopped the guest is on the console. -/// -/// **One wording for all three waits**, so a summary line, a redlist row and an -/// issue file quote the same words wherever the wait was — and so that nothing -/// in it is a measurement of the host. The silence that proved the panic was -/// fatal is deliberately not in the sentence: it differs by a poll interval -/// between two runs of one panic, and `alone_line` compares those two sentences -/// to decide whether a re-run reproduced the defect or found a second one. fn kernel_died_here(line: &str) -> String { format!( "kernel panic: {} — the guest went quiet because every CPU is halted, not because it \ @@ -568,8 +551,8 @@ fn kernel_died_here(line: &str) -> String { /// The heading a verdict puts the guest's own account under. /// -/// One spelling, so an issue file, a redlist row and a CI log all quote the -/// same words when they quote a report. +/// One spelling, so an issue file and a CI log both quote the same words when +/// they quote a report. pub const DIED_SAYING: &str = "--- what the kernel said as it died ---"; /// The heading a verdict puts a never-announced test's window under. @@ -641,11 +624,6 @@ impl WaitVerdict { } /// The sentence, without the account under it. - /// - /// What `headline` in `tests/toyos.rs` reads off a red's reason and what - /// `alone_line` compares two runs of one defect on — so the first line is - /// still one line, and the report below it can differ between two boots of - /// the same panic without reading as a second defect. pub fn sentence(&self) -> &str { self.0.lines().next().unwrap_or_default() } @@ -776,13 +754,10 @@ pub fn ceiling_self_check() -> Result<(), String> { if early >= CEILING { return Err(String::from("staged the panic after the ceiling, so it proves nothing")); } - // The same panic on a later boot, differing only in its stamp, is the same - // sentence — or the lone re-run of a reproducible panic reads as a second, - // different defect. `alone_line` is what compares the two. let again = "[kernel 1.503 cpu7] PANIC: panicked at kernel/src/sched/reserve.rs:812:9:"; if ceiling_verdict(Some(again), early, CEILING, quiet, 40).as_deref() != Some(panic.as_str()) { return Err(format!( - "one panic on two boots gives two sentences, so a re-run reads as a second defect:\n\ + "one panic on two boots gives two sentences:\n\ {panic}\n{:?}", ceiling_verdict(Some(again), early, CEILING, quiet, 40) )); @@ -925,9 +900,6 @@ pub fn ceiling_self_check() -> Result<(), String> { )); } } - // The sentence is still one line and still the sentence — `headline` in - // `tests/toyos.rs` reads it off a red's reason and `alone_line` compares two - // runs of one defect on it, so a report under it must not become part of it. if carried.sentence() != died_verdict { return Err(format!( "the report changed the sentence a summary quotes:\n{}\n{died_verdict}", diff --git a/tests/test-durations b/tests/test-durations index 86acf871b21..d76f46dec0c 100644 --- a/tests/test-durations +++ b/tests/test-durations @@ -135,7 +135,6 @@ abuse_thread_name 1 abuse_tls_alloc 15 acpi_table_inventory 4601 allocator_stress 787 -alone_line_reports_the_alone_run 0 apps_and_home_are_one_filesystem 7263 audio_idle_suspend 1013 audio_tone (smp=1) 8450 @@ -185,9 +184,7 @@ empty_dir_stat 12 endowment_denied 10 esp_filesystem 10123 exit_wait_storm 178 -quarantine_entries 0 -quarantine_exit_status 0 -quarantine_verdicts 0 +run_exit_status 0 fat_backing_revoked 8226 fault_gates 31 foreign_disk_untouched 4688 diff --git a/tests/toyos.rs b/tests/toyos.rs index 1e413c4835a..bc6a49851a1 100644 --- a/tests/toyos.rs +++ b/tests/toyos.rs @@ -18,7 +18,7 @@ use common::{ use toyos_build::bootlog::{self, boot_millis}; use toyos_build::heartbeat; use toyos_build::testargs::{self, Shard, SUITE}; -use toyos_build::redlist::{self, Quarantined}; +use toyos_build::redlist; use toyos_build::tiers::Tier; struct TestDef { @@ -433,13 +433,9 @@ const RUST_SKIP: &[&str] = &[ // each other and never against the binaries the registry discovers. // `check_no_collisions` closes that, and this is what it found. // - // Two verdicts under one name is not extra coverage, it is a name that - // cannot be read: `retry_task` searches the shared registry first, so a - // machine test of one of these that failed wide was re-run *as the shared - // binary* and its `ALONE:` line was about a different test. What the shared - // copy adds is the binary exiting 0 on a boot that gives it nothing to - // measure — `cache_eviction` in 132 ms against the 22.5 s its own device - // shape costs (run `31247206462`). + // What the shared copy adds is the binary exiting 0 on a boot that gives it + // nothing to measure — `cache_eviction` in 132 ms against the 22.5 s its + // own device shape costs (run `31247206462`). // // `cache_eviction` needs the small NVMe that makes the cache evict at all. "cache_eviction", @@ -1606,19 +1602,12 @@ const MACHINE_TESTS: &[(&str, Sched, Tier)] = &[ // verdict it does not have. ("suspend_detector", Sched::Parallel, Tier::Fast), ("suspend_invalidates_a_verdict", Sched::Parallel, Tier::Fast), - // Same again: whether a red that is a blown liveness guard still reads as - // one by the time it reaches the summary, and whether the `ALONE:` line - // under a red is about the run it claims to be about. ("stall_is_not_a_verdict", Sched::Parallel, Tier::Fast), - ("alone_line_reports_the_alone_run", Sched::Parallel, Tier::Fast), // Same: whether two guests can still be handed one lane's NVMe image, which // is what a shared-boot reboot did to itself. ("nvme_image_is_held_by_one_guest", Sched::Parallel, Tier::Fast), - // Same: the quarantine list asking whether it still refuses the - // things it exists to refuse. - ("quarantine_verdicts", Sched::Parallel, Tier::Fast), - ("quarantine_exit_status", Sched::Parallel, Tier::Fast), - ("quarantine_entries", Sched::Parallel, Tier::Fast), + // Same: what a whole run exits with. + ("run_exit_status", Sched::Parallel, Tier::Fast), // Same: the control-register verdict, against the machine this tree // actually booted before `arch/x86_64/control_regs.rs`. ("control_regs_verdict", Sched::Parallel, Tier::Fast), @@ -11552,14 +11541,13 @@ fn run_machine_test( }; let mut qemu = QemuInstance::boot_with_options(test_config, c_bins, rust_bins, options); - // A liveness ceiling, not a pace: a loaded shard once took past a - // fixed 500 ms drain to run iod's probe (run 33246638742, alone-green). - // The T14's readback needs no drain at all — the whole boot's records - // are on the stick — so the wait is here and the predicate is shared. - let log = qemu.boot_log().to_string() - + &qemu.drain_until(Duration::from_secs(10), |l| { - l.contains("sysret-ss: reloaded") || l.contains("sysret-ss: NOT reloaded") - }); + // iod's probe reports on either side of the ready marker and the + // drain reads only lines after it, so the boot log is asked first; + // the drain's ceiling is a liveness bound on a report still owed. + let mut log = qemu.boot_log().to_string(); + if !log.lines().any(sysret_ss_reported) { + log += &qemu.drain_until(Duration::from_secs(10), sysret_ss_reported); + } sysret_ss(&log) } "fsync_failed_commit" => common::volumes::fsync_failed_commit(test_config, c_bins, rust_bins), @@ -12975,11 +12963,8 @@ fn run_machine_test( "suspend_detector" => common::clock::self_check(), "suspend_invalidates_a_verdict" => suspend_invalidates_a_verdict(), "stall_is_not_a_verdict" => stall_is_not_a_verdict(), - "alone_line_reports_the_alone_run" => alone_line_reports_the_alone_run(), "nvme_image_is_held_by_one_guest" => nvme_image_is_held_by_one_guest(), - "quarantine_verdicts" => quarantine_verdicts(), - "quarantine_exit_status" => quarantine_exit_status(), - "quarantine_entries" => quarantine_entries(), + "run_exit_status" => run_exit_status(), "control_regs_verdict" => control_regs_verdict(), "i8042_quarantine_verdict" => idle_trip_verdict(), "suite_split" => suite_split(), @@ -14777,12 +14762,7 @@ fn run_machine_test( let start = toyos_build::build::boot_start(&config.join("system.toml")); for program in &start { let line = format!("init: started {program}"); - if !log.contains(&line) { - return Err(format!( - "the shipped `[boot] start` names {program} and init never said \ - {line:?}\n{log}" - )); - } + qemu::await_marker(&mut qemu, &mut log, &line, &line)?; } serial::Serial::named("boot console", log.as_str()).must_be_clean()?; eprintln!( @@ -17005,18 +16985,35 @@ fn window_held(before: &str, during: &str) -> Result<(), String> { Ok(()) } +/// `kernel/src/arch/x86_64/hw.rs`'s probe, when it could not run. +const SYSRET_SS_UNARMED: &str = "sysret-ss: probe could not arm"; +/// The probe, when a switch refreshed SS from null. +const SYSRET_SS_RELOADED: &str = "sysret-ss: reloaded"; +/// The probe, when SS stayed null across a switch. +const SYSRET_SS_NOT_RELOADED: &str = "sysret-ss: NOT reloaded"; + +/// Whether `line` is the probe's last word: each of its three outcomes. +fn sysret_ss_reported(line: &str) -> bool { + [SYSRET_SS_RELOADED, SYSRET_SS_NOT_RELOADED, SYSRET_SS_UNARMED] + .iter() + .any(|end| line.contains(end)) +} + /// The context switch reloads SS from null before a `sysretq` can see it. /// /// Text in, a verdict out: every line it reads is a kernel record, so the /// T14's readback and a QEMU boot log are judged by this one predicate. fn sysret_ss(log: &str) -> Result<(), String> { - if log.contains("sysret-ss: NOT reloaded") { + if log.contains(SYSRET_SS_UNARMED) { + return Err(format!("the SS-reload probe could not arm, so it measured nothing:\n{log}")); + } + if log.contains(SYSRET_SS_NOT_RELOADED) { return Err(format!( "the switch did not reload SS — a sysretq here would hand userland an \ unusable one:\n{log}" )); } - if !log.contains("sysret-ss: reloaded") { + if !log.contains(SYSRET_SS_RELOADED) { return Err(format!( "the SS-reload probe never reported — iod may not have run it:\n{log}" )); @@ -18626,52 +18623,30 @@ struct Outcome { /// What the suite may conclude from one outcome. #[derive(PartialEq, Debug)] enum Verdict { - /// `Some` when the name is quarantined: still a pass, and still worth a - /// line, because one green of a quarantined test closes nothing. - Pass(Option<&'static Quarantined>), - /// Red. `Some` when the name is quarantined and this is not the failure its - /// row quotes: the row excuses one failure, and this is another. - Fail(Option<&'static Quarantined>), + Pass, + Fail, /// The host stopped in the middle of it. Neither a pass nor a fail: the /// guest, QEMU's virtual clock and every wall-clock margin the test's /// assertion rests on all jumped by however long the lid was closed, so the /// run measured something and it was not this tree. Invalid, - /// Quarantined, and failed the way its row quotes. Not red — and reported by - /// name, with its issue, on every run. - Quarantined(&'static Quarantined), } impl Outcome { fn verdict(&self) -> Verdict { - self.verdict_against(redlist::QUARANTINE) + if self.suspended >= common::clock::SUSPENDED_AT_LEAST { + return Verdict::Invalid; + } + match self.reason { + None => Verdict::Pass, + Some(_) => Verdict::Fail, + } } /// Whether this red is a blown liveness guard rather than an answer. - /// - /// Deliberately *not* a [`Verdict`] arm. A stall is red on exactly the same - /// terms as any other red — the exit code, the quarantine lookup and - /// the alone re-run all have to treat it identically, and an arm would make - /// each of those a place where somebody could decide otherwise. What it - /// changes is only what the reader is told, which is the whole complaint: - /// the run establishes nothing about this tree, so nobody should bisect it. fn stalled(&self) -> bool { self.reason.as_deref().is_some_and(|r| r.contains(STALLED)) } - - /// The table is a parameter so the gates can state a case rather than - /// depend on what the tree happens to quarantine today. - fn verdict_against(&self, quarantine: &'static [Quarantined]) -> Verdict { - if self.suspended >= common::clock::SUSPENDED_AT_LEAST { - return Verdict::Invalid; - } - let listed = quarantine.iter().find(|q| q.test == self.name); - match (&self.reason, listed) { - (None, listed) => Verdict::Pass(listed), - (Some(reason), Some(row)) if row.excuses(reason) => Verdict::Quarantined(row), - (Some(_), listed) => Verdict::Fail(listed), - } - } } /// The one line of a failure that names it. @@ -18684,188 +18659,6 @@ fn headline(reason: Option<&str>) -> String { reason.unwrap_or("check failed").lines().next().unwrap_or("check failed").to_string() } -/// What a red says when its name is quarantined and its failure is not the one -/// the row quotes. -fn quarantined_for_something_else(row: &Quarantined) -> String { - format!( - "{} is quarantined for something else, so its row does not cover this failure: {} \ - excuses {:?}", - row.test, row.issue, row.says - ) -} - -/// What the isolated re-run of one red is allowed to say about it. -/// -/// **A red-again arm quotes the alone run's own failure**, and says so when it -/// is not the failure the wide run found. `red again — the defect is real` used -/// to be the whole line, and the `failures:` summary beside it always carries -/// the *wide* run's message: on PR #22's run `31424496450` the wide run failed -/// `xhci_hid_break`'s endpoint count and the alone re-run failed its pointer -/// delivery three minutes later, and the job said `red again` over the wide -/// run's sentence — so an adjudicator read one assertion's evidence for -/// another's. Two different assertions in one job is not a weaker finding than -/// one twice; it is a different and larger one, and the line now says which it -/// was. -/// -/// Which of the two it is is not a text comparison — an assertion that prints -/// what it measured writes a different sentence every time it fires, so -/// `toyos_build::alone::same_failure` decides it and both readings are quoted. -/// -/// The green arms are untouched. They are a classification the issue files are -/// written against, and nothing about them was wrong. -/// -/// Pure, and every input a parameter, so [`alone_line_reports_the_alone_run`] -/// can stage the divergence rather than wait for CI to produce one. -fn alone_line(name: &str, wide: &str, shared_the_host: bool, alone: Option<&Outcome>) -> String { - let Some(outcome) = alone else { - return format!(" ALONE {name}: the lone run reported nothing about it"); - }; - match outcome.verdict() { - // **Two different findings, and which one it is depends on whether the - // first run shared the host** — the parallel phase's width, never the - // run's, because the serial tail is one guest at any width. Beside other - // guests, a green retry says this one was not, which is a classification - // defect. - Verdict::Pass(_) if shared_the_host => format!( - " ALONE {name}: GREEN — it fails only beside other guests, so its \ - Sched::Parallel is wrong. The run stays red on the classification." - ), - // Alone both times, nothing differed that the harness controls: it - // failed once and passed once, which is a *rate* and says nothing about - // `Sched`. CI runs one lane per machine, so every one of its retries is - // the second kind. - Verdict::Pass(_) => format!( - " ALONE {name}: GREEN, and it was alone both times — nothing the harness \ - controls differed, so it failed once and passed once. That is a rate and \ - not a classification." - ), - Verdict::Fail(_) | Verdict::Quarantined(_) => { - let said = headline(outcome.reason.as_deref()); - // One decision and one classifier: byte equality is the case - // `same_failure` already answers, so asking it first would be a - // second rule nothing tests. - if toyos_build::alone::same_failure(&said, wide) { - format!( - " ALONE {name}: red again, the same failure both times — the defect is \ - real.\n wide: {wide}\n alone: {said}" - ) - } else { - format!( - " ALONE {name}: red again on a DIFFERENT failure — it failed twice, on two \ - assertions, so this is not one defect reproduced and the divergence is \ - itself the finding.\n wide: {wide}\n alone: {said}" - ) - } - } - Verdict::Invalid => format!(" ALONE {name}: the host was suspended during the retry too"), - } -} - -/// Whether the `ALONE:` line still reports the run it is a line about. -/// -/// The staged pair is the one that was mis-reported: a wide failure and an -/// alone failure that are not the same sentence. A gate rather than a comment -/// because the defect is invisible from inside a green run — every arm prints -/// *a* plausible line, and only the quoted text says which run it came from. -fn alone_line_reports_the_alone_run() -> Result<(), String> { - const WIDE: &str = "3 endpoint(s) were found Running after the break, want 2"; - const OTHER: &str = "input never came back: no pointer event moved by (2560, -1920)"; - let red = |reason: &str| Outcome { - name: "a_test".to_string(), - reason: Some(format!("{reason}\n[kernel 2.639 cpu0] a whole capture nobody diffs")), - elapsed: Duration::from_secs(9), - suspended: Duration::ZERO, - }; - let green = Outcome { - name: "a_test".to_string(), - reason: None, - elapsed: Duration::from_secs(9), - suspended: Duration::ZERO, - }; - - // The two greens, byte for byte what they have always been: the issue - // files quote these, and a reworded classification would silently - // invalidate the record rather than add to it. - let wide_green = alone_line("a_test", WIDE, true, Some(&green)); - if !wide_green.contains( - "GREEN — it fails only beside other guests, so its Sched::Parallel is wrong. \ - The run stays red on the classification.", - ) { - return Err(format!("the shared-host green arm has changed wording:\n{wide_green}")); - } - let lone_green = alone_line("a_test", WIDE, false, Some(&green)); - if !lone_green.contains( - "GREEN, and it was alone both times — nothing the harness controls differed, so it \ - failed once and passed once. That is a rate and not a classification.", - ) { - return Err(format!("the alone-both-times green arm has changed wording:\n{lone_green}")); - } - for line in [&wide_green, &lone_green] { - if line.contains(WIDE) { - return Err(format!("a green quotes the wide run's failure:\n{line}")); - } - } - - // Red again on the same assertion: still "the defect is real", now with the - // sentence the *alone* run produced under it. - let same = alone_line("a_test", WIDE, false, Some(&red(WIDE))); - if !same.contains("red again, the same failure both times") || !same.contains(WIDE) { - return Err(format!("a reproduced failure does not say so, or does not quote it:\n{same}")); - } - if same.contains("[kernel 2.639") { - return Err(format!("the line pasted the whole capture into the summary:\n{same}")); - } - - // One assertion at two readings: still a reproduction, and it quotes both - // rather than picking one. - const MEASURED: &str = "the controller started at 0.303 s, past the 0.3 s the ports are held \ - empty for"; - const AGAIN: &str = "the controller started at 0.300 s, past the 0.3 s the ports are held \ - empty for"; - let twice = alone_line("a_test", MEASURED, false, Some(&red(AGAIN))); - if !twice.contains("red again, the same failure both times") { - return Err(format!("one assertion at two readings reads as two assertions:\n{twice}")); - } - // The labels are the whole content of the line: two sentences under swapped - // labels is the mis-attribution this gate exists to stop, said in the other - // direction. - for (label, sentence) in [("wide: ", MEASURED), ("alone: ", AGAIN)] { - if !twice.contains(&format!("{label}{sentence}")) { - return Err(format!("{sentence:?} is not the line's {label:?} run:\n{twice}")); - } - } - - // And the case the old line could not tell apart from it. - let diverged = alone_line("a_test", WIDE, false, Some(&red(OTHER))); - if !diverged.contains("DIFFERENT failure") { - return Err(format!("two different failures read as one reproduced:\n{diverged}")); - } - for both in [WIDE, OTHER] { - if !diverged.contains(both) { - return Err(format!("the divergent line drops {both:?}:\n{diverged}")); - } - } - if diverged.find(WIDE) > diverged.find(OTHER) { - return Err(format!("the divergent line reads alone-then-wide:\n{diverged}")); - } - - // The host stopping during the retry is neither, and a retry that never - // reported is not a verdict about anything. - let suspended = Outcome { - suspended: common::clock::SUSPENDED_AT_LEAST, - ..red(OTHER) - }; - let asleep = alone_line("a_test", WIDE, false, Some(&suspended)); - if !asleep.contains("the host was suspended during the retry too") { - return Err(format!("a suspended retry reads as a verdict:\n{asleep}")); - } - let missing = alone_line("a_test", WIDE, false, None); - if !missing.contains("the lone run reported nothing about it") { - return Err(format!("a retry that reported nothing reads as a verdict:\n{missing}")); - } - Ok(()) -} - /// One live guest holds its lane's NVMe image, and the next one may not. /// /// **The overlap this stages is the one the shared-boot reboot used to @@ -18966,13 +18759,13 @@ fn stall_is_not_a_verdict() -> Result<(), String> { } // Red is red. A stall that stopped failing the run would be a gate that // reports and enforces nothing. - let red = matches!(outcome.verdict_against(&[]), Verdict::Fail(_)); + let red = outcome.verdict() == Verdict::Fail; if red != reason.is_some() { return Err(format!("{what} is red={red}, and a reason is always red")); } } - let mut tally = Tally::new(&[]); + let mut tally = Tally::new(); tally.record(Outcome { name: "a_stalled_test".to_string(), reason: Some(format!("{STALLED} waiting for nothing at all — it went quiet")), @@ -19015,7 +18808,7 @@ fn stall_is_not_a_verdict() -> Result<(), String> { fn nightly_tier_is_announced() -> Result<(), String> { let held: [&str; 2] = ["desktop_window_child", "sshd_fail_closed"]; let announced = - Tally::new(&[]).holding_back(&held).summary(1, Duration::ZERO, Duration::ZERO); + Tally::new().holding_back(&held).summary(1, Duration::ZERO, Duration::ZERO); for want in [ "not run — the nightly tier", "desktop_window_child, sshd_fail_closed", @@ -19026,7 +18819,7 @@ fn nightly_tier_is_announced() -> Result<(), String> { return Err(format!("a run holding tests back never says {want:?}:\n{announced}")); } } - let whole = Tally::new(&[]).summary(1, Duration::ZERO, Duration::ZERO); + let whole = Tally::new().summary(1, Duration::ZERO, Duration::ZERO); if whole.contains("nightly") || whole.contains("held back") { return Err(format!("a run that held nothing back says it did:\n{whole}")); } @@ -19048,12 +18841,12 @@ fn suspend_invalidates_a_verdict() -> Result<(), String> { .checked_sub(Duration::from_millis(1)) .expect("SUSPENDED_AT_LEAST must be at least 1ms for this case to mean anything"); let cases: [(&str, Option<&str>, Duration, Verdict); 6] = [ - ("a pass on a host that stayed up", None, awake, Verdict::Pass(None)), - ("a fail on a host that stayed up", Some("the guest said no"), awake, Verdict::Fail(None)), + ("a pass on a host that stayed up", None, awake, Verdict::Pass), + ("a fail on a host that stayed up", Some("the guest said no"), awake, Verdict::Fail), ("a pass across a suspend", None, slept, Verdict::Invalid), ("a fail across a suspend", Some("timed out"), slept, Verdict::Invalid), - ("a pass across clock jitter", None, jitter, Verdict::Pass(None)), - ("a fail across clock jitter", Some("the guest said no"), jitter, Verdict::Fail(None)), + ("a pass across clock jitter", None, jitter, Verdict::Pass), + ("a fail across clock jitter", Some("the guest said no"), jitter, Verdict::Fail), ]; for (what, reason, suspended, want) in cases { let outcome = Outcome { @@ -19077,18 +18870,12 @@ fn suspend_invalidates_a_verdict() -> Result<(), String> { /// only reports, and what the process exits with. [`Tally::exit_code`] and /// [`Tally::summary`] are that arithmetic, and both are gated. struct Tally { - quarantine: &'static [Quarantined], passed: usize, failures: Vec<(String, String)>, /// The subset of `failures` whose guard expired rather than whose assertion /// failed, by name. Red like any other — and named apart, because a run /// that never got the guest going has measured the host and not the tree. stalls: Vec, - /// A quarantined test that failed. Reported, never red. - fired: Vec<&'static Quarantined>, - /// A quarantined test that passed. Not red, and reported: one green of a - /// quarantined test closes nothing. - quiet: Vec<&'static Quarantined>, invalid: Vec<(String, Duration)>, /// What the tier held back, by name. Not a verdict and never red — it is the /// one thing a reader of the last line cannot infer from anything else in @@ -19098,14 +18885,11 @@ struct Tally { } impl Tally { - fn new(quarantine: &'static [Quarantined]) -> Self { + fn new() -> Self { Tally { - quarantine, passed: 0, failures: Vec::new(), stalls: Vec::new(), - fired: Vec::new(), - quiet: Vec::new(), invalid: Vec::new(), relegated: Vec::new(), } @@ -19118,35 +18902,21 @@ impl Tally { } fn record(&mut self, outcome: Outcome) { - match outcome.verdict_against(self.quarantine) { - Verdict::Pass(None) => self.passed += 1, - Verdict::Pass(Some(row)) => { - self.passed += 1; - self.quiet.push(row); - } - Verdict::Fail(listed) => { + match outcome.verdict() { + Verdict::Pass => self.passed += 1, + Verdict::Fail => { if outcome.stalled() { self.stalls.push(outcome.name.clone()); } let said = headline(outcome.reason.as_deref()); - let said = match listed { - None => said, - Some(row) => format!("{said} — {}", quarantined_for_something_else(row)), - }; self.failures.push((outcome.name, said)); } - Verdict::Quarantined(row) => self.fired.push(row), Verdict::Invalid => self.invalid.push((outcome.name.clone(), outcome.suspended)), } } - /// **Three statuses, and a quarantined failure is none of them.** - /// - /// It never reaches this function, which is the statement: a run whose only - /// reds were quarantined is exit 0, and a failure on no list is exit 1. - /// - /// 2 keeps its existing meaning untouched: the run established nothing, - /// because the host stopped in the middle of it. + /// **Three statuses**: 0 green, 1 red, and 2 when the run established + /// nothing, because the host stopped in the middle of it. fn exit_code(&self) -> i32 { if !self.failures.is_empty() { return 1; @@ -19160,9 +18930,7 @@ impl Tally { /// Everything the run has to say, as one block, ending in the result line. /// /// A string rather than a pile of `eprintln!`s so that the gate can read - /// what an agent reads. **The result line names every quarantined test the run - /// judged, failed or green**: the whole hazard of this mechanism is a run that - /// looks clean, or a row that looks needed, because nobody scrolled up. + /// what an agent reads. fn summary(&self, total: usize, elapsed: Duration, suspended: Duration) -> String { let mut out = String::new(); let mut say = |line: String| { @@ -19212,26 +18980,6 @@ impl Tally { self.stalls.len(), self.stalls.join(", ") )); - say( - " The guest stopped making progress, so the run established nothing \ - about this tree and there is nothing in it to bisect. Re-run; if one \ - recurs with the host to itself, the guest really is stopping." - .to_string(), - ); - say(String::new()); - } - if !self.fired.is_empty() { - say("quarantined — known defects this run reproduced:".to_string()); - for row in &self.fired { - say(format!(" {} {}", row.test, row.issue)); - } - say(String::new()); - } - if !self.quiet.is_empty() { - say("quarantined and green this run — one green closes nothing:".to_string()); - for row in &self.quiet { - say(format!(" {} {}", row.test, row.issue)); - } say(String::new()); } if !self.invalid.is_empty() { @@ -19251,18 +18999,6 @@ impl Tally { say(String::new()); } - let named = |what: &str, rows: &[&Quarantined]| { - if rows.is_empty() { - return String::new(); - } - let names: Vec<&str> = rows.iter().map(|row| row.test).collect(); - format!(", {} {what}: {}", rows.len(), names.join(", ")) - }; - let quarantined_note = format!( - "{}{}", - named("quarantined", &self.fired), - named("quarantined and green", &self.quiet) - ); // **In the result line, because that is the line a shard's job summary // extracts and the line anybody reads.** A count of what ran means // something different depending on how much was not attempted. @@ -19273,7 +19009,7 @@ impl Tally { }; match self.exit_code() { 1 => say(format!( - "test result: FAILED. {} passed, {} failed{quarantined_note}, {} invalidated, \ + "test result: FAILED. {} passed, {} failed, {} invalidated, \ {total} total ({elapsed:.1?}){held}", self.passed, self.failures.len(), @@ -19281,7 +19017,7 @@ impl Tally { )), 2 => { say(format!( - "test result: INVALID. {} passed{quarantined_note}, {} invalidated by a \ + "test result: INVALID. {} passed, {} invalidated by a \ host suspend of {suspended:.0?}, {total} total ({elapsed:.1?}){held}", self.passed, self.invalid.len(), @@ -19292,13 +19028,8 @@ impl Tally { .to_string(), ); } - _ if !self.fired.is_empty() => say(format!( - "test result: ok, NOT clean. {} passed{quarantined_note}, {total} total \ - ({elapsed:.1?}){held}", - self.passed, - )), _ => say(format!( - "test result: ok. {} passed{quarantined_note}, {total} total ({elapsed:.1?}){held}", + "test result: ok. {} passed, {total} total ({elapsed:.1?}){held}", self.passed )), } @@ -19306,264 +19037,37 @@ impl Tally { } } -/// Every claim [`redlist::QUARANTINE`] makes that only the registry can check. -/// -/// `runnable` is the whole registry rather than the two const lists, because the -/// shared boot's C and Rust tests are discovered and a name that only exists -/// there must still be listable. -fn check_quarantine( - quarantine: &'static [Quarantined], - runnable: &BTreeSet<&str>, -) -> Result<(), String> { - for row in quarantine { - if !runnable.contains(row.test) { - return Err(format!( - "{} is quarantined and no list registers it — a renamed or deleted test must \ - take its row with it, or the quarantine is waiting for whatever gets that \ - name next", - row.test - )); - } - } - Ok(()) -} - -/// What the quarantine decides about one outcome, and what it must not. -/// -/// The negative controls are the point: a quarantined test failing any way its -/// row does not quote is an ordinary red, so is a failure of a name on no list -/// — a name that merely extends a listed one is on no list — and a suspend -/// invalidates a quarantined verdict like any other. The match is the row's -/// fragment inside the *whole* failure text, spelled as the row spells it: a -/// reason that says nothing is excused by nothing, a fragment reached only past -/// the headline still excuses, and a fragment that differs only in case is a -/// different fragment. -fn quarantine_verdicts() -> Result<(), String> { - static LISTED: &[Quarantined] = &[Quarantined { - test: "known_to_red", - says: &["the guest never answered", "the shell never answered again"], - issue: "i.md", - }]; - let awake = Duration::ZERO; - let slept = common::clock::SUSPENDED_AT_LEAST + Duration::from_secs(120); - let row = &LISTED[0]; - let cases: [(&str, &str, Option<&str>, Duration, Verdict); 12] = [ - ( - "a quarantined test failing the way its row quotes", - "known_to_red", - Some("round 2: the shell never answered again:\n"), - awake, - Verdict::Quarantined(row), - ), - ( - "a quarantined test whose failure says nothing at all", - "known_to_red", - Some(""), - awake, - Verdict::Fail(Some(row)), - ), - ( - "a quote the reason carries below its headline, as a shared boot's does", - "known_to_red", - Some("exit code Some(101)\nthe guest never answered"), - awake, - Verdict::Quarantined(row), - ), - ( - "a quote the failure repeats in another case", - "known_to_red", - Some("The Guest Never Answered"), - awake, - Verdict::Fail(Some(row)), - ), - ( - "the row's second alternative", - "known_to_red", - Some("the guest never answered"), - awake, - Verdict::Quarantined(row), - ), - ("a quarantined test passing", "known_to_red", None, awake, Verdict::Pass(Some(row))), - ( - "a quarantined test failing some other way", - "known_to_red", - Some("the client binary was not built"), - awake, - Verdict::Fail(Some(row)), - ), - ( - "an unlisted test failing the same way", - "some_other_test", - Some("round 2: the shell never answered again:\n"), - awake, - Verdict::Fail(None), - ), - ("an unlisted test passing", "some_other_test", None, awake, Verdict::Pass(None)), - ( - "an unlisted name that extends a listed one, failing the same way", - "known_to_red_controls", - Some("round 2: the shell never answered again:\n"), - awake, - Verdict::Fail(None), - ), - ( - "an unlisted name that extends a listed one, passing", - "known_to_red_controls", - None, - awake, - Verdict::Pass(None), - ), - ( - "a quarantined test across a host suspend", - "known_to_red", - Some("the shell never answered again"), - slept, - Verdict::Invalid, - ), - ]; - for (what, name, reason, suspended, want) in cases { - let outcome = Outcome { - name: name.to_string(), - reason: reason.map(str::to_string), - elapsed: Duration::from_secs(3), - suspended, - }; - let got = outcome.verdict_against(LISTED); - if got != want { - return Err(format!("{what} is {got:?}, and it has to be {want:?}")); - } - } - Ok(()) -} - /// What a whole run exits with, and what its last line says. /// /// Driven through [`Tally`] rather than asserted about it: the property that /// matters is what `--land`'s gate reads off the process, and that is the exit /// code after `record` has seen every outcome. -fn quarantine_exit_status() -> Result<(), String> { - static LISTED: &[Quarantined] = &[Quarantined { - test: "a_test_pending_on_a_defect", - says: &["the shell never answered again"], - issue: "issues/nowhere.md", - }]; - let outcome = |name: &str, reason: Option<&str>| Outcome { +fn run_exit_status() -> Result<(), String> { + let outcome = |name: &str, reason: Option<&str>, suspended: Duration| Outcome { name: name.to_string(), reason: reason.map(str::to_string), elapsed: Duration::from_secs(3), - suspended: Duration::ZERO, + suspended, }; - let fired = || outcome("a_test_pending_on_a_defect", Some("the shell never answered again:\n")); - - let mut only_quarantined = Tally::new(LISTED); - only_quarantined.record(outcome("something_else", None)); - only_quarantined.record(fired()); - let text = only_quarantined.summary(2, Duration::from_secs(9), Duration::ZERO); - if only_quarantined.exit_code() != 0 { - return Err(format!( - "a run whose only red was quarantined exits {}, and it has to be 0:\n{text}", - only_quarantined.exit_code() - )); - } - // The whole hazard is a run that reads as clean. The result line is the one - // line every reader and every log-scraper looks at, so it is the line that - // has to carry it. - let result = text.lines().last().unwrap_or_default(); - for wanted in ["a_test_pending_on_a_defect", "1 quarantined", "NOT clean"] { - if !result.contains(wanted) { - return Err(format!("the result line does not say {wanted:?}: {result}")); - } - } - if !text.contains("issues/nowhere.md") { - return Err(format!("the report never points at the issue that owns the defect:\n{text}")); - } - - // A quarantined test going green is reported and is not red. - let mut went_green = Tally::new(LISTED); - went_green.record(outcome("a_test_pending_on_a_defect", None)); - let text = went_green.summary(1, Duration::from_secs(9), Duration::ZERO); - if went_green.exit_code() != 0 { - return Err(format!("a quarantined test passing exits {}:\n{text}", went_green.exit_code())); - } - if !text.contains("one green closes nothing") { - return Err(format!("a quarantined test passing is not reported:\n{text}")); - } - let result = text.lines().last().unwrap_or_default(); - if !result.contains("1 quarantined and green: a_test_pending_on_a_defect") { - return Err(format!("the result line does not name the quarantined test that passed: {result}")); - } - - // Negative control: a red on no list is still an ordinary red, and a - // quarantined one beside it does not soften the status. - let mut real_red = Tally::new(LISTED); - real_red.record(outcome("something_else", Some("the disk came back short"))); - real_red.record(fired()); - if real_red.exit_code() != 1 { - return Err(format!( - "a run with an unlisted red exits {}, and it has to be 1", - real_red.exit_code() - )); - } - - // And a listed test failing some other way: the row must not reach it. - let mut wrong_failure = Tally::new(LISTED); - wrong_failure.record(outcome("a_test_pending_on_a_defect", Some("the client was not built"))); - let text = wrong_failure.summary(1, Duration::from_secs(9), Duration::ZERO); - if wrong_failure.exit_code() != 1 { - return Err(format!( - "a listed test failing another way exits {}, and it has to be 1:\n{text}", - wrong_failure.exit_code() - )); - } - if !text.contains("a_test_pending_on_a_defect is quarantined for something else") { - return Err(format!("the report does not say why the row did not cover it:\n{text}")); - } + let slept = common::clock::SUSPENDED_AT_LEAST + Duration::from_secs(120); - // The same two arms over the table the suite really runs under: every row - // excuses the failure it quotes, and none excuses a failure nobody has seen. - for row in redlist::QUARANTINE { - let mut quoted = Tally::new(redlist::QUARANTINE); - let quote = row.says.first().copied().unwrap_or_default(); - quoted.record(outcome(row.test, Some(&format!("round 2: {quote}:\n")))); - if quoted.exit_code() != 0 { - return Err(format!( - "{} failing the way its row quotes exits {}, and it has to be 0", - row.test, - quoted.exit_code() - )); - } - let mut unseen = Tally::new(redlist::QUARANTINE); - unseen.record(outcome(row.test, Some("a wholly different assertion, never seen before"))); - if unseen.exit_code() != 1 { - return Err(format!( - "{} is quarantined for {:?}, and failing on a text matching none of them exits \ - {}: it has to be 1", - row.test, - row.says, - unseen.exit_code() - )); - } + let mut red = Tally::new(); + red.record(outcome("a_red", Some("the disk came back short"), Duration::ZERO)); + red.record(outcome("a_suspended_one", None, slept)); + if red.exit_code() != 1 { + return Err(format!("a run with a red exits {}, and it has to be 1", red.exit_code())); } - // Exit 2 keeps its meaning: a suspended run establishes nothing. - let mut suspended = Tally::new(LISTED); - suspended.record(Outcome { - name: "something_else".to_string(), - reason: None, - elapsed: Duration::from_secs(3), - suspended: common::clock::SUSPENDED_AT_LEAST + Duration::from_secs(120), - }); + let mut suspended = Tally::new(); + suspended.record(outcome("a_suspended_one", None, slept)); if suspended.exit_code() != 2 { - return Err(format!( - "a suspended run exits {}, and it has to be 2", - suspended.exit_code() - )); + return Err(format!("a suspended run exits {}, and it has to be 2", suspended.exit_code())); } // The clean case, so that none of the above is passing because everything // reds. - let mut clean = Tally::new(LISTED); - clean.record(outcome("something_else", None)); + let mut clean = Tally::new(); + clean.record(outcome("a_green", None, Duration::ZERO)); let text = clean.summary(1, Duration::from_secs(9), Duration::ZERO); if clean.exit_code() != 0 { return Err(format!("a clean run exits {}, and it has to be 0", clean.exit_code())); @@ -19845,60 +19349,6 @@ fn driven_binaries(sources: &[String]) -> BTreeSet { found } -fn quarantine_entries() -> Result<(), String> { - static NAMED_NOTHING: &[Quarantined] = - &[Quarantined { test: "a_test_that_was_renamed", says: &["x"], issue: "issues/i.md" }]; - static GOOD: &[Quarantined] = - &[Quarantined { test: "a_real_test", says: &["x"], issue: "issues/i.md" }]; - let runnable: BTreeSet<&str> = ["a_real_test", "another_real_test"].into_iter().collect(); - match check_quarantine(NAMED_NOTHING, &runnable) { - Ok(()) => return Err("a row for a test that no longer exists was accepted".to_string()), - Err(refusal) if !refusal.contains("no list registers it") => { - return Err(format!("a stale row was refused, but for {refusal:?}")) - } - Err(_) => {} - } - // The negative control: the check is refusing that and not refusing - // everything put in front of it. - check_quarantine(GOOD, &runnable).map_err(|e| format!("a well-formed row was refused: {e}"))?; - check_quarantine(&[], &runnable).map_err(|e| format!("an empty list was refused: {e}")) -} - -/// The task that would run `name` again, by itself. -/// -/// **Every red from the parallel phase is re-run alone**, and the two possible -/// answers are both findings. Same verdict: the defect is real and the width had -/// nothing to do with it. Green: the test is red only when it shares the host, -/// which makes its [`Sched::Parallel`] wrong — a bug in this file, not in the -/// kernel, and one the suite has no other way to notice. -/// -/// **A green retry does not turn the run green.** A rerun-only pass counting as -/// a pass is selective test running by the back door; the failure line -/// says which of the two it was and the run stays red until somebody fixes the -/// classification. That is the whole safety argument for widening the parallel -/// phase: getting a scheduling answer wrong costs a red run, never a quiet one. -/// -/// A group member is re-run **as its group**, not on its own, so that the only -/// thing that changed between the two attempts is how many guests the host had. -fn retry_task<'a>(name: &str, all_tests: &[&'a TestDef]) -> Option> { - if let Some(def) = all_tests.iter().find(|t| t.name == name) { - return Some(Task::Shared(vec![def], shared_kernel(name))); - } - if let Some((registered, _, _)) = SCREEN_TESTS.iter().find(|(n, _, _)| *n == name) { - return Some(Task::Screen(registered)); - } - let (registered, _, _) = MACHINE_TESTS.iter().find(|(n, _, _)| *n == name)?; - let names = match group_of(registered) { - None => vec![*registered], - Some(group) => MACHINE_TESTS - .iter() - .filter(|(n, _, _)| group_of(n) == Some(group)) - .map(|(n, _, _)| *n) - .collect(), - }; - Some(Task::Machine(names)) -} - /// Which of the two shared boots a name belongs on — a *kernel build*, because /// `SYS_DEBUG` is compiled in or it is not, and never a boot parameter. fn shared_kernel(name: &str) -> &'static [&'static str] { @@ -20061,8 +19511,8 @@ fn run_task(task: Task<'_>, bins: &Bins<'_>, report: &std::sync::mpsc::Sender], known: &BTreeMap) { fn report_line(outcome: &Outcome) { let reason = || outcome.reason.as_deref().unwrap_or("check failed"); match outcome.verdict() { - Verdict::Pass(None) => eprintln!(" PASS {} ({:.0?})", outcome.name, outcome.elapsed), - Verdict::Pass(Some(row)) => eprintln!( - " PASS {} ({:.0?}) — quarantined, and one green closes nothing: {}", - outcome.name, outcome.elapsed, row.issue - ), - Verdict::Fail(listed) => { + Verdict::Pass => eprintln!(" PASS {} ({:.0?})", outcome.name, outcome.elapsed), + Verdict::Fail => { eprintln!("FAIL {}: {}", outcome.name, reason()); - let other = listed - .map(|row| format!(" — {}", quarantined_for_something_else(row))) - .unwrap_or_default(); if outcome.stalled() { eprintln!( " STALL {} ({:.0?}) — the guard expired, so this says nothing about \ - the tree{other}", + the tree", outcome.name, outcome.elapsed ); } else { - eprintln!(" FAIL {} ({:.0?}){other}", outcome.name, outcome.elapsed); + eprintln!(" FAIL {} ({:.0?})", outcome.name, outcome.elapsed); } } - Verdict::Quarantined(row) => { - // The reason in full, exactly as a red would print it. A - // quarantined failure is still a defect reproducing, and the run - // that reproduced it is the only place its evidence exists. - eprintln!("XFAIL {}: {}", outcome.name, reason()); - eprintln!( - " XFAIL {} ({:.0?}) — quarantined, {}", - outcome.name, outcome.elapsed, row.issue - ); - } Verdict::Invalid => eprintln!( " INVL {} ({:.0?}) — the host was suspended for {:.0?} while it ran", outcome.name, outcome.elapsed, outcome.suspended @@ -20755,12 +20188,6 @@ fn check_registration() { /// The half [`check_registration`] could not ask: the shared boot's tests are /// *discovered* from the binaries in `tests/toyos-rust-tests` and `tests/c`, so /// nothing declared can be compared against them until they exist. -/// -/// A name in both places is two tests reporting one name, and the damage is not -/// a duplicate line. [`retry_task`] searches the shared registry first, so a -/// machine test of that name which failed wide is re-run *as the other test* and -/// its `ALONE:` verdict is about neither. Four names were doing this and the -/// suite had never been able to see them. fn check_no_collisions(shared: &[TestDef]) { let mut shared_seen = BTreeSet::new(); let shared_twice: Vec<&str> = shared @@ -20784,11 +20211,24 @@ fn check_no_collisions(shared: &[TestDef]) { assert!( clash.is_empty(), "{clash:?} name both a binary on the shared boot and a test that declares its own \ - machine — two verdicts under one name, and `retry_task` takes the shared one. Add \ - each to RUST_SKIP with the reason its own test exists, or rename one of the two." + machine — two verdicts under one name. Add each to RUST_SKIP with the reason its \ + own test exists, or rename one of the two." ); } +/// `redlist::DISABLED` against every name `all_tests` plus the three declared +/// registries could produce a verdict for, before any boot on any entry point. +fn check_redlist(all_tests: &[TestDef]) -> Result<(), String> { + let runnable: BTreeSet<&str> = all_tests + .iter() + .map(|t| t.name.as_str()) + .chain(AUDIO_TESTS.iter().map(|(name, _)| *name)) + .chain(SCREEN_TESTS.iter().map(|(n, _, _)| *n)) + .chain(MACHINE_TESTS.iter().map(|(n, _, _)| *n)) + .collect(); + redlist::check(redlist::DISABLED, |name| runnable.contains(name), &compile::repo_root()) +} + fn main() { let args: Vec = std::env::args().skip(1).collect(); @@ -20802,6 +20242,14 @@ fn main() { std::process::exit(1); } }; + // The one selection every entry point below takes, the metal's included: a + // disabled test runs nowhere, and every run names each one with its issue. + for row in redlist::DISABLED { + eprintln!("[toyos] disabled: {} — {}", row.test, row.issue); + } + let keep = |name: &str| { + filter.is_none_or(|f| name.contains(f)) && redlist::disabled(redlist::DISABLED, name).is_none() + }; let debug_mode = SUITE.present(&args, &testargs::DEBUG); let list_mode = SUITE.present(&args, &testargs::LIST); @@ -20905,9 +20353,14 @@ fn main() { eprintln!("[toyos] Compiling {} C tests for the corpus boot...", c_names.len()); let c_bins = compile_c_tests(&c_names); check_metal_only_unshared(&rust_bins, &c_bins); + let c_compiled: Vec = c_bins.iter().map(|(n, _)| n.clone()).collect(); + if let Err(refusal) = check_redlist(&build_test_registry(&rust_bins, &c_compiled)) { + eprintln!("[toyos] src/redlist.rs: {refusal}"); + run.exit(1); + } let selected: Vec<(&str, &'static metal::Metal)> = METAL .iter() - .filter(|(name, _)| filter.is_none_or(|f| name.contains(f))) + .filter(|(name, _)| keep(name)) .map(|(name, decl)| (*name, decl)) .collect(); // The shared boots carry names no registration holds, so an empty @@ -20928,7 +20381,6 @@ fn main() { &dir, &selected, &{ - let keep = |n: &str| filter.is_none_or(|f| n.contains(f)); let mut boots = shared_metal(&rust_bins, keep); boots.push(c_corpus_metal(&c_bins, keep)); boots @@ -20959,10 +20411,17 @@ fn main() { let rust_bins = qemu::build_toyos_bins(&rust_tests_dir); toyos_build::build::build_host_judges(&common::compile::repo_root(), !nocapture && !debug_mode); + // Every name this process could produce a verdict for, before `--list`, + // `--debug` and `--audio-gate` can return without ever reaching it. + let all_tests = build_test_registry(&rust_bins, &c_compiled); + if let Err(refusal) = check_redlist(&all_tests) { + eprintln!("[toyos] src/redlist.rs: {refusal}"); + run.exit(1); + } + // --list: print test names and exit if list_mode { - let tests = build_test_registry(&rust_bins, &c_compiled); - for t in &tests { + for t in &all_tests { println!("{}", t.name); } for (name, _) in AUDIO_TESTS { @@ -20986,7 +20445,7 @@ fn main() { let mut audio_to_run: Vec<&str> = AUDIO_TESTS .iter() .map(|(name, _)| *name) - .filter(|n| filter.is_none_or(|f| n.contains(f))) + .filter(|n| keep(n)) .collect(); assert!(!audio_to_run.is_empty(), "no audio test matches filter {filter:?}"); // Sharded too, and this is the tier it buys the most for: the thorough @@ -21026,7 +20485,6 @@ fn main() { return; } - let all_tests = build_test_registry(&rust_bins, &c_compiled); check_no_collisions(&all_tests); check_metal_only_unshared(&rust_bins, &c_bins); // Every row against the catalogue before any boot, so a name the suite did @@ -21035,25 +20493,10 @@ fn main() { qemu::carrying(&c_bins, &rust_bins, names.iter().copied()); } check_shard_partition(&all_tests); - // Every name this process could produce a verdict for, which is what a - // quarantine row has to be one of. Taken before the filter, so a filtered - // run cannot make a stale row look well-formed. - let runnable: BTreeSet<&str> = all_tests - .iter() - .map(|t| t.name.as_str()) - .chain(AUDIO_TESTS.iter().map(|(name, _)| *name)) - .chain(SCREEN_TESTS.iter().map(|(n, _, _)| *n)) - .chain(MACHINE_TESTS.iter().map(|(n, _, _)| *n)) - .collect(); - if let Err(refusal) = check_quarantine(redlist::QUARANTINE, &runnable) { - eprintln!("[toyos] src/redlist.rs: {refusal}"); - run.exit(1); - } - let keep = |name: &str| filter.is_none_or(|f| name.contains(f)); // The tier filter, and it is not conditional on the name filter: a rule with // an exception for filtered runs is two rules, and the second one is the one - // nobody remembers. `cargo test -- desktop_window_child` refuses below and + // nobody remembers. `cargo test -- screen_diag_boot` refuses below and // says what to type instead, which is the same information a silent skip // would have withheld. let in_tier = |tier: Tier| tier.selected(nightly, shard.is_some()); @@ -21127,18 +20570,13 @@ fn main() { --nightly to run them." ); } else { - eprintln!("No tests match filter {filter:?}"); + eprintln!("No enabled test matches filter {filter:?}"); } run.exit(1); } - for row in redlist::QUARANTINE { - // Before anything boots, so that the run reads as what it is from its - // first line: a suite carrying quarantined names is not a clean suite. - eprintln!("[toyos] quarantined: {} — {}", row.test, row.issue); - } let test_config = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/testcases"); - let mut tally = Tally::new(redlist::QUARANTINE).holding_back(&held_back); + let mut tally = Tally::new().holding_back(&held_back); let suite_start = common::clock::mark(); let bins = Bins { @@ -21170,9 +20608,6 @@ fn main() { } let (mut parallel, mut serial) = build_tasks(&tests_to_run, &machine_to_run, &screen_to_run); - // Every red the wide phase produced, re-run by itself before anything is - // believed about it. See [`retry_task`] for why both answers are findings - // and why neither turns the run green. let known = load_durations(); // After the phases are decided and before either is ordered: what a shard // divides is the work, and a task's answer to `Sched` is a property of the @@ -21228,29 +20663,12 @@ fn main() { eprintln!("\nrunning {total} tests\n"); let mut timed: Vec<(String, Duration)> = Vec::new(); - // Every red, with whether it had the host to itself when it happened and - // *what it said* — the third field, because the re-run below has to be able - // to answer whether the two runs failed the same way, and by the time it - // runs this outcome has been moved into the tally. Reds only: a test the - // host slept through has no verdict to confirm, and re-running it would put - // a second guess beside the first — and a quarantined failure has already - // been answered by its row, which names the issue that owns it. - let mut reds: Vec<(String, bool, String)> = Vec::new(); - let mut collect = |outcomes: &[Outcome], shared_the_host: bool| { - reds.extend( - outcomes - .iter() - .filter(|o| matches!(o.verdict(), Verdict::Fail(_))) - .map(|o| (o.name.clone(), shared_the_host, headline(o.reason.as_deref()))), - ); - }; if !parallel.is_empty() { longest_first(&mut parallel, &known); eprintln!(" --- parallel, {width} wide ---"); let started = std::time::Instant::now(); let outcomes = run_phase(parallel, width, &bins, &slots); eprintln!(" --- parallel done in {:.1?} ---", started.elapsed()); - collect(&outcomes, width > 1); timed.extend(outcomes.iter().map(|o| (o.name.clone(), o.elapsed))); outcomes.into_iter().for_each(|o| tally.record(o)); } @@ -21259,31 +20677,11 @@ fn main() { let started = std::time::Instant::now(); let outcomes = run_phase(serial, 1, &bins, &slots); eprintln!(" --- serial done in {:.1?} ---", started.elapsed()); - // **The serial tail is one guest whatever `--jobs` says**, which is - // exactly why its reds were never re-run: the loop below was written for - // the parallel phase and read the *run's* width. So the two - // `Sched::Serial` reds of run `31252989653` — `screen_pager_keys` and - // `usb_transport_break` — carried no `ALONE:` line at all and nobody - // could say whether either was reproducible. - collect(&outcomes, false); timed.extend(outcomes.iter().map(|o| (o.name.clone(), o.elapsed))); outcomes.into_iter().for_each(|o| tally.record(o)); } qemu::set_width(1); - if !reds.is_empty() { - eprintln!(" --- re-running {} failure(s) alone ---", reds.len()); - for (name, shared_the_host, wide) in &reds { - let Some(task) = retry_task(name, &tests_to_run) else { - eprintln!(" ALONE {name}: no way to run it by itself; verdict stands"); - continue; - }; - let outcomes = run_phase(vec![task], 1, &bins, &slots); - let alone = outcomes.iter().find(|o| &o.name == name); - eprintln!("{}", alone_line(name, wide, *shared_the_host, alone)); - } - } - // Gate A, alone. `tests/audio-baseline.toml`'s numbers were recorded with // one QEMU on the host and no concurrent agents, so a run beside anything // else is not the instrument they describe — which makes this a @@ -21331,9 +20729,8 @@ fn main() { save_durations(known, &timed); } - // Three exit statuses, because there are three things a run can establish, - // and a quarantined failure is deliberately none of them — see - // [`Tally::exit_code`], which is where the whole decision now lives. + // Three exit statuses, because there are three things a run can establish — + // see [`Tally::exit_code`], which is where the whole decision now lives. // // A green run is a claim that this tree passed, and `--land`'s gate consumes // exactly this number. A run that spanned a suspend did not establish that: