Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
---
status: expected-red
kind: tooling
opened: 2026-09-28
---

# `quiesce_leaves_the_volume_whole` passes only when the refused flush closes inside the stop's `PARK`, which a host stall takes away

The verdict needs `fsync: … durable on attempt 9` on the console before
`Syncing filesystems...`. In other words, the `quiesce-fsync-refuse` ladder has
to close before `quiesce::stop` spends `PARK`. The
kernel's own `const` assert in `fat32_adapter::mirror_refuse` only covers the
parks: 1270 ms of `RETRY_SOONEST` doubling against `PARK`. The I/O of nine
attempts, and any time the guest is not running, are not covered by anything.
`PARK`'s own doc says a thread can outlast it and that the record then names
the shortfall. So the test asserts an outcome the kernel does not promise, and
under TCG the guest clock runs with the host's.

The one red, the Fast tier for PR #562 at `2a9c77ee`:

```
refusals 1..8 at 630 643 658 681 723 809 974 1304 ms (the nominal ladder)
{0.649 init} init: power: the machine stops …
[kernel 2.703 cpu0] Syncing filesystems...
[kernel 2.705 cpu1 tid=1] fsync: /log/quiesce-fsync.bin durable on attempt 9 after 2073ms
stop: 6 of 7 userland thread(s) stopped … in 2037 ms of a 2010 ms budget
```

Attempt 9 was due at about 1944 ms (a 640 ms park after 1304). It closed at
2705 ms, two milliseconds after the stop's own deadline wake, which fired
27 ms late. A single stall of the whole guest from before 1944 ms to past
2676 ms explains both late wakes firing together. Nothing in the guest was
waiting on the other. Nothing here is PR #562's doing either: on that branch
the kernel paths this boot runs (`quiesce.rs`, `block.rs`, `fat32_adapter.rs`)
are the same as on `main`. The one change to `quiesce_fsync.rs` removes a
deadline the guest only reached on a hang.

## Exit condition

The verdict no longer rests on guest time. One way: a stop whose budget ran out
over the parked update is read as its own outcome, with the volume judged whole
or not by the checker. Another: the actuator holds the ladder open until the
stop has swept, rather than for a fixed ladder of parks. Then this file and its
`src/redlist.rs` row are deleted.

## Owner

`tests/common/volumes.rs` `quiesce_leaves_the_volume_whole`, the
`quiesce-fsync-refuse` actuator. Nobody holds it.
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
---
status: expected-red
kind: defect
opened: 2026-09-28
---

# `copy-meets-a-remap` holds a CPU with `IF` clear, waiting for a thread that may be queued behind it on that CPU

`kernel/src/user_ptr.rs`'s `remap_race::hold` spins inside the copier's
syscall with interrupts off until the racing process maps again, and panics
after `BOUND` (10 s). The thread that has to map is the program's main thread.
Nothing puts it on another CPU: `CpuHandles::place` puts the copier on the
least-loaded CPU that is answering, and a steal is a `StealRequest` that only
the victim's own pass answers (`SchedPass::answer_steal_requests`). A CPU
spinning with `IF` clear takes no pass. If the main thread is queued on that
CPU, no other CPU can take it, and the kernel panics.

In the orchestrator's Fast tier for PR #562 at `2a9c77ee` (a two-CPU guest):

```
[kernel 10.738 cpu1] sched: cpu=1 ready=0 dying=0 stopped=0 parked=4 current=None trips=91
[kernel 11.142 cpu0 tid=1] PANIC: panicked at src/user_ptr.rs:402:13:
copy-meets-a-remap: pid 6 held a copy 10000ms and never mapped again
cpu0 is on ctx … pid=6 tid=1 (the copier, in `copy_out::<SchedInfo>` → `remap_race::hold`)
cpu1 is on ctx … pid=3 tid=0
```

cpu1 was idle with nothing ready, and the main thread (pid 6 tid 0) was on
neither CPU. Queued behind the spin on cpu0 is the reading that fits all of
this. Nothing in the capture proves it: no line says which queue held the
thread.

Also seen in the orchestrator's nightly for PR #536 at its head `069722c3`,
the same panic site, `src/user_ptr.rs:402:13`.

## Exit condition

The hold cannot strand the thread it waits for. For example: the racing thread
is placed on a different CPU from the copier before the cue, or the hold waits
with the CPU able to run passes. `user_copy_races_munmap` is then green, and a
mutation that puts both threads on one CPU reds with a line saying so rather
than with this panic. Then this file and its `src/redlist.rs` row are deleted.

## Owner

`kernel/src/user_ptr.rs` `remap_race`, `tests/toyos-rust-tests/src/bin/copy_out_races_munmap.rs`. Nobody holds it.
21 changes: 10 additions & 11 deletions issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md
Original file line number Diff line number Diff line change
@@ -1,24 +1,24 @@
---
status: open
status: expected-red
kind: tooling
opened: 2026-08-23
---

# `syscall_window_nmi` under-counts window arrivals on a contended host

Seen once in passing, on a dev host running a second worktree's 12-wide suite
against the same twelve guest slots (2026-08-23):

```
FAIL syscall_window_nmi: 44 window arrivals against 572 in Ring 3. Every
iteration passes through both exactly once, so they are of one order; a 10x
shortfall says the arrivals are not being classified where they land
```

Green in the same session's alone re-run (4 s) and green again on a quiet
re-run of the same tree. `cargo run -- --known-red syscall_window_nmi` says
`NOT ON THE LIST`, so no rate has ever been written down for it and this is the
first datum rather than a regression against one.
Also seen in the orchestrator's own runs at a main-level head, `c5d09bb6`:

```
FAIL syscall_window_nmi: 24 sprayed window arrivals against 515 in Ring 3. Every
iteration passes through both exactly once, so they are of one order; a 10x
shortfall says the arrivals are not being classified where they land
```

Two readings and nothing here separates them: the storming CPU genuinely lands
in the three-instruction window less often when the host is oversubscribed —
Expand All @@ -31,12 +31,11 @@ rate measured on a host whose company is recorded (`tests/CLAUDE.md`).
Not `Sched::Parallel` being wrong. The harness suggests that on every alone-green
red, and re-classifying a red whose mechanism is unknown answers nothing.

**2026-08-25, promoted to `defect`.** A test that reds on a loaded host with no
**2026-08-25.** A test that reds on a loaded host with no
rate written down is an unadjudicated red, and CLAUDE.md's rule is that such a
red is fixed at its owner rather than re-run away. The act is a measurement: a
window-arrival rate taken across widths on hosts whose company is recorded, so
the host reading can be excluded before the classification reading is
investigated. Until that exists nothing can decide whether the assertion bounds
this kernel or the dev host. Owed by whoever next runs a load sweep on this
instrument; `cargo run -- --known-red syscall_window_nmi` still answers `NOT ON
THE LIST`.
instrument.
12 changes: 12 additions & 0 deletions src/redlist.rs
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,10 @@ pub const DISABLED: &[Disabled] = &[
test: "partition_claim_departure",
issue: "issues/boot-media/partition-claim-departure-exits-clean-with-none-of-its-refusals-said.md",
},
Disabled {
test: "quiesce_leaves_the_volume_whole",
issue: "issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md",
},
Disabled {
test: "quiesce_stops_the_machine",
issue: "issues/kernel/a-quiesce-writers-first-pass-outlasts-the-jobs-five-second-spin-up.md",
Expand Down Expand Up @@ -88,10 +92,18 @@ pub const DISABLED: &[Disabled] = &[
test: "swap_netd",
issue: "issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md",
},
Disabled {
test: "syscall_window_nmi",
issue: "issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md",
},
Disabled {
test: "usb_transport_break",
issue: "issues/kernel/a-held-disk-waits-for-a-pass-no-cpu-takes-when-every-cpu-is-in-a-call-on-it.md",
},
Disabled {
test: "user_copy_races_munmap",
issue: "issues/kernel/copy-meets-a-remap-holds-a-cpu-the-thread-it-waits-on-may-be-queued-behind.md",
},
];

/// The row of `rows` that disables `test`, matched by the whole name.
Expand Down
Loading