Main's nightly: a vanished disk is refused at ROOT's hold instead of panicking, and the reds #506 and #527 left - #535
Conversation
`usb_transport_break` has been red since #506: its `transport_gives_up` boot leaves the gate's disk offline and still registered, `hold_source` asked `gpt::claimable` for ROOT's partition, the offline disk's table read failed, `claimable` answered `Unusable` before it had looked at the boot stick, and `hold_source` panicked the boot on it. A device crashed the kernel. `gpt::seek` now reads every disk and names the ones that did not answer instead of stopping at the first. `hold_source` holds ROOT's span when a disk that answered carries it: a silent disk that answers later either lacks it, or carries it again and makes every claim of it `Ambiguous`. Every other outcome (on no disk that answered, carried twice, refused by its table, no span a view can hold) withholds ROOT's GUID from every claim through `gpt::withhold`, which `claimable` answers `KernelDriven`, the same answer a claim of the held span gets. Neither path refuses the boot. `claimable` keeps its answers: a silent disk still makes it `Unusable`. `transport_gives_up` now asserts the hold line names the disk the gate left offline. Measured: - fix: `cargo test --test toyos-build -- --nightly usb_transport_break` EXIT=0, three runs. - negative control, the whole fix reverted onto 1ce7183 as a checked patch: EXIT=1, `PANIC: panicked at src/rootfs.rs:177:19: boot: the partition ROOT was read from, ..., cannot be held: Unusable`, wide and alone. - mutation, a silent disk not recorded (`silent.clear()`): EXIT=1 on the new assertion, wide and alone. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
`screen_diag_boot` has been red since #506. #506 moved storage behind init's spawn, which left the i8042 probe in the peripherals phase, before NVMe, xHCI, the USB disks, ROOT's hold and the mounts. The first `i8042:` line then sat 86 rows above `Boot: complete`. The diagnostic boot's panel shows the log's tail, and the T14's panel holds 67 rows, so the line that answers "why is the keyboard dead" was off the flashed machine's screen. QEMU's 3-page log no longer showed it on its last page either. Before #506 the probe ran after storage (run 36111884575's diag boot: 57 rows above the end). It now runs first in the device phase, which is after storage again. Nothing between the two positions reads a key: no task runs before `smp::set_ready`. Measured: - before: `cargo test --test toyos-build -- --nightly screen_diag_boot` EXIT=1, `"i8042:" is not on screen five seconds after the boot finished`, wide and alone. - after: EXIT=0, the first `i8042:` line 16 rows above the end. - the suites the probe's position can move, after: `--nightly i8042` EXIT=0 (13 tests), `screen_` EXIT=0 (22), `pre_idle` EXIT=0, `keyboard` EXIT=0, `console_locale` EXIT=0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
…ry flush's `home_budget_refusal_retried` and `log_flush_retry` timed out waiting for `===READY===` on the nightly (run 36285169430). Both arm `fsync-budget-spent`, which ran the first attempt of every `until_answered` run under an operation already over. Since logd owns `/log`, logd flushes after every round it wrote a line in, and each refused-then-retried flush commits four kernel records. Those records are the next round's lines, so the next flush is refused too. The storm never ends. In CI it reached `_0003.log` by 28 s and `===READY===` never reached the console. The actuator now refuses once per run: a file's `SYS_FSYNC` keyed by its `FileId`, and a claimed partition's read, write and flush keyed by device, unique GUID and kind. That matches what the actuator stands in for: a loaded host whose budget ran out once, not on every flush forever. Each of its tests still has its own first attempt refused. The loop itself is a logd defect on a device whose every flush overruns its budget. It is filed as issues/filesystem/logd-flushes-the-records-its-own-refused-flush-made.md. `home_budget_refusal_retried` now judges the storm's absence: no flush retried in the 2 s after the guest's. Measured: - negative control, the kernel half reverted as a checked patch with the new judge kept: `cargo test --test toyos-build -- --nightly home_budget_refusal_retried` EXIT=1, `220 flush(es) retried in the 2 s after the guest's` wide and `131` alone. - fix: EXIT=0. - every arm of the actuator, after: `--nightly log_flush_retry` EXIT=0, `partition_claim` EXIT=0 (3), `fsync` EXIT=0 (6), `quiesce` EXIT=0 (5), `home_` EXIT=0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
…he last word Since #527, a console holder's line (every program's, through logd) waits in `log::console`'s queue for `klogd`. `klogd` takes a chunk of records and then a chunk of the queue per hold of the wire. The stop never drained that queue. A line queued just before the stop was written either after `Rebooting.` by the power-off's `flush_final`, or not at all if the reset came first. This is issues/kernel/a-holders-queued-line-can-reach-the-console-after-the-last-word.md. It is also the shape of the nightly's `metal_job_reboot` and `quiesce_wakes_on_the_last_exit` reds: `===READY===` never reached, or a drain after `===READY===` with no kernel line in it, because the job's own kernel lines had gone out ahead of the queued marker. `quiesce()` now drains every record and queued line under the wire right after `quiesce::stop()` has stopped every holder, before `Syncing filesystems...`. `console-queue-at-the-stop` is the deterministic stimulus. It queues one line once every holder is stopped, and keeps `klogd` off the queue from the stop's claim on. `quiesce_stops_the_machine` arms it and judges that line above the last word. Measured: - negative control, the drain disabled as a checked patch that builds (`if false { ... }`): `cargo test --test toyos-build -- --nightly quiesce_stops_the_machine` EXIT=1, `1 line(s) reached the console after the boot's last word: console: a holder's line, queued once the stop had stopped every holder`, wide and alone. - with the drain: EXIT=0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Both were red on main's nightly at 1ce7183 (run 36290616312), wide and alone. Each judge's premise was the log before #527. - `usb_reset_hands_devices_back`: `nothing_after_the_last_word` took the text's last `Rebooting.`. Since #527 the next loader pass prints the boot's newest records under `log-tail:`, newest first. The last `Rebooting.` in the text is that copy, and the older records under it read as spawns after the last word ("3 of 4 reset path(s) unmet: a process started after "Rebooting."... "| log-tail: ... spawn: ..."). The judge now takes the last `Rebooting.` that is not a `log-tail:` line, and reads to the next loader pass as before. - `usb_flush_optional`: it wanted `Shutting down.` in `/log`. Since #527 `/log` ends at init's stop line, because the stop stops logd with every other thread, so the kernel's last word is on the console alone. The judge now wants init's stop line (`bootlog::stopping_line`). Measured, `cargo test --test toyos-build -- --nightly <name>`: - `usb_reset_hands_devices_back`: the old judge restored as a checked patch EXIT=1 (`3 of 4 reset path(s) unmet`), the new one EXIT=0. - `usb_flush_optional`: before EXIT=1 (`the shutdown's last line never reached the file`, wide and alone), after EXIT=0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Four update tests stalled on main's nightly at 1ce7183 (run 36290616312), wide and alone: `update_boots_the_new_kernel`, `update_falls_back_from_a_dying_kernel`, `update_floor_is_the_images_own` and `update_refusals_boot_the_other_slot`. #527 moved every stop onto init's `power` port, so the `reboot` applet (`toyos::power::stop`) asks init, which has logd make the log whole and then stops the machine. `tests/updatecase/system.toml` still gave toybox `syscap = ["power"]` and no `power` connector. The host's `reboot` over ssh was "accepted" and the machine never went down. toybox now receives `power`, as it does in `system.toml`. The `SysCap` power right goes with the change, because no applet in this image uses it any more. Measured: `cargo test --test toyos-build -- --nightly update_boots_the_new_kernel` before, EXIT=1 (STALLED after the reboot, wide and alone). After, `--nightly update_` EXIT=0 (7 of 7). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
… not fix Run 36290616312, read after #527: - gate A's `audio_tone_load.smp1` median against the dev host's TCG sample: a new sighting added to issues/audio/gate-a-has-no-runner-baseline.md. It is the instrument. Harm was null, and the same lane passed on this branch's nightly and on the one before #527. - `wake_storm_cost`: a third sighting added to its issue, green alone. - `swap_crash_rolls_back` and `i8042_health_cadence`: each red once wide and green alone twice, filed as findings. - `log_ring_keeps_the_owners_slots`, seen on the dev host: a ring's owner is named only when logd reads init's registration, so a child that floods first takes the owner's slots. Filed with the mechanism and the rates. The closing fix is in `toyos/src`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Five more interleaved rounds after the merge of 16d2e64: branch 0 of 5, origin/main 2 of 5, each red the owner-word race. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
|
Review of #535 at 64c4636 (reviewer.md, round 1) Readiness: PR CI at 64c4636 is green (run 36312296089: abi-split success; host is skipped on pull requests by design). Nightly 36297455432 ran on c271588. The two commits after it touch only Net lines, BLOCKER
NOTE
REMOVE
SEND BACK |
… cheap The review sent #535 back because nothing tested the withhold path: deleting the `WITHHELD` check in `gpt::claimable`, or `crate::gpt::withhold(guid)` in `rootfs::withhold`, survived every test. - `partclaim-root-withheld` (new actuator): device block 0 of every NVMe disk refuses reads across `rootfs::hold_source` alone, then answers again. On `InternalDisk` the one disk carries ROOT, so the hold finds it silent and withholds the GUID; `partition_claim_gives_up`'s third boot then claims ROOT's GUID, which the disk now answers for and nothing holds, and wants `PermissionDenied` (KernelDriven's word) plus the kernel's `partclaim: ... withholds it` line. - `gpt::seek` answers its own `Unnamed { Ambiguous, Unusable }`, so `hold_source`'s match has no arm for a claim error no table read makes. - `transport_gives_up` wants the hold line to name the gate's disk alone (`16 + index`, the index off `usb-gate: disk N designated`). - `nothing_after_the_last_word` moves to `bootlog` with a host `#[test]` over a crafted log-tail copy and a spawn after the real word. - `ssh_fire` refuses any answer but accepted/closed/silent; the client's `fire` reports `exited <n>` when the program came back, which a `reboot` that ended the machine never does. - The run key `fsync-budget-spent` refuses once is a closure, evaluated only in actuator kernels: a shipping partition transfer no longer takes `described`'s lock to build a key nothing reads. - The stop drains the console queue only; the record backlog stays `klogd`'s, so the sync is not spent behind a slow wire. - The false "the queue only shrinks from here" comment and the issue's contradicted paragraph are deleted; the storage-phase i8042 ordering is filed as issues/panic-path/a-storage-phase-panic-reads-its-key-off-an-unconfigured-i8042.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
|
Review of #535 at a4f68c5 (reviewer.md, round 2; last reviewed head 64c4636) Readiness: PR CI at a4f68c5 is green: run 36314453901 has abi-split success, and host is skipped on pull requests by design. The added Net lines,
Since 64c4636: kernel +74/-39, tests and src +164/-54. Round-1 BLOCKER
Round-1 NOTEs
BLOCKERNone. NOTE
REMOVE
READY |
hold_source computed one Result per arm and called withhold from four sites, so the only tested arm was Ok(None); the other three could drop their withhold silently and nothing would notice. Fold the found partition and its view into one Result<(found, view), &'static str> keyed on why it failed, and call withhold once from the single Err arm the existing test already exercises. The log text is unchanged. Also: drop the false and the branch-scoped lines from the console queue issue (main has had the queue since #527, and the branch name rots at the merge), and align the `put` row in toyos_ssh's usage table with its neighbours. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Both are sent back by the review of #542, and the owner rejects the first. - The red-streak gate (`src/ci.rs`'s `nightly-red` second step, `RED_STREAK` and everything that read the run history, and nightly.yml's `actions: read`) goes whole. A red is a red: a flaky test is disabled at once with its issue, so nothing waits two nights to find out that it reproduces. - `tests/common/update.rs` waits on the machine under `qemu::GUEST_WEDGED` again, as every other wait does. `Spans`, `A_BOOT`, `A_STOP`, `A_REBOOT`, `A_DEATH`, `QemuInstance::boot_ceiling` and `take_pending` go. `A_STOP` was a timing verdict and not a hang bound. It left out the stop's sync, which only `block::DEADMAN` bounds (120 s), and it counted `quiesce::PARK`, which that module calls a budget and not a bound. On the one-wide KVM lane it came to a flat 12 s. - `since_the_reboot` goes with them. It existed to show a refused `reboot`, and #535's `ssh_fire` now refuses one by name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Nightly 36314576406 at a4f68c5 reddened lan_swap and xhci_flap in guest (1) and handle_kill_policy in guest (8). None of them is this branch's. Each is filed where its code lives, with the recommendation that it goes on #542's disabled list when that lands. - xhci_flap: a lost wake in main's driver. Slot_gone's Teardown arm leaves the port Settled with the device in it and CSC unacknowledged, and poll steps no port unless one is dirty or outstanding. The same sentence was red on wt/toyos-lld at a55d62c (run 36287592139). On the dev host, QEMU 11.1.1 TCG, a printed serial shows every other collapse stuck about 700 ms until the next cycle's edges. The committed four-cycle gate is green by parity. At CYCLES = 3 it is red (EXIT=1), and adding `self.ports_dirty = true;` after torn_down() makes it green (EXIT=0). All measurement patches were applied checked and reverted, and the tree is clean. - lan_swap: the redial ceiling that is already recorded (issues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.md), reached on the 82574 bench. The branch changes nothing on that path. - handle_kill_policy: the same failure text, byte for byte, on main's nightly at 16d2e64. Consistent with the deferred release this branch does not touch. Named runs, dev host, one at a time, branch 877b8c9 / main 16d2e64: xhci_flap 0/0, lan_swap 0/0, handle_kill_policy 0/0, log_reserve_window_negative 0/0. The shard-8 `log-gate: FAILED` line is that negative control's owed refusal, printed by a passing test on both trees. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
One conflict, the ssh client's usage block: main's `fire` answer `exited <n>` and this branch's `probe` line, both kept. `tests/updatecase/system.toml` carries main's side. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
page_cache.rs stays deleted. #535's `partclaim-root-withheld` moves off its read-fault injector onto `block::unanswered`, which gains `answer`: the kernel drives only USB disks here, so the actuator refuses block 0 of each of those across `rootfs::hold_source` alone. main.rs keeps the branch's `gpt::probe_usb_disks` and drops the kernel mounts main still carries; main's move of `platform_devices` into the device phase merges as it is. partclaim.rs: both arms. `root_withheld` boots the partclaim config off its USB stick (`Profile::UsbDisk`) with the crafted disk beside it, because an NVMe boot here puts ROOT on a disk the kernel does not drive and would withhold it with no actuator armed at all. power.rs: the branch's `OTHERS`, with main's `console-queue-at-the-stop` actuator and its judge. storage.rs: `home_budget_refusal_retried` stays deleted with the kernel's NVMe fsync path. #535's claim that a refused flush is refused once and not on every `logd` flush moves to `log_flush_retry`'s first boot, where `/log`'s flush is fsd's claim flush: no `partclaim: a flush durable on attempt` line in the 2 s after the guest's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch works through main's nightly reds. It starts from #534's run (36285169430) and then re-measures after #527 on main's run at 1ce7183 (36290616312). #527 changed which tests were red. This PR fixes everything in both runs that has a fix here, and files the rest.
What changed, per decision
1. ROOT's hold: a disk that does not answer is named, never a panic (
kernel/src/rootfs.rs,kernel/src/gpt.rs)usb_transport_breakhas been red since #506. Itstransport_gives_upboot leaves the gate's disk offline but still registered.hold_sourceaskedgpt::claimablefor ROOT's partition, and the offline disk's table read failed.claimableansweredUnusablebefore it had looked at the boot stick, andhold_sourcepanicked. A device crashed the kernel. The issue file namedusb-port-goneas the cause. The measured cause is thetransport_gives_upsub-boot: in the main boot the gone disk is never probed.gpt::seekreads every disk and names the ones that did not answer, instead of stopping at the first. It answers its ownUnnamed { Ambiguous, Unusable }, sohold_sourcehas no arm for a claim error that no table read makes.Ambiguous.gpt::withhold, andclaimableanswersKernelDriven, the same answer a claim of the held span gets. The other cases are: on no disk that answered, carried twice, refused by its table, or no span a view can hold. Neither path refuses the boot.claimablekeeps its answers: a silent disk still makes itUnusable.transport_gives_upasserts that the hold line's silent list is exactly the gate's disk,[16 + index], with the index read offusb-gate: disk N designated.partclaim-root-withheldmakes device block 0 of every NVMe disk refuse reads acrossrootfs::hold_sourceonly, and then answer again. OnInternalDiskthe one disk carries ROOT, so the hold finds that disk silent and withholds the GUID. The third boot ofpartition_claim_gives_upthen claims ROOT's GUID. The disk now answers for that GUID and nothing holds its span. The boot wantsPermissionDenied(KernelDriven's word) and the kernel'spartclaim: <ROOT> is where ROOT was read from, and the kernel withholds it.2. The i8042 probe runs first in the device phase (
kernel/src/main.rs)screen_diag_boothas been red since #506, which moved storage behind init's spawn. That left the i8042 probe 86 rows aboveBoot: complete. The T14's panel holds 67 rows, so the line that answers "why is the keyboard dead" was off the flashed machine's screen. Before #506 it ran after storage: 57 rows up in run 36111884575. It now runs first in the device phase, which puts it after storage again, 16 rows up.This move has a cost, which is recorded rather than fixed here. A panic in the storage phase now meets the i8042 as firmware left it, and the panic panel reads the key that retires its reset bound off port
0x60. The panics in question are NVMe or xHCI init,hold_source, and the DATA and FAT mounts. The ordering is the one that shipped before #506. Moving the probe back would bring back thescreen_diag_bootred. The cost is filed asissues/panic-path/a-storage-phase-panic-reads-its-key-off-an-unconfigured-i8042.md.3.
fsync-budget-spentrefuses each run's first attempt once (kernel/src/object/ops.rs,kernel/src/syscall/device.rs)home_budget_refusal_retriedandlog_flush_retrytimed out on #534's nightly. The actuator refused the first attempt of everyuntil_answeredrun. logd flushes after every round it wrote a line in, and each refused-then-retried flush commits four kernel records. Those records are the next round's lines, so the next flush is refused too, and the loop never stops. Measured: 220 retried/logflushes in 2 s on the dev host. In CI the storm reached_0003.logby 28 s, and===READY===never reached the console.FileId, and a claimed partition's read, write and flush keyed by device, GUID and kind.described's lock to build a key that nothing reads.allow(dead_code)onRunandallow(unused_variables)onuntil_answeredstay: withoutboot-actuatorsthe closure is passed and never called.home_budget_refusal_retriedjudges that the storm is absent, over a fixed window. An absence has no event to wait on, and the site says so.issues/filesystem/logd-flushes-the-records-its-own-refused-flush-made.md.4. The stop drains the console queue before the last word (
kernel/src/syscall/machine.rs,kernel/src/log/console.rs)Since #527, a program's console line waits in
log::console's queue forklogd, and the stop never drained that queue. A line queued just before the stop was written afterRebooting.by the power-off'sflush_final, or not at all. This isissues/kernel/a-holders-queued-line-can-reach-the-console-after-the-last-word.md. It also has the shape of #534'smetal_job_rebootandquiesce_wakes_on_the_last_exitreds.quiesce()now drains every queued line under the wire right afterquiesce::stop(), beforeSyncing filesystems....klogd, so a deep backlog on a slow UART no longer runs ahead of the sync. As a result, at the stop a queued line can reach the wire ahead of records committed before it. Records and queued lines were already interleaved only chunk by chunk.console-queue-at-the-stopis the deterministic stimulus. It queues one line once every holder is stopped, and keepsklogdoff the queue from the stop's claim on.quiesce_stops_the_machinearms it and judges the line above the last word.5. Two stop judges read the log as #527 left it (
tests/common/power.rs,tests/common/usb.rs,src/bootlog.rs)usb_reset_hands_devices_back: the judge took the text's lastRebooting.. Since Logging: records from every producer, and a kernel that waits on nobody #527 that is the next loader pass'slog-tail:copy, and the copy is printed newest first. The judge now takes the lastRebooting.that is not alog-tail:line. It lives inbootlog::nothing_after_the_last_word, with a host#[test]over a crafted log-tail copy and a spawn after the real word.usb_flush_optional: the judge wantedShutting down.in/log./lognow ends at init's stop line, so the judge wants that line (bootlog::stopping_line).6.
updatecase'srebootasks init's power port (tests/updatecase/system.toml,tests/common/ssh.rs,tests/ssh-client-host)Four update tests stalled on main's nightly. #527 moved every stop onto init's
powerport, but this config still gave toyboxsyscap = ["power"]and no connector. The host'srebootover ssh was "accepted" and the machine never went down. toybox nowreceives = ["power"], as insystem.toml.firenow reportsexited <n>when the program comes back. Arebootthat ended the machine never comes back.ssh_firerefuses every answer except accepted, closed or silent, so a refused reboot fails its test by name instead of stalling it.metaltalk's own judge already accepts only those three.Each red, and where it stands
usb_transport_breakscreen_diag_boothome_budget_refusal_retriedlog_flush_retryquiesce_leaves_the_volume_wholequiesce_wakes_on_the_last_exitmetal_job_rebootsoundd_log_stallusb_reset_hands_devices_backusb_flush_optionalupdate_boots_the_new_kernel,update_falls_back_from_a_dying_kernel,update_floor_is_the_images_own,update_refusals_boot_the_other_slotaudio_tone_load.smp1medianwake_storm_cost,swap_crash_rolls_back,i8042_health_cadenceportability-windowscontinue-on-errorNothing was quarantined.
Gates, each the command's own exit code (dev host, TCG)
At a4f68c5, after the review. Each named QEMU test was run alone, one at a time. The full tier was not run here, and the orchestrator schedules it.
cargo test --test toyos-build -- --nightly partition_claim_gives_up(with the newwithheldboot): EXIT=0.--nightly usb_transport_break: EXIT=0.--nightly partition_claim(3): EXIT=0.--nightly quiesce(6): EXIT=0.--nightly update_(7): EXIT=0.home_budget_refusal_retried,log_flush_retry,metal_job,fsync: EXIT=0 each.cargo test --lib(toyos-build, 397 tests): EXIT=0.cargo test --workspace --exclude toyos-build: EXIT=0.Before the review, at 64c4636:
--nightly usb_transport_break: base EXIT=1 (therootfs.rs:177panic, wide and alone), fix EXIT=0 on three runs.--nightly screen_diag_boot: base EXIT=1, fix EXIT=0.usb_flush_optional: base EXIT=1, fix EXIT=0.--nightly update_: baseupdate_boots_the_new_kernelEXIT=1 (stalled), fix EXIT=0 (7 of 7).i8042(13),screen_(22),pre_idle,keyboard,console_locale,console,reboot,shutdown,stop,usb_reset_hands_devices_back: EXIT=0.Negative controls and oracles (high-risk: the ROOT hold, the stop path, a block-layer actuator)
--nightly partition_claim_gives_upwas run and the tree was restored clean.if WITHHELD.lock().contains(&target) { ... }block deleted fromgpt::claimable: MUT_EXIT=1, wide and alone. The guest saidROOT, withheld when its disk did not answer the boot's hold,: expected PermissionDenied, and the claim was minted.crate::gpt::withhold(guid);inrootfs::withhold. The bare deletion does not build, because-D dead-codefires ongpt::withholdbeing unused. Asif false { crate::gpt::withhold(guid); }it builds and gives MUT_EXIT=1, wide and alone, with the same guest line.gptcrate and that the kernel's parser did not choose.PANIC: ... cannot be held: Unusable.silent.push(0)ingpt::seekgives MUT_EXIT=1 onusb_transport_break, wide and alone:disks that did not answer: [0]against[17]. Oracle: the recorded real failure, main's nightly reds since The loader puts ROOT in memory and the kernel mounts it from there #506.if false { crate::log::console::drain_for_the_stop(); }, on a4f68c5),--nightly quiesce_stops_the_machineMUT_EXIT=1, tree restored clean. Alone:1 line(s) reached the console after the boot's last word: console: a holder's line, queued once the stop had stopped every holder. Wide it also caught a real program line,{0.801 tid=6 pid=6 test-runner} quiesce-writer: 5 2. The same control on the earlier whole-backlog drain was EXIT=1 too. Oracle: the recorded red in the issue (618e68e,quiesce_stops_the_machine).fsync-budget-spentonce per run. Negative control: the kernel half reverted with the new judge kept gives EXIT=1 with220 flush(es) retried in the 2 s after the guest'swide and131alone. Oracle: CI run 36285169430's console, the storm to_0003.log.match None::<&&str> {inbootlog::nothing_after_the_last_wordbuilds and givescargo test --lib bootlog::tests::a_spawn_afterexit 101, red on the spawn-after-the-word assertion. The old judge restored gives EXIT=1 onusb_reset_hands_devices_back.rebootover ssh (6).updatecaseput back tosyscap = ["power"]as a checked patch gives--nightly update_boots_the_new_kernelMUT_EXIT=1 in 4 s, wide and alone:`reboot` over ssh answered "exited 1", so it did not end the machine. Before this change, the same config stalled the test to its timeout.What I am unsure of
log_ring_keeps_the_owners_slotsis red beside the otherlog_guests on both arms: 3 of 14 on this branch and 2 of 13 on origin/main, in interleaved--nightly log_runs. Each red is the owner-word race: a ring's owner is named only when logd reads init's registration. Filed asissues/kernel/a-log-rings-owner-is-named-only-when-logd-reads-its-registration.md. The closing fix is intoyos/src, which this brief may not touch.issues/audio/gate-a-has-no-runner-baseline.md, which needs a per-host baseline, a schema decision. On this lane harm was null: 0 dropouts, 0 underruns, 0 ceiling breaches. The same lane passed on this branch's nightly at dbf4ace and on the one before Logging: records from every producer, and a kernel that waits on nobody #527.Round-2 review fixes (877b8c9)
rootfs::hold_sourcecalledwithholdfrom four arms, and the only test reaching any of them wasOk(None). Restructured so the found partition and its view are computed as oneResult<(found, view), &'static str>keyed on the refusal reason, andwithhold(guid, why, &sought.silent)is called from the singleErrarm of the outer match. The log text is byte-identical. The review's named mutation, replacing an arm's ownreturn withhold(...)withreturn, no longer has a line to apply to: each arm now yields a plainErr(why)value, and the one remainingwithholdcall is the one the existingpartition_claim_gives_upboot already exercises.issues/kernel/a-holders-queued-line-can-reach-the-console-after-the-last-word.md: main has carried the console queue since Logging: records from every producer, and a kernel that waits on nobody #527, and thenightly-green2branch name rots at the merge.putrow intoyos_ssh's usage doc-comment with its neighbours.cargo test --test toyos-build -- --nightly partition_claim_gives_up: EXIT=0.cargo test --workspace --exclude toyos-build: EXIT=0.Nightly on this branch
Run 36297455432 on c271588 (this branch merged with main at 16d2e64). It predates the review round.
audio_tone_load.smp1 wake lateness: median 5765 -> 6556 (z=4.07). Harm was null: dropouts 0/60, underruns 0, ceiling breaches 0/60. The runner's sample did not move across Logging: records from every producer, and a kernel that waits on nobody #527. Each post-Logging: records from every producer, and a kernel that waits on nobody #527 sample against the pre-Logging: records from every producer, and a kernel that waits on nobody #527 one gives Mann-Whitney z=1.40 (main 1ce7183), -1.20 (this branch at dbf4ace, where gate A passed) and 0.84 (this run). The gate compares against the dev host's TCG sample, and its verdict flips at this median. That isissues/audio/gate-a-has-no-runner-baseline.md: it needs a per-host baseline, a schema decision, and it is not fixed here.continue-on-error).An earlier run on dbf4ace (36292135439) was cancelled once the merge landed. Before cancelling, guest (12) had shown the
usb_reset_hands_devices_backjudge red, which is fixed in 15408f8. Its audio (2) passed.Nightly 36314576406
Run on a4f68c5, two commits before the head. None of the reds is this branch's. Named runs on the dev host (QEMU 11.1.1, TCG), one at a time: branch at 877b8c9, and
mainat 16d2e64 in its own worktree. CI ran QEMU 11.1.0.xhci_flap(guest 1)EXIT=0EXIT=0lan_swap(guest 1)EXIT=0(12 dials turned away)EXIT=0(11)handle_kill_policy(guest 8)EXIT=0EXIT=0log-gate: FAILEDline (guest 8)EXIT=0EXIT=0xhci_flap:issues/hardware/a-collapsed-replug-is-enumerated-only-when-another-port-event-arrives.md. When a collapsed replug's teardown completes,slot_goneleaves the portSettledwith the device still in it and CSC unacknowledged.pollsteps no port unless one is dirty or outstanding, so the device waits for an unrelated event. The same sentence was red onwt/toyos-lldat a55d62c, an ancestor of main (run 36287592139). On the dev host, a printed serial shows every other collapse stuck for about 700 ms until the next cycle's edges arrive. That makes the committed 4-cycle gate green by parity. Checked patches, each reverted in the same script:CYCLES = 3reds (EXIT=1).CYCLES = 3plusself.ports_dirty = true;aftertorn_down()is green (EXIT=0).EXIT=0).QEMU's
hcd-xhci.cis byte-identical at v11.1.0 and v11.1.1. The fix belongs to the driver's owner and is not in this PR. Recommend Test suite, first pass: a disabled list replaces the quarantine, sysret waits on its report, update_* reboot through the power connector #542's disabled list until it lands.lan_swap:issues/build/lan-swap-redial-spent-its-ceiling-on-a-nightly-shard.md. The host's redial was turned away 64 times and gave up. The guest's own console completed the swap (in serviceat 6.141 s), andlogdlistened again at 1.176 s but admitted no reader after that. This is the ceilingissues/diagnostics/a-swaps-redial-asks-again-with-no-event-to-wait-on.mdrecords, on the 82574 bench.swap_crash_rolls_backhit it on main's nightly at 1ce7183. This branch changes nothing on the swap path:Ssh::swap,metalswap,Stream::redial, logd, netd and init are untouched, and thefirechange reaches onlyssh_fireandSsh::fire. Recommend Test suite, first pass: a disabled list replaces the quarantine, sysret waits on its report, update_* reboot through the power connector #542's disabled list.handle_kill_policy:issues/kernel/handle-kill-policy-census-grew-one-sharedmem-on-two-nightlies.md. Main's nightly at 16d2e64 (run 36306830048, guest 8) failed the same way, byte for byte:[("SharedMem", 9, 10)], with identical first-census counts. It was green on the nightlies at 1ce7183 and at c271588, which already carries 16d2e64, so it is a rate. It is consistent withissues/kernel/deferred-release-outlives-its-syscall.md: the settle takes two readings 10 ms apart, and the red boot reports TLB shootdown waits of up to 13667 us. Which process held the tenth region is not shown. Recommend disabling it via Test suite, first pass: a disabled list replaces the quarantine, sysret waits on its report, update_* reboot through the power connector #542's list.log-gate: FAILED: cpu7 seq 517 …: printed bylog_reserve_window_negative, which passed. That test is the negative control that removes the reserve bracket, and this line is the refusal it owes ([log] unbracketed: …). Main's nightly printed the same shape (cpu0 seq 726) under the same passing test. Both named runs print it and pass.audio (2): known,issues/audio/gate-a-has-no-runner-baseline.md, and not fixed here.portability-windows: the declared frontier.No code changed after this run: 00ec52e adds only the three issue files.
cargo run -- --known-redanswers NO for all four names.🤖 Generated with Claude Code
https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK