Kernel: a kill never waits on its victim — the last thread out tears its process down - #549
Conversation
…ishes it Two processes holding each other's handle and killing at once deadlocked: each killer sat in `retire_task` on a thread that was itself inside the other kill, reached no safe point, and the kernel panicked at the 10 s tripwire (audit K1). TRANSFER makes it reachable from userland. `kill_process` now claims, posts every retire (`scheduler::post_retire`) and hands the rest to `reaper`, a kernel thread no process can kill, which waits the threads out (`scheduler::await_released`) and runs the same teardown tail as exit. The reaper stops with userland at a shutdown (`quiesce::exempt`): it runs userland's teardowns, so a stop waits for one in progress and stops it when parked, as it did the killer. A process's end is pollable: its handle's read watch is the object's watch and it is readable once the exit is published; closing one handle ends no poll. `toyos::process::Process::watch_end` is the SDK side. toyos-proclife's model now makes a retire a wait — a thread inside a scripted op reaches no safe point until it returns — adds the reaper op, a deadlock check and L6 (every claimed teardown publishes), and the KillEachOther cases. `mutate-kill-waits-for-its-victims` restores the base's inline kill and reds them. Guest tests: `mutual_kill`, and a poll arm in `process_lifecycle`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ling is a killer The audit's K1 names the sibling form too (A2 kills B while B1 kills A). 19 schedules, every one ends; red under mutate-kill-waits-for-its-victims. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the stop names The reaper stops with userland (quiesce::exempt), so the stop's exact thread count is one higher. Measured: 11 of 11 before this, and green after. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review of 61c0a7f: CI is red at this head. Run 36330513554 NOT READY FOR REVIEW |
|
Review of PR #549 at 61c0a7f (high-risk: process lifecycle, scheduler) CI:
No CI job runs QEMU, so BLOCKER
NOTE
REMOVE
Brief questions
SEND BACK |
… asserts it Answers the review of #549 at 61c0a7f. The reaper no longer pops its queue in order and waits for the head. Every thread release now also posts `sched::payload::ANY_RELEASED`, and the reaper waits on that (or on its own `WORK` when nothing is owed). It takes the oldest owed kill whose threads are all released, so a slow victim no longer delays another kill's release. The tripwire is per kill and counts from when the kill was owed (`scheduler::RETIRE_GIVE_UP`, the same 10 s). The teardowns still run one at a time on one thread. That is filed as issues/kernel/the-reaper-runs-every-kills-teardown-one-at-a-time.md. `fold_retired` (was `retire_threads`) asserts that every thread it folds is released before any teardown frees what they ran on. Both paths wait before they reach it: exit through `await_released`, and the reaper through its choice. Mutation: the reaper takes a kill whose threads are not all released (`killed.released() || true`). mutual_kill then goes red with "teardown: a thread of the process is still on the scheduler", both wide and alone. Deleted: - `scheduler::retire_task`. The exit path posts every sibling's retire first, then awaits each one, so siblings die concurrently. - `retire_threads`' `retire` parameter. - `reaper::REAPER`/`is` and the `kthread::spawn` return change. A kernel thread now declares `OnStop::Runs` or `OnStop::Stops` in its row at spawn, and `quiesce` reads the row (`kthread::runs_through_the_stop`). - The self-kill branch in dispatch. A self-kill posts its own retire, dies at its Ring 3 boundary, and the reaper publishes 137. mutual_kill now checks that. Also: - `block::counted` reads the same row, so a block operation the reaper opens is counted as the stop's. - The stop record's thread count loses "userland": it now counts the reaper too. - `Shootdown::serve` raises `flushed` with `fetch_max`, so a serve nested inside another cannot move it backwards (audit K2, now reachable from the reaper's IF=1 teardown). kernel-loom's `a_nested_serve_is_not_undone_by_the_one_it_interrupted` reds under `store`. - process_lifecycle: closing another handle to a running process leaves a poll on it empty. It reds under `KObjectRef::Process(_) => true`. - kill_while_blocked arm 4 kills from the observer. The separate killer existed only because a kill used to wait. - The model's reaper takes a released kill, and its exit posts every retire before it waits. - Stale `retire_task` citations in toyos-sched follow the rename. The sim paragraph that said the kernel never batches one process's retires is deleted, because it now does. Filed: issues/kernel/a-manage-holder-can-kill-a-kernel-thread.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review of PR #549 at a7457a3, round 2 (high-risk: process lifecycle, scheduler) CI at a7457a3: Round-1 BLOCKERs
BLOCKER
NOTE
REMOVE
SEND BACK |
…hoice is the model's A kernel thread's pid names no process a handle could hold: `process::process_object`, the one door from a pid to a `Process` object, refuses it, so `SYS_PROCESS_OPEN` answers NotFound and a MANAGE holder can no longer kill iod, usbd, klogd or the reaper into the tripwire's halt. The control is a second verdict under `process-reopen-selftest`: after the last kthread is spawned, every row's pid is asked for its object. The reaper's choice of owed kill is `toyos_proclife::teardown::next_released`, which the kernel and the model both call; the model's own copy is deleted. The model's reaper now parks on ANY_RELEASED the way the kernel's does, a kill's posts, its owe and its return are separate sections, and a new law (L7) holds at every state: the reaper never sleeps while a released kill is owed. The explorer skips a state it has reached before, since every law is a property of the state alone; without that the three-kill chain took 78 s. ANY_RELEASED is posted only on a killed thread's release. The kill mark and the release share one word on the TaskHandle, so a release and a kill posted at once decide by one read-modify-write each: `Hw::release` holds no `TaskShared` whose kill bit it could read. The idle reaper waits on ANY_RELEASED too, which deletes `WORK`, the idle branch and the re-check. The teardown issue is deleted: every handle drop in a teardown either only queues work (a file's write-back) or is a deferred row released from the zero queue, so no stimulus can hold one. The kthread issue is deleted with its fix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review of PR #549 at 1c5ca17, round 3 (high-risk: process lifecycle, scheduler) CI at 1c5ca17: run 36336743687, Round-2 BLOCKERs
BLOCKER
NOTE
REMOVE
SEND BACK |
…ire wait A kill claims its victim, posts every thread's retire and returns. Each thread leaves by its own hand (`process::leave`): at its Ring 3 exit boundary when killed, in `exit` or `thread_exit` otherwise. It leaves its address space, folds its blocked-time accounting into the process's and marks itself out under the table lock, and the one whose leaving empties a claimed process frees the process and publishes its exit on its own stack. Its kernel stack goes as every stack does, in `Hw::release` on the pass after the switch. Nobody waits for another thread's release, so a kill cannot deadlock on its victim by construction. Deleted: `kernel/src/reaper.rs` and its kernel thread, `ANY_RELEASED`, the owed queue, `next_released`, the `OnStop` row, `scheduler::await_released` and `RETIRE_GIVE_UP` with the reaper's tripwire, `TaskHandle`'s released flag and accounting copy, `ProcessAccounting::child_threads_cpu_ns`, and the teardown's thread-by-thread fold. `issues/kernel/retire-tripwire-is-not- queue-shaped.md` closes: its constant is deleted and nothing replaces it. The decisions are `toyos-proclife`'s and the kernel calls them: `claim_teardown` now carries the exit code, `retire_set` is the claim's retire set, and `leave` answers whether the thread leaving is the last one out. A poisoned thread's death is its leaving, done for it by the idle loop: a poisoned main thread claims its process and retires the rest, and a poisoned last thread publishes the claim's code. A join does not collect a thread of a process being torn down: that thread's TLS is still mapped where its siblings may run. The model drops the reaper and scripts every thread's way out as its own steps. Laws at every state: one claim and one teardown per process, nothing freed or published while a thread is still in, a stack freed only after the switch, no mapped TLS dropped under a running sibling, and no kill or exit waiting on another process's teardown; at every leaf, every claimed process published and every killed thread gone. New controls: `mutate-first-out-tears-down` and `mutate-join-collects-in-a-teardown`. `kill_ends_every_wait` kills a child parked in a futex, a poll ring, a process wait, a thread join and a sleep, and waits for each to end. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review of PR #549 at fa1e325, round 4. The design is new (last thread out) and is judged whole. This is high-risk code: the process lifecycle, the scheduler and the quiesce stop. CI at fa1e325: run 36341735456, Earlier BLOCKERs
BLOCKER
NOTE
REMOVE
SEND BACK |
|
PR #553's round-5 review corrected its "What #549 must drop" section. Posting the corrected list here since it is a merge instruction for this PR, not part of #553's own record:
Not a drop, but shared ground: |
main's #553 deletes panic recovery: every kernel panic halts. This branch's poison work goes with it: `toyos-proclife/src/poison.rs`, `Op::Poison` and its two scripts, `World::poison` and its `poisoned` set, `process::PoisonWake` and `zombify_poisoned`, the idle loop's poison bank in `reap_finished`, the TLS issue a poisoned sibling raised, and the "poison path" wording in `teardown.rs` and `ProcessEntry::teardown_code`. `kthread::open_selftest` reads `ROWS` as `[AtomicU64]`. `mark_thread_zombie`, which main kept and this branch's `leave` replaced, goes. main.rs's `usbd::start()` context is gone with usbd. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
The machine's stop counts a zombie as nothing left to stop. The last thread
out was marked a zombie by `proclife::leave`, and `note_progress` posted,
before its teardown ran; a stop taken then returned while that teardown's
`close_all` still dropped files, and the shutdown's `drain_all` and
`sync_all` ran beside it. A dropped file's write-back was queued after the
drain, and its dirty pages were not durable at power-off.
`proclife::leave` now marks every thread but the last one out, and answers
`Last { code, mark }`; `proclife::torn_down` marks that one in
`teardown_bookkeeping`, after its records and releases, and `note_progress`
posts after that mark. The model scripts the mark as its own step, and its
L3 now reads: a teardown in flight keeps its thread in the process. L3's
old sentence, a stack freed only after the switch, checked the model's own
step order; toyos-sched-sim's I11 is that check, and the model's copy and
its hand-run teeth test go.
`mutate-last-out-leaves-before-its-teardown` marks the last one out before
its teardown, as the head before this did; it reds
`the_last_one_out_is_in_its_process_until_its_teardown_is_done` on L3 and
`only_the_thread_that_empties_a_claimed_process_tears_it_down` on its
location assertion. `quiesce-last-teardown` holds the last thread out of a
child of `quiesce_last` between its leaving and its teardown until the stop
counts it alone; `quiesce_wakes_on_the_last_teardown` judges that boot.
The exit boundary calls `process::leave` at its own preempt depth: nothing
the teardown runs parks, and a park there asserts. `current_data` and
`process_data` panic on a missing entry rather than ending the thread
without leaving: an entry outlives every thread that has not left it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Nothing but `process_lifecycle`'s own arm polled it: the `Process` arms of `read_watch`, `has_data` and `close_ends_polls`, `ProcessObject`'s `Arc<Watch>`, `toyos::process::Process::watch_end` and `its_end_completes_a_poll` go back to main's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
A kill returns before its victim ends, so the killer can observe the end itself: arm 4 kills the spinner and waits for its exit code. The in-guest `ENDS_WITHIN` deadline and the timed stdout reader go; a spinner the exit boundary misses never ends, and the harness's hang ceiling says so. The deferred-release issue loses the sentence whose only reason was `retire_task`'s park. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
`a_nested_serve_is_not_undone_by_the_one_it_interrupted` ran the nesting as one written-out schedule, one execution. The nested issue and serve now run on a spawned thread, so loom places them at every point of the outer serve: 11 executions at a preemption bound of 2, and the model reds when `serve` stores what it owes instead of raising to it. Filed: `sync.rs`'s lock spin says it runs with `IF` clear, and kernel threads and the idle loop spin there with it set, which is where one serve nests inside another. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
|
Review of PR #549 at 0c48340, round 5. This is high-risk code: the process lifecycle, the scheduler and the quiesce stop. CI at 0c48340: run 36388092522 ( Earlier BLOCKERs (round 4, fa1e325)
The orchestrator's questions
BLOCKER
NOTE
REMOVE
SEND BACK |
Takes #560, #541, #565, #563, #569 and #570. `src/ci.rs` and `tests/toyos.rs` merge without conflict; `rust` takes main's pin, 1b236638, since this branch carries no fork commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…lone `hold` released on the first sweep that counted one thread running, which need not be the held one. With the last thread out marked a zombie before its teardown, no sweep counts it; a sweep that counted some other thread still in Ring 3 released the hold, and `quiesce_wakes_on_the_last_teardown`, which judged on `sweeps >= 2`, passed with the teardown unwaited. `hold` now claims the held thread's id, `sweep` notes whether it counted that thread as running, and the hold ends only on a sweep that counted exactly one running thread, the held one. It then logs `<actuator>: the stop counts quiesce-last alone`. The harness requires that line before the `stop:` record for all three `quiesce-last-*` boots, and for the teardown boot also the held process's `exit: test_rs_quiesce_last pid=... code=0` record between the two. The Teardown hold sits in `process::leave`, which the Ring 3 exit boundary also reaches, at depth 0, where `yield_now` asserts. `hold` refuses by name there (`scheduler::may_yield`), logs that the thread left outside a syscall, and lets the stop go on unheld; the harness then reds on the missing line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…h IF set `kernel_exit_to_user_check` is entered with IF clear, so `process::leave` at the exit boundary ran a killed process's whole teardown -- `close_all`, the address-space drop, the records -- with interrupts masked on that CPU. A syscall's exit runs the same teardown with IF set. `leave_ring3_if_due` now sets IF around `leave` and clears it again before the exit pass, the bracket `kernel_exit_to_user_check` already puts around `do_preempt`. The preempt depth is unchanged: 0, `BASELINE_IRQ_EXIT`, which is `do_preempt`'s own. What IF adds is a nested interrupt, which on a Ring 0 return reaches no scheduler entry, and a `preempt::enable` at depth 0 calling `do_preempt`, which asserts that same baseline. A park asserts at depth 0 with IF either way, and nothing the teardown runs parks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
`check_rust_result` printed only the kernel's account when the harness ended a test at its ceiling, though the streamed stdout is in the result; it now prints it, as the exit-code branches do. `kill_ends_every_wait` prints `<wait>: killing` before each kill, so a wait the kill cannot end is the last line of that red. The race by which an arm passes without reaching its wait moves out of the module doc into `issues/kernel/a-kill-ends-every-wait-arm-can-pass-without-reaching-its-wait.md`, and `quiesce::last`'s in-guest `STAGED` deadline is filed as `issues/diagnostics/the-quiesce-last-staging-dies-on-ten-seconds-of-guest-clock.md`. `teardown`'s doc loses a clause naming a wake `teardown_bookkeeping` does not make. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
|
Review of PR #549 at 751e36d, round 6. This is high-risk code: the process lifecycle, the scheduler and the quiesce stop. CI at 751e36d: run 36405909561, headSha 751e36d, concluded success, and its one job Earlier findings (round 5, 0c48340)
The orchestrator's questions
BLOCKER
NOTE
REMOVE
SEND BACK |
…nd sleep leads the wait arms `quiesce::hold` claimed the one hold slot before checking whether the thread may yield, and left it claimed on refusal, so the thread a boot actually stages could never take it. The slot is now released on refusal. `kill_ends_every_wait`'s `process-wait` and `thread-join` arms both sleep underneath (the waited process's `nanosleep`, the joined thread's `std::thread::sleep`), so a mutation that breaks sleep surfaced under one of their names. `sleep` now runs first among the three. `stopped_boot`'s gave-up error was the one error in the function that omitted the boot's whole log, so a sighting of a stop giving up named no threads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…ark's disabled issue Its Fast-tier sighting on PR #549 at 751e36d is the same shape: two threads the stop never counts down to, one of them the held thread by construction. The teardown boot's own extra thread beyond quiesce_last's usual five is the defect the issue already tracks, not this branch's teardown change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Under the sleep mutation (`sys_nanosleep` parking uncancellably) the sleep arm still passed: its child exited 137, and the ceiling hang came only from the process-wait arm, whose waited process runs the same `nanosleep(u64::MAX)` through the same `sys_nanosleep`. The arm and the mutation were right; the kill was early. The parent killed on reading the child's `parked in <wait>` marker, which the child writes before the syscall that parks, so a kill that lands while the child is still returning from that write is honoured at the write's `kernel_exit_to_user_check` and the named wait is never entered. The waited process of the process-wait arm had a whole second spawn's time to park, and hung as the mutation says it should. `spawn` now also polls the `SysCap` roster until the child's main thread is `SCHED_BLOCKED`, under a five-second hang guard. Between the marker's write and the named wait's syscall the child runs nothing else that parks: its code is paged from `/system`, which is the memory image. That is the filed issue's exit condition, so it is closed. `tests/testcases/system.toml` no longer counts the binaries that read the roster; this one makes five and the next landing moves it again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
A QEMU test's only clock is the harness's hang ceiling, so the five-second in-guest guard on the roster poll goes. The poll sleeps 10 ms between reads, as futex_wake_counts' wait_until_parked does, and prints `<wait>: waiting for the roster to show it parked` before it starts, so a ceiling red names where it stopped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Review, round 8, at 5163e9bCI at 5163e9b: Earlier findings (round 6, 751e36d)
The orchestrator's questions
BLOCKERNone. NOTE
REMOVE
LAND AFTER NAMED CHANGES |
…ss ceiling, on a header over 256 entries or a main thread gone/zombied; file the missing roster-entry decoder as its own issue; name the teardown claim the quiesce issue leaves dark. - kill_ends_every_wait.rs asserts the roster header's entry count is at most the 256-entry buffer, and its poll panics by name when the child's main thread is absent from the roster or zombied, rather than spinning to the harness's ceiling. - issues/design-debt/toyos-abi-decodes-only-the-roster-header-not-its-entries.md: toyos-abi decodes the sysinfo header but not an entry, and four readers hand-spell its offsets 8 and 9 instead. - issues/kernel/quiesce-wakes-on-the-last-park-gave-up-on-one-thread-beside-the-held-one.md: quiesce_wakes_on_the_last_teardown is the only guest check that a stop waits for a teardown in flight, named in both Exit and what no enabled test checks; the unproven sixth-thread attribution is removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
#564 deleted `quiesce_wakes_on_the_last_exit`, `quiesce_dump_holds_the_stopped`, their actuators `quiesce-last-exit` and `quiesce-dump`, `quiesce_last`'s exiting thread, and collapsed `quiesce::last::Last` to the park. This branch had edited all of them only to carry its new `quiesce-last-teardown` beside them; those edits have no purpose without their subjects, so the deletions are taken. `quiesce-last-teardown` and `quiesce_wakes_on_the_last_teardown` are this branch's own, not something #564 deleted: they are the guest check that the stop waits for a teardown in flight (the last one out stays in its process until `torn_down`). So they stay, and with them `Last` stays an enum of two (Park, Teardown), `hold` keeps its `last` argument and `sys_nanosleep` passes `Last::Park`, and `woken_by_the_held_thread` stays a helper with two callers. - kernel/src/actuator.rs: `quiesce_last_exit` and `quiesce_dump` go (main); `quiesce_last_teardown` stays (branch). - kernel/src/quiesce.rs: `last` keeps the branch's claimed-id hold, `ALONE` and the `may_yield` refusal; `Last::Exit` goes. - kernel/src/syscall/proc.rs: main's `hold()` becomes `hold(Last::Park)`. - tests/common/power.rs: `quiesce_wakes_on_the_last_exit` and `quiesce_dump_holds_the_stopped` go (main); the park and teardown verdicts keep the branch's `counts … alone` and teardown `exit:` checks. - tests/toyos-rust-tests/src/bin/quiesce_last.rs: the exiting thread goes (main); the teardown child stays (branch). - tests/toyos.rs: `quiesce_wakes_on_the_last_teardown` registered Nightly beside `quiesce_wakes_on_the_last_park`, which #564 moved to Nightly: same binary, same verdict, same disabling issue, and a new row has no catch record to put it anywhere else. Its CARRIES row and dispatch arm stay; the exit and dump rows go. - toyos-quiesce `LAST_THREAD`, tests/quiescelastcase/system.toml and the two quiesce issues: name the teardown actuator beside the park, not the exit; the "in a Fast tier" of the gave-up issue's exit conditions is deleted, since both its tests are Nightly now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Review, merge of #564, at 4e67c19Scope: Accounting: of the 44 files the branch changes against ec0a91a, 38 carry the branch's delta onto 807f456 unchanged (sorted Net against BLOCKERNone. NOTE
REMOVE
LAND AFTER NAMED CHANGES |
…ad of removing The merge onto 807f456 rewrote comments that named specific tests or actuators where a canonical list already exists elsewhere (kernel/src/actuator.rs, tests/common/power.rs, CARRIES in tests/toyos.rs): toyos-quiesce's LAST_THREAD doc, quiescelastcase/system.toml's header, quiesce_last.rs's module doc, and tests/toyos.rs's RUST_SKIP entry all drop their repeated lists rather than carry them forward again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Brings in #572 (host QEMU's edk2), #580 and #579 (disabled reds), #549 (a kill never waits on its victim) and #566 (metaltalk redial). - src/redlist.rs: one row each for user_copy_races_munmap and quiesce_leaves_the_volume_whole, which both sides added. main's new rows stay (netd_refused_accept, quiesce_wakes_on_the_last_teardown, root_chunk_refused_on_a_usb_stick, syscall_window_nmi). The rows for tests or issues this branch deleted go (hda_tone, doom_sound_flood, latency_wake, sched_check_build), and so does lan_swap, whose issue main deleted with swap_netd's and swap_crash_rolls_back's rows. - The two issue files both sides added take main's text. - tests/common/power.rs: main's woken_by_the_held_thread, shared by the new quiesce_wakes_on_the_last_teardown, without the two clock verdicts this branch took off QEMU (stopped_the_machine in stopped_boot, and woken_by_its_threads). - tests/common/qemu.rs: qemu_command takes main's firmware_vars and has no audio_wav, so profile_argv passes six paths. The too_many_arguments allow goes, because seven parameters do not trigger it. - kill_while_blocked.rs: main's text. After #549 a kill does not park in retire_task, so this branch's doc for arm 4 was false. main's arm also has no clock. - tests/toyos.rs check_rust_result: this branch's single-print form, which already carries the stdout main added. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Main's rust pin has not moved since the last merge, so the fork is unchanged. Every conflict, and how it was resolved: Modify/delete, main deleted: - issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md: #566 fixed the defect and deleted the issue. This branch had added one sighting to it, and a sighting of a fixed defect has no home, so the file stays deleted. - src/heartbeat.rs: #562 deleted `kernel_heartbeat`'s CPU-mask and gap verdicts with the file. This branch had given its done-line table blockd and fsd rows. The table goes with the verdict it served. - tests/doomcase/system.toml: #562 moved the doom audio tests to metal and deleted their QEMU config. This branch had added blockd and fsd rows to it. Nothing boots it now. Modify/delete, this branch deleted: - tests/toyos-rust-tests/src/bin/ftruncate_flush_race.rs, tests/toyos-rust-tests/src/bin/quiesce_fsync.rs, issues/build/ftruncate-flush-race-reds-intermittently-and-nothing-says-why.md, issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md and issues/kernel/a-root-metadata-read-refused-on-budget-is-not-retried.md: main's hunks remove timing from them or note its own runs. They are about the kernel FAT flush, the stop's kernel sync and the kernel's metadata read, which this branch deletes, so they stay deleted. Content: - kernel/src/actuator.rs: main's `quiesce_last_teardown` (#549) is kept. The kernel FAT actuators `fat_flush_meta_refuse`, `resize_evict_window` and `resize_fault_refuse` stay deleted. `process_reopen_selftest` stays where this branch has it, with main's doc (#549 also opens every kernel thread's pid). - src/redlist.rs: both conflicted rows go. `doom_sound_flood` left QEMU with #562, and this branch deletes `ftruncate_flush_race`. - tests/common/gpt.rs: this branch's `device_saying` and decoy `boot` are kept. Main drops the `drain_serial` window, so its `qemu` binding is no longer `mut`. - tests/common/inspect.rs: main's "nothing plays audio" (#562 deleted `inspect_plays`) is taken, with this branch's clause on the boot stick. - tests/common/iommu.rs: main's `panic-reboot-fast` and its wait for the fatal path's reset are kept. This branch's `iommu_empty_domain` reads the xHCI's DCBAAP over QMP, and QEMU has exited by the time that reset is seen. So `fault_boot` now takes a `holding` read, which it runs after the fault line and before it waits for the reset, while the fatal path holds its panel. `iommu_context_absent` reads nothing there. - tests/common/origin.rs: main's judgement of `log_ring_keeps_the_owners_slots` is taken whole: init says it waited a flush out, or its stop line is missing. That drops the millisecond inference between two records, whose record this branch had changed from `Syncing filesystems...` to the stop record (#562: no QEMU test measures time). - tests/common/volumes.rs: main's timing edit to `ftruncate_flush_race` goes with the test. - tests/logstallcase/system.toml: main drops `power` and the `shutdown` symlink, since the metal row reads `/log` without a stop. This branch's blockd and fsd rows are kept, because fsd holds `/log`. - tests/toyos-rust-tests/src/bin/blockd_io.rs: main's `claim_when_free`, now generic and with no deadline, is taken inside this branch's `if let Some(syscap)`. `bench` is this branch's blockd-only arm with main's timing removed: no MiB/s, and the line says only how many Flushes each run took. The module doc's "timed" goes. - tests/toyos-rust-tests/src/roster.rs (add/add): both sides wrote one roster decoder. Main's is taken whole, because five binaries read it and it has no deadline (#562). This branch's copy had a 5 s give-up. - tests/toyos-rust-tests/src/bin/process_lifecycle.rs: main's is taken whole. This branch's only change to it was the move onto its own roster.rs. - tests/toyos-rust-tests/src/bin/process_stats.rs: main's `refused_calls_are_counted` and its roster wait for the held child are kept, and so are this branch's two connection arms. The system capability is taken once in `main` and passed to the three arms that read the roster, since a second take of the label finds nothing. The connection arms now wait on main's `threads_of` for the child's main thread to be blocked, with no deadline. - tests/toyos-rust-tests/src/bin/quiesce_twice.rs: main's `Duration`-only import. This branch deletes the owed file, so `File` and `Write` go. - tests/toyos.rs: - RUST_SKIP: main's audio rows are taken. `audio_tone_load` goes, since main deleted it. `log_volume_reread` goes, since this branch deletes it. - MACHINE_TESTS: `quiesce_leaves_the_volume_whole` stays deleted. `quiesce_wakes_on_the_last_teardown` comes from main with main's comment. `blockd_serves_nothing` is kept. `hda_tone` and `hda_client_stall` went to metal with #562, and `hda_two_live_refused` takes main's comment. - CARRIES and dispatch: the same. - `nvme_wide_sector`: this branch's blockd arm, which already had no drain window. - toyos-quiesce/src/lib.rs: this branch's `FILES_MS`, `FLUSH_MS` and `SYNC_MS` are kept, with main's `LAST_THREAD` doc, which names both quiesce-last actuators. - userland/logd/src/policy.rs: this branch deletes the module doc and the `LOG_WRITE_BUDGET` paragraphs main edited one line of, so they stay deleted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two processes that held each other's handle and killed each other at the same time deadlocked the kernel, which then panicked (audit K1). Each killer waited in
retire_taskon a thread that was itself inside the other kill. With this change a kill never waits on its victim: every thread of a killed process leaves by its own hand, and the last one out tears the process down. No kernel thread is created for it.What changed, and why
process::kill_process,scheduler::post_retire).waitlearn of the end from the process object's exit, which only the teardown publishes.exitclaims the same way, and retires the calling thread's siblings.retire_task, its 10 s tripwire andTaskHandle's released flag, accounting copy and release post are deleted.process::leave).scheduler::leave_ring3_if_due). Any other thread leaves inexitorthread_exit.leavewithIFset (scheduler::leave_ring3_if_due).kernel_exit_to_user_checkis entered withIFclear. A killed process'sclose_all, address-space drop and records would otherwise all run with interrupts masked on that CPU; a syscall's exit runs the same teardown withIFset.kernel_exit_to_user_checkalready puts arounddo_preempt: set beforeleave, cleared before the exit pass.BASELINE_IRQ_EXIT, which isdo_preempt's own. WhatIFadds is a nested interrupt, whose Ring 0 return reaches no scheduler entry, and apreempt::enableat depth 0 callingdo_preempt, which asserts that same baseline. The teardown mints noParkableand calls noyield_now: it closes handles by dropping them, and the VFS lock it can reach is a spinlock.proclife::leaveanswersLast { code, mark }without marking the thread.teardown_bookkeepingmarks it withproclife::torn_down, after the teardown'ssyscalls:,memory:andexit:records and afterclose_all.quiesce::note_progressposts after that mark.close_allstill dropped files. The shutdown'sdrain_allandsync_allthen ran beside the teardown, and a dropped file's write-back was queued after the drain.teardown's doc states its invariant: it waits on nothing, because it runs on a killed thread whose one cancel may be spent, and at a depth where a park asserts.process::current_data,process::process_data). An entry outlives every thread that has not left it, andproclife::leaveandroute_thread_exitassert the same.toyos-proclife's, and the kernel calls them.claim_teardowncarries the exit code.retire_setis the set a claim retires.leaveanswers whether the thread leaving is the last one out, andtorn_downmarks that one.join::collect_zombie). A thread that left keeps its TLS mapped until the last one out. Dropping its entry would return frames its siblings may still run in.ProcessAccounting::child_threads_cpu_nsis deleted.process::process_object), so userland can kill no kernel thread.process_reopen_selftest'sprocess-open-kthreadverdict holds it.Shootdown::serveraisesflushedwithfetch_max. Kernel threads and the idle loop spin inLock::lockwithIFset, so the 0xFE IPI can nest one serve inside another. A store would moveflushedbackwards, and the initiator would spin toACK_TIMEOUT's panic.mutual_killends with a process killing itself.quiesce-last-*holds its thread until a sweep counts it, and it alone, as running (quiesce::last).holdclaims the held thread's id.sweepnotes whether it counted that thread as running, and the hold ends only on a sweep that counted exactly one running thread, the held one. It then logs<actuator>: the stop counts quiesce-last alone.quiesce-last-teardownholds the last thread out of a child ofquiesce_lastbetween its leaving and its teardown. Its hold is inprocess::leave, which the exit boundary also reaches at depth 0, whereyield_nowasserts. Thereholdrefuses by name (scheduler::may_yield): it logs<actuator>: quiesce-last left outside a syscall and is not heldand lets the stop go on unheld.woken_by_the_held_thread) requires thecounts … aloneline before thestop:record, for bothquiesce-last-*boots. For the teardown boot it also requires the held process'sexit: test_rs_quiesce_last pid=… code=0record between the two.kill_while_blockedarm 4 waits for its spinner's exit. The in-guestENDS_WITHINdeadline and the timed reader are deleted. A spinner the exit boundary misses never ends, and the harness's hang ceiling says so.check_rust_resultprints the streamed stdout on its error branch, as its exit-code branches do.kill_ends_every_waitprints<wait>: killingbefore each kill.issues/kernel/retire-tripwire-is-not-queue-shaped.mdcloses, because its constant is deleted.issues/kernel/a-killed-shutdown-caller-panics-on-its-second-drain-backoff.md.issues/kernel/the-lock-spins-shootdown-poll-says-if-is-clear-and-two-callers-spin-with-it-set.md.issues/diagnostics/the-quiesce-last-staging-dies-on-ten-seconds-of-guest-clock.md.kernel/src/quiesce.rs,hold). The slotHELD.compare_exchangeclaims is released back toNOBODYon themay_yieldrefusal, so the thread the boot stages can still take it.kill_ends_every_wait'ssleeparm runs beforeprocess-waitandthread-join. Both of those arms also sleep underneath (the waited process'snanosleep, the joined thread'sstd::thread::sleep), so a sleep mutation would otherwise surface under their name instead of its own.kill_ends_every_waitkills a child only once the kernel's roster shows its main thread parked (spawn,main_thread_parked).parked in <wait>marker is written before the syscall that parks. A kill landing while the child still returns from that write is honoured at the write'skernel_exit_to_user_check, so the named wait is never entered, and the arm passes whatever the wait does with a kill.sleeparm passed, and the hang came fromprocess-wait, whose waited process runs the samenanosleep(u64::MAX)through the samesys_nanosleepbut had a whole second spawn's time to park.spawnnow polls theSysCaproster (Rights::ROSTER, which test-runner endows) until the child's main thread isSCHED_BLOCKED, sleeping 10 ms between reads asfutex_wake_counts'swait_until_parkeddoes. It has no bound of its own: the harness's ceiling is the test's only clock, and the<wait>: waiting for the roster to show it parkedline printed before the poll names where a ceiling red stopped. Between the marker's write and the named wait's syscall the child runs nothing else that parks; its code faults in from/system, the memory image.issues/kernel/a-kill-ends-every-wait-arm-can-pass-without-reaching-its-wait.mdcloses on its exit condition.tests/testcases/system.tomlno longer counts the binaries that read the roster.quiesce_wakes_on_the_last_teardownis disabled, behind the same issuequiesce_wakes_on_the_last_parkalready is (src/redlist.rs,issues/kernel/quiesce-wakes-on-the-last-park-gave-up-on-one-thread-beside-the-held-one.md). Its Fast-tier failure is the same shape: two threads the stop never counted down to alone, one of them the held thread by construction.tests/common/power.rs's gave-up error now appends the boot's whole log, so the next sighting names both.The model (
toyos-proclife)Every thread's way out is scripted as its own steps: leave; then, for the last one out, free, mark and publish; then post and go.
Laws checked at every state:
Laws checked at every leaf:
A state where nothing can move is a deadlock.
KillEachOtherreaches 457 states, and every schedule ends.kernel-loom'sa_nested_serve_is_not_undone_by_the_one_it_interruptedruns the nested serve on a thread of its own and explores 11 executions at a preemption bound of 2.Every wait a killed thread can be in
kernel/src/syscall/io.rs:72kernel/src/syscall/io.rs:158kill_while_blockedarms 1–2kernel/src/syscall/io.rs:173,:188kernel/src/syscall/io.rs:203,:218kernel/src/syscall/ipc.rs:309kill_while_blockedarm 3kernel/src/syscall/proc.rs:76kill_ends_every_waitkernel/src/syscall/proc.rs:179kill_ends_every_waitkernel/src/syscall/proc.rs:205kill_ends_every_waitkernel/src/scheduler.rs:474kill_ends_every_waitkernel/src/inbox/mod.rs:415kill_ends_every_waitkernel/src/block.rs:88, fromobject/ops.rsblock::OPERATIONkernel/src/file_backing.rs:96read_block_retryingBudgetExpiredfor up toblock::DEADMANunder the process-data lock without reading the kill bit, and the kill lands when the fault returnswriteback::drain_allbackoffkernel/src/block.rs:88, fromwriteback.rsissues/kernel/a-killed-shutdown-caller-panics-on-its-second-drain-backoff.mdWIREsleep lockkernel/src/sleeplock.rs:100kernel/src/quiesce.rs:189quiesce::PARK; the caller is the shutdownpark_foreverkernel/src/quiesce.rs:381,arch/x86_64/hw.rs:401,watch.rsiod,klogd,iod's drainiod.rs:40,:47,log/console.rs:475,writeback.rs:52sync.rsticket lock;arch/x86_64/tlb.rsshootdown ack;drivers/virtio.rsused ring; disk waits under driver locksblock::OPERATION, and the kill lands afterkill_while_blockedarm 4,mutual_killA thread banded by the machine's stop is never dispatched again:
SafePoint::StopoutranksExit, and the machine is going down.Host gates, at ae06655;
cargo test --lib,--listand--build-onlyagain at 5163e9b, EXIT=0 each (5163e9b changes only the guest testkill_ends_every_wait)cargo test --libcargo test --workspace --exclude toyos-buildcargo test --manifest-path kernel-loom/Cargo.tomlcargo run -- --clippyclippy: 10 invocations cleancargo run -- --ci hostHost: 53 step(s), all green;mutate-last-out-leaves-before-its-teardown: 2 verdicts reachedcargo test --test toyos-build -- --listkill_ends_every_waitamong them, and lists itcargo run -- --build-onlyred-wait-sleep.patch(kernel/src/syscall/proc.rs:sys_nanosleepparks inwait_uncancellableuntil its deadline) applies withgit apply --checkat ae06655,kernel/cargo checkandcargo check --features boot-actuatorsexit 0 on the mutant, andgit apply -Rleft the tree clean.Also at ae06655:
toyos-proclife, inside the workspace gate, 32 passed. Thekernel/check variants are inside--clippy: x86_64 and aarch64, each plain, withboot-actuators, and withboot-actuators,test-actuators.Guest runs, run by the orchestrator, at 4d612d6
kill_ends_every_waitred-wait-futextimed out after 300s, stdout ending onfutex: killingred-wait-polltimed out after 300s, stdout ending onpoll: killingred-wait-process-waittimed out after 300s, stdout ending onprocess-wait: killingred-wait-thread-jointimed out after 300s, stdout ending onthread-join: killingred-wait-sleeptimed out after 300s, but stdout ending onprocess-wait: killingaftersleep: a kill ended it: the early kill aboveEarlier, at 751e36d:
mutual_kill,process_lifecycle,process_reopen_selftest,quiesce_stops_the_machine,quiesce_refuses_a_second_shutdown, andkill_while_blockedwith its redlist row lifted, EXIT=0 each.At 5163e9b:
red-wait-sleep(cargo test --test toyos-build -- --nightly kill_ends_every_waitwith the patch applied) EXIT=1,timed out after 300s, stdout ending onsleep: killing; the Fast tier EXIT=0, 396 of 396 passed.Negative controls
Host, in
--ci hostat ae06655:mutate-last-out-leaves-before-its-teardown: the last one out is marked before its teardown.the_last_one_out_is_in_its_process_until_its_teardown_is_donefails onpid 1 tid 0 is tearing its process down and the table has it out, so a stop would not wait for it.mutate-kill-waits-for-its-victims,mutate-first-out-tears-down,mutate-join-collects-in-a-teardown,mutate-claim-teardown-always-winsandmutate-spawn-skips-the-insert-recheck.Host, at 751e36d as a checked patch, restored clean:
Shootdown::servestores instead of raising:cargo test --manifest-path kernel-loom/Cargo.toml --test tlb_shootdownEXIT=101.a_nested_serve_is_not_undone_by_the_one_it_interruptedfails onthe serve it interrupted published an older generation over it.Guest, by the orchestrator:
kernel/,toyos-proclife/andtoyos-sched/at the merge base withorigin/main, tests kept, at 751e36d:mutual_killEXIT=1,retire_task: task not released after 10000msnamingRunning(CpuId(0))andRunning(CpuId(1)), a kernel panic.kill_ends_every_waitarm: the futex, poll, process-wait and thread-join arms red on their own<wait>: killingline at 4d612d6; the sleep arm red on its ownsleep: killingline at 5163e9b, above.quiesce-last-teardown's hold released on any sweep, not one that counted the held thread, at 751e36d:red-teardownEXIT=1, noquiesce-last-teardown: the stop counts quiesce-last aloneline before the stop's record, which reads5 of 5 … over 2 sweep(s).Independent oracles
mutual_kill, which the whole-change revert reproduces.toyos-sched-sim's I11, written independently of this change, for the free-after-switch the teardown relies on.Lines
git diff --shortstat origin/main...HEAD, at 5163e9b:kernel/src: +219/−340, net −121.toyos/srcunchanged.toyos-proclife: +657/−520.kernel-loom: +30.src/: +18.Merge of #564 (the measured schedule), at 4e67c19
quiesce_wakes_on_the_last_exit,quiesce_dump_holds_the_stopped, thequiesce-last-exitandquiesce-dumpactuators andquiesce_last's exiting thread. This branch's edits to them only carriedquiesce-last-teardownbeside them, so the deletions are taken.quiesce-last-teardownandquiesce_wakes_on_the_last_teardownare this branch's, and stay: they are the guest check that the stop waits for a teardown in flight. Soquiesce::last::LastisParkandTeardown,sys_nanosleepholds withLast::Park, andwoken_by_the_held_threadkeeps two callers.quiesce_wakes_on_the_last_teardownis registered Nightly (still disabled), besidequiesce_wakes_on_the_last_park, which tests: the measured schedule — Fast is every PR, then Nightly, then Weekly; 15 never-caught tests and the kernel code only they armed deleted #564 moved to Nightly: same binary, same verdict, same issue.kill_ends_every_waitandmutual_killare shared-boot tests, so Fast.cargo test -p toyos-build --lib(398 passed, 2 ignored),cargo test --test toyos-build -- --list,cargo run -- --clippy(clippy: 10 invocations clean),cargo run -- --build-only.Guest runs, run by the orchestrator, at 4e67c19
kill_ends_every_wait: EXIT=0mutual_kill: EXIT=0--nightly quiesce_refuses_a_second_shutdown: EXIT=0--weekly process_reopen_selftest: EXIT=0🤖 Generated with Claude Code
https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j