Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
8864321
kernel: a kill posts its victim's retires and returns; the reaper fin…
Japabu Sep 27, 2026
b88bf98
proclife: KillEachOther gives each process a second thread, and a sib…
Japabu Sep 27, 2026
db57dea
tests: quiesce_stops_the_machine counts the reaper among the threads …
Japabu Sep 27, 2026
61c0a7f
Merge origin/main
Japabu Sep 27, 2026
1ae6558
kernel: the reaper takes whichever kill is released, and the teardown…
Japabu Sep 27, 2026
a7457a3
Merge origin/main
Japabu Sep 27, 2026
c88bb4a
Merge remote-tracking branch 'origin/main' into wt/toyos-killfix
Japabu Sep 27, 2026
1c5ca17
review #549 round 2: no kernel thread's pid opens, and the reaper's c…
Japabu Sep 27, 2026
b4d7ef6
Merge remote-tracking branch 'origin/main' into wt/toyos-killfix
Japabu Sep 27, 2026
fa1e325
kernel: the last thread out tears its process down; no reaper, no ret…
Japabu Sep 27, 2026
456d8cb
Merge origin/main into wt/toyos-killfix
Japabu Sep 28, 2026
a6aaf3e
kernel: the last thread out is in its process until its teardown is done
Japabu Sep 28, 2026
57c89b5
kernel: a process's end is not pollable
Japabu Sep 28, 2026
179f41e
kill_while_blocked: arm 4 waits for its spinner's exit, with no clock
Japabu Sep 28, 2026
0c48340
kernel-loom: the nested serve is a thread of its own
Japabu Sep 28, 2026
393a4f2
Merge origin/main into wt/toyos-killfix
Japabu Sep 28, 2026
63d0b08
quiesce-last: the hold ends on a sweep that counted the held thread a…
Japabu Sep 28, 2026
a592498
scheduler: a killed thread's teardown at the Ring 3 boundary runs wit…
Japabu Sep 28, 2026
751e36d
kill_ends_every_wait: a ceiling red names the arm; the arm race is filed
Japabu Sep 28, 2026
009fa2a
A refused hold releases its slot, a gave-up stop names its threads, a…
Japabu Sep 28, 2026
ce95b55
Merge origin/main into wt/toyos-killfix
Japabu Sep 28, 2026
4d612d6
quiesce_wakes_on_the_last_teardown shares quiesce_wakes_on_the_last_p…
Japabu Sep 28, 2026
0f28ea0
kill_ends_every_wait kills a child only once the roster shows it parked
Japabu Sep 28, 2026
ae06655
Merge origin/main into wt/toyos-killfix
Japabu Sep 28, 2026
5163e9b
kill_ends_every_wait's roster wait has no clock of its own
Japabu Sep 28, 2026
82f917b
Round 8 review fixes: the roster poll fails by name, not by the harne…
Japabu Sep 28, 2026
80276cb
Merge remote-tracking branch 'origin/main' into wt/toyos-killfix
Japabu Sep 28, 2026
4e67c19
Merge origin/main (#564, the measured schedule) into wt/toyos-killfix
Japabu Sep 28, 2026
4c2e052
Round of #549 review: delete four stale lists the merge rewrote inste…
Japabu Sep 28, 2026
254f1fe
Merge remote-tracking branch 'origin/main' into wt/toyos-killfix
Japabu Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
---
status: open
kind: defect
opened: 2026-09-28
---

# `toyos-abi` decodes only the roster header, not its entries

`SysinfoHeader::decode` (`toyos-abi/src/syscall.rs:85`) is "the one spelling of
the header's offsets, so a reader takes a field rather than an index" — but
that one spelling stops at the header. `SYSINFO_ENTRY_SIZE`
(`toyos-abi/src/syscall.rs:103`) documents the entry's byte layout in prose
only, and every reader of an entry hand-spells its offsets instead of calling
a decoder:

- `tests/toyos-rust-tests/src/bin/kill_ends_every_wait.rs:140-141` —
`entry[9] == 0`, `entry[8]`.
- `tests/toyos-rust-tests/src/bin/process_lifecycle.rs:258` —
`buf[pos + 9] != 0, buf[pos + 8]`.
- `tests/toyos-rust-tests/src/bin/abuse_thread_name.rs:62` — `entry[9]`.
- `userland/toybox/src/ps.rs:64-65` — `buf[pos + 8]`, `buf[pos + 9] != 0`.

Four readers, each free to walk off by a field the way the header's own doc
comment warns against. `kernel/src/syscall/machine.rs:309-310` is the one
writer (`entry[8] = state; entry[9] = is_thread;`), so the four are already
one silent renumbering away from reading the wrong column.

**Exit**: a `SysinfoEntry::decode` (or equivalent) in `toyos-abi`, the four
readers above converted to call it, and no roster-entry offset left
hand-spelled outside that one function.

Owner: `toyos-abi/src/syscall.rs`, whoever next touches the roster ABI.
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
---
status: open
kind: tooling
opened: 2026-09-28
---

# The `quiesce-last-*` staging dies on ten seconds of guest clock

`kernel/src/quiesce.rs`'s `last::STAGED` is a 10 s in-guest deadline. `hold`
panics the boot when the stop has not counted the held thread alone within it,
and `await_the_held_thread` panics when no thread named `quiesce-last` has
reached its syscall within it. `quiesce-last-park` and
`quiesce-last-teardown` all rest on it, and `quiesce_twice`'s `WAITS_WITHIN` is
priced inside it. A QEMU test's only clock is the harness's ceiling: a guest
slower than the deadline dies with a staging panic that names no defect.

Owner: orchestrator. Exit condition: both sides of the staging wait on the
other's event with no deadline, and a staging that never arrives is the
harness's ceiling.
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
---
status: open
kind: defect
opened: 2026-09-27
---

# A killed shutdown caller panics on its second drain backoff

`writeback::drain_all` (`kernel/src/writeback.rs`) retries a refused drain
through `block::between_attempts`, whose park is cancellable, and discards what
it answers. A thread whose kill bit is set gets `Cancelled` from that park at
once, retries, and parks again; `TaskHandle::take_cancel` asserts on the second
cancel reported to one thread, so the kernel panics. `ops::until_answered`
reads the kill bit before it parks and does not have this; `drain_all` does not.

The caller is the shutdown syscall (`syscall/machine.rs`), which a kill posted
before `quiesce::stop` can still reach, and it takes a second backoff only when
the device refuses two drain attempts on budget. Not reproduced.

Exit condition: `drain_all`'s retry either stops parking once its caller is
killed or parks uncancellably, with a test that kills a caller parked in a
refused drain.
6 changes: 0 additions & 6 deletions issues/kernel/deferred-release-outlives-its-syscall.md
Original file line number Diff line number Diff line change
Expand Up @@ -196,12 +196,6 @@ at all.** It belongs with that track, not beside it.

### What "give the batch an owner" costs, worked out 2026-08-20

An owner has to be the *thread*, not the CPU. `kill_process` phase 2 calls
`scheduler::retire_task`, which parks the killer until the victim's record is
released, so the killing thread can be moved to another CPU between the
`close_all` that queues its objects and the syscall exit that would drain them —
a per-CPU list would strand exactly the batch it was added to own.

A per-thread list cannot live behind `ThreadData`'s lock either:
`teardown_resources` holds `ProcessData` across `close_all`, and its own first
line is that the two locks are never held together. So the list has to be on the
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,10 @@ FAIL quiesce_wakes_on_the_last_park: the stop gave up on 2 thread(s) that never
- PR #555 at `d2656765`.
- PR #559 at `ac948e6a`.

`quiesce_wakes_on_the_last_teardown`'s boot shares the same shape, in the Fast
tier: PR #549 at `751e36d9`, `4 of 6 … 2010 ms of a 2010 ms budget over 2
sweep(s)`. That boot normally counts 5 userland threads.

The earliest is PR #510 at `98e803cb`, recorded in
`issues/build/quiesce-wakes-on-the-last-park-lost-its-serial-ready-beside-other-guests.md`:
`stop: 4 of 7 userland thread(s) stopped ... in 2010 ms of a 2010 ms budget`.
Expand Down Expand Up @@ -49,7 +53,10 @@ enabled test checks any of these:
- that a band, a park or an exit wakes the stop, rather than its deadline;
- `in_flight == 0` with `begun > 0`;
- the thread census;
- the `console-queue-at-the-stop` drain.
- the `console-queue-at-the-stop` drain;
- that the stop waits for a teardown in flight —
`quiesce_wakes_on_the_last_teardown`'s only claim, and the only enabled
guest check of it, disabled by this same issue.

`quiesce_refuses_a_second_shutdown` stays green over a lost post. It judges the
stop only by `stopped_the_machine`, so a stop that spends its budget and then
Expand All @@ -60,7 +67,9 @@ finds everything stopped passes it.
- an instrument that names each thread still running when the stop gives up,
with its name, tid, cpu and scheduler state;
- the mechanism it names fixed;
- this test green in a Fast tier beside other guests;
- this test green beside other guests;
- `quiesce_wakes_on_the_last_teardown` — the stop waits for a teardown in
flight — green beside other guests;
- an enabled guest test checking each claim listed above.

Owner: the stop path, `kernel/src/quiesce.rs`; held by the orchestrator.
92 changes: 0 additions & 92 deletions issues/kernel/retire-tripwire-is-not-queue-shaped.md

This file was deleted.

11 changes: 3 additions & 8 deletions issues/kernel/scheduler-pass-blocks-in-xhci.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,20 +23,15 @@ sits on the wrong side of the boundary the budget describes.
What a CPU inside that recovery holds is *every message addressed to it*: an
`Adopt` carrying a task, a `Wake` for a parked thread, a `Retire`. Nothing in the
scheduler can shorten it — every reap and every wake is bounded by the owning
CPU's pass latency by design, which is exactly why the design is sound. The one
thing in the tree that notices is `scheduler::retire_task`'s 1 s guard, and it
notices by panicking:
CPU's pass latency by design, which is exactly why the design is sound.

```
retire_task: task not released after 1s: InTransit(CpuId(1))
```

That panic fired on the owner's T14 at 949.792 s of uptime with doom exiting. The
*balance*-path half of it is fixed: `hand_off` reaps a killed task rather than
handing it on, gated by simulator invariant I14. This half is not,
and it would produce the same panic with `Blocked(CpuId(n))` in the message
instead — the guard cannot tell a lost message from a busy CPU, which is what it
is written as if it could.
handing it on, gated by simulator invariant I14. This half is not.

The second instance of the same shape used to be the idle loop's log flush;
the idle loop touches no filesystem now — the log is logd's file — and the
Expand All @@ -48,7 +43,7 @@ that `drain_irqs` only ever does work it can finish: drain the event ring,
dispatch HID reports, note that a port or an endpoint owes work. The debounce and
the port reset were already moved off this path for exactly this reason (CLAUDE.md,
USB hotplug); the control transfers inside `configure` and `recover_endpoints`
were not. Until then, `retire_task`'s bound is measuring the USB bus.
were not.

**And the budget cannot see it.** `cpu::MAX_PASS_NS` is measured against by
`SchedPass::finish`, from the `now` the pass was entered with to the end of
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
---
status: open
kind: defect
opened: 2026-09-28
---

# The lock spin's shootdown poll says `IF` is clear, and two of its callers spin with it set

`Lock::lock` (`kernel/src/sync.rs`) polls TLB shootdowns inside its spin with the
comment "this spin runs with `IF` clear". Two callers spin there with `IF` set:

- a kernel thread, which `kernel_start` (`arch/x86_64/entry.rs`) enters with `sti`;
- the idle loop, whose `reap_finished` takes `PROCESS_TABLE.lock()`.

There the 0xFE IPI can land inside `tlb::poll`'s own `serve_if_owed`, so one
CPU runs a serve nested inside another. `Shootdown::serve` raises `flushed` with
`fetch_max` for exactly that case, and `kernel-loom`'s
`a_nested_serve_is_not_undone_by_the_one_it_interrupted` reds when it stores.
The comment gives the poll a reason that holds only for syscall context.

**Exit**: the comment states when the spin runs with `IF` set and when clear, and
why the poll is needed in each.
30 changes: 30 additions & 0 deletions kernel-loom/tests/tlb_shootdown.rs
Original file line number Diff line number Diff line change
Expand Up @@ -234,3 +234,33 @@ fn an_initiator_answers_while_it_waits() {
assert!(s.wait_turn(0, 1, g0, || {}));
});
}

/// A serve that an interrupt nests inside another publishes the later
/// generation, and the outer serve, finishing after it, must not take it back.
///
/// The nested serve runs on a thread of its own, so loom places it at every
/// point of the outer one, the nesting among them. Reds when `serve` stores
/// what it owes instead of raising to it.
#[test]
fn a_nested_serve_is_not_undone_by_the_one_it_interrupted() {
model(|| {
let s = Arc::new(Shootdown::new());
let first = s.issue();
let nested = {
let s = Arc::clone(&s);
loom::thread::spawn(move || {
let later = s.issue();
s.serve(1, || {});
later
})
};
s.serve(1, || {});
let later = nested.join().unwrap();
assert!(s.served(1, first));
assert!(
s.served(1, later),
"cpu 1 flushed for the nested shootdown, and the serve it interrupted \
published an older generation over it",
);
});
}
5 changes: 4 additions & 1 deletion kernel/src/actuator.rs
Original file line number Diff line number Diff line change
Expand Up @@ -119,6 +119,9 @@ actuators! {
/// Hold the thread named `toyos_quiesce::LAST_THREAD` inside `SYS_NANOSLEEP`, and the shutdown until it is held there, until the stop waits on it alone: its park is then the stop's last transition.
quiesce_last_park = "quiesce-last-park";

/// The same for the last thread out of its process, between its leaving and its teardown: that teardown is then the stop's last transition.
quiesce_last_teardown = "quiesce-last-teardown";

/// Refuse the second directory-entry write of the file `writeback_durability` stages for the retry gate — the first is that file's own seed being made durable — as a budget expiry, so a flush fails at its metadata write with its pages already written and settled.
fat_flush_meta_refuse = "fat-flush-meta-refuse";

Expand Down Expand Up @@ -524,7 +527,7 @@ actuators! {
/// Judged by `partition_claim_gives_up`.
partclaim_root_withheld = "partclaim-root-withheld";

/// Reopen init by pid once it is spawned, the way `SYS_PROCESS_OPEN` does.
/// Reopen init by pid once it is spawned, and open every kernel thread's pid, the way `SYS_PROCESS_OPEN` does.
process_reopen_selftest = "process-reopen-selftest";

/// Offer the block layer a second device claiming a registered `DeviceId`, and report what it did with it.
Expand Down
8 changes: 2 additions & 6 deletions kernel/src/arch/x86_64/hw.rs
Original file line number Diff line number Diff line change
Expand Up @@ -466,12 +466,8 @@ impl Hw for KernelHw {
}

/// Reached once per task, from a later pass running on another stack, so dropping `payload`
/// here never frees the stack this call stands on; `publish_released` must be last — a
/// retirer's wait ends only once this drop has happened.
/// here never frees the stack this call stands on.
fn release(&self, _key: TaskKey, payload: KernelPayload, acct: TaskAccounting) {
let handle = payload.handle.clone();
handle.finalize(acct);
drop(payload);
handle.publish_released();
payload.handle.finalize(acct);
}
}
6 changes: 6 additions & 0 deletions kernel/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -656,6 +656,12 @@ pub(crate) unsafe extern "C" fn kernel_main(kernel_args: &KernelArgs) -> ! {
log::console::start();
iod::start();

// Here: the last kernel thread is spawned.
#[cfg(feature = "boot-actuators")]
if actuator::process_reopen_selftest() {
sched::kthread::open_selftest();
}

smp::set_ready();

// After the release, because a shootdown waits on CPUs that are not
Expand Down
Loading
Loading