Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
c14fc9a
issues: the kernel still creates threads (track)
Japabu Sep 27, 2026
f63ac32
Kernel: delete usbd, and the kernel-thread panic policy with it (K1)
Japabu Sep 27, 2026
5a7964d
CLAUDE.md: the kernel creates no thread but the per-CPU idle loop
Japabu Sep 27, 2026
020a2ef
CLAUDE.md: no kernel threads is a principle, not a snapshot fact
Japabu Sep 27, 2026
653008e
Kernel: a Ring 0 fault outside a syscall halts; QEMU's reset judges a…
Japabu Sep 27, 2026
bb8c1e9
tests: QEMU's stop reason is klogd_*_halts' verdict, not the guest's …
Japabu Sep 27, 2026
fbd31a1
tests: klogd_*_halts stop at klogd's spawn line, the last one both ar…
Japabu Sep 27, 2026
dd83890
Merge remote-tracking branch 'origin/main' into wt/toyos-nokthread
Japabu Sep 27, 2026
7eb6b51
tests: klogd's spawn line is not the last line of a recovered boot
Japabu Sep 27, 2026
894a85e
Kernel: every kernel panic halts, and panic recovery is deleted
Japabu Sep 27, 2026
6897bff
tests: the syscall-death boots take their kernel from the parameter; …
Japabu Sep 27, 2026
af5737c
issues: a sysroot cloned during a toolchain rebuild never gets its cargo
Japabu Sep 27, 2026
ede2220
Merge remote-tracking branch 'origin/main' into wt/toyos-nokthread
Japabu Sep 27, 2026
73e00d7
Point the citations of the deleted `recover_or_halt` at `fatal_except…
Japabu Sep 27, 2026
f9700b9
CLAUDE.md: the Kernel and Userspace daemons paragraphs are two paragr…
Japabu Sep 27, 2026
69e8dff
Panic path: reset without waiting on the console wire
Japabu Sep 27, 2026
c645c04
syscall_fault_halts reads a live window; panic_recovery folds into fa…
Japabu Sep 27, 2026
8e525da
Merge remote-tracking branch 'origin/main' into wt/toyos-nokthread
Japabu Sep 27, 2026
355e0d6
syscall_fault_halts reads an untouched .bss window; its red carries t…
Japabu Sep 27, 2026
eb19f5b
Merge remote-tracking branch 'origin/main' into wt/toyos-nokthread
Japabu Sep 27, 2026
37d9826
Round 5 review fixes: dead IrqGuard trait items go, the null-read che…
Japabu Sep 27, 2026
fe7c7a6
Merge remote-tracking branch 'origin/main' into wt/toyos-nokthread
Japabu Sep 27, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ A subdirectory `CLAUDE.md` loads when a file in that subtree is `Read`, and not
- **Zero silent debt.** Dead code is deleted; every abstraction earns its place. A discovered compromise has exactly two legal outcomes: remove it, or record it with ownership, evidence and an exit condition — and it stays a present-state weakness until removed.
- **Fail fast, trust nothing.** Panics over silent degradation; exhaustive matches; the unimplemented dies loudly. Input that crossed a trust boundary is never trusted and never panics the kernel — it is refused.
- **The kernel never crashes from userland.** A kernel bug crashes loudly; a userland bug never reaches it.
- **No kernel threads.** The kernel creates no thread but the per-CPU idle loop: kernel work runs, bounded, on the thread or interrupt that caused it and is charged to it; long-running work with no owner is a userland server's.
- **Rust is first class.** Not POSIX, not C. Unrepresentable is best: prefer compile-time safety over runtime checks over tests.
- **Existing Rust just works.** A program that builds for other operating systems builds and runs on ToyOS unchanged; the ecosystem gains ToyOS support through forks carried upstream, never through ToyOS-specific replacement crates.
- **Development ergonomics above all.** Iteration speed beats feature count; tooling comes first.
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
---
status: open
kind: tooling
opened: 2026-09-27
---

# A guest with no virtio device keeps no serial when its run reds

`tests/common/lane.rs`'s `keep_serial` copies a red run's `uart-*.log` files to
`target/red-run-serial`. A profile with no virtio device (`qemu_command`'s
`shape.virtio.present()` false: `Profile::Metal`, and every `power.rs` guest
built on `panicked()`) routes its 16550 to QEMU's stdio instead, so its only
record is the reader thread's memory, and the kept lane directory is empty. The
console stream of a virtio guest is not kept either; only its UART is.

**Evidence:** `syscall_fault_halts` red at `8e525da4`: the run printed
`this red run's serial logs are kept at target/red-run-serial/toyos-tmp-45425-0`,
and that directory's `lane-0` is empty. The check's own error carried no
capture, so what the kernel said was lost.

**Exit condition:** every guest's console, whichever device carries it, is
written under its lane as it is read, and a red run keeps it; a red
`panic_reboots` run shows a non-empty kept lane.
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
---
status: open
kind: tooling
opened: 2026-09-27
---

# A sysroot cloned during a toolchain rebuild never gets its cargo

The primary checkout's `rust/build/aarch64-apple-darwin/stage2/bin` was
recreated at 22:00 on 2026-09-27 and, read while a `toyos-build --build-only`
ran in the primary, held `rustc` and `rustdoc` and no `cargo`; the `cargo` link
appeared at 22:08. At 22:03:59 a `cargo test --test toyos-build` in the
`wt/toyos-nokthread` worktree rebuilt std and published
`rust/build/sysroots/5dc157f7fac727be` cloned from that `bin/`, so it has
`rustc` and `rustdoc` and no `cargo`. The key is found again on every later
run and nothing re-provisions it: every harness run from that worktree since
panics at `tests/toyos.rs:2970` with `the toyos toolchain at
.../sysroots/5dc157f7fac727be/bin is missing cargo` on the C corpus, before any
test runs.

**Evidence:** the two directory listings and the harness log above, read on
the dev host; not reproduced on purpose.

**Exit condition:** a sysroot is never published without the provisioned
`cargo`, and a run that finds a published one without it provisions it or
refuses by name at the toolchain step rather than inside the corpus.
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
---
status: open
kind: tooling
opened: 2026-09-27
---

# `idle_stack_guard`'s "the read succeeded" arm reads a line nothing writes

`tests/common/faults.rs`'s `idle_stack_guard` reds with *the page below the idle
stack is still mapped* when its capture contains `debug syscall returned`. The
only guest it drives, `tests/toyos-rust-tests/src/bin/test_panic_child.rs`,
writes `SYS_DEBUG {action} returned {rc:#x}` when the syscall comes back. The
arm can never fire: a guard page that is still mapped reaches the drain's
ceiling and reds instead on the missing `#PF UNHANDLED` line, naming the wrong
cause.

**Evidence:** read from both sources; no mutation run.

**Exit condition:** the arm and the child share one constant for the line, and
a mutation that maps the guard page reds with the arm's own message.
25 changes: 25 additions & 0 deletions issues/isolation/interrupt-entry-keeps-a-ring-3-ac-flag.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
---
status: open
kind: defect
opened: 2026-09-27
---

# Interrupt entry keeps a Ring 3 thread's AC flag

SMAP binds a Ring 0 access only while `RFLAGS.AC` is clear, and a Ring 3 thread
can set `AC` with `popf` (`CR0.AM` is clear, so it costs the thread nothing).
The syscall entry clears it through `IA32_FMASK`
(`kernel/src/arch/x86_64/syscall.rs`), but an interrupt or trap gate does not:
Intel SDM Vol. 3A §6.12.1.3 names TF, VM, RF and NT, and IF for an interrupt
gate. `arch::entry::ring3_naked_asm` prepends only `cld`, and
`kernel/src/arch/x86_64/control_regs.rs` says the boot `clac` is the only one
the kernel needs.

So an interrupt or exception taken from a Ring 3 thread that set `AC` runs its
handler with SMAP off: a kernel bug there that touches a user address reads or
writes it silently.

**Evidence:** read from the code and the SDM; no test stages it.

**Exit condition:** every Ring 0 entry from Ring 3 runs with `AC` clear, gated by
a guest program that sets `AC` and a handler that must fault on a user address.
Original file line number Diff line number Diff line change
Expand Up @@ -25,15 +25,6 @@ disk. On the T14 that is `logd`'s `SYS_FSYNC`, which reaches
kernel::arch::syscall::gate::syscall_entry
```

The panic lands in a syscall, so `percpu::in_syscall` makes it recoverable:
`try_recover_from_panic` ends that thread and returns to the scheduler, where
`deadline::wedge_if_staged` folds the CPU into the wedge. `apic::halt_all_cpus`
is therefore never reached, and neither is `deadline::stand_down` — which is why
the boot deadline and not the panic path's own bound ends the machine. That
composition is correct; what is not is that a boot staged to measure a *device*
silently loses its log writer, and the only account of it is a `PANIC` record in
a ring tail nobody was draining.

Two things are true and neither is decided here:

- The detector firing is right in general — a lock held for ever by a CPU that
Expand All @@ -46,7 +37,7 @@ Two things are true and neither is decided here:

Any actuator that stops a CPU inside a driver — today the `usb-wedge-*` arms,
which are QEMU registrations and reach no flashed image. It does not change
their verdicts, but it adds a panic, a dead `logd`, and a page of dropped
their verdicts, but it adds a panic and a page of dropped
records to every one of them.

## Exit condition
Expand Down
4 changes: 1 addition & 3 deletions issues/kernel/deferred-release-outlives-its-syscall.md
Original file line number Diff line number Diff line change
Expand Up @@ -237,9 +237,7 @@ So the two constraints meet. **The hook queue may not park, and the row that
would want to is not on the hook queue but in a `Drop` that also may not.**
Moving `File` to `deferred` swaps one illegal site for another. The second shape
above is still right for the `deferred` rows, and by itself it reaches neither
`File` nor `close_all` — that one is also called from `recover_or_halt`'s
`Blame::Process` arm (`arch/idt/exceptions.rs:348`), which has no syscall to
return through.
`File` nor `close_all`.

The track carries this as **wall 4**, with the three shapes the owner has to
choose between. Nothing here should be built before that choice, because all
Expand Down
13 changes: 2 additions & 11 deletions issues/kernel/every-wait-in-this-kernel-is-a-spin.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,10 +92,6 @@ both before any lock conversion; the order is forced, not preferred.
- **The sleep lock.** A sleep-lock holder stays preemptible and raises no
preempt count, so the baseline assertion keeps meaning exactly "a spinlock is
held".
- **`usbd` and `iod` on the existing kernel-thread machinery.** No housekeeping
thread's wait can stop another's, and a panic inside one is recoverable rather
than a halted machine. Three threads, not one, because a stuck USB enumeration
must not stop the log.
- **xHCI async, and the four lock conversions.** Inseparable. A CPU never waits
for a device: the lock is dropped before the park, and a completion is matched
to its asker by identity, never by arrival order.
Expand Down Expand Up @@ -160,10 +156,7 @@ both before any lock conversion; the order is forced, not preferred.
`PROCESS_TABLE.lock()` at `loader/mod.rs:694`. `release_process` and
`kill_process` bracket `teardown_resources` between two table acquisitions and
hold neither across it. So `PROCESS_TABLE` does not have to convert for the
park to be legal, and converting it anyway would buy the exception-recovery
path a `try_lock` with no answer for its failure: `recover_or_halt`'s
`Blame::Process` arm reaches `process::exit` from a CPU exception
(`arch/idt/exceptions.rs:348`), and nothing on that path mints a `Parkable`.
park to be legal.
What *does* have to convert is `Lock<ProcessData>`, and wall 5 is why.

**Six, and the count above is the xHCI chunk's rather than the machine's.**
Expand Down Expand Up @@ -384,9 +377,7 @@ means everywhere, not only here. Three shapes:
anywhere.
3. **Give the batch an owner** — `deferred-release-outlives-its-syscall`'s own
second shape — and make `File` deferred. That buys a parkable release site on
the syscall path and does *not* cover `close_all` reached from
`recover_or_halt`'s `Blame::Process` arm, which has no syscall to return
through; and it needs `drain_zero_handles`'s two scheduler sites to stop
the syscall path, and it needs `drain_zero_handles`'s two scheduler sites to stop
running hooks that can park, which is a redesign of that queue rather than a
use of it.

Expand Down
14 changes: 6 additions & 8 deletions issues/kernel/nothing-charges-kernel-memory-to-a-process.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,8 +20,8 @@ is not a value with a destructor" is unrepresentable; the descriptor type and
its un-refcounted clone are gone; the file cache installs a real budget and a
budget that was never installed is now a loud kernel bug; and the unbounded
user-string copy has a ceiling. What survives is the *accounting*, which nothing
has touched, plus two items: panic recovery still runs no teardown, and peak
memory is written by two paths that overwrite each other.
has touched, plus one item: peak memory is written by two paths that overwrite
each other.

Blocked on nothing. Two things worth knowing before it is restarted, because
both cost a day to discover:
Expand All @@ -34,9 +34,7 @@ both cost a day to discover:
=`, `drop()` and burial in a collection all pass silently. A drop bomb is the
state of the art, and `Unmapped<T>` is already exactly such an obligation.

**The terminal state of every unbounded grower is this entry's.** The allocation
failure itself reports cleanly — a failed kernel allocation takes `alloc`'s
no_std default handler and panics with the size, the layer and the call site
named — so what is left at the end of a grower is *who dies*: whichever thread
happened to allocate, not the one that exhausted the heap. That is what charging
fixes, and nothing else does.
**The terminal state of every unbounded grower is a halted machine.** A failed
kernel allocation panics with the size, the layer and the call site named, and
every kernel panic halts, so a grower userland can drive is refused at its bound
or it ends the machine.
54 changes: 0 additions & 54 deletions issues/kernel/the-capability-end-state-is-twelve-answers.md
Original file line number Diff line number Diff line change
Expand Up @@ -349,60 +349,6 @@ Four decisions, in one place, for the owner:

Question 12 is not in this set: its track already holds it.

## Kernel-resident workers

**Kernel-resident workers are a control-flow boundary, not a memory boundary** —
each one is audited periodically, exists only where independent blocking
progress, fault containment of execution flow, or latency isolation requires it,
and a new one needs explicit architectural justification; work moves to
userspace when the IPC/wait machinery makes that an isolation gain rather than
overhead.

**Census of 2026-08-20.** `sched::kthread` caps a shipping machine at three and
dies naming a fourth (`kernel/src/sched/kthread.rs:50`, `:102`). All three
exist, and all three are started at the end of `kernel_main`
(`kernel/src/main.rs:696`-`703`):

- **`klogd`** — the machine's only console drainer, "one thread where every idle
CPU used to drain". `OnPanic::Halt`, because "a machine whose only console
drainer has been killed goes silent with nothing left able to say so"
(`kernel/src/log/console.rs:1`, spawn at `:116`).
- **`usbd`** — owns the xHCI port machine so USB work runs in a context of its
own instead of whichever thread trapped: "a stuck USB enumeration must not
stop the log". Spawned on every machine including one with no controller, at
one kernel stack, so the machine has one answer to how many kernel threads it
has. `OnPanic::Recover`, because every loss it causes is visible
(`kernel/src/drivers/xhci/usbd.rs:1`, spawn at `:47`). Its body is one park
today — nothing posts to it yet.
- **`iod`** — owns the deferred write-back queue, because `OpenFileState::drop`
must flush under a lock that is becoming a sleep lock and a `Drop` impl cannot
hold a `Parkable`. `OnPanic::Recover`, because a killed `iod` costs deferred
write-back and both `SYS_FSYNC`'s error path and `/system/bin/logd`'s give-up policy
can see that (`kernel/src/iod.rs:1`, spawn at `:52`). Its body is one park
today — nothing pushes yet. **One `iod` machine-wide is a decision with a
measurement owed** at the 128-core target, recorded at its own site
(`kernel/src/iod.rs:24`).

Two more exist only on a `boot-actuators` kernel and are stimulus rather than
workers: `lognest`, one thread (`kernel/src/log/nested.rs:68`), and `logstorm`,
one per log shard (`kernel/src/log/storm.rs:115`) — which is why the cap is
`3 + MAX_LOG_SHARDS` on that build and 3 on a shipping one
(`kernel/src/sched/kthread.rs:50`).

**Verdict:** all three are justified at their site, each by independent blocking
progress or fault containment of execution flow, and each states its panic
policy. None is a memory boundary. Nothing is owed to userspace yet; the sleep
locks the two idle bodies are waiting for are
`issues/kernel/every-wait-in-this-kernel-is-a-spin.md`.

**PID-backed pseudo-processes for kernel workers** are pragmatic today — a
process-table row is what makes a kernel thread nameable in `ps`, in
`sched::dump` and in a crash report, and every field of it is the empty value
rather than a plausible one (`kernel/src/sched/kthread.rs:291`). The moment that
representation leaks misleading user-process semantics into policy,
observability, lifecycle or APIs, identity/accounting separates from
user-process semantics rather than preserving the abstraction for convenience.

## The rest of the review's standing rules

- **The adversarial handle-lifecycle suite** the review lists — stale-handle
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -21,8 +21,7 @@ times:
`Source` enum with two hand-written dispatches, and ad-hoc wake paths.
- **USB with three concurrency models:** a state machine inside the scheduler
pass, disk I/O spinning with interrupts off for up to 4.75 s under one global
lock, and boot discovery calling a blocking bind from a scheduler pass. The
thread reserved to own the controller, `usbd`, only parks.
lock, and boot discovery calling a blocking bind from a scheduler pass.
- **Storage done busy-waiting under spinlocks:** one global VFS lock, NVMe with
one command outstanding and polled, and a 2 s operation budget "with
preemption off" that budgets audio stalls. That budget drags a refusal chain
Expand Down Expand Up @@ -138,7 +137,7 @@ times:
`/log` survives usbd killed mid-batch, the keyboard keeps working while
a stick misbehaves, and Ctrl+Alt+D on the machine's own keyboard files
the dump with usbd killed.
5. **USB owned by its thread, then by userland**, with discovery and recovery
5. **USB by userland**, with discovery and recovery
written once as straight-line code. **Exit**: no interrupts-off window
longer than a register access, and keyboard input keeps flowing while a
stick misbehaves.
Expand Down
34 changes: 34 additions & 0 deletions issues/kernel/the-kernel-still-creates-threads.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
---
status: open
kind: track
opened: 2026-09-27
---

# The kernel still creates threads

The owner's ruling: the kernel creates no thread but the per-CPU idle loop.
Kernel work runs, bounded, on the thread or interrupt that caused it and is
charged to it; long-running work with no owner is a userland server's. The
ruling holds only if the whole kernel-thread machinery is deleted —
`sched/kthread.rs`, `kthread::spawn`, the row table and every policy a row
carries, `is_kernel_task`/`current_is_kernel_thread`, and every special case that exists
only for kernel threads. Stopping new uses is not enough.

**Exit condition:** `kernel/src/sched/kthread.rs` does not exist, and no kernel
code creates a schedulable task other than the per-CPU idle loop.

**Evidence:** `git grep -n 'kthread::spawn(' -- kernel/src`.

**Stages:**

- **K2:** the reaper PR #549 introduces becomes last-thread-out: the victim's
last thread tears down its own process on its way out of the kernel, and the
scheduler frees that thread's kernel stack after switching away. Blocked on
#549 landing.
- **K3:** the test-only `logstorm`/`lognest` producers are deleted if the log
gate does not need kernel-context producers. Blocked on nothing.
- **K4:** `klogd` goes: the owner-approved driver-model design moves the
console to logd, and this track owns that move.
- **K5:** `iod` goes with the kernel's write-back queue; met only when #536
lands with no new `kthread::spawn`.
- **K6:** delete the machinery named above. Blocked on K2–K5.
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
---
status: open
kind: defect
opened: 2026-09-27
---

# A crash report can name a syscall that already ended

The crash report prints `Syscall:` and a user backtrace only while
`percpu::in_syscall()` holds (`kernel/src/arch/x86_64/idt/exceptions.rs`). That
compares this CPU's recorded syscall task with the current task, and only
`leave_syscall` clears the word, on the CPU where the syscall ends. `Hw::switch`
in `kernel/src/arch/x86_64/hw.rs` moves the current identity and never the
bracket. The number, the user `rip` and the `rbp` the user backtrace walks from
are per-CPU copies too (`syscall_entry` in `kernel/src/arch/x86_64/syscall.rs`).

A syscall that parks on CPU A and finishes on CPU B is wrong both ways:

- **A keeps naming it.** A's word still names the thread. When that thread next
runs in Ring 3 on A before any other syscall enters there, an interrupt
handler's death on A reports the finished syscall's number, user `rip` and
user backtrace as the context it died in.
- **B forgets another.** `leave_syscall` on B clears whatever B's word named,
even a second thread still parked inside a syscall it entered on B. When that
thread resumes on B and the kernel dies inside its syscall, the report prints
no `Syscall:` line.

`process::handle_fault` (`kernel/src/process.rs`) prints `percpu::syscall_num()`
with no bracket at all, so a handle fault in a syscall that migrated names the
syscall that last entered on the CPU it faulted on.

`syscall_entry` already pushes the user `rsp`, the user `rip` and the number at
the top of the thread's own kernel stack, which moves with the thread; reading
them there would delete every per-CPU copy. The frame alone cannot say whether
it is a syscall's: an interrupt from Ring 3 puts `SS` and `CS` in the slots
where a syscall's frame holds the user's `rsp` and `rdi`, and userland chooses
both.

**Evidence:** read from the code; no test stages it.

**Exit condition:** the syscall context a crash report or a handle fault prints
is the current thread's, whichever CPU it entered on, and a guest test stages
both migrations: an interrupt-context death on the first CPU whose report
carries no `Syscall:` line, and a death inside a syscall that parked while
another thread's syscall ended on its CPU, whose report carries its own.
Loading
Loading