Skip to content

Kernel: every kernel panic halts and panic recovery is deleted; delete usbd; open the no-kernel-threads track (K1) - #553

Merged
Japabu merged 22 commits into
mainfrom
wt/toyos-nokthread
Sep 28, 2026
Merged

Japabu merged 22 commits into
mainfrom
wt/toyos-nokthread

Conversation

@Japabu

@Japabu Japabu commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Every kernel panic now halts the machine, including a panic inside a syscall, and all of the kernel's panic recovery is deleted. The branch also opens the owner's no-kernel-threads track and does its first stage (K1): usbd is deleted.

The ruling this lands

The owner ruled:

  • A kernel panic means a kernel invariant broke, so it always halts, then reboots or holds the panel under the existing panic policy.
  • A Ring 0 fault is a kernel bug and halts, whatever syscall or thread is current.
  • A Ring 3 fault kills only its own process, as before.
  • "The kernel never crashes from userland" is kept by refusing bad input at the boundary, not by surviving panics.

This reverses the old rule that a recoverable panic ended only the process that caused it.

What is deleted, and why

  • The panic handler's decision. The percpu::in_syscall() branch is gone. So are try_recover_from_panic (x86 and the aarch64 stub) and recover_or_halt. The handler now always captures, flushes and calls halt_all_cpus.

  • Poisoned threads. Deleted:

    • kernel/src/sched/poison.rs, the per-CPU POISONED bank and poison_tid;
    • schedule_no_return;
    • process::PoisonWake and zombify_poisoned;
    • toyos-proclife/src/poison.rs and Watch::thread, whose one caller was the poison path;
    • the poison-overwrite feature, together with its loom model kernel-loom/tests/poison_set.rs and its red in src/ci.rs.

    The idle loop's reap_poisoned becomes reap_finished. It only collects exits that have been published, behind the same ReapGate.

  • The scheduler's (Ready(_), Dead) edge. Its one stated user was schedule_no_return. The only production edge into Dead is Running → Dead.

  • The panel's survived-panic path. Deleted: panic_console::discard_capture, CaptureAccess::discard, CaptureLatch::owned_by, and the loom model of a discard. The latch's release stays, because capture_into still uses it. Its loom model is renamed a_released_latch_hands_the_snapshot_to_the_next_captor.

  • blame.

    • toyos_userbound::{blame, Blame, Faulted} are deleted. The ring of a trap now decides whose fault it is. That ring is the opaque Ring, which is built only from cs.
    • fatal_exception kills the process on a Ring 3 fault and halts on anything else.
    • A recursive fault always halts. A Ring 3 frame can only arrive with this CPU already in Fatal/Panic if the kernel left that state set, and that is a kernel bug.
  • Ring 0 demand paging. page_fault_handler now sends a not-present fault to process::handle_page_fault only for a Ring 3 frame: if is_user && process::handle_page_fault(..). It used to do that for any fault taken while a thread was current. A Ring 0 fault falls through to fatal_exception, which halts before anything is mapped into the current process. handle_page_fault's kernel-thread refusal became unreachable and is deleted.

  • SYS_DEBUG action 2's one-shot arming. It existed to refuse a second call into a lock that a recovered panic had stranded.

  • deaf_window's unconditional sti. It now holds an IrqGuard. The unconditional enable was there because panic recovery could leave IF clear. Machine::irq_guard/type IrqGuard, the trait method deaf_window does not use, are deleted together with their five impls (kernel/src/arch/x86_64/hw.rs, kernel/src/arch/aarch64/hw.rs, toyos-sched/src/cpu.rs, toyos-sched/sim/src/hw_impl.rs, toyos-sched/loom/tests/loom_retire.rs); deaf_window calls the concrete crate::arch::IrqGuard::close() directly and never went through the trait. The review named only the first three; the aarch64 kernel and the loom build are what caught the other two.

  • The aarch64 percpu::in_syscall stub. It has no caller now.

  • The kernel-thread panic policy. OnPanic, a row's recoverable word and panic_recovers_here are deleted.

  • K1. usbd is deleted: kernel/src/drivers/xhci/usbd.rs, its start(), the usbd-panic actuator, and usbd in the test lists. The track is issues/kernel/the-kernel-still-creates-threads.md (kind: track), with stages K2–K6.

The panic path's reset does not wait on the console wire

panic_reboot::reboot_now resets through acpi::reset_now, not acpi::reboot. The log is already drained by that point, and acpi::reboot opens with serial::flush_final, which spins PANIC_LOCK_SPIN_LIMIT (100,000,000 iterations of try_lock and pause) on the console wire. A CPU this panic stopped may hold that wire, and a stopped CPU never releases it, so that wait could only run out. The limit being a count and not a time is filed as panic-path/the-panic-lock-spin-limit-is-a-count-its-comment-calls-a-second.md.

This change is measured green, but what it fixes is not proven. At f9700b90, syscall_panic_halts went red once: the guest printed returning this machine to firmware, QEMU reported no reset in the harness's wait, and the test took 98 s at a 2.96x liveness width. The held-wire arm below does not reproduce that. With the wire held and the old acpi::reboot path, the row stays green (EXIT=0). It takes 27 s, against 7–8 s for the fixed path at the same head, so the held wire's spin costs about 19–20 s at a 1.50x width and still resets inside the 37 s budget. A held wire alone therefore does not account for that red, and nothing measured here explains it. No deterministic red arm exists on the old path: flush_final's spin is bounded, so a held wire delays the reset and does not stop it.

The syscall-death rows also subscribe to QMP before they write the command (power::watch_the_bound), the way machine_reboot does. QMP delivers no event that was emitted before its client connected.

Tests

  • Four Nightly rows, all run through power::syscall_death_resets:

    Row SYS_DEBUG action Must say
    syscall_panic_halts PANIC the panic, Syscall: num=92, the user backtrace
    syscall_fault_halts NULL_READ on the 2 MiB-aligned window inside a 4 MiB .bss array the child never touches KERNEL PANIC: read unmapped address at <that address>, Syscall: num=92, the user backtrace
    lock_across_switch_halts LOCK_ACROSS_SWITCH the tripwire, with check_tripwire_attribution scoping its panicked at to syscall/dispatch.rs
    heap_over_ceiling_halts HEAP_OVER_CEILING exceeds MAX_HEAP_ALLOC

    Each row boots panic-reboot-fast with QMP and runs test_rs_test_panic_child <action>. The verdict is QEMU's guest-reset inside the fast bound plus the reset allowance, the same reading klogd_death_resets takes (klogd_panic_halts, klogd_fault_halts). The two share died_and_reset. When a row's own check fails, its error carries the guest's capture. These guests put the 16550 on stdio, so no uart-*.log exists for a red run to keep.

  • syscall_fault_halts carries the demand-paging claim.

    • NULL_READ is a Ring 0 read of the address its caller passes. The kernel records SYS_DEBUG: a Ring 0 read of <addr> first, and check_ring0_read_unmapped reads the address from that record. It is a record and not the child's own line because the panic path drains records, while a program's line is still in userland when the read ends the machine. check_ring0_read_unmapped also refuses address 0: a child that regressed to calling syscall::debug(action) for every action, dropping the window argument, reads address 0 and must red rather than pass.
    • The address has to be demand-paged. An anonymous mmap does not qualify: sys_mmap maps its pages when it makes it, and the region is RegionKind::Mapped. A Ring 0 read of one takes SMAP's fault on a present page and says protection violation. RegionKind::Anonymous regions are the loader's, a segment's pages past its file bytes. So the child reads the aligned window inside static mut UNTOUCHED: [u8; 4 MiB]. In the built child, UNTOUCHED is the first 0x400000 bytes of .bss at 0x7e2c8, in a PT_LOAD segment with 5656 file bytes and 4200196 memory bytes, so the window holds no other object and no file bytes.
    • A kernel that demand-paged that window for Ring 0 would re-execute the read into SMAP's fault (+smap on both x86-64 CPU models, src/arch.rs) and say protection violation instead. This row replaces the read of address 0.
  • panic_recovery is folded into fault_gates. fault_gate_child gains a pf arm (read_null). check_fault_gates requires its SEGFAULT tid= header and fault_gate_child::read_null in the backtrace. Deleted: panic_recovery.rs, check_panic_recovery, its timeout and duration rows, and segfault_child, which nothing else ran.

  • heap_ceiling_recovery is renamed heap_ceiling_bounds. Its over-ceiling arm moved to heap_over_ceiling_halts. Its "the heap still works after recovery" arm is deleted.

  • screen_recoverable_untouched and screen_survived_panic_not_blamed are deleted. They tested a panic that survives, and none can now.

  • test_panic_child needs an action. Every caller names one.

  • Harness prose that relied on recovery is cut:

    • serial::Died::Kernel;
    • qemu::ceiling_verdict's doc;
    • await_guest's scoping comment;
    • Ppm::identical_to's doc;
    • needs_actuators and suite_split.
  • Unpriced in tests/test-durations: syscall_panic_halts, syscall_fault_halts, lock_across_switch_halts, heap_over_ceiling_halts, klogd_panic_halts, klogd_fault_halts and heap_ceiling_bounds. The committed profile is measured on a CI runner (committed_durations_path), and a dev host's TCG times are not written into it. fault_gates is priced at 31 from before it had the eighth arm. panic_recovery's row is deleted.

Issues

  • Deleted, because the defect went with its code:
    • panic-path/panic-holding-process-table-hangs.md;
    • panic-path/the-discards-refusal-branch-is-exercised-by-nothing.md.
  • Filed:
    • panic-path/a-crash-report-can-name-a-syscall-that-already-ended.md. It covers both migration directions and process::handle_fault's unbracketed syscall_num();
    • panic-path/the-panic-lock-spin-limit-is-a-count-its-comment-calls-a-second.md;
    • build/idle-stack-guards-returned-arm-reads-a-line-nothing-writes.md;
    • build/a-sysroot-cloned-during-a-toolchain-rebuild-never-gets-its-cargo.md;
    • build/a-guest-with-no-virtio-keeps-no-serial-when-its-run-reds.md: lane::keep_serial keeps only uart-*.log, so a red run of a guest whose 16550 is on stdio keeps an empty lane. The harness owns it;
    • isolation/interrupt-entry-keeps-a-ring-3-ac-flag.md;
    • panic-path/a-syscall-panic-reset-the-guest-announced-and-qemu-never-saw.md: the f9700b90 red above, filed once the wire-held measurement refuted the only proposed mechanism.
  • Edited:
    • nothing-charges-kernel-memory-to-a-process.md: an unbounded grower now ends the machine;
    • a-deliberate-wedge-…md: the recovery paragraph is deleted;
    • the two wall-4 issues: the clauses citing the deleted recovery arm are deleted;
    • the-capability-end-state-is-twelve-answers.md: its kernel-resident-workers section is deleted, since the track states the opposite rule;
    • the-kernel-is-small-interrupts-post-and-threads-wait.md: usbd is gone from it.

in_syscall stays

in_syscall stays, because only the crash report's Syscall: lines read it now. Reading the syscall frame at the top of the thread's own kernel stack instead does not work: that frame cannot tell a syscall's frame from a Ring 3 interrupt's without trusting words userland chose. An interrupt puts SS and CS in the slots where a syscall's frame holds the user's rsp and rdi. The per-thread bit that would be left needs new per-thread state written on every syscall, not one rule in one function. The issue above records this.

Gates

At fe7c7a6c, which is this branch after merging origin/main at 41ad548c:

Gate Exit
cargo run -- --clippy, including the three aarch64 kernel invocations EXIT=0, "10 invocations clean"
cargo test --lib EXIT=0 (379 passed, 1 ignored)
cargo test --workspace --exclude toyos-build (toyos-sched, its sim and loom included) EXIT=0
cargo test --test toyos-build -- --list (builds the C corpus and every Rust test binary, checks the redlist, boots nothing) EXIT=0
cargo run -- --build-only EXIT=0

Measured by the orchestrator at fe7c7a6c: syscall_fault_halts --nightly EXIT=0; with the demand-paging mutant EXIT=1, "expected KERNEL PANIC: read unmapped address at 0x10000200000: the read did not fault as unmapped"; with the child's window multiplied by 0 (the address-0 arm) EXIT=1, "expected the demand-paged window, not the null read"; dump_nmi_probe --nightly EXIT=0; the Fast tier 393 passed, 1 failed, the one being lan_mdns_answer (a macOS socket path over SUN_LEN in the harness, fixed by its own PR). Earlier heads' guest results below stand where this round did not touch their paths.

Negative controls and the independent oracle (run by the orchestrator)

The independent oracle is QEMU's own SHUTDOWN event, read through qemu::QmpShutdown. It is the hypervisor reporting that the guest reset, not the guest's serial.

Each patch is applied with git apply --check then git apply, and reversed with git apply -R. Each mutated tree passes cargo run -- --clippy (EXIT=0).

Arm Patch Command Result
Ring 0 demand paging demand-paging-ring0.patch: if (is_user || percpu::current_tid().is_some()) && process::handle_page_fault(..) cargo test --test toyos-build -- --nightly syscall_fault_halts EXIT=1 at eb19f5b5 (553r5-demandpaging-syscall_fault_halts.log). The kernel says KERNEL PANIC: read protection violation at 0x10000200000; the check says expected \KERNEL PANIC: read unmapped address at 0x10000200000`: the read did not fault as unmapped at 0x10000200000`. The machine still resets, so the red comes from the named assertion.
The held wire on the old path wire-held-red.patch: reboot_now goes back to acpi::reboot, with the wire held first (core::mem::forget(serial::try_wire())) cargo test --test toyos-build -- --nightly syscall_panic_halts EXIT=0 at 8e525da4, 27 s. A measurement of the wait, not a red arm; see above.
The held wire with the fix wire-held-green.patch same EXIT=0 at 8e525da4, 8 s.
R2: a Ring 3 fault halts R2-ring3-halts.patch: if false && ctx.ring().is_user() in fatal_exception cargo test --test toyos-build -- --nightly fault_gates EXIT=1 at 8e525da4.

Green arms, orchestrator-run: at 8e525da4, cargo test (Fast) — 393 of 394 passed, the one red lan_mdns_answer (path must be shorter than SUN_LEN), a harness socket-path defect #560 fixes and red the same way on #536, #541, #554, #555 and #559 (#560's own run of it is EXIT=0); and, under cargo test --test toyos-build -- --nightly <row>, each of syscall_panic_halts, lock_across_switch_halts, heap_over_ceiling_halts, klogd_panic_halts, klogd_fault_halts, heap_ceiling_bounds, panic_reboots, panic_key_holds, fault_gates and disk_backtrace EXIT=0 (553r4-*.log); syscall_fault_halts alone was EXIT=1 there (553r4-syscall_fault_halts.log), fixed by 355e0d64. At eb19f5b5, rerun under the same nightly command: syscall_fault_halts, syscall_panic_halts, lock_across_switch_halts and heap_over_ceiling_halts each EXIT=0 (553r5-green-*.log).

Unsure / left outside the fence

  • What made syscall_panic_halts red at f9700b90 is unknown, as the section above says.
  • NULL_READ keeps its name while it reads any address. A rename changes toyos-abi/src and so the sysroot key, which is outside this brief.
  • The recursive-fault path changed. A recursive fault on a Ring 3 frame now halts where it used to kill the process. No test stages it.

Net against origin/main (41ad548c): +630 −1782. Production +130 −1061 (kernel +102 −503, toyos-userbound +8 −226, toyos-proclife +5 −158, kernel-loom +8 −136, toyos-sched +4 −31, toyos-abi +2 −2, src +1 −5); tests +243 −568; issues +256 −153; CLAUDE.md +1, the orchestrator's.

🤖 Generated with Claude Code

Japabu and others added 2 commits September 27, 2026 20:08
The owner's ruling: the kernel creates no thread but the per-CPU idle
loop, and the ruling holds only once the whole kernel-thread machinery
is deleted. The track names its exit condition, the spawn sites left
once K1 lands, and stages K1-K6 with what unblocks each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
usbd's body only parked. It goes with its start() call, its
`usbd-panic` actuator, and every test list that named it.

It was also the only thread that ever walked `OnPanic::Recover`. The
one other claimant, iod, was never measured recovering, and #536
deletes iod. So the policy goes rather than moving its actuator to iod:
`OnPanic`, the row's `recoverable` word, `panic_recovers_here`, and
the Release/Acquire pair that published the word before the identity.
A kernel thread's panic now halts the machine, because
`percpu::in_syscall` compares the running task's identity with the one
that entered the syscall, and a kernel thread never enters one. The
panic handler asks only that.

A row is now one word. Every reader learns the id it compares through
something ordered after the publish: the table lock held across it, or
the run queue the task is dispatched from. Relaxed is enough for that.

klogd_panic_halts's verdict was the ready marker's absence. A recovered
klogd takes the console with it, so that verdict held with kernel-thread
panics made recoverable (EXIT=0 on the mutated kernel). The verdict is
now the line only halt_all_cpus writes, awaited as an event, and the
same mutation reds it (EXIT=1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu
Japabu marked this pull request as ready for review September 27, 2026 18:09
The owner's ruling, placed where every agent reads it before it knows
which subsystem it is in.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Review of 5a7964dd against origin/main. PR CI host at this head: success (run 36339673529). The guest tests are measured in the PR body at f63ac32d, and the head differs from it only in CLAUDE.md. Net: kernel +34 −133, tests +23 −86, issues +45, CLAUDE.md ±1.

Panic path, every context. A user thread is unchanged from main: panic_recovers_here() returned None for any non-kthread, so the handler already fell to in_syscall(). The same holds for early boot (PERCPU_READY), for idle (no pid), for an IRQ taken from Ring 3 or idle, and for a Ring 3 fault. A Ring 3 fault never reaches the panic handler; blame returns Process and the process is killed. The only contexts that change are the ones the branch intends. A panic on a kernel thread, or in an IRQ handler that runs while one is current, now halts. On main, an IRQ-handler panic that landed while iod/usbd was current poisoned that thread and the machine carried on. in_syscall() compares identity, so a kernel thread running while a user syscall is parked on the same CPU does not inherit that syscall's recovery. The line is right for panics. It is not the line the fault path draws; see the first NOTE.

BLOCKER

  • PR body "Independent oracle: none" / tests/toyos.rs:13395: a high-risk change to the panic path names no independent oracle, and CLAUDE.md requires one. A cheap one is already in the tree. Boot klogd-panic with panic-reboot-fast and qmp: true, then require QEMU's own reset through qemu::QmpShutdown inside PANIC_FAST_SECS + RESET_ALLOWANCE, the way tests/common/power.rs panic_reboots does. That is the hypervisor's report, not the guest's serial. M1 (|| sched::kthread::current_is_kernel_thread() at kernel/src/main.rs:181) must turn it red too.

NOTE

  • toyos-userbound/src/fault.rs:102, kernel/src/arch/x86_64/idt/exceptions.rs:120. This predates the branch and is outside its fence. A Ring 0 #PF on a user-half address is ProcessThroughKernel whenever any tid is current, and it then goes to try_recover_from_panic. A kernel thread's null dereference (cr2=0, is_user_addr(0) holds) therefore poisons the thread and the machine carries on. For klogd that is exactly the silent mute its old Halt row claimed to prevent. An IRQ handler's null dereference while a user thread is current kills that innocent process instead of halting. Panics are judged by in_syscall() and faults by on_a_thread: two answers to one question. File it as a defect. The candidate fix is to feed percpu::in_syscall() in as the blame input, after measuring that no Ring 0 user-address access happens outside a syscall.
  • Negative control: I accept M1 in place of the whole revert. The whole revert is green by construction, because klogd's row was Halt on main. Only iod's outcome moved, and no test can see it. But the branch makes a per-thread panic policy unrepresentable: the handler at kernel/src/main.rs:181 reads no per-thread state, and bringing one back means adding a type, which review sees. M1 red, and M1 plus the old verdict green, are the right pair.
  • tests/test-durations:257: still prices klogd_panic_halts at two boots (16658), and it is now one. Re-price it, then re-tier it under src/tiers.rs or say why it stays Nightly.
  • tests/toyos.rs:13424: this is a second copy of screen_late_panic's arm-line derivation (tests/toyos.rs:5941). Make it one fn.
  • issues/kernel/the-kernel-still-creates-threads.md:43 (K5): Storage: file servers for DATA, the log and the boot volume; the kernel's NVMe and FAT go #536 does not remove a kernel thread; it swaps one. Its head (wt/toyos-fsd, kernel/src/main.rs:592) spawns a new kthread::spawn("probe", probe, 0, OnPanic::Recover) to host sched_operation_nesting and sysret_ss_probe, which ride iod today (kernel/src/iod.rs:20-29). That spawn also conflicts with this branch. The track owes a stage for those two probes' host.
  • issues/kernel/the-kernel-still-creates-threads.md:40 (K4): nothing plans to move the console. issues/kernel/every-driver-is-still-in-the-kernel.md never mentions the console or klogd, and its exception criterion ("stays in the kernel only if the kernel needs it while userspace is dead") keeps the console in. K4 is blocked on a decision nobody holds, which makes it an owner question.
  • c14fc9af carries the usbd.rs deletion and does not compile alone, so a bisect through the merge hits a broken commit. It cannot be fixed without rewriting history; recorded here only.
  • CLAUDE.md:39: under "A snapshot", the orchestrator's sentence states the ruling as present fact while klogd, iod, lognest and logstorm exist. This is the orchestrator's call.

REMOVE

  • kernel/src/sched/kthread.rs:4-6 "A panic inside one halts the machine: …": rewritten module prose. It misleads, because a fault inside one does not halt (first NOTE), and the contract is the handler.
  • kernel/src/sched/kthread.rs:113 "Before enqueue_new: from that call the task can run and panic.": its reason went stale with this branch. No row reader is on the panic path any more.
  • kernel/src/main.rs:667 "After klogd so its spawn log has a drainer.": a rewritten comment that is not load-bearing. Records commit to their shards whoever drains them.
  • tests/toyos.rs:1234: a rewritten row comment that no longer says why the row is Nightly.
  • tests/toyos.rs:13420 "The verdict is the line only halt_all_cpus writes.": false. panic_reboot::arm is also called from the reentry and early-panic branches (kernel/src/main.rs:135, :149).
  • issues/kernel/the-kernel-still-creates-threads.md:23-28: the pasted grep output. Its line numbers move with every landing, and Storage: file servers for DATA, the log and the boot volume; the kernel's NVMe and FAT go #536 and Kernel: a kill never waits on its victim — the last thread out tears its process down #549 each change it. Keep the command.
  • issues/kernel/the-kernel-still-creates-threads.md:32: K1 is done when this merges.
  • issues/kernel/the-kernel-still-creates-threads.md:37-38 "move to a userland test program whose threads emit through the syscall path, or": refuted for lognest by kernel/src/log/nested.rs:45 (IF is clear for a whole syscall).
  • issues/kernel/the-kernel-still-creates-threads.md:40-42 "under issues/kernel/every-driver-is-still-in-the-kernel.md … Blocked on that track moving the console.": that track carries no such move.
  • issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md:24-25 "The thread reserved to own the controller, usbd, only parks.": that thread is deleted.
  • issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md:141 "owned by its thread, then": the kernel-thread half of stage 5, which the ruling forbids. Lines 65, 115 and 128-140 name step 10's userland usbd ("the whole xHCI moves", "usbd killed mid-batch"). That usbd is consistent with the ruling, and those lines stay.
  • issues/kernel/every-wait-in-this-kernel-is-a-spin.md:95-98: this bullet's invariant, that a kernel-thread panic is recoverable and there are three threads, is now false. :31 stays: it is a true record of what completions: the duration kinds, the one park site, the sleep lock, and the daemons on it — pipeline 2's first half #91 landed.
  • issues/kernel/the-capability-end-state-is-twelve-answers.md:352-405: the whole "Kernel-resident workers" section. Its rule (:354-359) contradicts the ruling. Its census and verdict (:361-396) name usbd, OnPanic and three threads. Its PID-backed paragraph (:398-404) describes the machinery the track deletes. The track is now the one home for this.

SEND BACK

Japabu and others added 5 commits September 27, 2026 20:27
Four kernel threads still exist; the rule is what new work follows,
so it sits among the principles the tree does not yet meet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… kernel thread's death

Answers the first review of #553.

Faults and panics now ask one question. `blame` took "a thread is
current" as its input, so a Ring 0 page fault on a user address was the
process's whenever any tid was current: a kernel thread's null
dereference poisoned `klogd` and the machine carried on silent, and an
interrupt handler's null dereference killed whichever user thread it
landed on. `blame` now takes `percpu::in_syscall()`, the input the panic
handler already reads, so a Ring 0 fault outside a syscall is the
kernel's and halts, and one inside a syscall stays the process's.

The audit behind it: every user-memory access in the kernel goes through
the direct map (`user_ptr::window`, the futex word, the loader's
`KernelSlice`s, the inbox and shm pages, the crash dump's hand walk);
the only Ring 0 dereference of a user-half address is `SYS_DEBUG`'s
staged `NULL_READ`, inside a syscall. Three defects the audit found are
filed rather than fixed here:
issues/panic-path/the-syscall-bracket-outlives-a-migrated-syscall.md,
issues/isolation/interrupt-entry-keeps-a-ring-3-ac-flag.md and
issues/isolation/a-ring-0-page-fault-demand-pages-user-memory.md.

`klogd-fault` reads address zero on klogd's first instruction.
`klogd_panic_halts` and the new `klogd_fault_halts` share
`power::klogd_death_resets`: boot on `panicked()`'s guest with
`panic-reboot-fast`, and require QEMU's own `guest-reset` inside the
bound, the way `panic_reboots` does, which now shares
`resets_inside_the_bound` with them. The two-boot price of
`klogd_panic_halts` is dropped from tests/test-durations; the name is
unmeasured on a runner until the next profile.

Review REMOVEs applied: the kthread module-doc and publish comments,
main.rs's spawn-order comment, the MACHINE_TESTS row comment, the false
arm-line claim; the track loses K1, the pasted grep, K3's userland
option and K4's pointer, and K4/K5 are rewritten as the orchestrator
ruled; the stale usbd prose in three issue files is deleted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…arm line

The first cut waited for the arm line as its ready marker, so a machine
that recovered went red on the guest's silence before QEMU was asked.
The boot now stops at a line both a halting and a recovering kernel
write (`PANIC: panicked at`, `#PF UNHANDLED: cr2=0x0`), and the verdict
is `QmpShutdown`'s `guest-reset` inside the bound; the report's lines and
the arm line are read from the boot log plus the drain after it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ms write

Under the old blame input a kernel thread's fault is recovered and
nothing after klogd's spawn line reaches the wire: the fault's own
`#PF UNHANDLED` record is committed and never drained. So a report line
as the boot's marker still let the guest's silence decide. The marker is
now `kthread: klogd pid=`, written before klogd runs, and QEMU's stop
reason is the only verdict on halt versus carry on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under M1 the recovered machine goes on writing; the spawn line is only
one both arms write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu Japabu changed the title Kernel: delete usbd and the kernel-thread panic policy; open the no-kernel-threads track (K1) Kernel: a Ring 0 panic or fault outside a syscall halts; delete usbd; open the no-kernel-threads track (K1) Sep 27, 2026
Japabu and others added 5 commits September 27, 2026 21:39
The owner's ruling: a kernel panic means a kernel invariant broke, and the
kernel does not keep running on state it can no longer trust. A panic inside
a syscall halts like any other; a Ring 0 fault is a kernel bug and halts; a
Ring 3 fault still kills only its process. "The kernel never crashes from
userland" is kept by refusing bad input at the boundary, not by surviving
panics. This reverses the rule that a recoverable panic ends only the
offending process.

Deleted:
- the panic handler's `in_syscall()` branch, `try_recover_from_panic` (x86
  and the aarch64 stub) and `recover_or_halt`;
- `sched/poison.rs`, the per-CPU poison bank, `poison_tid`,
  `schedule_no_return`, `process::PoisonWake`/`zombify_poisoned`,
  `toyos_proclife::poison`, `Watch::thread`, the `poison-overwrite` feature
  and its loom model and CI red; the idle loop's `reap_poisoned` is now
  `reap_finished` and only collects published exits;
- `panic_console::discard_capture`, `CaptureAccess::discard`,
  `CaptureLatch::owned_by` and the loom model of a discard;
- `toyos_userbound::{blame, Blame, Faulted}`: whose fault a trap is is now
  its ring alone, and `fatal_exception` kills a Ring 3 fault's process and
  halts on anything else;
- Ring 0 demand paging: `page_fault_handler` resolves a not-present fault
  only for a Ring 3 frame, so a Ring 0 fault halts before anything is
  mapped into the current process; `handle_page_fault`'s kernel-thread
  refusal went with it;
- `SYS_DEBUG` action 2's one-shot arming, which existed to refuse a second
  call into a lock a recovered panic stranded;
- the aarch64 `percpu::in_syscall` stub. The x86 one stays: the crash
  report reads it to print the syscall a death happened inside.

Tests:
- `panic_recovery` keeps its Ring 3 arm only (a user segfault kills its
  process and the system lives), and leaves `ACTUATOR_TESTS`.
- `syscall_panic_halts`, `syscall_fault_halts`, `lock_across_switch_halts`
  and `heap_over_ceiling_halts` each drive one `SYS_DEBUG` death on its own
  boot and take the verdict from QEMU's `guest-reset` through QMP, through
  `power::syscall_death_resets`, which shares `klogd_death_resets`'s tail.
- `heap_ceiling_recovery` is `heap_ceiling_bounds`: its over-ceiling arm
  and the heap-still-works arm after it are gone.
- `screen_recoverable_untouched` and `screen_survived_panic_not_blamed`
  are deleted: there is no survived panic to paint or not paint.

Issues: the stale-bracket, Ring-0-demand-paging, stranded-PROCESS_TABLE and
discard-refusal defects are gone with the code they were about. What is left
of the stale bracket is a crash report naming a finished syscall, filed
narrowly; `idle_stack_guard`'s inert "the read succeeded" arm is filed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the wedge issue loses its recovery paragraph

`panic-reboot-fast` already selects the test kernel, which carries
`test-actuators`, and the harness refuses a boot that also names the build.
The wedge arms' DEADLOCK panic in `logd`'s fsync now halts like any other;
`usb_reset_records_the_phase_it_cut` stays green on this tree, so only the
paragraph describing the recovery composition goes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Found while measuring this branch's Ring-3 red arm: the run published a
sysroot without `cargo`, and every harness run from the worktree has panicked
on the C corpus since.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion`

A Ring 3 fault's `kill_process(-1)` now sits in `fatal_exception` itself.
The sysroot issue says only what was read, not when the primary's build ran.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu Japabu changed the title Kernel: a Ring 0 panic or fault outside a syscall halts; delete usbd; open the no-kernel-threads track (K1) Kernel: every kernel panic halts and panic recovery is deleted; delete usbd; open the no-kernel-threads track (K1) Sep 27, 2026
@Japabu

Japabu commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator Author

Review of 73e00d78 against origin/main. Gate: PR CI host at this head is green (run 36347325470, conclusion success). --ci host boots no guest. The only QEMU runs are 14 targeted tests at 6897bff9, taken before the merge of c5518949. Reviewed anyway because the brief asks for the arms owed. Net: production +105 −1002, tests +179 −482, issues +145 −148.

Round 1 BLOCKER, no independent oracle for the kernel-thread halt: CLOSED. klogd_panic_halts and klogd_fault_halts take QEMU's guest-reset over QMP. M1 is red: nokthread-k1-r2/arm-m1.log, EXIT=1, "QEMU never reported stopping". MB is red: arm-mb.log, EXIT=1.

BLOCKER

  • (no guest run at this head) — The branch turns two things into a machine halt: every Ring 0 fault on a user address, and every panic reachable from a syscall. PR CI runs no guest test, and only 14 hand-picked tests ran, on a pre-merge tree. A shared-boot test that stayed green only because the kernel killed its caller now halts the shared boot. Only the whole suite finds that. Owed at the next head, with EXIT and log: cargo test, then cargo test --test toyos-build -- --nightly syscall_panic_halts syscall_fault_halts lock_across_switch_halts heap_over_ceiling_halts klogd_panic_halts klogd_fault_halts heap_ceiling_bounds panic_recovery.
  • kernel/src/arch/x86_64/idt/exceptions.rs:504 — R2 (a Ring 3 fault halts) is unmeasured, and it is the other half of the ruling. Apply nokthread-r3/R2-ring3-halts.patch (if false && ctx.ring().is_user()) onto the head. cargo test --test toyos-build -- panic_recovery fault_gates must go red. Reverse the patch and it must go green.
  • kernel/src/arch/x86_64/idt/exceptions.rs:447 — On high-risk code, no test can fail on the claim "a Ring 0 fault maps nothing into the current process". NULL_READ reads address 0, which lies in no region. The minimal test:
    • Turn DA::NULL_READ into a Ring 0 read_volatile of the address the caller passes (debug_with(action, addr), reading a2).
    • test_panic_child takes a 4 MiB anonymous mapping it never touches and prints reading {addr:#x} for base + 2 MiB, a window inside a live region. Then it calls the action.
    • syscall_fault_halts requires KERNEL PANIC: read unmapped address at {that addr:#x}, Syscall: num=92 and the reset.
    • Negative control: - if is_user && process::handle_page_fault(fault_addr, frame.error_code) { / + if (is_user || percpu::current_tid().is_some()) && process::handle_page_fault(fault_addr, frame.error_code) { must turn it red.
    • Under that mutant the machine still resets. The window gets mapped, and the re-executed read takes SMAP's #PF PRESENT on this +smap CPU. So the reset alone stays green, and the assertion must name unmapped and the address.
    • This row replaces address 0, so one row carries both claims.

NOTE

  • toyos-sched/src/task.rs:278 — Delete (Ready(_), Dead). The branch deleted its one stated reason (schedule_no_return), and the only production edge into Dead is Running → Dead (task.rs:1168). A legality table that admits an edge nobody takes is a weaker check on the scheduler.
  • tests/common/power.rs:950 — syscall_death_resets writes the command before died_and_reset connects QmpShutdown (:901). The guest's 5 s fast bound is the only margin before the SHUTDOWN it must not miss. Connect before the write, the way machine_reboot does.
  • kernel/src/arch/x86_64/percpu.rs:762 (brief item 3) — Keeping in_syscall is acceptable here: nothing decides on it any more, and it is recorded. The filed issue has half the defect, though:
    • leave_syscall on the CPU a migrated syscall finishes on clears whatever bracket that CPU holds, even another parked thread's. So a genuine syscall death afterwards prints no Syscall: line.
    • process::handle_fault (process.rs:1676) prints the per-CPU syscall_num() with no bracket at all.
    • There is a way that makes the issue go away. syscall_entry already pushes the user rsp, the user rip (rcx) and the number (rdi) at the top of the thread's own kernel stack, the one kernel_rsp names and switch moves with the thread. Reading them there deletes OFF_SYSCALL_RIP/OFF_SYSCALL_NUM and the identity compare. The one per-thread bit left is whether the top frame is a syscall's. Add both directions and handle_fault to the issue.
  • tests/toyos-rust-tests/src/bin/panic_recovery.rs — What is left is a subset of fault_gates plus disk_backtrace's symbol checks. Add a pf arm to fault_gate_child/fault_gates, then delete panic_recovery.rs, check_panic_recovery and its runtime row (tests/toyos.rs:18290). If it stays, its name and "all panic recovery tests passed" are false.
  • kernel/src/sched/dump.rs:387 — deaf_window's unconditional sti has lost its one reason, which was panic recovery leaving IF clear. Either an IrqGuard is right now, or there is a reason the code has to show.
  • kernel/src/sched/driver.rs:715 — metal-panic-probe fired from the idle loop only because a syscall panic used to recover (the comment the branch deleted). Nothing keeps it there now. It is outside the fence, so leave it.
  • tests/test-durations — klogd_panic_halts, klogd_fault_halts, the four *_halts syscall rows and heap_ceiling_bounds are unpriced.
  • CLAUDE.md:40 — The branch deletes the blank line between Kernel and Userspace daemons, which joins the two paragraphs. This is the orchestrator's file.

REMOVE

  • toyos-abi/src/syscall.rs:793 "Armed once per boot." — the arming is deleted.
  • kernel/src/drivers/panic_console/mod.rs:261 "until recovery"
  • kernel/src/drivers/panic_console/mod.rs:292 "whose fall-through skips recovery — the only path that never paints."
  • tests/toyos-rust-tests/src/bin/process_lifecycle.rs:8 "or panic recovery" — no such teardown exists.
  • tests/common/power.rs:913-915 "a recovered klogd … a halting and a recovering kernel both write" — describes a kernel that cannot exist.
  • issues/isolation/interrupt-entry-keeps-a-ring-3-ac-flag.md:20 "instead of faulting into blame" — blame is deleted.
  • tests/toyos-rust-tests/src/bin/heap_ceiling.rs:114-116 "Measured against the old code: …" — chronology in a paragraph the branch rewrote.
  • kernel-loom/tests/reap_gate.rs:29-30 "Verified 2026-08-17, both ways round." — a date in a rewritten sentence, verifying a message that changed.
  • kernel/src/sched/dump.rs:387 — a comment rewritten to a bare assertion.
  • toyos-sched/src/hw.rs:104-107 — rewritten so that "both exits" now carries one exit's reason.
  • kernel/src/quiesce.rs:28-29 "klogd and iod" — a rewritten list that K3–K5 move.
  • issues/kernel/deferred-release-outlives-its-syscall.md:240-241, issues/kernel/every-wait-in-this-kernel-is-a-spin.md:160-162, :384-385 — 73e00d78 corrected these citations instead of deleting them.
  • PR body "## Kept from the earlier rounds" — review chronology in main's record.
  • PR body "Unsure", first bullet — the sysroot narrative (22:08, "every harness run from this worktree") belongs to its issue, not to main's record.

SEND BACK

Japabu and others added 4 commits September 27, 2026 22:29
…aphs again

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`syscall_panic_halts` went red at f9700b9: the guest printed "returning
this machine to firmware" and QEMU reported no reset for the rest of the
harness's wait. The QMP connection was up (a late connect to an exited
QEMU fails loudly instead), and the 20 s drain after it ran to its end,
so QEMU was still running.

The mechanism, read from the code and that run's serial:
- `klogd` was stopped after one 16-byte UART burst of a line
  (`[kernel 0.573 cp`). The halt IPI is a fixed vector, so a CPU stops only
  with interrupts on, and the one console lock held with interrupts on is
  the wire.
- `reboot_now` reset through `acpi::reboot`, whose `serial::flush_final`
  spins `PANIC_LOCK_SPIN_LIMIT` (100,000,000 `try_lock`+`pause`) on that
  wire. The limit is a count, not a time bound.

`reboot_now` now resets through `acpi::reset_now`: the log is already drained
by then, and nothing on the panic path waits on a lock a stopped CPU may
hold. The count itself is filed:
`panic-path/the-panic-lock-spin-limit-is-a-count-its-comment-calls-a-second.md`.

The syscall-death rows also subscribe to QMP before they write the command,
as `machine_reboot` does (`watch_the_bound`). QMP delivers no event emitted
before its client connected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ult_gates

- **The demand-paging row.** `SYS_DEBUG` `NULL_READ` is now a Ring 0 read of
  the address its caller passes, recorded first as `SYS_DEBUG: a Ring 0 read
  of <addr>` (a record, because the panic path drains records and a program's
  own line is still in userland when the machine ends). `test_panic_child`
  maps 4 MiB it never touches and passes base + 2 MiB. `syscall_fault_halts`
  requires `KERNEL PANIC: read unmapped address at <that addr>`: a kernel that
  demand-paged the window for Ring 0 re-executes into SMAP's protection fault
  and says `protection violation` instead. This row replaces the read of 0.
- **`panic_recovery` is folded into `fault_gates`**: a `pf` arm in
  `fault_gate_child` whose report `check_fault_gates` requires. Deleted:
  `panic_recovery.rs`, `check_panic_recovery`, its timeout and duration rows,
  and `segfault_child`, which nothing else ran.
- **`(Ready(_), Dead)` is not a legal edge.** The only production edge into
  `Dead` is `Running -> Dead`.
- **`deaf_window` holds an `IrqGuard`**; the unconditional `sti` had no
  reason left.
- **`in_syscall` stays.** The stack-top frame cannot tell a syscall's from a
  Ring 3 interrupt's without reading words userland chose. The issue gains
  the other migration direction and `process::handle_fault`.
- The review's REMOVE lines are deleted, as are the `fatal_exception`
  citations in the two wall-4 issues; the recovering-`klogd` sentence went
  with the previous commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Japabu and others added 2 commits September 27, 2026 23:53
…he capture

- **The window was never demand-paged.** An anonymous `mmap` is mapped when
  it is made (`sys_mmap` allocates its pages and maps them with
  `alloc_and_map`/`map_range`, and the region is `RegionKind::Mapped`), so
  base + 2 MiB of one was present, and the Ring 0 read took SMAP's fault on a
  present page: `read protection violation`, which the row refused. The
  kernel's report was right; the test's region was wrong. Demand-paged
  (`RegionKind::Anonymous`) regions are the loader's: a segment's pages past
  its file bytes. `test_panic_child` now reads the 2 MiB-aligned window
  inside a 4 MiB `static mut` it never touches, so the fault is not-present
  and must say `unmapped`, and a kernel that fills it for Ring 0 re-executes
  into SMAP's `protection violation`.
- **A syscall-death check's error carries the capture.** These guests put the
  16550 on stdio, so no `uart-*.log` exists for the red run to keep, and the
  check's bare message was all that was left. The harness-wide gap is filed:
  `issues/build/a-guest-with-no-virtio-keeps-no-serial-when-its-run-reds.md`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Round 5 review of eb19f5b5 against origin/main (e5ffe950).

Gate.

  • PR CI host at eb19f5b5 is queued and has no conclusion yet (run 36353717547). The previous CI run, at 8e525da4, was cancelled.
  • The local host gates at eb19f5b5 all exited 0 (nokthread-r5/gates.summary): clippy, --lib, --workspace --exclude toyos-build, --list and --build-only.
  • Every guest result below is the orchestrator's run.
  • Net: +582 −1737. Production +118 −902, tests (tests/, kernel-loom/tests) +246 −682, issues +217 −153, CLAUDE.md +1.

Round-3 BLOCKERs

BLOCKER

(none)

NOTE

  • PR body, the "Ring 0 demand paging" row and the "Green arms" paragraph — both still say "owed at eb19f5b5" — the measurements now exist (the 553r5 lines). The claim stands only on the body, so the body must carry each command, exit code and log.
  • PR CI host at eb19f5b5 — no conclusion yet — it must be success before the merge.
  • syscall_panic_halts went red once at f9700b90: the guest announced its reset and QEMU saw no reset — the cause is unexplained, and it is recorded only in the PR body. File it under issues/panic-path/ with that run's log as evidence and an exit condition. Until it is explained, it is a panic-path defect.
  • PR body, "What Kernel: a kill never waits on its victim — the last thread out tears its process down #549 must drop" — the list is not exact:
  • toyos-sched/src/hw.rs:89,104-105 — type IrqGuard and fn irq_guard, re-documented as "Has no caller in either world" — with the panic-recovery reason gone, this is dead code. Delete it together with its impls at kernel/src/arch/x86_64/hw.rs:36,53-55, toyos-sched/src/cpu.rs:2501,2507 and toyos-sched/sim/src/hw_impl.rs:129,156.
  • tests/toyos.rs:3192 check_ring0_read_unmapped accepts the address 0 — so the row loses its power if the test child regresses. For example, test_panic_child calling syscall::debug(action) for every action reads null and stays green, and so does demand-paging-ring0.patch on top of it. Refuse addr == 0; that mutant must then red on syscall_fault_halts.
  • kernel/src/sched/dump.rs:387 — no guest run covers deaf_window's IrqGuard — its one test is dump_nmi_probe (Nightly), and it is in none of the runs. Owed: cargo test --test toyos-build -- --nightly dump_nmi_probe at the head.
  • toyos-abi/src/syscall.rs:792 NULL_READ — the name is false now that the action reads the caller's address — a rename costs one sysroot rebuild, which is no reason to keep a false name.

REMOVE

  • issues/panic-path/the-panic-lock-spin-limit-is-a-count-its-comment-calls-a-second.md:16-27 — it gives, as the cause of the f9700b90 red, a mechanism the PR body calls unproven and the held-wire arm did not reproduce. The defect stands on the code.
  • same file :13-14 "and nobody has measured that on any machine this kernel boots" — false. wire-held-red measured it: 27 s against 7–8 s.
  • same file :29 "the run above, and" — it dangles once :16-27 go.
  • kernel/src/drivers/xhci/stop.rs:103 "written from acpi::reboot" and :638 "Called from acpi::reboot and acpi::shutdown and nowhere else" — both false. The panic path now reaches before_reset through reset_now alone.
  • kernel/src/drivers/panic_console/mod.rs:291-292 "Called from the idle loop." — this only narrates the caller.
  • tests/common/metal.rs:722 "and panic_recovery" — it cites a test the branch deleted.
  • tests/toyos.rs:3141-3146 — the "N = 12 rounds of fault_gates + panic_recovery" provenance. It cites the deleted test, and measurement provenance does not belong in source.
  • PR body, "Unsure": the 69e8dff5/c14fc9af bullet and the eb19f5b5 co-author bullet — commit chronology, not main's record.
  • PR body, "Unsure": the metal-panic-probe bullet — it is neither filed nor load-bearing.

LAND AFTER NAMED CHANGES

Japabu and others added 2 commits September 28, 2026 01:04
…ck refuses address 0, and three stale comments the review named are cut

`Machine::irq_guard`/`type IrqGuard` have no caller in either world — `deaf_window`
calls the concrete `crate::arch::IrqGuard::close()` directly and never went
through the trait. Deleted from the trait and its five impls (kernel x86_64
and aarch64, the sim, the unit-test double, and the loom model), the last two
of which the review's own list missed and clippy's aarch64 kernel invocation
and the loom build caught in their place.

`check_ring0_read_unmapped` accepted address 0, so a `test_panic_child`
regression to `syscall::debug(action)` for every action — dropping the
demand-paged window argument — would read null and stay green. It now refuses
`addr == 0` by name.

Three comments the review found false once panic recovery left: `xhci/stop.rs`
no longer says its summary is written from `acpi::reboot` or that `before_reset`
has exactly two callers (`reset_now` is the only path now); `panic_console`'s
`probe_due` doc drops "called from the idle loop", which only narrated the
caller; the spin-limit issue drops the disproven "nobody has measured that"
claim and the `f9700b90` paragraph the wire-held measurement already refuted,
along with the `tests/common/metal.rs` and `tests/toyos.rs` comments citing the
deleted `panic_recovery` test.

Filed `issues/panic-path/a-syscall-panic-reset-the-guest-announced-and-qemu-never-saw.md`
for the `f9700b90` red the spin-limit issue used to (wrongly) explain.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu
Japabu added this pull request to the merge queue Sep 27, 2026
Merged via the queue into main with commit 1ec6daa Sep 28, 2026
1 check passed
Japabu added a commit that referenced this pull request Sep 28, 2026
Brings in #537, #553 (every kernel panic halts, usbd and poison deleted,
the no-kernel-threads track), #559 and #561. The kernel now spawns one
thread outside the actuator build, klogd: this branch deleted iod and
main deleted usbd, and neither side's replacement is a kthread.

Conflicts, each resolved against both sides' hunks:

- kernel/src/drivers/nvme.rs, kernel/src/iod.rs (modify/delete): deleted.
  Main's changes to them were the `mm::policy::MmioPolicy` import rename
  and dropping `OnPanic::Recover` from iod's spawn, adaptations to code
  this branch removes; neither carries behaviour to move elsewhere.
- kernel/src/sched/kthread.rs: main's row table without panic policy.
  MAX_KERNEL_TASKS is 1 (klogd), and 2 + MAX_LOG_SHARDS in the actuator
  build (klogd, lognest, one logstorm per shard, which can run in one boot).
- kernel/src/main.rs: neither `usbd::start()` nor `iod::start()`; main's
  panic handler without recovery, this branch's storage phase.
- kernel/src/quiesce.rs: main's header, which names no kernel thread.
- kernel-loom/src/lib.rs: neither `poison` (main) nor `durability` (here).
- tests/toyos.rs: `heap_ceiling_bounds` (main's rename) without
  `cache_eviction` (deleted here); `blocked_dump` and `klogd_hosted` ask
  for klogd alone; `klogd_panic_halts` is main's, whose usbd arm and
  recover row are gone with usbd and poison.
- issues/build/a-lane-s-tap-socket-path-outgrows-sun-len-on-the-dev-host.md:
  main's body and `kind: tooling`, this branch's two-shapes measurement;
  `status: open`, since no disabled row names `lan_mdns_answer` on either
  side and `--known-red` answers NO, so the quarantine paragraph goes.

Beyond the conflicts: toyos-inventory gets the `description` main's
hostws gate now requires of every workspace package, and the
no-kernel-threads track loses K5, which this branch meets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Japabu added a commit that referenced this pull request Sep 28, 2026
main's #553 deletes panic recovery: every kernel panic halts. This branch's
poison work goes with it: `toyos-proclife/src/poison.rs`, `Op::Poison` and
its two scripts, `World::poison` and its `poisoned` set, `process::PoisonWake`
and `zombify_poisoned`, the idle loop's poison bank in `reap_finished`, the
TLS issue a poisoned sibling raised, and the "poison path" wording in
`teardown.rs` and `ProcessEntry::teardown_code`.

`kthread::open_selftest` reads `ROWS` as `[AtomicU64]`. `mark_thread_zombie`,
which main kept and this branch's `leave` replaced, goes. main.rs's
`usbd::start()` context is gone with usbd.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu
Japabu deleted the wt/toyos-nokthread branch September 28, 2026 09:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant