Kernel: every kernel panic halts and panic recovery is deleted; delete usbd; open the no-kernel-threads track (K1) - #553
Conversation
The owner's ruling: the kernel creates no thread but the per-CPU idle loop, and the ruling holds only once the whole kernel-thread machinery is deleted. The track names its exit condition, the spawn sites left once K1 lands, and stages K1-K6 with what unblocks each. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
usbd's body only parked. It goes with its start() call, its `usbd-panic` actuator, and every test list that named it. It was also the only thread that ever walked `OnPanic::Recover`. The one other claimant, iod, was never measured recovering, and #536 deletes iod. So the policy goes rather than moving its actuator to iod: `OnPanic`, the row's `recoverable` word, `panic_recovers_here`, and the Release/Acquire pair that published the word before the identity. A kernel thread's panic now halts the machine, because `percpu::in_syscall` compares the running task's identity with the one that entered the syscall, and a kernel thread never enters one. The panic handler asks only that. A row is now one word. Every reader learns the id it compares through something ordered after the publish: the table lock held across it, or the run queue the task is dispatched from. Relaxed is enough for that. klogd_panic_halts's verdict was the ready marker's absence. A recovered klogd takes the console with it, so that verdict held with kernel-thread panics made recoverable (EXIT=0 on the mutated kernel). The verdict is now the line only halt_all_cpus writes, awaited as an event, and the same mutation reds it (EXIT=1). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The owner's ruling, placed where every agent reads it before it knows which subsystem it is in. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review of Panic path, every context. A user thread is unchanged from main: BLOCKER
NOTE
REMOVE
SEND BACK |
Four kernel threads still exist; the rule is what new work follows, so it sits among the principles the tree does not yet meet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… kernel thread's death Answers the first review of #553. Faults and panics now ask one question. `blame` took "a thread is current" as its input, so a Ring 0 page fault on a user address was the process's whenever any tid was current: a kernel thread's null dereference poisoned `klogd` and the machine carried on silent, and an interrupt handler's null dereference killed whichever user thread it landed on. `blame` now takes `percpu::in_syscall()`, the input the panic handler already reads, so a Ring 0 fault outside a syscall is the kernel's and halts, and one inside a syscall stays the process's. The audit behind it: every user-memory access in the kernel goes through the direct map (`user_ptr::window`, the futex word, the loader's `KernelSlice`s, the inbox and shm pages, the crash dump's hand walk); the only Ring 0 dereference of a user-half address is `SYS_DEBUG`'s staged `NULL_READ`, inside a syscall. Three defects the audit found are filed rather than fixed here: issues/panic-path/the-syscall-bracket-outlives-a-migrated-syscall.md, issues/isolation/interrupt-entry-keeps-a-ring-3-ac-flag.md and issues/isolation/a-ring-0-page-fault-demand-pages-user-memory.md. `klogd-fault` reads address zero on klogd's first instruction. `klogd_panic_halts` and the new `klogd_fault_halts` share `power::klogd_death_resets`: boot on `panicked()`'s guest with `panic-reboot-fast`, and require QEMU's own `guest-reset` inside the bound, the way `panic_reboots` does, which now shares `resets_inside_the_bound` with them. The two-boot price of `klogd_panic_halts` is dropped from tests/test-durations; the name is unmeasured on a runner until the next profile. Review REMOVEs applied: the kthread module-doc and publish comments, main.rs's spawn-order comment, the MACHINE_TESTS row comment, the false arm-line claim; the track loses K1, the pasted grep, K3's userland option and K4's pointer, and K4/K5 are rewritten as the orchestrator ruled; the stale usbd prose in three issue files is deleted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…arm line The first cut waited for the arm line as its ready marker, so a machine that recovered went red on the guest's silence before QEMU was asked. The boot now stops at a line both a halting and a recovering kernel write (`PANIC: panicked at`, `#PF UNHANDLED: cr2=0x0`), and the verdict is `QmpShutdown`'s `guest-reset` inside the bound; the report's lines and the arm line are read from the boot log plus the drain after it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ms write Under the old blame input a kernel thread's fault is recovered and nothing after klogd's spawn line reaches the wire: the fault's own `#PF UNHANDLED` record is committed and never drained. So a report line as the boot's marker still let the guest's silence decide. The marker is now `kthread: klogd pid=`, written before klogd runs, and QEMU's stop reason is the only verdict on halt versus carry on. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under M1 the recovered machine goes on writing; the spawn line is only one both arms write. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The owner's ruling: a kernel panic means a kernel invariant broke, and the
kernel does not keep running on state it can no longer trust. A panic inside
a syscall halts like any other; a Ring 0 fault is a kernel bug and halts; a
Ring 3 fault still kills only its process. "The kernel never crashes from
userland" is kept by refusing bad input at the boundary, not by surviving
panics. This reverses the rule that a recoverable panic ends only the
offending process.
Deleted:
- the panic handler's `in_syscall()` branch, `try_recover_from_panic` (x86
and the aarch64 stub) and `recover_or_halt`;
- `sched/poison.rs`, the per-CPU poison bank, `poison_tid`,
`schedule_no_return`, `process::PoisonWake`/`zombify_poisoned`,
`toyos_proclife::poison`, `Watch::thread`, the `poison-overwrite` feature
and its loom model and CI red; the idle loop's `reap_poisoned` is now
`reap_finished` and only collects published exits;
- `panic_console::discard_capture`, `CaptureAccess::discard`,
`CaptureLatch::owned_by` and the loom model of a discard;
- `toyos_userbound::{blame, Blame, Faulted}`: whose fault a trap is is now
its ring alone, and `fatal_exception` kills a Ring 3 fault's process and
halts on anything else;
- Ring 0 demand paging: `page_fault_handler` resolves a not-present fault
only for a Ring 3 frame, so a Ring 0 fault halts before anything is
mapped into the current process; `handle_page_fault`'s kernel-thread
refusal went with it;
- `SYS_DEBUG` action 2's one-shot arming, which existed to refuse a second
call into a lock a recovered panic stranded;
- the aarch64 `percpu::in_syscall` stub. The x86 one stays: the crash
report reads it to print the syscall a death happened inside.
Tests:
- `panic_recovery` keeps its Ring 3 arm only (a user segfault kills its
process and the system lives), and leaves `ACTUATOR_TESTS`.
- `syscall_panic_halts`, `syscall_fault_halts`, `lock_across_switch_halts`
and `heap_over_ceiling_halts` each drive one `SYS_DEBUG` death on its own
boot and take the verdict from QEMU's `guest-reset` through QMP, through
`power::syscall_death_resets`, which shares `klogd_death_resets`'s tail.
- `heap_ceiling_recovery` is `heap_ceiling_bounds`: its over-ceiling arm
and the heap-still-works arm after it are gone.
- `screen_recoverable_untouched` and `screen_survived_panic_not_blamed`
are deleted: there is no survived panic to paint or not paint.
Issues: the stale-bracket, Ring-0-demand-paging, stranded-PROCESS_TABLE and
discard-refusal defects are gone with the code they were about. What is left
of the stale bracket is a crash report naming a finished syscall, filed
narrowly; `idle_stack_guard`'s inert "the read succeeded" arm is filed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the wedge issue loses its recovery paragraph `panic-reboot-fast` already selects the test kernel, which carries `test-actuators`, and the harness refuses a boot that also names the build. The wedge arms' DEADLOCK panic in `logd`'s fsync now halts like any other; `usb_reset_records_the_phase_it_cut` stays green on this tree, so only the paragraph describing the recovery composition goes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Found while measuring this branch's Ring-3 red arm: the run published a sysroot without `cargo`, and every harness run from the worktree has panicked on the C corpus since. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion` A Ring 3 fault's `kill_process(-1)` now sits in `fatal_exception` itself. The sysroot issue says only what was read, not when the primary's build ran. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review of Round 1 BLOCKER, no independent oracle for the kernel-thread halt: CLOSED. BLOCKER
NOTE
REMOVE
SEND BACK |
…aphs again Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`syscall_panic_halts` went red at f9700b9: the guest printed "returning this machine to firmware" and QEMU reported no reset for the rest of the harness's wait. The QMP connection was up (a late connect to an exited QEMU fails loudly instead), and the 20 s drain after it ran to its end, so QEMU was still running. The mechanism, read from the code and that run's serial: - `klogd` was stopped after one 16-byte UART burst of a line (`[kernel 0.573 cp`). The halt IPI is a fixed vector, so a CPU stops only with interrupts on, and the one console lock held with interrupts on is the wire. - `reboot_now` reset through `acpi::reboot`, whose `serial::flush_final` spins `PANIC_LOCK_SPIN_LIMIT` (100,000,000 `try_lock`+`pause`) on that wire. The limit is a count, not a time bound. `reboot_now` now resets through `acpi::reset_now`: the log is already drained by then, and nothing on the panic path waits on a lock a stopped CPU may hold. The count itself is filed: `panic-path/the-panic-lock-spin-limit-is-a-count-its-comment-calls-a-second.md`. The syscall-death rows also subscribe to QMP before they write the command, as `machine_reboot` does (`watch_the_bound`). QMP delivers no event emitted before its client connected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ult_gates - **The demand-paging row.** `SYS_DEBUG` `NULL_READ` is now a Ring 0 read of the address its caller passes, recorded first as `SYS_DEBUG: a Ring 0 read of <addr>` (a record, because the panic path drains records and a program's own line is still in userland when the machine ends). `test_panic_child` maps 4 MiB it never touches and passes base + 2 MiB. `syscall_fault_halts` requires `KERNEL PANIC: read unmapped address at <that addr>`: a kernel that demand-paged the window for Ring 0 re-executes into SMAP's protection fault and says `protection violation` instead. This row replaces the read of 0. - **`panic_recovery` is folded into `fault_gates`**: a `pf` arm in `fault_gate_child` whose report `check_fault_gates` requires. Deleted: `panic_recovery.rs`, `check_panic_recovery`, its timeout and duration rows, and `segfault_child`, which nothing else ran. - **`(Ready(_), Dead)` is not a legal edge.** The only production edge into `Dead` is `Running -> Dead`. - **`deaf_window` holds an `IrqGuard`**; the unconditional `sti` had no reason left. - **`in_syscall` stays.** The stack-top frame cannot tell a syscall's from a Ring 3 interrupt's without reading words userland chose. The issue gains the other migration direction and `process::handle_fault`. - The review's REMOVE lines are deleted, as are the `fatal_exception` citations in the two wall-4 issues; the recovering-`klogd` sentence went with the previous commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he capture - **The window was never demand-paged.** An anonymous `mmap` is mapped when it is made (`sys_mmap` allocates its pages and maps them with `alloc_and_map`/`map_range`, and the region is `RegionKind::Mapped`), so base + 2 MiB of one was present, and the Ring 0 read took SMAP's fault on a present page: `read protection violation`, which the row refused. The kernel's report was right; the test's region was wrong. Demand-paged (`RegionKind::Anonymous`) regions are the loader's: a segment's pages past its file bytes. `test_panic_child` now reads the 2 MiB-aligned window inside a 4 MiB `static mut` it never touches, so the fault is not-present and must say `unmapped`, and a kernel that fills it for Ring 0 re-executes into SMAP's `protection violation`. - **A syscall-death check's error carries the capture.** These guests put the 16550 on stdio, so no `uart-*.log` exists for the red run to keep, and the check's bare message was all that was left. The harness-wide gap is filed: `issues/build/a-guest-with-no-virtio-keeps-no-serial-when-its-run-reds.md`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Round 5 review of Gate.
Round-3 BLOCKERs
BLOCKER(none) NOTE
REMOVE
LAND AFTER NAMED CHANGES |
…ck refuses address 0, and three stale comments the review named are cut `Machine::irq_guard`/`type IrqGuard` have no caller in either world — `deaf_window` calls the concrete `crate::arch::IrqGuard::close()` directly and never went through the trait. Deleted from the trait and its five impls (kernel x86_64 and aarch64, the sim, the unit-test double, and the loom model), the last two of which the review's own list missed and clippy's aarch64 kernel invocation and the loom build caught in their place. `check_ring0_read_unmapped` accepted address 0, so a `test_panic_child` regression to `syscall::debug(action)` for every action — dropping the demand-paged window argument — would read null and stay green. It now refuses `addr == 0` by name. Three comments the review found false once panic recovery left: `xhci/stop.rs` no longer says its summary is written from `acpi::reboot` or that `before_reset` has exactly two callers (`reset_now` is the only path now); `panic_console`'s `probe_due` doc drops "called from the idle loop", which only narrated the caller; the spin-limit issue drops the disproven "nobody has measured that" claim and the `f9700b90` paragraph the wire-held measurement already refuted, along with the `tests/common/metal.rs` and `tests/toyos.rs` comments citing the deleted `panic_recovery` test. Filed `issues/panic-path/a-syscall-panic-reset-the-guest-announced-and-qemu-never-saw.md` for the `f9700b90` red the spin-limit issue used to (wrongly) explain. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings in #537, #553 (every kernel panic halts, usbd and poison deleted, the no-kernel-threads track), #559 and #561. The kernel now spawns one thread outside the actuator build, klogd: this branch deleted iod and main deleted usbd, and neither side's replacement is a kthread. Conflicts, each resolved against both sides' hunks: - kernel/src/drivers/nvme.rs, kernel/src/iod.rs (modify/delete): deleted. Main's changes to them were the `mm::policy::MmioPolicy` import rename and dropping `OnPanic::Recover` from iod's spawn, adaptations to code this branch removes; neither carries behaviour to move elsewhere. - kernel/src/sched/kthread.rs: main's row table without panic policy. MAX_KERNEL_TASKS is 1 (klogd), and 2 + MAX_LOG_SHARDS in the actuator build (klogd, lognest, one logstorm per shard, which can run in one boot). - kernel/src/main.rs: neither `usbd::start()` nor `iod::start()`; main's panic handler without recovery, this branch's storage phase. - kernel/src/quiesce.rs: main's header, which names no kernel thread. - kernel-loom/src/lib.rs: neither `poison` (main) nor `durability` (here). - tests/toyos.rs: `heap_ceiling_bounds` (main's rename) without `cache_eviction` (deleted here); `blocked_dump` and `klogd_hosted` ask for klogd alone; `klogd_panic_halts` is main's, whose usbd arm and recover row are gone with usbd and poison. - issues/build/a-lane-s-tap-socket-path-outgrows-sun-len-on-the-dev-host.md: main's body and `kind: tooling`, this branch's two-shapes measurement; `status: open`, since no disabled row names `lan_mdns_answer` on either side and `--known-red` answers NO, so the quarantine paragraph goes. Beyond the conflicts: toyos-inventory gets the `description` main's hostws gate now requires of every workspace package, and the no-kernel-threads track loses K5, which this branch meets. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
main's #553 deletes panic recovery: every kernel panic halts. This branch's poison work goes with it: `toyos-proclife/src/poison.rs`, `Op::Poison` and its two scripts, `World::poison` and its `poisoned` set, `process::PoisonWake` and `zombify_poisoned`, the idle loop's poison bank in `reap_finished`, the TLS issue a poisoned sibling raised, and the "poison path" wording in `teardown.rs` and `ProcessEntry::teardown_code`. `kthread::open_selftest` reads `ROWS` as `[AtomicU64]`. `mark_thread_zombie`, which main kept and this branch's `leave` replaced, goes. main.rs's `usbd::start()` context is gone with usbd. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Every kernel panic now halts the machine, including a panic inside a syscall, and all of the kernel's panic recovery is deleted. The branch also opens the owner's no-kernel-threads track and does its first stage (K1):
usbdis deleted.The ruling this lands
The owner ruled:
This reverses the old rule that a recoverable panic ended only the process that caused it.
What is deleted, and why
The panic handler's decision. The
percpu::in_syscall()branch is gone. So aretry_recover_from_panic(x86 and the aarch64 stub) andrecover_or_halt. The handler now always captures, flushes and callshalt_all_cpus.Poisoned threads. Deleted:
kernel/src/sched/poison.rs, the per-CPUPOISONEDbank andpoison_tid;schedule_no_return;process::PoisonWakeandzombify_poisoned;toyos-proclife/src/poison.rsandWatch::thread, whose one caller was the poison path;poison-overwritefeature, together with its loom modelkernel-loom/tests/poison_set.rsand its red insrc/ci.rs.The idle loop's
reap_poisonedbecomesreap_finished. It only collects exits that have been published, behind the sameReapGate.The scheduler's
(Ready(_), Dead)edge. Its one stated user wasschedule_no_return. The only production edge intoDeadisRunning → Dead.The panel's survived-panic path. Deleted:
panic_console::discard_capture,CaptureAccess::discard,CaptureLatch::owned_by, and the loom model of a discard. The latch's release stays, becausecapture_intostill uses it. Its loom model is renameda_released_latch_hands_the_snapshot_to_the_next_captor.blame.toyos_userbound::{blame, Blame, Faulted}are deleted. The ring of a trap now decides whose fault it is. That ring is the opaqueRing, which is built only fromcs.fatal_exceptionkills the process on a Ring 3 fault and halts on anything else.Fatal/Panicif the kernel left that state set, and that is a kernel bug.Ring 0 demand paging.
page_fault_handlernow sends a not-present fault toprocess::handle_page_faultonly for a Ring 3 frame:if is_user && process::handle_page_fault(..). It used to do that for any fault taken while a thread was current. A Ring 0 fault falls through tofatal_exception, which halts before anything is mapped into the current process.handle_page_fault's kernel-thread refusal became unreachable and is deleted.SYS_DEBUGaction 2's one-shot arming. It existed to refuse a second call into a lock that a recovered panic had stranded.deaf_window's unconditionalsti. It now holds anIrqGuard. The unconditional enable was there because panic recovery could leave IF clear.Machine::irq_guard/type IrqGuard, the trait methoddeaf_windowdoes not use, are deleted together with their five impls (kernel/src/arch/x86_64/hw.rs,kernel/src/arch/aarch64/hw.rs,toyos-sched/src/cpu.rs,toyos-sched/sim/src/hw_impl.rs,toyos-sched/loom/tests/loom_retire.rs);deaf_windowcalls the concretecrate::arch::IrqGuard::close()directly and never went through the trait. The review named only the first three; the aarch64 kernel and the loom build are what caught the other two.The aarch64
percpu::in_syscallstub. It has no caller now.The kernel-thread panic policy.
OnPanic, a row'srecoverableword andpanic_recovers_hereare deleted.K1.
usbdis deleted:kernel/src/drivers/xhci/usbd.rs, itsstart(), theusbd-panicactuator, andusbdin the test lists. The track isissues/kernel/the-kernel-still-creates-threads.md(kind: track), with stages K2–K6.The panic path's reset does not wait on the console wire
panic_reboot::reboot_nowresets throughacpi::reset_now, notacpi::reboot. The log is already drained by that point, andacpi::rebootopens withserial::flush_final, which spinsPANIC_LOCK_SPIN_LIMIT(100,000,000 iterations oftry_lockandpause) on the console wire. A CPU this panic stopped may hold that wire, and a stopped CPU never releases it, so that wait could only run out. The limit being a count and not a time is filed aspanic-path/the-panic-lock-spin-limit-is-a-count-its-comment-calls-a-second.md.This change is measured green, but what it fixes is not proven. At
f9700b90,syscall_panic_haltswent red once: the guest printedreturning this machine to firmware, QEMU reported no reset in the harness's wait, and the test took 98 s at a 2.96x liveness width. The held-wire arm below does not reproduce that. With the wire held and the oldacpi::rebootpath, the row stays green (EXIT=0). It takes 27 s, against 7–8 s for the fixed path at the same head, so the held wire's spin costs about 19–20 s at a 1.50x width and still resets inside the 37 s budget. A held wire alone therefore does not account for that red, and nothing measured here explains it. No deterministic red arm exists on the old path:flush_final's spin is bounded, so a held wire delays the reset and does not stop it.The syscall-death rows also subscribe to QMP before they write the command (
power::watch_the_bound), the waymachine_rebootdoes. QMP delivers no event that was emitted before its client connected.Tests
Four
Nightlyrows, all run throughpower::syscall_death_resets:SYS_DEBUGactionsyscall_panic_haltsPANICSyscall: num=92, the user backtracesyscall_fault_haltsNULL_READon the 2 MiB-aligned window inside a 4 MiB.bssarray the child never touchesKERNEL PANIC: read unmapped address at <that address>,Syscall: num=92, the user backtracelock_across_switch_haltsLOCK_ACROSS_SWITCHcheck_tripwire_attributionscoping itspanicked attosyscall/dispatch.rsheap_over_ceiling_haltsHEAP_OVER_CEILINGexceeds MAX_HEAP_ALLOCEach row boots
panic-reboot-fastwith QMP and runstest_rs_test_panic_child <action>. The verdict is QEMU'sguest-resetinside the fast bound plus the reset allowance, the same readingklogd_death_resetstakes (klogd_panic_halts,klogd_fault_halts). The two sharedied_and_reset. When a row's own check fails, its error carries the guest's capture. These guests put the 16550 on stdio, so nouart-*.logexists for a red run to keep.syscall_fault_haltscarries the demand-paging claim.NULL_READis a Ring 0 read of the address its caller passes. The kernel recordsSYS_DEBUG: a Ring 0 read of <addr>first, andcheck_ring0_read_unmappedreads the address from that record. It is a record and not the child's own line because the panic path drains records, while a program's line is still in userland when the read ends the machine.check_ring0_read_unmappedalso refuses address 0: a child that regressed to callingsyscall::debug(action)for every action, dropping the window argument, reads address 0 and must red rather than pass.mmapdoes not qualify:sys_mmapmaps its pages when it makes it, and the region isRegionKind::Mapped. A Ring 0 read of one takes SMAP's fault on a present page and saysprotection violation.RegionKind::Anonymousregions are the loader's, a segment's pages past its file bytes. So the child reads the aligned window insidestatic mut UNTOUCHED: [u8; 4 MiB]. In the built child,UNTOUCHEDis the first 0x400000 bytes of.bssat 0x7e2c8, in aPT_LOADsegment with 5656 file bytes and 4200196 memory bytes, so the window holds no other object and no file bytes.+smapon both x86-64 CPU models,src/arch.rs) and sayprotection violationinstead. This row replaces the read of address 0.panic_recoveryis folded intofault_gates.fault_gate_childgains apfarm (read_null).check_fault_gatesrequires itsSEGFAULT tid=header andfault_gate_child::read_nullin the backtrace. Deleted:panic_recovery.rs,check_panic_recovery, its timeout and duration rows, andsegfault_child, which nothing else ran.heap_ceiling_recoveryis renamedheap_ceiling_bounds. Its over-ceiling arm moved toheap_over_ceiling_halts. Its "the heap still works after recovery" arm is deleted.screen_recoverable_untouchedandscreen_survived_panic_not_blamedare deleted. They tested a panic that survives, and none can now.test_panic_childneeds an action. Every caller names one.Harness prose that relied on recovery is cut:
serial::Died::Kernel;qemu::ceiling_verdict's doc;await_guest's scoping comment;Ppm::identical_to's doc;needs_actuatorsandsuite_split.Unpriced in
tests/test-durations:syscall_panic_halts,syscall_fault_halts,lock_across_switch_halts,heap_over_ceiling_halts,klogd_panic_halts,klogd_fault_haltsandheap_ceiling_bounds. The committed profile is measured on a CI runner (committed_durations_path), and a dev host's TCG times are not written into it.fault_gatesis priced at 31 from before it had the eighth arm.panic_recovery's row is deleted.Issues
panic-path/panic-holding-process-table-hangs.md;panic-path/the-discards-refusal-branch-is-exercised-by-nothing.md.panic-path/a-crash-report-can-name-a-syscall-that-already-ended.md. It covers both migration directions andprocess::handle_fault's unbracketedsyscall_num();panic-path/the-panic-lock-spin-limit-is-a-count-its-comment-calls-a-second.md;build/idle-stack-guards-returned-arm-reads-a-line-nothing-writes.md;build/a-sysroot-cloned-during-a-toolchain-rebuild-never-gets-its-cargo.md;build/a-guest-with-no-virtio-keeps-no-serial-when-its-run-reds.md:lane::keep_serialkeeps onlyuart-*.log, so a red run of a guest whose 16550 is on stdio keeps an empty lane. The harness owns it;isolation/interrupt-entry-keeps-a-ring-3-ac-flag.md;panic-path/a-syscall-panic-reset-the-guest-announced-and-qemu-never-saw.md: thef9700b90red above, filed once the wire-held measurement refuted the only proposed mechanism.nothing-charges-kernel-memory-to-a-process.md: an unbounded grower now ends the machine;a-deliberate-wedge-…md: the recovery paragraph is deleted;the-capability-end-state-is-twelve-answers.md: its kernel-resident-workers section is deleted, since the track states the opposite rule;the-kernel-is-small-interrupts-post-and-threads-wait.md:usbdis gone from it.in_syscallstaysin_syscallstays, because only the crash report'sSyscall:lines read it now. Reading the syscall frame at the top of the thread's own kernel stack instead does not work: that frame cannot tell a syscall's frame from a Ring 3 interrupt's without trusting words userland chose. An interrupt putsSSandCSin the slots where a syscall's frame holds the user'srspandrdi. The per-thread bit that would be left needs new per-thread state written on every syscall, not one rule in one function. The issue above records this.Gates
At
fe7c7a6c, which is this branch after mergingorigin/mainat41ad548c:cargo run -- --clippy, including the three aarch64 kernel invocationscargo test --libcargo test --workspace --exclude toyos-build(toyos-sched, its sim and loom included)cargo test --test toyos-build -- --list(builds the C corpus and every Rust test binary, checks the redlist, boots nothing)cargo run -- --build-onlyMeasured by the orchestrator at
fe7c7a6c:syscall_fault_halts --nightlyEXIT=0; with the demand-paging mutant EXIT=1, "expectedKERNEL PANIC: read unmapped address at 0x10000200000: the read did not fault as unmapped"; with the child's window multiplied by 0 (the address-0 arm) EXIT=1, "expected the demand-paged window, not the null read";dump_nmi_probe --nightlyEXIT=0; the Fast tier 393 passed, 1 failed, the one beinglan_mdns_answer(a macOS socket path over SUN_LEN in the harness, fixed by its own PR). Earlier heads' guest results below stand where this round did not touch their paths.Negative controls and the independent oracle (run by the orchestrator)
The independent oracle is QEMU's own
SHUTDOWNevent, read throughqemu::QmpShutdown. It is the hypervisor reporting that the guest reset, not the guest's serial.Each patch is applied with
git apply --checkthengit apply, and reversed withgit apply -R. Each mutated tree passescargo run -- --clippy(EXIT=0).demand-paging-ring0.patch:if (is_user || percpu::current_tid().is_some()) && process::handle_page_fault(..)cargo test --test toyos-build -- --nightly syscall_fault_haltseb19f5b5(553r5-demandpaging-syscall_fault_halts.log). The kernel saysKERNEL PANIC: read protection violation at 0x10000200000; the check saysexpected \KERNEL PANIC: read unmapped address at 0x10000200000`: the read did not fault as unmapped at 0x10000200000`. The machine still resets, so the red comes from the named assertion.wire-held-red.patch:reboot_nowgoes back toacpi::reboot, with the wire held first (core::mem::forget(serial::try_wire()))cargo test --test toyos-build -- --nightly syscall_panic_halts8e525da4, 27 s. A measurement of the wait, not a red arm; see above.wire-held-green.patch8e525da4, 8 s.R2-ring3-halts.patch:if false && ctx.ring().is_user()infatal_exceptioncargo test --test toyos-build -- --nightly fault_gates8e525da4.Green arms, orchestrator-run: at
8e525da4,cargo test(Fast) — 393 of 394 passed, the one redlan_mdns_answer(path must be shorter than SUN_LEN), a harness socket-path defect #560 fixes and red the same way on #536, #541, #554, #555 and #559 (#560's own run of it is EXIT=0); and, undercargo test --test toyos-build -- --nightly <row>, each ofsyscall_panic_halts,lock_across_switch_halts,heap_over_ceiling_halts,klogd_panic_halts,klogd_fault_halts,heap_ceiling_bounds,panic_reboots,panic_key_holds,fault_gatesanddisk_backtraceEXIT=0 (553r4-*.log);syscall_fault_haltsalone was EXIT=1 there (553r4-syscall_fault_halts.log), fixed by355e0d64. Ateb19f5b5, rerun under the same nightly command:syscall_fault_halts,syscall_panic_halts,lock_across_switch_haltsandheap_over_ceiling_haltseach EXIT=0 (553r5-green-*.log).Unsure / left outside the fence
syscall_panic_haltsred atf9700b90is unknown, as the section above says.NULL_READkeeps its name while it reads any address. A rename changestoyos-abi/srcand so the sysroot key, which is outside this brief.Net against
origin/main(41ad548c): +630 −1782. Production +130 −1061 (kernel +102 −503, toyos-userbound +8 −226, toyos-proclife +5 −158, kernel-loom +8 −136, toyos-sched +4 −31, toyos-abi +2 −2, src +1 −5); tests +243 −568; issues +256 −153;CLAUDE.md+1, the orchestrator's.🤖 Generated with Claude Code