K3: delete the logstorm and lognest kernel threads - #586
Conversation
kernel/src/log/storm.rs and kernel/src/log/nested.rs existed only to exercise four guest tests (log_conservation_smp1, log_nested_emit, log_reserve_window, log_reserve_window_negative). Delete both files, their kthread::spawn call sites, their five actuators and cmdline tokens (log-storm, log-unbracketed-reserve, log-nested-emit, log-nested-reserve, log-shared-reservation), the x86-64 log_nest IDT gate and vector, the aarch64 LogNest vector, the userland test-runner log-gate builtin those tests alone drove, and every test registration and helper that served them — leaving klogd and iod as the kernel's only remaining kthread::spawn sites. Two real claims those tests alone checked — same-CPU interrupt reentrancy inside emit's IF/TF-off bracket, and SYS_LOG_READ's conservation law under concurrent multi-shard write load — are now unverified, recorded in issues/kernel/deleting-logstorm-and-lognest-left-two-log-claims-unverified.md rather than left silent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Review, round 1, at 4c35a94Gate: CI Net: +51 / -1565. Production ( BLOCKER
NOTE
REMOVE
SEND BACK |
Round 1 deleted the logstorm and lognest kernel threads together with the four guest tests they drove. The threads stay deleted; the checks come back, with producers that are not kernel threads. The nesting gate (log_nested_emit, log_reserve_window and the negative control log_reserve_window_negative) is restored unchanged except for where its body runs: log::nested::start_once now runs it inline in the first SYS_LOG_READ, between enable_interrupts() and disable_interrupts(). That is the kthread's context: Ring 0, IF=1, not preempted (the Ring 0 timer only sets need_resched), as the page-fault arm and deadline.rs already run. The log_nest gate, LOG_NEST_VECTOR, aarch64 Vector::LogNest (so Hda and VirtioSound keep 2 and 3), IrqGuard::unclosed, both injection points, the kernel-loom shims and the three actuators return with it. nested::inject is private now: its only outside caller was log-shared-reservation. The conservation law returns as log_conservation_smp2. Its producer is a std::thread of test-runner's new `log-storm` builtin, calling a new test-only SYS_DEBUG action, LOG_PATTERNED (21, under test-actuators), 1024 times; each call emits one `logstorm t=0 i=<n>` record through log::storm::emit_patterned, which the reader regenerates byte for byte. `concurrent` counts storm records taken by a read after which the producer's own counter was still below 1024, not batch order; the loop ends once that counter is 1024 and QUIET_READS reads were empty, so STORM_SETTLE and the logstorm start/done parsers go. --smp 2 rather than 1: a Ring 3 producer on the reader's only CPU interleaves only at 10 ms quantum ends, and one whose 1024 calls fit in one quantum would make every run vacuous. Deleted and not restored: the log-storm and log-shared-reservation actuators (the second was read by no test), storm::start_once/body, watch::park_forever, the boot-actuators kthread row budget, the Producer.shards/migrated= ledger nothing asserted, the host's a/b field parse that only migrated= used, and close_probe's always-empty params. test-durations drops log_conservation_smp1's number, measured for a different producer; the renamed test is unpriced until measured. issues/diagnostics/a-shards-timestamps-run-backwards-at-seq-517.md is deleted: its only evidence is the line `[log] unbracketed: log-gate: FAILED: cpu6 seq 517 ...`, which is log_reserve_window_negative's own eprintln on a pass — the designed refusal. 517 is 5 + 512: the burst is 512 records reserved ahead of the interrupted one on a shard holding four boot records, and one more boot record moved it to 518, which is the "fixed position" the issue read as a defect. The sched_stress red it names shared a run with that line and nothing more. issues/kernel/deleting-logstorm-and-lognest-left-two-log-claims-unverified.md is deleted: both claims are checked again. The K3 stage is deleted from the-kernel-still-creates-threads.md: the logstorm and lognest threads are gone and the checks moved to syscall-context producers. K6 is blocked on K2, K4 and K5. Eight test manifests justified test-runner's logread by the log gate, which none of those boots runs; the comments go and the grants are filed as issues/isolation/test-runner-holds-logread-where-no-log-builtin-runs.md. The echo-spawn and negative-control-timeout issues' exits name the restored sites. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts, every hunk of each side accounted for:
- tests/doomcase/system.toml: main deletes the case; this branch only
removed its logread comment, so the deletion stands, and the logread
issue no longer lists it (seven manifests, not eight).
- tests/doommusiccase/system.toml: main's shortened namespace comment,
without the logread comment this branch removed.
- tests/common/logread.rs: this branch's STORM_GATE beside main's
one-line CEILING comment.
- userland/test-runner/src/log_gate.rs: main deleted the guest's 30 s
CEILING ("no QEMU test measures time"); it goes here too, with its
elapsed check and the Instant/Duration imports. The host's 60 s
ceiling is what reds a gate that never finishes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lands inside it on every run At 8910046 the gate's non-vacuity check ("concurrent", storm records taken by a read while the producer had not finished) held only when the scheduler happened to interleave the reader and the producer. One TCG run at --smp 2 had the producer finish all 1024 calls before the reader reached a storm record, and the gate refused: a timing verdict, which main's #562 rules out of QEMU tests. The producer now stops after HANDOVER (64) records and waits on a channel until the reader has taken a storm record, then emits the rest. The reader sends only after it has loaded the producer's counter for that read, so that read counts as concurrent on every run: 64 is below the target, and the producer cannot move until the send. The wait is bounded by HANDOVER_WAIT (30 s, inside the host's 60 s) and fails loudly, naming the reader that never arrived. Controls, deterministic both: - join the producer before the first read: the producer times out at the handover and the gate reports it; - the same with the handover deleted: the producer finishes before any read and the concurrent check refuses, as at 8910046. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review, round 3, at daca039Gate: Net: +251 / -570. Production ( No kernel thread: Earlier blockers
BLOCKER
NOTE
REMOVE
SEND BACK |
…st clock decides a verdict Round 3's review of #586 found three holes in the storm gate. - The producer's 30 s `recv_timeout` at the handover was an in-guest deadline deciding a verdict, which #562 rules out. It is `recv` now: only a disconnect is an error, and the host's `CEILING` reds a reader that never arrives. - `concurrent` was at least one batch by construction: the handover read counted even while the producer was parked. A read now counts only if the producer's counter moved across it, and only after the lap. The producer emits until such a read has happened and the reader sets `stop`, so the overlap is waited on rather than sampled. The guest's "raced nothing" refusal could no longer fire and is deleted; the host still refuses `concurrent=0`. - `lost` was zero on two of three runs, so read.rs's `lost +=` was measured by nothing. After the handover the reader now blocks until the producer has emitted `shards * 512 + 1` more records, which puts more than a shard's worth into one shard whichever CPUs the producer ran on, and the host asserts `lost > 0`. `STORM_RECORDS` goes: the emitted count is the producer's counter. REMOVEs: the `LOG_PATTERNED` arm's comment, the "rather than on a kernel thread" clause in `log::user::read`, the unchecked "runs beside it" and "a read lands inside the storm" claims, `STORM_SETTLE` in the timing-verdicts issue, and the narration in the logread-grants issue. The echo-spawn and negative-control-timeout issues are back to main's text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review, round 4, at 32d1792Gate: Net: +260 / -577. Production ( Earlier blockers
BLOCKERNone. NOTE
REMOVE
LAND AFTER NAMED CHANGES |
One conflict, kernel/src/arch/aarch64/boot.rs, both hunks kept: main's (#583) new `report_counter_origin`, empty on AArch64 because no register says where the generic timer counts from, which main.rs's `report_power_on` calls on both architectures; and this branch's `timer()`, its doc and its body (the EL1 virtual timer, logged, stopped until the scheduler arms it) in place of main's `owed!` stub. #583's KernelArgs layout word needs nothing on the AArch64 side: the loader writes it in the portable bootloader/src/main.rs, the kernel refuses a foreign one in the portable `kernel_main`, which AArch64's `_start` reaches, and that `_start` reads its four fields by `offset_of!`, so the layout moves under it by construction. Auto-merged, each checked against the branch's own hunk: actuator.rs (main's layout actuator beside this branch's irq-storm and timer-floor), main.rs (main's layout refusal and power-on report beside this branch's `mod hw` and headless GOP), x86_64/boot.rs (main's UTC `clock` and `report_counter_origin` beside this branch's irq-storm/timer-floor refusal), aarch64/mod.rs (#586 drops log-shared-reservation's window from `percpu_fetch_add`; this branch's percpu.rs still calls `log::nested::reserve_window`, which main keeps), clock.rs, sched/kthread.rs, src/build.rs, src/sourcegate.rs, tests/toyos.rs. kernel/src/hw.rs is this branch's alone: main touched neither it nor either architecture's hw.rs. The rust gitlink takes main's 9c3eea44. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings #583, #586, #593, #597, #555, #600, #610 and #611. Every conflicted hunk, and where it went: - rust: 1471893e39c, which merges main's pin 9c3eea441d8 into this branch's 62fa74d7a50 with no conflict (main's three bootstrap commits and this branch's six std files do not meet). Pushed to ToyOSOrg/rust wt-toyos-fsd. - kernel/src/actuator.rs: #583 deletes `rtc-zone-east`, and that deletion stands. This branch's doc for `leak-rollback-selftest` stands, since the FAT reopen control went with the kernel's FAT adapter. - kernel/src/fat32_adapter.rs (modify/delete): the deletion stands. #583's hunk made FAT stamp UTC (`clock::utc_secs`) with the refusal reason at the site. FAT stamping is fsd's on this branch, so it is carried there: `local_secs`, which recovered a zone through `toyos_wallclock::resolve` (deleted by #583) and cited logd's `wall.rs` (deleted by #583), becomes `utc_secs`, `clock_epoch()` alone, with main's reason. fsd no longer depends on toyos-wallclock; userland/Cargo.lock drops the edge. - kernel/src/sched/kthread.rs: #586 deleted `lognest` and `log-storm`, and this branch deleted `iod`, so klogd is the one kernel thread in every build: MAX_KERNEL_TASKS = 1, with no feature split. - issues/kernel/the-kernel-still-creates-threads.md: #586 met K3 and deleted it; this branch meets K5 and deletes it. K6 is blocked on K2 and K4. - userland/logd/src/main.rs: #583's `boot_secs` rename, beside this branch's `Published::new()`, which takes no argument here. The `owed` and `retrying_since` fields stay deleted (this branch). - userland/logd/src/serve.rs: this branch's `serving` flag, with #583's `boot_secs`. - tests/test-durations: #586 deleted `log_conservation_smp1` and this branch deleted `log_backing_read_error`; neither row stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs
Track
issues/kernel/the-kernel-still-creates-threads.md, stage K3: thelogstormandlognestkernel threads are deleted, and the four log checksthey drove now run on producers that are not kernel threads.
What changed and why
The nesting gate runs inline in
SYS_LOG_READ.log::nested::start_onceno longer calls
kthread::spawn("lognest", …): it runs the body in the firstread's own syscall, between
arch::cpu::enable_interrupts()anddisable_interrupts(). That is the context the kthread had: Ring 0, IF=1, notpreempted, because the Ring 0 timer only sets
need_resched.trap_dispatch's page-fault arm anddeadline.rsalready run in it, andlog::user::readholds no lock at the hook. Everything the gate needs isrestored unchanged:
log_nestIDT gate andLOG_NEST_VECTOR;Vector::LogNest, soHdaandVirtioSoundkeep 2 and 3 and novector moves;
IrqGuard::unclosed, both injection points (shard.rsandpercpu::reserve_log_slot) and thekernel-loomshims;log-nested-emit,log-nested-reserveandlog-unbracketed-reserve;log_nested_emit,log_reserve_windowandlog_reserve_window_negative.nested::injectis private now, because its one outside caller waslog-shared-reservation.The conservation law comes back as
log_conservation_smp2, with a userlandproducer.
SYS_DEBUGaction,LOG_PATTERNED(21), compiled only undertest-actuators, so no shipping kernel carries it. It callslog::storm::emit_patterned(0, arg).log-stormbuiltin runslog_gatewith onestd::threadproducer, which makes that call once per record andbumps an
AtomicU64after each call.decides any of them:
HANDOVER(64) records, then blocks inrecvuntil the readerhas taken a storm record. Only a disconnect is an error. A reader that
never arrives hangs, and the host's 60 s
CEILINGreds it, as No QEMU test measures time, and audio is judged on metal only #562rules.
shards × 512 + 1and then blocks until the producerhas emitted that many more records. That puts more than a shard's worth
into at least one shard, whichever CPUs the producer ran on. So the
ring laps this reader's cursor on every run, and
lostis never zero.stop. The reader sets itafter the first read that takes storm records while the producer's
counter moved across it (the counter is loaded before and after
SYS_LOG_READ).concurrentcounts only such reads after the lap.emittedis the producer's counter, since the count now depends on theshard count and on when
stoplands.QUIET_READSreads came backempty.
STORM_SETTLE,STORM_RECORDSand thelogstorm start/doneparsers are deleted.
returns after a concurrent read. The host still refuses
concurrent=0,read=0and nowlost=0.kernel_features: TEST_KERNEL, since theproducer needs
SYS_DEBUGand arms nothing.LOG_PATTERNEDas a test-onlySYS_DEBUGaction. The owner delegated kernel-interface decisions to the
orchestrator.
Why
--smp 2: two CPUs keep two shards in the merge and the ledger, andgive the producer a CPU the reader is not on. The rename is carried through
tests/toyos.rsandtests/common/logread.rs.tests/test-durationsdropsthe old name's number, which was measured for a different producer.
Deleted and not restored:
log-stormandlog-shared-reservationactuators (no test ever readthe second);
storm::start_once/body/STARTED;watch::park_forever;boot-actuatorskthread row budget;Producer.shards/mark_shard/migrated=, which nothing asserted, and thehost's
a/bfield parse that onlymigrated=used;close_probe's always-emptyparams, now inlined.Issues.
and K5.
issues/diagnostics/a-shards-timestamps-run-backwards-at-seq-517.mdisdeleted. Its only evidence is
[log] unbracketed: log-gate: FAILED: cpu6 seq 517 …, which is the linelog_reserve_window_negativeprints when itpasses: the designed refusal. Seq 517 is 5 + 512, so the "fixed position"
the issue describes is the burst's width.
logreadby a gate theirboot never runs lose the comment. The grants stay, filed as
issues/isolation/test-runner-holds-logread-where-no-log-builtin-runs.md.STORM_SETTLE.Orchestrator's guest runs at 32d1792
Gates at
32d17926, this worktreecargo run -- --ci host: EXIT=0 ("Host: 54 step(s), all green").cargo run -- --build-only: EXIT=0.cargo run -- --build-only --boot-config tests/testcases --kernel-feature boot-actuators --kernel-feature test-actuators:EXIT=0. This compiles test-runner and the test kernel, which the plain
image build does not.
mut-join-first,mut-join-after-handover,mut-no-lostandmut-nest-no-sti: each passedgit apply --check, applied, and built thattest image with EXIT=0. Each was reverted with EXIT=0, leaving
git status --porcelainempty.High-risk checks
The log's interrupt-atomicity claim, and a new
SYS_DEBUGaction.git apply --check, applied, and built the test image, then was revertedwith
git status --porcelainleft empty.log_reserve_window_negativeis the control on theIrqGuardbracket.mut-nest-no-stideletes theenable_interrupts()line inlog/nested.rs. It must turnlog_reserve_window_negativered: with IFclear, the IPI pends until
sysret, the burst lands after the outerrecord, and the gate passes.
mut-join-firstjoins the producer before the first read. The producerparks at the handover, and
log_conservation_smp2must hang to thehost's
CEILING.mut-join-after-handoverjoins the producer right after the handoversend, so the reader takes nothing more until the storm has ended. Theproducer never sees
stop, andlog_conservation_smp2must hang to theCEILING.mut-no-lostdeletesself.lost += oldest.saturating_sub(want);fromkernel/src/log/read.rs. The lap makes the sequence numbers show a gapon every run, so
log_conservation_smp2must refuse with "conservationfailed".
IPI delivery, which the unbracketed control measures by breaking it. The
conservation law's other oracle is the text regenerated byte for byte from
t=/i=in userland, independent of the kernel's formatter.🤖 Generated with Claude Code