Skip to content

Storage: file servers for DATA, the log and the boot volume; the kernel's NVMe and FAT go - #536

Open
Japabu wants to merge 71 commits into
mainfrom
wt/toyos-fsd
Open

Japabu wants to merge 71 commits into
mainfrom
wt/toyos-fsd

Conversation

@Japabu

@Japabu Japabu commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Steps 5 to 7 of issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md under the owner's option B: /apps, /config, /home, /state, /log and /boot are served by /system/bin/fsd, one process per role (DATA, LOG, BOOT), and the kernel keeps ROOT in memory and /tmp. The kernel's NVMe driver, its read-write bcachefs adapter, its FAT32 adapter, every DATA, /boot and /log mount, and its write-back queue, iod and durability ledger are deleted.

The owner accepted the gap this leaves: C programs reach no served path until libc is a client of the SDK's file-server crate (issues/filesystem/c-programs-name-no-file-a-file-server-holds.md states the order).

Decisions

  • init never waits on a file server from its loop, which is also where a server that ended is started again. Its calls into the file servers run on one worker thread, and init waits on the call and on a service's end together. A launch's files are read there before its spawn: the package's resolution, and std's CommandExt::prepare, which finds the program, judges the working directory and reads a program on a file server into a toyos::process::Image. The spawn on the loop then calls no file server. A relative working directory is refused before resolution.
  • std connects to a file server outside its capability table's lock. The connect waits for the server's hello. Of two threads that connect at once, the first answer kept is the one both use.
  • A program on a file server is spawned from a memory object (SpawnArgs::image), which the kernel pages it from, and only an object of ordinary memory the kernel allocated: a device aperture is refused InvalidArgument before anything else of the spawn is read. Its libraries come from /system/lib alone (issues/filesystem/a-package-cannot-ship-its-own-libraries.md). dlopen from a memory object is not in the ABI, since nothing calls it until libc is a file-server client.
  • New in the ABI: toyos_abi::inventory::Record::Loaded (kind 6), the partition the loader named for a role (Role::Root, Boot or Log) by its unique GUID, and blockd's --running <guid>, the running ROOT it serves no session on, which init passes from that record.
  • The kernel's inventory is read whole or refused, by one reader (toyos::syscap::SysCap::records, used by init and inspect). A refused count, a refused read and a record that does not decode are each the answer, never a shorter list. A machine that grew between the count and the read is asked again, at most four rounds, and one that grew on every round is refused by name. init's slot grant and a storage row's start read through it, so a storage row whose inventory is refused does not start, and says why; blockd is never started without --running for want of it.
  • A write that would take fsd's block cache past DIRTY_LIMIT dirty blocks flushes first, and is refused whole when that flush is, so a disk that refuses leaves the cache no larger.
  • DATA keeps every write it accepts. A file's extents live in its one btree entry, at most 4064 bytes: 250 runs for home/a. On a volume whose free blocks are all apart, a file is refused at 250 blocks while blocks are free (issues/filesystem/a-file-on-data-is-refused-at-250-runs-while-blocks-are-free.md).
    • A file extends in place when the block after its last is free, and when another file took it, looks 256 blocks further on (SPREAD), so files growing in turn extend runs rather than alternate blocks.
    • A leaf splits into as many nodes as its entries fill in order, so a large entry landing between small ones takes a node of its own. A split that finds no block for a later sibling gives back the ones it took; a sibling whose write fails leaks its block, since nothing names it and the tree is left as it was.
    • A write whose runs the entry could not name is refused whole, ResourceExhausted, said by name, and its allocations go back.
    • sync writes every file it can and names the ones it could not, so an fsync fails only for its own file. A close whose entry would not write answers the error and keeps the node, and each later sync names it again until it is written. The write-back (fsd::writeback::WriteBack) stays due after a sync that left a file unwritten. FAT's volume follows the same contract: a file that will not level is named, its close refused and its node kept, and each later sync names it again until it levels.
  • blockd's partitions live inside Drive::Up, so nothing lists one without a controller. Service::listing and Service::place are the handshake's and the open's decisions: Drive::Absent lists nothing and opens NotFound, Drive::Unusable refuses both Unusable, Drive::ClaimRefused refuses both ClaimRefused, and fsd then says DATA is absent this boot rather than serving it from memory.
  • Nothing the kernel flushes reaches a device. /tmp's pages are the file and ROOT is read-only, so the write-back queue, iod, durability.rs and its kernel-loom model are deleted. So are the file cache's dirt, pin, shrink mark and flush plan, the VFS's flush, sync and close surface, writeback-stall, the stop's sync, block::OpenUpdate and the scheduler's MID_UPDATE bit.
    • A handle's drop releases its reference under the file cache's lock alone.
    • A governed page is never dirty, so eviction always finds one.
    • SYS_FSYNC answers 0 on a kernel file; on a partition claim it is unchanged.
    • A tmpfs file's mtime lives in the file cache, set by create, write and truncate.
    • The stop's own record (toyos_quiesce::Record) is what a reader of the stop orders against.
  • No kernel thread replaces iod. With main's usbd, lognest and log-storm gone too, klogd is the kernel's one thread in every build (MAX_KERNEL_TASKS = 1), and issues/kernel/the-kernel-still-creates-threads.md loses K5. sched-operation-nesting's task half and sysret-ss-probe need a task, and run once a boot on the first syscall, init's (syscall/dispatch.rs's task_probes): the nesting gate in that task's own deadline slot as site syscall, and the SS probe's park, which switches away and back before that syscall's sysretq. sysret_ss_reload judges the boot log alone, since init's first syscall precedes the ready marker.
  • A handle held across a file server's restart answers Gone (StaleNetworkFileHandle) and is never reopened by path. Exit: issues/filesystem/a-file-held-across-a-file-servers-restart-answers-gone.md.
  • std hands a directory's connection over in ticket order, so a thread that let go does not barge ahead of the one it woke.
  • The stop's bounds are declared once in toyos_quiesce: FILES_MS (init's file-call bound), FLUSH_MS and SYNC_MS. The kernel's staged hold is their sum plus quiesce::PARK.
  • toyos-fat32 has one refusal, the device's. BudgetExpired, RepairPending, RepairNotice and the repair episode had no producer or caller once the kernel's FAT adapter went; logd's WouldBlock flush retry goes with them.
  • A partition claim is named where it is blamed. Holder::Claim carries its GUID, and a claim's retried flush names its partition.
  • ROOT on no disk this kernel drives is said so, apart from a disk that did not answer.
  • iommu-userdev-foreign-dma stages a network function's first grant alone.
  • init names a log ring's owner at the spawn that creates it, and no writer is the owner before that. init lays each ring out owned by Pid::MAX, which no writer has, so every writer leaves CHILD_KEEP slots. It names the program once the spawn returns its pid, before the ring goes to logd. logd writes no ring.
  • logd's --stall holds the stalled program's registration until --stall-until instead of skipping its reads. So log_ring_keeps_the_owners_slots holds whatever logd has done.
  • fsd's test actuators are --end-on <path> (a write), --end-at-read <path> (the first read-only open, once a boot), --end-at-mount <role> (that role's first server, once a connection waits on it and before it accepts one, once a boot) and --let-go-at-read <path> (the first read-only open, once a boot, syncs, lets the partition go and is refused; that client's next request ends the server).
  • heap_ceiling_bounds' lowered bound is the machine's own live threads plus 16, counted at arming by the function sys_sysinfo counts its roster with, so it moves with the daemons a boot runs rather than with a constant.
  • A refused controller claim is not an absent controller. init tells a block service each device of its row the kernel refused with anything but NotFound (--claim-refused <name>, beside --running). blockd then serves Drive::ClaimRefused, whose listing and every open are refused with the new wire word ClaimRefused (6), and fsd serves DATA absent by name. NotFound is a machine without the controller, and stays Drive::Absent.
  • DATA is one partition, counted once over both sources (fsd::data::find). init claims every partition of a role the kernel's disks carry, and fsd counts those claims with blockd's TOYOS-DATA listing: none is memory, one is served, two or more are refused by name and DATA is absent. A listing blockd refuses leaves the count unknown, so DATA is absent then too, even beside a claim. A log or boot claim is by the loader's GUID, which the kernel refuses when two partitions carry it.
  • A wait on a connection is IPC, reading or writing. sys_read and sys_write charged a connection's wait to WaitClass::Pipe, and every file call is now a request on a connection and a wait for its reply. The class is chosen beside the pipe in ops::pipe_read and ops::pipe_write, whose matches name every object kind, so a new pipe-bearing kind does not compile without one. A bare pipe's wait stays Pipe, the one class watch-window holds, so the canary's window is staged as before.
  • A refused partition claim refuses that file server's start, by name, and the role is absent. A claim the kernel refuses and a log or boot role the loader named no partition for are each StartError::Partition in init's storage start. On a restart init closes the role's ports, as for any refused start; at boot it says the role did not start, closes its ports and goes on, where any other refused first start still panics init. Nothing serves the role's paths from memory, and fsd, never started with neither its claim nor its GUID, panics by name if it is.
  • The device judge reads each record from its writer. metaldevices::unmet takes logd's whole file and judges RECORDS against the kernel's records alone and BLOCKD_RECORDS (nvme, nvme-served) against blockd's own lines (bootlog::lines_of). A kernel line or another program saying blockd's words does not count, nor a program saying Boot: complete (. blockd_writes_the_lines_its_table_reads holds the table to userland/blockd/src/main.rs, where main prints both lines.
  • fsd stamps a file with UTC nanoseconds since the Unix epoch, off one wall-clock reader (fsd::volume::now_nanos), which meets A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587's contract: DATA keeps the nanosecond, FAT keeps the whole seconds of its level, and 0 is undated. The reader derives the kernel's anchor BOOT_SECS exactly, once: SYS_CLOCK_EPOCH is the anchor plus the whole seconds of the counter at the call, and the clock page is the same counter on the kernel's formula, so a call bracketed by two counter readings inside one second names it. A bracket that straddles a second is asked again, and three that straddle panic by name. FAT divides the same clock, and FAT stamps UTC. The kernel's /tmp stamps clock::mtime_now at create, write and ftruncate.

The std fork: ToyOSOrg/rust wt-toyos-fsd at c4c65e3e87a, which merges main's pin 90697f1401a into this branch's 1471893e39c with no conflict; over main's pin it is this branch's six std files alone.

The merge of origin/main (a7cd327) is 1f96997. Its message says where each hunk of the two deleted issues went. The merge of origin/main (69d1b53) is e93d8c2, with no conflict; main's rust pin had not moved, so the fork is unchanged. The merge of origin/main (ec06384) is b21e3dd. Main's rust pin had not moved, so the fork is unchanged there too. It had one modify/delete conflict: main disabled ftruncate_flush_race behind issues/build/ftruncate-flush-race-reds-intermittently-and-nothing-says-why.md and gave that issue a status, a sighting, an exit and an owner. This branch deletes that test, with the kernel flush it raced and the ftruncate-flush-stall actuator. So every one of those hunks is about a test that no longer exists: the issue stays deleted, and main's src/redlist.rs row for it goes too, because the harness refuses a row whose test nothing registers. That merge also brings #573's forget_another_compiler, which meets the exit of this branch's issues/build/a-worktrees-bootstrap-cache-outlives-a-compiler-rebuilt-at-the-same-version.md, so 069722c closes it. The merge of origin/main (ec0a91a) is 93435d5, with no conflict: tests/toyos.rs is the one file both sides change, and main's hunks there are metal_sim_hostile_clipboard and a check in metal_sim_client_death. Main's rust pin had not moved.

The merge of origin/main (807f456, #564's measured schedule) is ead11ef; its message accounts for every conflicted hunk. Where #564 deleted a test or the kernel code only that test armed (so_cache_refusals with so-cache-tiny and the so_cache_policy binary this branch had moved onto /tmp; quiesce_wakes_on_the_last_exit and quiesce_dump_holds_the_stopped with quiesce-last-exit and quiesce-dump), the deletion stands: none of this branch's edits to them served a surviving test. Where this branch deleted a test #564 retiered, this branch's deletion stands, and issues/filesystem/home-budget-refusal-retried-is-red-on-every-nightly.md stays deleted: its test is gone on both sides and the reproducer its exit names is the kernel's /home over the NVMe driver this branch deletes. Every surviving row takes #564's tier. The five rows this branch adds, fsd_restart, fsd_end_at_mount, fsd_claim_held, fsd_two_data and blockd_serves_nothing, keep the Fast tier they were added at, as #564 kept the tier of every row main added. Main's rust pin had not moved.

The merge of origin/main (0368861, which brings #549, #562, #566, #581, #582, #584, #585, #591, #594, #595, #596 and #598) is 76e45e3. Its message accounts for every conflicted hunk. Main's rust pin had not moved. Where main deleted what this branch had edited (src/heartbeat.rs, tests/doomcase, the swap-redial issue), main's deletion stands. Where this branch deleted what main edited (ftruncate_flush_race, quiesce_fsync, their issues, the root-metadata issue), this branch's deletion stands. Both sides wrote a roster decoder in tests/toyos-rust-tests/src/roster.rs. Main's is kept, since it has no deadline, and process_stats' two connection arms read it. #562's iommu_* fault boots now wait for the fatal path's reset, and QEMU has exited by then. So fault_boot takes a holding read that runs before that wait, and iommu_empty_domain reads the xHCI's DCBAAP over QMP there, while the panel holds. b607ce7 deletes the quiesce_leaves_the_volume_whole sighting from issues/build/parallel-tests-red-under-other-suites.md.

The merge of origin/main (d1d83f6, which brings #555, #583, #586, #593, #597, #600, #610 and #611) is 27b5641. Its message accounts for every conflicted hunk. rust is 1471893e39c, which merges main's new pin with no conflict. kernel/src/fat32_adapter.rs stays deleted, and #583's UTC stamp is carried to fsd (the decision above). #586 met K3 of the kernel-threads track and this branch meets K5, so both go, and K6 waits on K2 and K4.

The merge of origin/main (8ee3c51, which brings #587, #603, #614 and #615) is 5d4174f. Its message accounts for every conflicted hunk, and for the issues main brought that this tree makes false: #615's two-shared-members-assume-a-bcachefs-home-the-t14-does-not-have.md, the-last-handle-to-close-stamps-the-file-with-its-own-mtime.md and a-fat-files-mtime-reads-finer-through-its-writer-than-after-a-reopen.md are deleted.

Negative controls

Every patch the rows below name is in the PR comment of round 19 (issuecomment-5891368856). At 5337e79 each passes git apply --check, cargo run -- --build-only exits 0 under it, and git apply -R leaves the tree clean. Arms marked "run by the orchestrator" are guest runs this branch's author did not make.

change test green mutation red
a launch's files read on init's worker fsd_restart EXIT=0 at 7a65674 and 58ad95d the fix unmade: 2b2563c, actuator and test without it EXIT=1 wide and alone: DATA ends at the image's read and init never starts it again
std connects outside the capability lock fsd_end_at_mount EXIT=0 at b6bcd1e, by the orchestrator the fork's capability() connecting under the lock again EXIT=1 at b6bcd1e, by the orchestrator
DATA keeps every accepted write, whole change fsd host every_file_reads_back_as_a_plain_map_of_its_accepted_writes EXIT=0 bcachefs/src and userland/fsd/src at 9b281ce, the test added EXIT=101: home/a: 4096 bytes at 202574, of 200444 held, Err(ResourceExhausted) after the first write refused for want of blocks
files growing in turn extend their runs same EXIT=0 SPREAD = 2 EXIT=101: home/a: 4129 bytes at 1119843, refused past 250 pages
a run starts past the file's last block same EXIT=0 every run from the allocator's cursor EXIT=101: home/b: 4099 bytes at 1132859, refused
a refused write's blocks go back a_write_its_entry_could_not_name_is_refused_and_every_accepted_one_kept, and the map test EXIT=0 let _ = dropped; EXIT=101: "a refused write's blocks went back", free 166 against 167
a write the entry cannot name is refused whole a_write_its_entry_could_not_name_is_refused_and_every_accepted_one_kept EXIT=0 the fit check always true EXIT=101
sync writes every file it can an_entry_refused_costs_only_its_own_file EXIT=0 sync returns at the first refusal EXIT=101
a close that lost its writes is refused same EXIT=0 close logs, drops the node and answers OK EXIT=101
a node kept after a refused close is kept by the next failing sync same EXIT=0 sync releases an unheld node whose entry it could not write EXIT=101: "and kept again"
FAT: one file that will not level costs only itself fsd host a_file_that_will_not_level_costs_only_its_own_file EXIT=0 sync returns at the first unlevel file EXIT=101: Err(Io), not Ok([(1, Io)])
FAT: a close that did not level is refused same EXIT=0 return Err(e); deleted from close EXIT=101: "a close that left the entry behind is refused", Ok(())
FAT: a node kept after a refused close is kept by the next failing sync same EXIT=0 sync releases an unheld node it could not level (r1-fat-sync-releases.patch) EXIT=101: "and keeps its node"
a machine that grew between count and read is read grown toyos host syscap::tests::a_machine_that_grew_once_is_read_grown and a_machine_that_grows_every_round_is_refused_by_name EXIT=0 the ResourceExhausted => continue arm deleted (b2-no-retry.patch) EXIT=101, both red: Err(Read(ResourceExhausted)) against the grown list at syscap.rs:274, and against Err(Grew) at :289
a refused flush leaves the cache no larger, whole change fsd host at_the_dirty_limit_a_refused_flush_refuses_the_write_and_the_cache_grows_no_larger EXIT=0 Cache::write as at b6bcd1e (r3-cache-old-write.patch) EXIT=101: "a refused write left the cache no larger", 8195 dirty against 8192
an image's libraries come from /system/lib alone abuse_elf_loader (Fast) run by the orchestrator an image's libraries from its argv's directory (red-arm-image-libs-beside.patch) EXIT=1 at bfb1763 (536m-red-abuse.log:32–34), at fae0a5d (536r12-red-arm-image-libs-beside.log:39) and at 06c6195 (536r14-abuse.log:39), by the orchestrator: export_past_image as an image, Err(InvalidArgument) against Err(NotFound). The first assert ends the process, so tpoff_overflow_spawn's image arm under the mutation is unmeasured
the lowered sysinfo bound follows the machine's own threads heap_ceiling_bounds (nightly) EXIT=0 at fae0a5d and 06c6195, by the orchestrator the comparison made false && … (b1-control.patch) EXIT=1 at fae0a5d and 06c6195, by the orchestrator (536r12-b1-control.log:39, 536r14-heap-b1.log:39): "64 extra threads and sysinfo never refused"
a refused partition claim refuses the restart, and no server answers the role fsd_claim_held (Fast) EXIT=0 at 5235eda and 06c6195, by the orchestrator the claim's refusal said and the start carried on (claim-say-and-continue.patch) EXIT=1 at 5235eda and 06c6195, by the orchestrator (536r13-b3.log:47, 536r14-claim-say.log:46): the guest's second open answered NotFound, "not Gone", the restarted server having found no DATA and serving memory
a refused NVMe claim serves DATA absent by name iommu_virtio_platform (nightly), its no-unit arm EXIT=0 at 06c6195, by the orchestrator blockd's refused-claim path serves Drive::Absent (b1-refused-claim-absent.patch) EXIT=1 at 06c6195, by the orchestrator (536r14-b1.log:34): fsd's (Refused(ClaimRefused)); DATA is absent this boot never reached the boot console
two DATA partitions, one per source, are refused by name fsd_two_data (Fast) EXIT=0 at 06c6195, by the orchestrator fsd takes the kernel's one claim without counting blockd's (b2-kernel-claim-uncounted.patch) EXIT=1 at 06c6195, by the orchestrator (536r14-b2.log:33): this machine has 2 DATA partitions, 1 on the kernel's disks and 1 the block service serves never reached the console
a connection's wait is IPC, whole change process_stats (Fast), a_wait_on_a_connection_is_ipc EXIT=0 at 069722c and 0baa8b3, by the orchestrator ipc-class-reverted EXIT=1 at 069722c, by the orchestrator (536r15-ipc-reverted.log:45): "charged 0 ns to ipc and 2345422 ns to pipe"; EXIT=1 at 0baa8b3, by the orchestrator
a wait to write a full connection is IPC process_stats (Fast), a_wait_to_write_a_full_connection_is_ipc EXIT=0 at 0baa8b3, by the orchestrator write-class-pipe EXIT=1 at 0baa8b3, by the orchestrator: "a child that parked writing a full connection charged 0 ns to ipc…"
DATA counted over both sources fsd host data_is_one_partition_counted_over_both_sources EXIT=0 one claim beside any listing is Claimed (h1-find-guesses.patch) EXIT=101: "1 claims and 1 served located Claimed"
the stop syncs every writable file server home_overwrite_reads_back (Nightly) EXIT=0 at 06c6195 and 069722c, by the orchestrator the stop's self.sync_files() deleted (b3-no-stop-sync) at 06c6195 EXIT=1, by the orchestrator (536r14-b3.log): "the NVMe image does not mount on the host: BadMagic { … got: [0, 0, 0, 0] }". That red says nothing fsd wrote that boot reached the device, the volume's format included. An absent format would red the same way, so it did not name the lost file. The guest now fsyncs /home/overwrite-looped.bin before the pinned overwrite, so the format is on the device and only the stop's sync carries the pinned file there. At 069722c EXIT=1, by the orchestrator (536r15-b3.log:35): "reading home/overwrite-pinned.bin off the image: NotFound"
blockd: an unusable or unclaimed controller lists nothing and opens nothing blockd host an_unusable_or_unclaimed_controller_is_refused_and_an_absent_one_lists_nothing EXIT=0 listing's Unusable arm folded into Absent's; place's answering NotFound; listing's ClaimRefused arm answering Ok([]) (h2-unclaimed-lists-nothing.patch) EXIT=101 each: Ok([]) and Err(NotFound) against Err(Unusable); Ok([]) against Err(ClaimRefused)
fsd says DATA is absent behind an unusable controller nvme_wide_sector EXIT=0 at b6bcd1e, by the orchestrator listing's Unusable arm folded into Absent's EXIT=1 at b6bcd1e, by the orchestrator
the write-back stays due after an unwritten file fsd host a_sync_that_left_a_file_unwritten_is_due_again EXIT=0 synced forgets the write-back EXIT=101: None against Some(2s)
a split short of a sibling gives back its blocks bcachefs a_split_short_of_a_sibling_gives_back_the_ones_it_took EXIT=0 the give-back loop deleted EXIT=101: "the first sibling's block went back", 0 against 1
a leaf splits into as many nodes as it fills a_refused_rename_over_an_open_file_keeps_its_unsynced_writes EXIT=0 the largest fitting prefix and the rest EXIT=101
one blockd loop and handshake blockd_serves_nothing EXIT=0 at 1bd2ef4, and at b6bcd1e by the orchestrator serve_nothing restored EXIT=1 wide and alone: a listing that carries a payload was answered 5 None in 0 bytes, not 3 Some(Malformed)
a device aperture is no program blockd_dma_outside_the_lent (Weekly since #564) EXIT=0 at 1bd2ef4, and at b6bcd1e by the orchestrator || object.ram().is_none() deleted from SharedImage::over EXIT=1 wide and alone: a spawn from the register window was answered Err(BadAddress)
the fault blamed on netd's own slot userdev_dma_fault EXIT=0 the kernel blames slot 0 EXIT=1: owner=slot0 for a function on slot 1
the refused-once claim flush, judged at the stop log_flush_retry EXIT=0 staged_spent refuses every first attempt EXIT=1: 37 and 41 flushes of the log partition retried before the stop
a ring's owner named by init at the spawn log_ring_keeps_the_owners_slots EXIT=0 wide and alone at c86dd84: 1853 flood lines and the job's end line in /log the naming moved back to logd's open, and the placeholder removed EXIT=1 wide and alone: 1917 flood lines, and no ===TEST_END test_rs_log_flood exit=0===
the nesting gate's task half, on the first syscall operation_nesting EXIT=0 at e318afc Operation::begin stores the asked deadline unnarrowed when a task is current EXIT=1 wide and alone: syscall: level 3 asked for 4000000000 ns and the depth inside it recovered 4000000000 ns, against the 250000000 ns
the SS probe, on the first syscall sysret_ss_reload EXIT=0 at e318afc the mov ss in KernelHw::switch deleted EXIT=1 wide and alone: sysret-ss: NOT reloaded
a file on DATA is stamped to the nanosecond, whole change file_mtime (Fast), its /home arm owed userland/fsd/src as at 27b5641 (fsd-clock-reverted) owed: EXIT=1, "on /home a write after another is stamped N ns and the one before it N ns" or "/home stamps … whole seconds"
DATA keeps the nanosecond same owed home-whole-seconds owed: EXIT=1, as above
an undated file on DATA is undated file_mtime_undated (Nightly) owed fsd's undated case answers nanoseconds since boot (home-undated-boot-nanos) owed: EXIT=1, "/home/file-mtime-undated was written on a machine whose RTC never answered, and its mtime reads Ok(…)"
the kernel's /tmp stamps mtime_now file_mtime (Fast) owed site-114, m132-ops, m821-ops owed: EXIT=1 each, naming the create with truncation, the create of a missing file and the truncation
FAT answers nanoseconds of its whole seconds file_mtime (Fast), its /log arm owed m926-fat-seconds owed: EXIT=1, "/log/file-mtime reads back N ns after a reopen"
fsd stamps FAT in UTC wall_clock_utc (Weekly) owed the reader two hours east (utc-plus-7200) owed: EXIT=1, "this boot's FAT timestamp is 72NNs from the staged instant"
the anchor is the kernel's fsd host the_anchor_is_the_kernels, a_clock_that_always_straddles_is_refused_by_name EXIT=0 the straddle check deleted (anchor-no-retry) EXIT=101, both red: Some(2000000001) against Some(2000000000), and "test did not panic as expected"

Independent oracles:

  • a plain HashMap of what each accepted write made each file, against DATA's volume read back after a remount (every_file_reads_back_as_a_plain_map_of_its_accepted_writes): no line of it is the writer's;
  • the FAT host test is independent only in its input: its volume is built from fatgen103 by toyos-fat32-check's own fixture, but a_file_that_will_not_level_costs_only_its_own_file reads its verdict back through the driver under test (v.lstat). The independent reader of FAT this branch writes is log_flush_retry (fatfs and toyos-fat32-check over the log partition), a nightly the orchestrator runs;
  • SYS_SPAWN's order, image before argv: the same spawn from a RAM region answers BadAddress, so InvalidArgument from the BAR object is the image refused;
  • QEMU's own fault record beside the kernel's attribution (userdev_dma_fault);
  • the host's GPT reader for the log partition's GUID (usb_transport_break, log_flush_retry);
  • the recorded failure at 58ad95d (log_ring_keeps_the_owners_slots), where logd reported "1919 of its ring's 1919 records waiting". The green count is the ring's own arithmetic: 1853 = 1919 shared slots − 64 kept − test-runner's 2 lines.
  • the recorded failure at b6bcd1e (the orchestrator's Fast tier, abuse_elf_loader): the kernel's own lines, export_past_image: ELF: a defined symbol's value is outside the image on the path route and export_past_image: failed to load export_dep.so: not found on the image route, which is the answer the image arm now asserts;
  • the recorded failure at bfb1763 (heap_ceiling_bounds EXIT=1): quiescecase's stop census counts 19 userland threads (536m-fast.log:861) against 10 on Kernel: every kernel panic halts and panic recovery is deleted; delete usbd; open the no-kernel-threads track (K1) #553 r4, so a constant bound of 16 was spent before the test spawned anything;
  • the kernel's own inventory on every guest boot, read through SysCap::records by init's storage start and, in the update_* tests, by its slot grant;
  • for fsd's stamps: the instant the host stages with -rtc base=, read back off the image by the host's own bcachefs reader (file_mtime_survives_a_reboot) and by the host's FAT reader (wall_clock_utc); and the kernel's arithmetic for the anchor (the_anchor_is_the_kernels).

Gates, at 5337e79

  • cargo run -- --ci host EXIT=0, "Host: 55 step(s), all green". That covers the build system's lib tests, toyos-checks, the host workspace, the licences, clippy with warnings denied, and every userland crate's host suite, fsd and blockd included.
  • cargo run -- --build-only EXIT=0.
  • cargo test --test toyos-build -- --list EXIT=0: 252 Fast, 144 Nightly, 108 Weekly and 3 Local; 20 disabled.
  • git -C rust status is clean, and the rust gitlink is c4c65e3e87a, the commit pushed to wt-toyos-fsd.

Guest runs, by the orchestrator

At b6bcd1e: fsd_end_at_mount EXIT=0 and EXIT=1 under its fork red arm; nvme_large_device, quiesce_refuses_a_second_shutdown and blockd_serves_nothing EXIT=0; nvme_wide_sector EXIT=0 and EXIT=1 under its red arm; the nightly blockd_dma_outside_the_lent EXIT=0.

At bfb1763 (orch-runs/536m-*.log): EXIT=0 for the seven update_*, the four blockd_* nightlies and the six panic-path rows. The Fast tier EXIT=1, 395 passed and 3 failed, all three main's: lan_mdns_answer (path must be shorter than SUN_LEN, fixed by #560), and lan_swap and swap_crash_rolls_back (the redial turned away 64 times while the guest finished, which #565 disables). heap_ceiling_bounds EXIT=1, this branch's, fixed at 4aef9a0. The red arm EXIT=1, as the table says.

At fae0a5d (orch-runs/536r12-*.log): fsd_claim_held alone and the nightly heap_ceiling_bounds EXIT=0, and the three guest arms EXIT=1, as the table says. The Fast tier EXIT=1, 400 passed and 2 failed, both this branch's: fs_claim_held ran on the shared boot and exited 101, and suite_split named it driven by a machine test and also shared. 5235eda puts it on RUST_SKIP. partition_claim, boot_partition_identity and partition_claim_gives_up, which boot with a role's claim refused and so with that role absent, passed in it.

At 5235eda (orch-runs/536r13-*.log): fsd_claim_held and the nightly heap_ceiling_bounds EXIT=0, and the b3, b1 and abuse arms EXIT=1. Fast EXIT=1, 400 passed and 1 failed: blocking_read_window, main's filed flake (issues/build/blocking-read-window-reds-beside-other-guests.md). The full nightly EXIT=1, 513 passed and 2 failed: that flake, and iommu_virtio_platform, this branch's, which 06c6195 answers.

At 06c6195 (orch-runs/536r14-*.log, ab-brw-*.log): iommu_virtio_platform, fsd_two_data, home_overwrite_reads_back, fsd_claim_held and heap_ceiling_bounds EXIT=0, and every guest arm EXIT=1, as the table says. Fast EXIT=0, 401 of 401. The full nightly EXIT=1, 515 passed and 1 failed: syscall_window_nmi, "36 sprayed window arrivals against 557 in Ring 3". That red is main's:

  • under kernel/src/arch this branch changes only the VT-d unit (vtd/domain.rs, vtd/mod.rs). It does not touch the NMI gate (nmi_gate.rs), the syscall entry, the test's binary (nmi_window_spin.rs) or its judge in tests/common/faults.rs;
  • it is filed on main as issues/kernel/syscall-window-nmi-shortfalls-on-a-contended-host.md ("44 window arrivals against 572");
  • Clipboard copied once into a region the compositor made; copy-once is a type, the compositor forbids unsafe code #557's Fast (557r2-fast.log) is red on it too, with 25 in the window;
  • this branch's own Fast at the same head passed it, with 64 in the window.
    blocking_read_window was red 2 of 10 at 06c6195 against 0 of 10 on main at cd2e630, interleaved in one session (ab-brw-536-7.log:35: "only 27 of 500 round trips completed inside 3s").

At 069722c (orch-runs/536r15-*.log): process_stats and home_overwrite_reads_back EXIT=0, and their arms EXIT=1, as the table says. blocking_read_window was red 0 of 10 against 0 of 10 on main. Fast EXIT=0, 399 of 399. The full nightly EXIT=1, 512 of 514, both reds main's: syscall_window_nmi, whose storm boots without watch-window and whose arrivals nmi_gate::observe classes from the interrupted frame alone, so no wait class reaches it; and user_copy_races_munmap, the copy-meets-a-remap actuator's own bound, red on #562 too and disabled by #580. The orchestrator's heap_ceiling_bounds log keeps only the harness's PASS line, so the machine's live thread count at arming is not recorded.

At 0baa8b3 (536r16): process_stats EXIT=0, and EXIT=1 under each of write-class-pipe and ipc-class-reverted, as the table says. Fast EXIT=0, 400 of 400. The full nightly EXIT=0, 515 of 515.

At d5bbfbb (orch-runs/536m-*.log): process_stats (Fast) EXIT=0; quiesce_refuses_a_second_shutdown (nightly) EXIT=0; the eleven --weekly rows EXIT=0: home_overwrite_reads_back, blockd_dma_outside_the_lent, block_duplicate_id, late_storage_connect, log_partition_layout, log_partition_identity, root_candidate_malformed, root_named_but_absent, root_named_twice, sysret_ss_reload and blockd_lends_within_its_bound. Fast EXIT=0, 240 of 240 (536m-fast.log:937). The full nightly EXIT=0, 383 of 383 (536m-nightly.log:1102). The b3-no-stop-sync arm EXIT=1 (536m-b3.log:31): "reading home/overwrite-pinned.bin off the image: NotFound".

At 806da12: cargo test --lib -- metaldevices EXIT=0, 8 of 8. Two mutations of unmet's blockd reader, each shown to build: reading the kernel's records (the reading the T14 was judged with) EXIT=101, and reading the whole log unfiltered EXIT=101, both in each_record_has_teeth. The oracle is the recorded T14 readback (target/metal/metaldevicecase), run through both readings as a temporary test and restored: the old reading returns exactly the T14's one finding, nvme: no record carries "blockd: no NVMe controller this row names is on this machine", and the new one returns none. cargo run -- --ci host was not run at 806da12.

Owed at 5337e79, by the orchestrator:

  • file_mtime (Fast) EXIT=0, and EXIT=1 under each of fsd-clock-reverted, home-whole-seconds, site-114, m132-ops, m821-ops and m926-fat-seconds;
  • --nightly file_mtime_undated EXIT=0, and EXIT=1 under home-undated-boot-nanos;
  • --nightly file_mtime_survives_a_reboot EXIT=0 on fsd's DATA;
  • --weekly wall_clock_utc EXIT=0, and EXIT=1 under utc-plus-7200;
  • --nightly home_overwrite_reads_back EXIT=0, and EXIT=1 under b3-no-stop-sync;
  • process_stats (Fast) EXIT=0, and EXIT=1 under each of ipc-class-reverted and write-class-pipe;
  • Fast and the full nightly;
  • the T14's fs_large_file and home_backing_revoked exit 0, as at 27b5641, where Triage main's T14 metal reds: issues only, no redlist rows #615 had them red on main.

Unsure

  • iommu_empty_domain's QMP read of the xHCI's DCBAAP races the fatal path's panic-reboot-fast hold, which is 5 s on the guest's clock (kernel/src/panic_reboot.rs's FAST_BOUND). Three monitor round trips fit inside it on an idle host. If the reset comes first, QEMU is gone and the read fails loudly. Nothing waits on a guest event there.
  • tests/logstallcase has no power and no shutdown since No QEMU test measures time, and audio is judged on metal only #562, and its /log is fsd's now. soundd_log_stall's metal row reads its verdict off a /log that only fsd's write-back puts on the stick.
  • fsd's anchor assumes the counter agrees across CPUs, as the kernel's clock page does: a thread moved between its bracket's readings and the syscall reads another CPU's counter.
  • fs_turns' floor is one pass: the barging lock left the others at one, and a thread kept off-CPU for 31 of the first thread's turns reds it too (issues/filesystem/fs-turns-judges-the-scheduler-with-the-lock.md).
  • A tmpfs file's mtime moving under an in-place same-size write (file_cache::touch) is checked by no test.
  • A ROOT file's cache entry and clean pages outlive its last close until CLOCK evicts them.
  • Issue bodies that quote deleted names as their evidence (Syncing filesystems..., flush_file) are left as their authors wrote them.
  • The /boot view: every program keeps fs:/boot (read-only), because esp_files and the metal tests read it; the track wants it to be the updater's alone.
  • A ring's placeholder owner is judged by a host test at the ring's layout (a_placeholder_owned_ring_keeps_its_slots_from_every_pid_until_named, through Ring::lay_out_with_placeholder_owner, the one function init and the test call): EXIT=0, and EXIT=101 with the owner mutated to own(0).
  • The whole-change control for DATA reds on the give-back half first: at 9b281ce a write refused for want of blocks kept them, so the map test fails before any file reaches 250 runs. The extent half is reached by the SPREAD = 2 and cursor-only arms, each refused past 250 pages.
  • The stop's sync control rests on timing. Without sync_files, home_overwrite_reads_back is green only if fsd's 2 s write-back fires between the guest's last write and the stop. That write-back is due from the pinned file's first write, the first one after the guest's fsync. The test took 3 s whole at 06c6195, boot included; no test pins the window.
  • fsd_claim_held and the partition_claim and partition_claim_gives_up boots stage DATA on a stick and so give the NVMe controller a disk with no table (partclaim::tableless_nvme); the lane's blank NVMe image carries a second DATA, and two are now refused.
  • faults::refused_claim takes the other functions a machine refuses (beside), which only the no-unit arm names: 00:02.0, the NVMe controller tests/netcase's blockd row claims.
  • issues/filesystem/the-tmpfs-fallback-for-apps-and-home-has-no-judge.md still describes the kernel's open_data, which this branch deletes, and says two DATA partitions land on memory, which is now false. Left as its author wrote it.
  • fsd_claim_held orders the claim against init's restart through the server's own actuator: the server is gone from the claim before the guest asks for it, and ends only on that guest's next request. Another client's request in between is answered by the absent volume and does not end it.

Size

git diff --shortstat origin/main...HEAD at b607ce7: 275 files, +10425 −10032. By path, from git diff --numstat:

  • tests/ +2769 −3004;
  • the crate test directories (toyos-fat32/tests, kernel-loom) +28 −339;
  • src/ +35 −81;
  • issues/ +664 −673;
  • the lockfiles +24;
  • every other path +6905 −5935, in-crate #[cfg(test)] modules included. Of that, the kernel is +623 −5007 and userland +5019 −653.

The fork: +715 −189 over main's pin (git -C rust diff --shortstat 9c3eea441d8 1471893e39c).

Filed, known and still true

  • issues/filesystem/a-file-held-across-a-file-servers-restart-answers-gone.md
  • issues/filesystem/one-write-far-past-a-files-end-holds-data-s-server.md
  • issues/filesystem/a-file-server-maps-a-lent-window-at-the-size-it-expects.md
  • issues/filesystem/every-flush-of-a-served-file-syncs-its-whole-volume.md
  • issues/filesystem/create-new-on-a-kernel-path-is-not-exclusive.md
  • issues/filesystem/the-logs-server-ending-under-logds-append-is-unmeasured.md
  • issues/filesystem/a-served-directory-costs-each-process-a-two-mebibyte-window.md
  • issues/isolation/a-file-server-can-open-every-partition-blockd-serves.md
  • issues/isolation/one-program-can-take-every-file-servers-client-slot.md, held by the orchestrator, streams included
  • issues/filesystem/c-programs-name-no-file-a-file-server-holds.md
  • issues/filesystem/a-package-cannot-ship-its-own-libraries.md
  • issues/filesystem/a-file-server-killed-mid-sync-leaves-a-half-written-btree.md
  • issues/filesystem/a-link-on-a-kernel-mount-to-a-served-path-leads-nowhere.md
  • issues/filesystem/fs-turns-judges-the-scheduler-with-the-lock.md
  • issues/boot-media/an-image-on-a-disk-blockd-drives-cannot-write-its-slots.md
  • issues/kernel/two-completions-can-name-one-arrival-and-accept-parks.md
  • issues/hardware/metalprobes-usb-read-is-answered-from-fsds-cache.md
  • issues/filesystem/a-cached-read-from-a-file-server-is-measured-only-under-tcg.md
  • issues/filesystem/a-file-on-data-is-refused-at-250-runs-while-blocks-are-free.md
  • issues/filesystem/a-served-file-panics-when-asked-its-raw-fd.md
  • issues/filesystem/dir-symlink-skips-the-reconnect-every-other-path-call-makes.md
  • issues/filesystem/a-write-back-the-device-refused-waits-for-the-next-write.md
  • issues/build/needs-actuators-names-lower-sysinfo-bound-as-a-payload-action.md
  • issues/filesystem/a-foreign-data-partition-is-answered-with-memory.md: a TOYOS-DATA partition with neither our volume nor a designation stamp still serves DATA from memory
  • issues/kernel/watch-window-spins-out-a-hold-its-poster-is-queued-behind.md: a watch-window hold spins out its 50 ms when the task that would post is queued behind it on the same CPU
  • issues/filesystem/a-kernel-files-fstat-answers-its-handles-mtime.md
  • issues/filesystem/a-fat-files-mtime-through-its-handle-is-its-last-level.md
  • issues/filesystem/an-undated-file-on-fat-reads-back-as-1980.md, main's, now of fsd's FAT stamp

🤖 Generated with Claude Code

https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs

Japabu and others added 7 commits September 27, 2026 05:54
…VMe and FAT go (work in progress)

A checkpoint of the file-server branch before main is merged in: the
kernel's NVMe driver, page cache, read-write bcachefs adapter and FAT
adapter are deleted; fsd serves each role's directories as capabilities
init hands out; spawn and dlopen take an image a served file was read
into. Tests are still being moved onto the servers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
…client (work in progress)

The kernel tests whose subject was deleted go with it; the ones whose
claim survives are retargeted onto blockd, fsd and the kernel's USB
partition claims. fsd answers one request per client per wait and asks
an acceptor before it accepts: a duplicate completion parked the log's
server at boot. New: fs_escape, fsd_restart, and fsd's FAT read-error
host tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
…t-lld pin

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
… up to the tree

The nightly storage tests are retargeted onto fsd, blockd and the USB
partition claims, or deleted with their kernel subject (writeback's FAT
arms, the truncate race and its actuator). The claim path says when a
flush was answered on a retry. Every program keeps /boot in its view
until the rows declare their own. Issues whose subject left the kernel
are closed or re-pointed at fsd, and the weaknesses this change leaves
are filed: C programs, a link from a kernel mount, a half-written btree,
the slots of an NVMe-installed image, a file server's reach over blockd,
duplicate completions, the metal read span, and the cached read's cost.
The SDK crates take their minors: toyos-abi 0.17.0, toyos 0.19.0,
toyos-window 0.21.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
…e wizard its /config

Both boots had no fsd row, so /config was nobody's and the locale applet
was refused its write; locale_gate's wizard namespace now keeps
fs:/config beside the surface. The lockfiles take the SDK's new minors.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
@Japabu
Japabu marked this pull request as ready for review September 27, 2026 10:41
@Japabu

Japabu commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Review of #536 at 2aa73f1 (wt/toyos-fsd).

Gate, not met at 2aa73f1:

  • PR CI is red. Run 36313219131, job abi-split (cargo run -- --ci abi-split), exit code 1. toyos/Cargo.lock and tests/iced-counter/Cargo.lock still lock toyos 0.18.0, toyos-abi 0.16.0 and toyos-window 0.20.0. This head declares 0.19.0, 0.17.0 and 0.21.0.
  • No whole fast tier has run at this head. The last run was at 193600c: EXIT=1, 402 passed and 3 failed. The locale fix in 2aa73f1 was checked only by targeted runs.
  • lan_mdns_answer is not a declared known red. cargo run -- --known-red lan_mdns_answer in the worktree printed "lan_mdns_answer: NO, not quarantined — its failure fails the suite." (EXIT=0). Its second fast-tier failure, QEMU exiting before READY with IOMMU translation faults on 00:03/00:04, has no A/B test in the body. This branch moves IOMMU/DMA actuators, so that signature cannot be blamed on something outside the diff.

NOT READY FOR REVIEW

Japabu and others added 3 commits September 27, 2026 12:45
The ABI bump declared toyos 0.19.0, toyos-abi 0.17.0 and toyos-window
0.21.0, but toyos/Cargo.lock and tests/iced-counter/Cargo.lock still
locked 0.18.0, 0.16.0 and 0.20.0, and `abi-split` refused them (run
36313219131). Moved with `cargo update --offline -p ...` against each
lockfile's own manifest; every other tracked lockfile already held the
new versions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
`cargo test --lib` in the root had 20 reds at 334fbd8, all this
branch's, and no gate run on the branch had reached them:

- `no_diag_program_claims_the_screen`: 2aa73f1 gave the diag boot
  blockd, which claims the NVMe controller, and the diag image's whole
  guarantee is a config that declares no `devices`. The diag boot now
  runs fsd for the log role alone: the log partition is the kernel's
  partition claim on the boot stick, so logd keeps its /log and nothing
  in the image claims a device. screen_diag_boot and screen_log_absent
  EXIT=0 after it.
- `every_shipped_boot_config_is_covered`: tests/fsdrestartcase joins
  ALL_CONFIGS, and with it every per-config gate that list drives.
- the 18 `heartbeat::tests`: tests/metalcase now starts blockd and fsd,
  and DONE had no done line for either. DONE carries one line per
  process a row starts: blockd's `NVMe up`, and fsd's `<Role> serving`
  for each of its three roles. The recorded nightly captures were
  booted without either program, so the tests replay them against that
  start list's lines, and the table is held against metalcase's own
  start list in `a_program_the_done_table_does_not_know_is_refused`.
  kernel_heartbeat EXIT=0 on a metalcase boot after it.

`cargo test --lib`: 396 passed, EXIT=0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
lan_mdns_answer's two fast-tier failures are one defect of the dev host's
scratch path, measured by an A/B at 16d2e64 against this branch, two runs
per arm per $TMPDIR, one at a time: the default $TMPDIR reds on both with
`path must be shorter than SUN_LEN` (a 104-byte socket path, which QEMU
binds and the host's connect refuses); a $TMPDIR three bytes longer reds
on both with QEMU refusing its own -chardev and exiting 1 before READY,
the shape the fast tier showed on lane 11 (105 bytes); a $TMPDIR under
target/ passes on both, with no IOMMU line from either guest. The IOMMU
translation faults beside the fast-tier failure were other tests'
deliberate foreign-DMA actuators interleaved into the shared log.

The first shape is quarantined at the issue that owns it; the second
cannot be, since its text is only the harness's boot-death framing. With
the row, the default $TMPDIR run is XFAIL, EXIT=0.

The cached-read issue now carries the measurement of where a request
goes, fsd's READ arm timed from inside the server beside the kernel's
/tmp in the same guest: a 104 us round trip against 2.6 us, and at
256 KiB the two volatile word copies through the window at ~87%.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
@Japabu

Japabu commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Review of #536 at 7e2e104 (wt/toyos-fsd), round 2: a full review after the gate stop at 2aa73f1.

Gate at 7e2e104. PR CI is green: run 36315170451, abi-split SUCCESS. host was SKIPPED, which ci.yml does on every pull request. The orchestrator asked for a full review, so it proceeds past these three gaps:

  • There is no T14 reading. This head changes what the T14 boots: tests/metalcase now starts blockd and fsd, src/metaldevices.rs rewrites the internal-disk safety record to rest on NVME_UNDRIVEN, and the log volume is reached through fsd's partition claim. QEMU is not the hardware.
  • The whole fast tier has not run at this head. Nightly run 36315314096 is the orchestrator's to judge.
  • The body gives no green run, with a command and an exit code, for the two tests this branch adds: fs_escape in the shared block, and fsd_restart on the Fast tier. It gives only their negative controls.

Net lines

  • Whole branch (git diff --shortstat origin/main...HEAD): 204 files, +7102 −7293.
  • Production: +5061 −4057, net +1004. That includes about 370 lines of host tests inside fsd's own files.
  • Tests: +1592 −2582, net −990.
  • The rest: src/ net −12, issues/ net −212, lockfiles net +19.
  • Outside the tree: the std fork adds +632 −187 (100806e1).

BLOCKER

  • userland/init/src/main.rs:397,438 and 1431–1442 (resolve), install, read_binary, declared, 1612 (make_dir): init can deadlock and wedge the machine.
    • init makes blocking std file calls into DATA's fsd, with no time bound, from its only loop. That same loop is the only caller of restart_ended (line 831).
    • Case 1: DATA's fsd ends at boot before it answers make_session_home's hello.
    • Case 2: DATA's fsd ends while init is serving a launch. A launch token is handled before TOKEN_WAKE in the same pass.
    • In both cases Dir::reconnect queues a new connection on a port whose acceptor init itself holds, and receive (toyos/src/fs.rs:285) waits forever. The machine is wedged, and the power port is dead with it.
    • Test to add: in tests/fsdrestartcase, end DATA's server before init makes the session home, or launch an /apps path in the same pass as an end. The boot must reach READY and the launch must be answered. By the code both hang today.
    • Fix direction: init never waits unbounded on a server it supervises. Run the file work on a bounded thread, as sync_files already does, or restart before serving launches.
  • userland/fsd/src/main.rs:812: any program can crash DATA's server, and then close it for the whole boot.
    • The client chooses the STREAM offset. It is stored unchecked (line 696), so stream.offset += n as u64 overflows. [profile.toyos] sets overflow-checks = true, so fsd panics.
    • Four such panics in 10 s close /apps, /config, /home and /state to every process for the rest of the boot. Every program holds fs:/home, so every program can do this.
    • Test to add: a raw client, shaped like fs_escape's wire_open, opens a file for write, sends STREAM at offset u64::MAX - 1 and writes 16 bytes into the pipe. Then assert that no ended; started again line appears and that a new open succeeds. Red today.
    • Patch: refuse r.offset > MAX_FILE_BYTES at STREAM, and use checked_add in drain.
  • rust fork library/std/src/sys/fs/toyos.rs:419 (at 20436ec2): after a restart, a held handle reopens a different file and writes into it.
    • A handle held across a restart reopens by rel, with no identity check.
    • If the file was renamed away and a new one created at the path, or another file was renamed over it, the held handle now writes into the other file at its old offset. That is silent corruption.
    • This contradicts the body ("a file no longer at its path … answers Gone") and data.rs's own rule ("renamed over answers Gone").
    • Test to add: in fs_restart, hold ACROSS open for write, write a new file and rename it over ACROSS, end the server, then write through the held handle. It must be StaleNetworkFileHandle; today it succeeds into the replacement.
    • Fix direction: OPEN answers a node identity, and the reopen compares it.
  • userland/fsd/src/data.rs:613: a rename-over that fails loses the unsynced writes of the file it tried to replace.
    • self.orphan(to) runs before self.fs.rename(from, to).
    • If the format refuses the rename (NoSpace, NodeOverfull, a device error), to's holders are already marked Gone. persist skips gone nodes, so their unsynced length and extents are dropped while to still names the file.
    • This is a sibling of code that already does it right: fat.rs:359 orphans after the rename, and data.rs:584's own unlink comment states that rule.
    • Patch: move self.orphan(to) after the ?. Add a host test: a rename over an open, dirty to that the volume refuses leaves to readable at its written length.
  • toyos/src/fs.rs:154–186 and userland/fsd/src/main.rs:627–633: the 15× cached-read regression is this branch's own code, so it is fixed here, not filed.
    • toyos/src/fs.rs is a new file in this branch. The body's claim that the fix is "outside this change" is false, and the issue records a defect this branch introduces, not a compromise it found.
    • Is volatile needed? No, not for soundness. The protocol hands the window back and forth:
      • fsd owns it from receiving a request until it sends the reply.
      • The client owns it from receiving the reply until it sends its next request.
      • Honest peers never overlap, and the send and receive syscalls order the two sides.
      • Against a hostile peer, the one defence that matters is single-fetch: copy once into private memory, then act only on the copy. Both ends already do this.
      • A volatile scalar loop fetches each byte once, exactly as memcpy does. read_volatile is not atomic, so it does not make a racing peer write defined either.
    • The correct bulk copy:
      • One bounds-checked core::ptr::copy_nonoverlapping between the window's raw pointer and private memory, never forming a &[u8] or &mut [u8] over the shared bytes.
      • fsd copies the cache's blocks into the window once. It should not allocate a fresh zeroed vec![0; len] per request and copy again from there.
      • Volatile stays where something is polled or must not be elided: rings, doorbells, DMA.
    • Exit: the author's own issue exit condition, measured by the same interleaved runs.
  • tests/toyos-rust-tests/src/bin/writeback_reopen.rs:20 and writeback_spawn.rs:35: these tests can no longer fail on the claim they make.
    • They arm the kernel's writeback-stall but write /home, which fsd now serves. So they cannot fail on the kernel write-back claim they name.
    • kernel/src/writeback.rs, iod.rs and file_cache.rs are still live for /tmp, and now nothing guards them.
    • The resize-fault-refuse and resize-evict-window actuators (kernel/src/file_cache.rs:493–503, actuator.rs) are armed by nothing: their host test went with tests/common/volumes.rs.
    • writeback_durability.rs:70 still sleeps 200 ms for a kernel iod drain that no longer touches /log.
    • Either retarget these onto /tmp with each actuator's firing asserted, or delete the kernel code they guard in this PR.
  • userland/fsd/src/main.rs:52,359: past 128 connections a new client hangs forever, and 128 is a guess.
    • MAX_CLIENTS = 128 counts (process, directory) connections machine-wide, and std holds each one for the process's life.
    • Past the limit, fsd stops watching its acceptors. The next client waits in the port queue with no bound (toyos/src/fs.rs:285) and hangs silently. With four DATA directories, 33 processes can reach it.
    • Nothing measured it. Either refuse the 129th by name (accept it, answer its hello ResourceExhausted), or measure the desktop's connection count.
    • Test to add: one guest opens 129 connections, and the last is refused, not hung.

NOTE

  • Brief question 2, the isolation boundary.
    • The resolver refuses .., absolute paths on the wire, and escapes through symlinks. Handles cannot be forged across views, because fids are per connection.
    • Naming a capability a program does not hold cannot be tested today: every program holds every directory.
    • Names containing NUL are accepted, but they are no escape.
    • The one place fsd trusts a client number unchecked is the STREAM offset (blocker above).
    • Of the author's controls, the resolver clamp covers the resolver, the no-op fsync covers FSYNC, and the Gone check covers a stale generation. None of them covers identity across a restart (blocker above).
  • Brief question 3, the oracles.
    • The host "bcachefs reader" is the same in-tree crate fsd writes with. It is independent of fsd's cache, blockd and the transport, but not of the format or of data.rs's dir/ marker convention.
    • fatfs with fatgen103, and QEMU's NVMe trace, are independent.
    • Where they run: fsd_restart, home_overwrite_reads_back, apps_and_home_are_one_filesystem, partition_claim, internal_disk_boot and nvme_large_device are on the Fast tier. esp_filesystem, kernel_log_file, fat_backing_revoked, fs_rename_durable and blockd_* are Nightly.
  • Brief question 5, the rust fork pin.
    • 20436ec2 is the head of wt-toyos-fsd and has main's pin fbf6ad14 as an ancestor.
    • The delta is five toyos files. It has zero delta in alloc and core. The one cross-platform touch is a re-export inside the existing target_os = "toyos" arm of sys/fs/mod.rs.
    • One exception: os/toyos/fs.rs adopt_namespace(u32) is a safe fn that takes ownership of a raw handle. std's I/O-safety convention makes that unsafe, so it is not of upstream quality.
  • Brief question 6, the redlist row.
    • src/redlist.rs:87 is narrow: std's own SUN_LEN error text cannot match a guest failure. The A/B puts the red on main.
    • The issue says status: open, and issues/README requires expected-red for a quarantined defect.
    • The 105-byte shape still reds the dev-host tier depending on which lane the test lands on.
  • Brief question 4, actuator hooks for deleted code. kernel/src/arch/x86_64/vtd/domain.rs:167 and vtd/mod.rs STAGED/staged() put an actuator hook into IOMMU attach that is compiled into the production kernel. Put them under #[cfg(feature = "boot-actuators")].
  • kernel/src/file_backing.rs:91: up to 256 MiB of kernel pages per spawn is charged to no one. The comment says so and no issue records it, which is silent debt. Record it or bound it.
  • kernel/src/syscall/dispatch.rs:344–354: SYS_DLOPEN copies up to 256 MiB before sys_dlopen checks whether the name is already loaded. Check the name first.
  • The ABI change (SpawnArgs.image_*, SYS_DLOPEN's fourth word) cites no owner discussion. CLAUDE.md requires one. The orchestrator confirms.
  • userland/init/src/main.rs storage_endowment: BOOT's fsd receives a write-capable partition claim on the ESP.
    • main refused /boot writes in the kernel (KernelOnly); now only fsd's promise refuses them.
    • issues/isolation/a-file-server-can-open-every-partition-blockd-serves.md covers the blockd path only. Extend it to the claim path.
  • userland/fsd/src/main.rs:95–100,206: Capability.writable is always role ≠ Boot, which is also volume.writable(). One of the two is dead.
  • toyos/src/fs.rs:376–378: Dir::write repeats fid_call's generation check. Delete it.
  • userland/fsd/src/main.rs:804–829: drain loops until WouldBlock, so a writer that keeps the pipe full holds every other client off. pump's one-request fairness does not cover streams.
  • userland/fsd/src/main.rs:679: FSYNC syncs the whole volume, and the fork's File::flush is an fsync (toyos.rs:633). So every BufWriter flush on a served file syncs all of DATA, at 3.5 ms p50 as measured.
  • userland/fsd/src/main.rs:420–443: the probe before accept works around a kernel defect that is recorded in an issue, in one server. blockd, init and logd still park on accept. The fix belongs in the kernel poller.
  • userland/blockd/src/main.rs:540–548: a controller that fails Controller::open ends up with fsd logging "this machine has no DATA partition" and putting DATA in memory. That reports a failure as an absence. Answer MSG_LIST with Unusable, so fsd says Absent instead.
  • fsd_restart ends only DATA's server. Ending LOG's server while logd is appending, where the append is not retried and logd gets StaleNetworkFileHandle, is unmeasured.
  • A 2 MiB window per (process, directory) connection: the body does not measure the memory cost.
  • tests/toyos-rust-tests/src/bin/esp_files.rs:130–141: the /tmp/evil link to the loader now resolves to nothing, because the kernel has no /boot. The check passes without testing anything. Retarget it or delete it.
  • Brief question 7, cut tests that were weakened checks:
    • resize-*, writeback_reopen and writeback_spawn (blocker above).
    • esp_files' /tmp attack (note above).
    • cache_eviction: no guest test drives fsd's cache past CLEAN_LIMIT, so eviction round trips run only on the host.

REMOVE

  • The PR body's "the fix is in toyos/src/fs.rs and toyos/src/volatile.rs, outside this change": false.
  • kernel/src/file_cache.rs:498: "Mirrored in …/writeback_durability.rs".
  • tests/toyos-rust-tests/src/bin/writeback_durability.rs:1–15,42: kernel iod claims, and a citation of the deleted kernel/src/fat32_adapter.rs.
  • tests/toyos-rust-tests/src/bin/writeback_reopen.rs:12–15: "/home, which is NVMe-backed".
  • tests/toyos.rs:11553: "not the NVMe /home device".
  • kernel/src/actuator.rs:443: "NVMe's admin completion queue page". The victim is the xHCI DCBAA now.
  • kernel/src/drivers/panic_console/mod.rs:1140: "beside irq: and nvme:".
  • kernel/src/drivers/usb_storage.rs:16: "must stay clear of NVMe's range". Nothing NVMe registers any more.
  • userland/init/src/main.rs:308–310: "DATA keeps no directories, and the VFS's own record".
  • toyos-manifest/src/lib.rs:590: the device-class test's doc comment now sits above the new role test.

SEND BACK

Japabu and others added 5 commits September 27, 2026 13:50
`DataVolume::rename` marked every holder of `to` gone before asking the
format for the rename, so a rename the format refused (EntryTooLarge,
NoSpace, a device error) had already dropped `to`'s unsynced length and
extents: `persist` skips a gone node, and `to` still named the file. The
orphan now follows the format's answer, as fat.rs's rename and data.rs's
own unlink already order it.

Host test `a_refused_rename_over_an_open_file_keeps_its_unsynced_writes`:
a rename refused EntryTooLarge (240 extents fit `from`'s short name and
not `to`'s 300-byte one) leaves `to` readable and its length on the
volume after the next sync. With the orphan moved back before the rename
(a checked patch): EXIT=101, the read answers Gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
…ounds what a client names

The review of #536 at 7e2e104, blockers 1, 2, 3, 5 and 7:

- init made its file calls (the session home, a service's home, a launch's
  `/apps` package) from its one loop, which is also where a file server that
  ended is started again: a call a server's end left in the port's queue
  waited for ever. They run on one worker thread now (`Worker`), and
  `Init::files` waits on the call and on a service's end together, restarting
  while it waits; a server alive and silent costs `FILES_BOUND` (30 s, policy).
  fsd's `--end-at-hello <dir>` actuator ends the server at the first hello on
  `dir` this boot; `tests/fsdrestartcase` arms it on `/apps`, and
  `fs_restart` launches an `/apps` path first: init's resolution is that
  hello, and the launch is answered.
- `STREAM` refuses an offset past `MAX_FILE_BYTES` (`fs_stream_offset`), a
  drain ends its stream by name at that bound, and a drain takes one read per
  wake, `pump`'s fairness. Every other client number fsd keeps was already
  bounded where it is kept.
- A handle held across a restart reopens by path and keeps the file only when
  the server answers the identity it last saw (`Stat::ident`, a hash of what
  the entry records: DATA's first block, length and mtime; FAT's short name,
  creation stamp, first cluster and length). `fs_restart` renames a file over
  a held one, ends the server, and the held handle's write is Gone; the host
  reads the replacement's bytes back untouched.
- The window: one bounds-checked `copy_nonoverlapping` each way, never a
  reference over it, with who owns it when at the site; fsd's `READ` has the
  cache copy each block into the window once (`Out`, `Cache::visit`), and a
  `WRITE` lands in a kept scratch.
- A client past `MAX_SERVED` is answered `ResourceExhausted` at its hello and
  let go by name; acceptors wait only on `MAX_HANDSHAKES`, each answered or let
  go within `HANDSHAKE_TIMEOUT` (`fs_client_bound`).

Also: `Capability.writable` and `Dir::write`'s second generation check are
gone; a blockd whose controller would not open answers `Unusable`, and fsd
says DATA is absent rather than putting it in memory; `fs_cache_eviction`
reads a file longer than the cache keeps back off the disk.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
…es the kernel copies

Owner-approved revision of the image ABI this branch added. `SpawnArgs`
carries `image` (a handle to a shared memory object carrying `MAP`) and
`image_len` in place of a pointer and a length, and `SYS_DLOPEN`'s fourth
word points at an `ImageRef { handle, len }`. The caller reads the program
into an object it owns and is charged for; the kernel pages the child from
that object (`SharedImage`), copying each page once as it is read, and
refuses an object that is no memory it allocated or a length the object does
not hold. `SYS_DLOPEN` answers a name the process already holds before it
looks at the object. `ImageBacking`, which copied up to 256 MiB into kernel
pages charged to nobody, and `MAX_IMAGE_BYTES` go.

`spawn_image_object`: a program in an object runs; an object without `MAP`
is refused `PermissionDenied` and a length past it `InvalidArgument`; an
object overwritten with `hlt` right after the spawn ends its child and the
next spawn runs; and a held name's `dlopen` is answered through a handle it
could not read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Blocker 6 of the review. The kernel's write-back queue is still the teardown
path of every closed kernel file, and `/tmp` is the one writable directory
the kernel serves; it cannot move to a file server while C programs name no
served file (`issues/filesystem/c-programs-name-no-file-a-file-server-holds.md`).
So the tests guard it there:

- `writeback_reopen` and `writeback_spawn` stage `/tmp` files with `iod`
  parked, and the harness holds `iod`'s one line saying `writeback-stall`
  parked it.
- `resize-fault-refuse` and `resize-evict-window` are deleted with their
  hooks: the device read a shrink made has no device under it any more.
- `writeback_durability`, which nothing ran, is deleted with its sleeps.
- `esp_files` stages the loader attack as a link on `/home`, absolute and
  climbing, where it resolves to something: both are refused for writing.

The IOMMU actuators' staging (`STAGED`, `staged`, the xHCI class) is compiled
only with `boot-actuators`. The review's REMOVE lines are deleted: NVMe in
the foreign-DMA actuator's doc, the panel census and the USB id base, and a
test doc left above the wrong test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
- a held file is known across a restart by what its entry records, which a
  same-length replacement in the freed first block within the second passes;
- a file server maps a lent window at the size it expects;
- every flush of a served file syncs its whole volume;
- one write far past a file's end holds DATA's server;
- the log's server ending under logd's append is unmeasured;
- a served directory costs each process a 2 MiB window, unmeasured;
- the boot volume's server is endowed a claim that writes (extends the
  isolation issue on blockd sessions).

`lan_mdns_answer`'s issue is `expected-red`: `src/redlist.rs` quarantines it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
Japabu and others added 4 commits September 27, 2026 15:18
…the deadlock was

`--end-at-hello` ended the server at init's first hello on `/apps`, whose
connect then failed and was answered: it never reached the retry that
reconnects and waits in the port's queue, so init resolving from its loop
passed it too. `--end-at-request` ends the server at the first request after
a hello, on a connection init holds; its retry reconnects into the queue.
With init's resolution moved back onto its loop (a checked patch),
`fsd_restart` wedges the machine: fsd's end is the last line, and init
starts nothing again (EXIT=1, 322 s, twice). With the worker: EXIT=0.

The mark that makes the actuator once a boot is looked for and then made:
the std fork's `create_new` on a kernel path is not exclusive, which two
servers took as two first times; filed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014iqcj4jDKpaiDX8B7CMvmK
…, and what load showed, filed

`quiesce_stops_the_machine` counts every userland thread the stop names;
init's file worker is one more.

The cached read, measured by the same interleaved guest runs as before (an
uncommitted bench and a probe in fsd's `READ`, applied and restored), one
guest, 256 KiB requests on a 16 MiB file in DATA's cache against the kernel's
`/tmp`:

  before (the volatile loops, a zeroed buffer per request):
    fsd 1952-1975 us a request (127-128 MiB/s), 945-954 us of it in the server
  after:
    fsd 988-997 us (251-253 MiB/s), 107 us in the server
  the kernel: 99.5-99.8 us

2 MiB requests went from 367-402 MiB/s to 1159-1392 MiB/s. What is left at
256 KiB is the client's one copy out of the window: a 256 KiB
`copy_nonoverlapping` out of a shared mapping costs 2.86 ns a byte under TCG
in the same guest, against 0.23 between two heap buffers. The issue that
recorded the copies as a compromise goes; one asking for the measurement on
metal takes its place.

Filed from load-coincident reds on a loaded dev host, each green alone:
threads of one process starving on one directory's connection
(`quiesce_stops_the_machine`), the stop's flush and syncs outlasting
`quiesce-last`'s 10 s hold (`quiesce_wakes_on_the_last_exit`), and
`swap_crash_rolls_back`'s log stream redial spending its ceiling inside the
swap's probation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The four guest binaries this round adds to the shared block
(`fs_stream_offset`, `fs_client_bound`, `fs_cache_eviction`,
`spawn_image_object`) cut its T14 list in three, and the metal profile refuses
a list nobody priced. `shared-3`'s rows carry `shared-2`'s ceilings, from the
same constants, and no `measured`: no run has taken that boot yet. With them,
`--metal --metal-readback` stages all 25 images (EXIT=2, staged and not
judged, the machine untouched).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 2, at e00e669

Not ready. There is no CI result for this head. gh pr view 536 shows an empty statusCheckRollup and mergeable: CONFLICTING against origin/main a637f5c (#535). The last ci run on the branch is 36315170451, at 7e2e104, which was round 1's head. Everything this round changed (7e2e104..e00e669: 62 files, +1736 −945, including the spawn/dlopen ABI and init's Worker) has no CI conclusion. The PR body also says two other things are missing: the fast tier has not run at this head, and there is no T14 reading, although metalcase now boots blockd and fsd.

Because of that, the round-1 BLOCKERs B1–B7, the ABI's completeness and the size are not judged. The mutation EXITs in the PR body are local runs on a head that cannot land as it stands: the merge below rewrites the same kernel boot path and the same tests.

What has to happen before the next round

1. Merge origin/main and resolve these conflicts. I ran git merge-tree --write-tree origin/main HEAD and it exited 1. Account for every hunk on #535's side:

  • kernel/src/page_cache.rs: this branch deletes it and Main's nightly: a vanished disk is refused at ROOT's hold instead of panicking, and the reds #506 and #527 left #535 modified it. Main's nightly: a vanished disk is refused at ROOT's hold instead of panicking, and the reds #506 and #527 left #535 added partclaim-root-withheld to the read-fault injector, plus answer_table_reads. That actuator is the only stimulus for Main's nightly: a vanished disk is refused at ROOT's hold instead of panicking, and the reds #506 and #527 left #535's ROOT-withhold path. The path's code, rootfs.rs withhold and gpt.rs withhold, merges in with no textual conflict. So the actuator has to move onto the branch's block::unanswered. If it is simply dropped, that path is left with no test.
  • kernel/src/actuator.rs: the same actuator, sitting against this branch's deletion of pc_unbind_selftest and partclaim_table_unanswered.
  • kernel/src/main.rs: main arms and disarms partclaim-root-withheld around rootfs::hold_source after fat32_adapter::probe_boot_disks. The branch has gpt::probe_usb_disks followed by block::unanswered::refuse. Separately, main's move of arch::boot::platform_devices into the device phase (dbf4ace) merges automatically and still has to be checked against the branch's storage phase.
  • tests/common/partclaim.rs: main's root_withheld arm conflicts with the branch's crafted-USB-stick arm. Both have to survive.
  • tests/common/power.rs: in quiesce_stops_the_machine, main has OTHERS = 4, the console-queue-at-the-stop actuator and the judge for "queued line above the last word". The branch has OTHERS = 13 and neither of the other two. The merge needs the branch's count together with main's actuator and judge.
  • tests/common/storage.rs: main extended home_budget_refusal_retried with bf4fb06's check that each file is refused once, not on every flush. The branch deletes that test along with the kernel's NVMe fsync path. Main's nightly: a vanished disk is refused at ROOT's hold instead of panicking, and the reds #506 and #527 left #535 made a claim about logd's flush storm (issues/filesystem/logd-flushes-the-records-its-own-refused-flush-made.md), and /log's flush now goes through fsd. The merge has to say where that claim is now tested, or delete it and give the reason.
  • Four files merge without a textual conflict but are semantically coupled: kernel/src/gpt.rs (branch −263 lines, main +94 for withhold), kernel/src/rootfs.rs (main's withhold(guid, why, &sought.silent) against the branch's disk inventory), kernel/src/object/ops.rs (refusal once per file and per partition and kind) and kernel/src/syscall/machine.rs (drain_for_the_stop before Syncing filesystems..., set against init's new file-server syncs). Only a build and CI of the merge can say whether these work.

2. Settle each of the three wide reds against main, with an A/B under the same load. Running them alone and getting green is a re-run; it does not adjudicate them.

  • swap_crash_rolls_back is main's. main's nightly at 1ce7183 (run 36290616312) shows the same signature, "the stream's redial was turned away 64 time(s), its ceiling of 64", and main records it in issues/build/swap-crash-rolls-back-redial-turned-away-once-on-mains-nightly.md, where --known-red answered NO. Disable it in src/redlist.rs against main's issue before this lands. The branch's issues/build/swap-crash-rolls-back-gives-up-its-log-stream-inside-the-probation.md duplicates that issue: move its probation analysis into main's issue and delete it.
  • quiesce_wakes_on_the_last_exit is this branch's. The mechanism the branch itself filed is the sync_files stop step that userland/init/src/main.rs adds. Together with logd's 5.016 s write to /log through fsd, it runs past the kernel's unchanged 10 s STAGED hold. main's stop has no file-server syncs. Fix it before landing: the staging bound must come from the stop's own budget, as issues/kernel/the-stops-flush-and-file-syncs-can-outlast-quiesce-lasts-hold.md states. Do not disable it.
  • quiesce_stops_the_machine is this branch's. Its failure here was "one of six writers reached its loop in 5 s", which is starvation on the std fork's per-directory Mutex<Dir>, and this branch adds that mutex. main's red on this test had a different signature (lines after the last word), and Main's nightly: a vanished disk is refused at ROOT's hold instead of panicking, and the reds #506 and #527 left #535's drain closed it. Six threads writing to /log is a regression against main. Fix it before landing (a lock that hands over, or one connection per thread), with the starvation test the issue's exit names.

3. The T14 images do not reflect what will land. The 25 images were staged after 06a6077. Production code last changed at 28094fd, so they may match e00e669's code. But the PR names no directory (<dir>), and I found no head stamp tying them to e00e669. More importantly, the landing head is the merge with a637f5c, which changes main.rs, rootfs.rs, gpt.rs, ops.rs and machine.rs, all on the T14's boot path. Restage from the merged head, then run the T14. Running it on the current images would spend a hardware run on a head that will not land.

B3's identity, decided now so it does not cost a round

The design doc is right: reopen answers Gone until the format carries an object id and a generation, and Stat::ident's hash is deleted. Hashing the first block, length and mtime is not an identity. The collision it admits is the allocator's normal behaviour: its own issue says the allocator "hands back the most recently freed first". So deleting a file and writing a same-length file within the second puts the held handle's writes into another file. That is wrong behaviour with data loss, and filing it does not make it legal. Gone is the fail-fast answer, and it removes the ident field from the protocol, fsd's two ident implementations and std's comparison.

The ABI's syscall numbers

This branch neither adds nor retires a syscall number. SYS_DLOPEN's fourth word changes meaning in place (it was 0, now it points to an ImageRef), and SpawnArgs gets bigger, under toyos-abi 0.17.0. No retired number is involved. Completeness across the std fork, libc and the SDK is judged once CI is green at the merged head.

In the brief's two-way verdict this counts as SEND BACK.

NOT READY FOR REVIEW

Japabu and others added 4 commits September 27, 2026 16:31
page_cache.rs stays deleted. #535's `partclaim-root-withheld` moves off its
read-fault injector onto `block::unanswered`, which gains `answer`: the kernel
drives only USB disks here, so the actuator refuses block 0 of each of those
across `rootfs::hold_source` alone. main.rs keeps the branch's
`gpt::probe_usb_disks` and drops the kernel mounts main still carries; main's
move of `platform_devices` into the device phase merges as it is.

partclaim.rs: both arms. `root_withheld` boots the partclaim config off its USB
stick (`Profile::UsbDisk`) with the crafted disk beside it, because an NVMe boot
here puts ROOT on a disk the kernel does not drive and would withhold it with
no actuator armed at all.

power.rs: the branch's `OTHERS`, with main's `console-queue-at-the-stop`
actuator and its judge.

storage.rs: `home_budget_refusal_retried` stays deleted with the kernel's NVMe
fsync path. #535's claim that a refused flush is refused once and not on every
`logd` flush moves to `log_flush_retry`'s first boot, where `/log`'s flush is
fsd's claim flush: no `partclaim: a flush durable on attempt` line in the 2 s
after the guest's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The table-read injector's module doc, `root_withheld`'s doc and the storm
check's comment were rewritten in the merge; they go, and the check keeps one
clause for its fixed window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tined

The signature main's nightly at 1ce7183 left, "the stream's redial was
turned away", is quarantined against main's issue, and what the branch's
duplicate measured moves there. The duplicate's deletion went in with
506bbdc, as did the deletion of the hash identity's issue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ory's connection is handed over in turn

B3, the orchestrator's ruling: a hash of what an entry records is no
identity, since the allocator hands a freed block back first and a
same-length file written within the second passed it. `Stat::ident` goes
from the protocol, fsd's two implementations and `toyos-fat32`'s accessor,
and std (fork 2b2fb74a303) no longer reopens a held file by path: a handle
whose server ended answers StaleNetworkFileHandle. `fs_restart` now expects
Gone from a handle held on an unchanged file across an end, reads the
flushed file through a new open, and still reds if a handle on a file
renamed over is reopened into it. The exit is
issues/filesystem/a-file-held-across-a-file-servers-restart-answers-gone.md.

std's futex mutex let the thread that let go of a directory's connection
take it back before the waiter it woke ran; `quiesce_stops_the_machine`'s
six writers on `/log` starved five on a loaded host. Callers take a ticket
and are served in order. `fs_turns` has six threads write one directory
until one made 32 passes, and every one must have made 16.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 13, at 5235eda (rounds 12 and 13, and the merge of a7cd327)

Gate

  • CI: green at 5235eda. Run 36412284926, job host success; steps cargo run -- --ci host and --ci gate-stage success.
  • The orchestrator's guest runs at 5235eda (orch-runs/536r13-*.log, summary.txt:414–420, porcelain clean on every green arm):
    • EXIT=0: fsd_claim_held and --nightly heap_ceiling_bounds.
    • EXIT=1 under their arms: b3, b1 and the red abuse arm.
    • Fast: EXIT=1, 400 passed and 1 failed.
    • The full nightly: EXIT=1, 513 passed and 2 failed.
  • Missing: the T14 reading, at any head. The orchestrator schedules it.
  • The landing head does not exist yet. main is at 69d1b53. git merge-tree --write-tree origin/main HEAD exits 0 (tree 5fa3f54c) with no conflict. main's new commits share five files with the branch: xhci/mod.rs, xhci/wait/boot.rs, tests/common/iommu.rs (Disable handle_basic behind the deferred-release-outlives-its-syscall issue, and run userdev_dma_fault on a binary with no census #571's userdev_dma_fault binary swap sits next to the branch's slot_of), tests/common/usb.rs and tests/toyos.rs. None of these conflict. Every guest measurement is owed again at the merge.

Net lines (git diff --shortstat origin/main...HEAD): 273 files, +10094 −10180.

part added removed net
kernel +595 −4989 −4394
userland (lockfile included) +4953 −653 +4300
other production +1249 −274 +975
production +881
tests/ +2590 −3245 −655
crate test directories +28 −339 −311
src/ +71 −93 −22
issues/ +608 −587 +21

Production grew 153 lines since r12 (+728). That is the SDK reader (+172) and the start refusal, less toyos-inventory (−109) and inspect's reader (−29). Accepted.

Earlier BLOCKERs

  • r12-1, heap_ceiling_bounds: CLOSED. It is EXIT=0 at 5235eda (536r13-green-heap_ceiling_bounds.log:33). Under b1-control.patch it is EXIT=1 with "64 extra threads and sysinfo never refused" (536r13-b1.log:36). It passes in the full nightly as well.
  • r12-2, one inventory reader: CLOSED. toyos-inventory is deleted, and init and inspect both call SysCap::records. With b2-no-retry.patch applied, both growth tests fail, EXIT=101 (fsd-r12/b2-mutate-head.out).
  • r12-3, a refused partition claim still starting the server: CLOSED for the kernel-disk claim. fsd_claim_held is EXIT=0. Under b3-say-and-continue.patch it is EXIT=1: the guest's second open was answered "NotFound … not Gone" (536r13-b3.log:47). Its sibling on the block-service side is still open, as BLOCKER 1.
  • r12 owed list:
    1. Done at 5235eda.
    2. Two nightly reds, adjudicated below. One of them is this branch's.
    3. Fast has one red, adjudicated below.
    4. Done.
    5. T14: open.
  • r12 NOTEs and REMOVEs: all done, except tests/CLAUDE.md, which is the orchestrator's to handle.

The two reds

  • iommu_virtio_platform: this branch's, deterministic. tests/netcase/system.toml:71 gives netcase a blockd with pci:1b36:0010.
    • On the no-unit arm the kernel refuses that claim, so the judge's list gains 00:02.0 and faults.rs:409 reds. That is 536r13-nightly.log:419, and the list reads ["00:02.0", "00:03.0", "00:03.0"].
    • The two 00:03.0 entries are main's: the refusals of netd's claim and of test-runner's.
    • The same boot shows the defect behind the red (BLOCKER 1).
  • blocking_read_window: main's filed flake, not a lost wake.
    • The branch leaves the wake path alone:
      • watch.rs, the pipe path and scheduler.rs are unchanged apart from comments;
      • the MID_UPDATE removal is read only by stop_if_blocked.
    • The unstaged canary blocking_read_stress passed in the same Fast run, in 119 ms (536r13-fast.log:619).
    • The red run's own numbers say the machine was running, not parked:
      • 28 round trips in 3 s;
      • cpu=1582ms and cpu=1577ms for the two ends. That is about 28×2 lapsed 50 ms watch-window holds, and every wait then completed.
    • It has the same signature as issues/build/blocking-read-window-reds-beside-other-guests.md, which records it on main (21–27 of 500, about 1.5 s of CPU per end). It is not on the redlist.
    • It went 2 of 2 at this head. The identical guest code passed at fae0a5d, bfb1763 and r10. The count at this head is the reason for the note on this red under NOTE.

BLOCKER

  1. A refused NVMe claim serves DATA from memory.
    • Where: userland/blockd/src/main.rs:530 → userland/fsd/src/main.rs:323.
    • What happens: init says pci:1b36:0010 is on this machine and could not be handed over, then starts blockd without its claim. blockd says "no NVMe controller this row names is on this machine" (false) and serves Drive::Absent. fsd then says "this machine has no DATA partition" (false) and puts /apps, /config, /home and /state in RAM. This is 536r13-nightly.log:605–611 on the no-unit boot.
    • Why: every write is lost at reboot on any machine with no DMAR, which includes VT-d disabled in firmware. main drove that disk, and main's own rule (bcachefs_adapter::open_data: a candidate no driver claimed is Absent) forbade the tmpfs. This is r12-3's fallback one layer down, against the branch's own fsd/src/main.rs:290–291. PR body l.18 claims Unusable covers it, but only a failed Controller::open reaches Unusable.
    • Fix: a refused claim reaches fsd as a refusal, so DATA is Absent by name and never RAM. diskless_boot's no-controller arm stays as it is.
    • Test: iommu_virtio_platform's no-unit arm must:
      • accept 00:02.0's refusal by name;
      • assert blockd's refusal line and fsd's "DATA is absent this boot";
      • must_not_say the in-memory line.
    • Mutation: blockd's refused-claim path sent back to Drive::Absent must turn it EXIT=1.
  2. Two DATA partitions are guessed between.
    • Where: userland/init/src/main.rs:2276 counts only kernel-disk partitions, and userland/fsd/src/main.rs:315 takes a claim without asking blockd. fsd/src/main.rs:324 answers two blockd partitions with RAM.
    • What happens: a machine with DATA on the stick and on NVMe serves the stick's and silently ignores the NVMe's. fsd_claim_held's own boot is that machine: 536r13-green-fsd_claim_held.log shows blockd: partition 20CE8E90… at block 256 beside the stick's 7E2B4C6D…, and 536r13-b3.log:42 shows the NVMe copy formatted. main refused two, "rather than guessed between" (bcachefs_adapter::open_data).
    • Why: one condition gets three answers (refused on kernel disks, RAM on blockd, a guess across both), and the RAM arm is BLOCKER 1's harm.
    • Fix: count both sources once and refuse two by name as Absent. Give fsd_claim_held a machine with one DATA.
    • Mutation: a test with DATA on both disks must turn red when init's count goes back to kernel disks alone.
  3. The stop's sync of the file servers has no negative control.
    • Where: userland/init/src/main.rs:1234, self.sync_files(), replaces the kernel's sync. The branch deleted quiesce_leaves_the_volume_whole, main's test of the stop leaving the volume whole, and names no control for its replacement (filesystem, high-risk).
    • Patch: delete line 1234.
    • The test it must turn red: home_overwrite_reads_back. The guest writes /home without fsync, runs run shutdown about 0.3 s later against a 2 s write-back, and the host reads the result off the image.
    • Record both arms. If it stays green, add a test that can fail.

NOTE

REMOVE

  • PR body: "At 5235eda, a7cd327 is still main's tip, so there was nothing new to merge." It is false: main is at 69d1b53.
  • PR body: the "Not yet run: any guest test at 5235eda, the whole --nightly tier at any head …" paragraph. It is false.
  • PR body: "The --nightly run at fae0a5d is void: the host disk filled during it." It is chronology.
  • tests/toyos-rust-tests/src/bin/heap_ceiling.rs:11 and :35: "plus 16". It restates the kernel's LOWERED_SYSINFO_HEADROOM.
  • userland/blockd/src/main.rs:506: "or None on a machine that has none its row names". A refused claim is also None.

Owed before LAND, at the landing head (after the merge with 69d1b53)

  1. BLOCKERs 1–3, each with its green arm and its red arm.
  2. Fast and the full nightly: EXIT=0, or each red on --known-red with its issue.
  3. The T14 reading.

SEND BACK

Japabu and others added 2 commits September 28, 2026 14:14
No conflict. main's rust pin is where the branch's fork merge left it
behind (1b236638, an ancestor of the branch's 62fa74d7), so the branch's
pin stands. No line main added cites a file the branch deletes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…er memory

A controller blockd's row names that is on the machine and whose claim the
kernel refused was served as no controller at all: blockd said none was on
the machine, fsd said the machine had no DATA partition, and /apps, /config,
/home and /state went into memory on every machine with no DMAR. init now
tells a block service which of its row's claims were refused with anything
but NotFound (`--claim-refused <name>`); blockd serves Drive::ClaimRefused,
whose listing and opens are refused with the new wire word ClaimRefused, and
fsd serves DATA absent by name.

DATA's partition was decided three ways: two on the kernel's disks refused
the start, two on blockd's went to memory, and one on each took the kernel's
without asking blockd. init now claims every partition of a role the
kernel's disks carry, and fsd::data::find counts those claims and blockd's
TOYOS-DATA listing together once: none is memory, one is served, two or more
are refused by name and DATA is absent, and a listing refused leaves the
count unknown, so DATA is absent then too. A log or boot claim is by the
loader's GUID, which the kernel refuses when two partitions carry it, so that
arm is unchanged.

iommu_virtio_platform's no-unit arm, which boots tests/netcase with its
blockd row, accepts 00:02.0's refusal by name, and asserts init's, blockd's
and fsd's lines and that the in-memory line is never said. fsd_two_data
boots DATA on a stick and on NVMe and asserts the refusal, the absent
volume, no format and the stick unchanged byte for byte. fsd_claim_held and
the three partclaim boots, which stage DATA on a stick, get an NVMe disk
with no table so the machine has one DATA.

The two "plus 16" restatements of the kernel's headroom in heap_ceiling.rs
are deleted. The issue on a worktree's bootstrap cache is restored: the pull
request that fixes it is not merged. A foreign DATA partition still answered
with memory is filed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Japabu added a commit that referenced this pull request Sep 28, 2026
…ion, claims as sentences

The rename's init exclusions become one rule: a match stays only where
init means initialise, initial or the CPU's INIT signal, or sits in
third-party text; every other match names the program. The search is
bounded so CamelCase joins hit: no lowercase letter follows, and a
letter precedes only a capital I. The review's literal "no lowercase
letter on either side" misses TimerInit, which occurs 6 times outside
issues/ at origin/main cd2e630; this bound catches it.

Stage 2's exit now requires each decision deleted from
userland/supervisor, which calls the crate for it. Stage 4's claims are
sentences. The tier-and-redlist clause covers every guest test the
track names. "The fork's delta" is forkcheck's definition. soundd's and
blockd's new names reopen with the owner beside fsd's: mixer and disks
are words the tree already uses.

Removed: the stop on #536, the manifest's current order, the file
manager's name, and the power-broker clause citing the deleted applet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Japabu and others added 3 commits September 28, 2026 16:58
…le calls

Every file call on this branch is a request on a connection to a file
server and a blocking read of its reply, and sys_read/sys_write charged a
connection's wait to WaitClass::Pipe. watch-window holds every pipe-class
waiter in a Ring 0 spin of up to 50 ms with no preemption point, so under
the actuator each of logd's writes through fsd, and fsd's through blockd,
became such a spin beside blocking_read_window's canary. A connection's
wait is now WaitClass::Ipc, which is also what the blocked-time breakdown
should have said of it; a bare pipe's stays Pipe and stays held.

process_stats gains the arm that names it: a child parked reading a
connection, answered once the roster says it is blocked, charges its park
to blocked_ipc_ns.

home_overwrite_reads_back's guest fsyncs a file before the pinned
overwrite, so the volume's format is on the device and the stop's sync is
the only thing that carries the pinned file there: without that sync the
host's read names the lost file rather than an unformatted volume. The
comment naming SYS_SHUTDOWN's drain, which no longer exists, is deleted.

Filed: watch-window spins out a hold whose poster is queued behind it on
the same CPU, the mechanism the canary's 50 ms-a-half-trip reds fit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
One modify/delete conflict: main disabled ftruncate_flush_race behind
issues/build/ftruncate-flush-race-reds-intermittently-and-nothing-says-why.md
and gave that issue an expected-red status, a sighting, an exit condition
and an owner. This branch deletes the test with the kernel flush it raced
and the ftruncate-flush-stall actuator, so every one of those hunks is about
a test that no longer exists: the issue stays deleted and main's redlist row
for it is removed, since the harness refuses a row whose test nothing
registers. main's rust pin did not move, so the fork merge stands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…mpiler switch

The issue's exit was a worktree's bootstrap cache keyed on, or cleared
with, the compiler that builds it. The merge of ec06384 brought
src/sysroot.rs's forget_another_compiler: build_std records the compiler's
identity in <fork>/build/toyos-std/compiled-by, and on any other identity
empties that directory of everything but bootstrap's downloads (cache,
<host>/ci-llvm, <host>/rustfmt), so rust/build/toyos-std/bootstrap, where
the stale serde rlibs lived, goes with it. The rule is stated at that
function; nothing else cites the issue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 15, at 069722c

Gate. CI host COMPLETED SUCCESS at 069722c (run 36441953279). Guest runs by the orchestrator are in orch-runs/, summary.txt between QUEUE55-DONE and QUEUE56-DONE. Fast EXIT=0, 399 of 399 (536r15-fast.log:1581). Nightly EXIT=1, 512 of 514 (536r15-nightly.log:1606); both reds are judged below. The T14 reading is still missing.

Net lines (git diff --shortstat origin/main...HEAD): 279 files, +10465 −10222.

  • Production is +993: kernel −4384, userland +4389, other +988. That is +112 since r13, all of it r14's refusals plus r15's +18/−8 in io.rs. Accepted.
  • Tests: tests/ −477 and the crate test directories −311.

Earlier BLOCKERs

  • r13-1, a refused NVMe claim serving DATA from memory: CLOSED.
    • iommu_virtio_platform EXIT=0 at 06c6195, and again inside the 069722c nightly (536r15-nightly.log:500).
    • Under b1-refused-claim-absent.patch it is EXIT=1 (536r14-b1.log:34): fsd's Refused(ClaimRefused)… DATA is absent line is never said.
  • r13-2, two DATA partitions guessed between: CLOSED.
    • fsd_two_data EXIT=0 at 06c6195 and at 069722c, in Fast and in the nightly.
    • Under b2-kernel-claim-uncounted.patch it is EXIT=1 (536r14-b2.log:33).
    • The host arm h1-find-guesses is EXIT=101.
  • r13-3, no control on the stop's sync: CLOSED.
    • home_overwrite_reads_back EXIT=0 at 069722c.
    • Under b3-no-stop-sync.patch it is EXIT=1, "reading home/overwrite-pinned.bin off the image: NotFound" (536r15-b3.log:35). The red now names the lost file, not the unformatted volume of 536r14-b3.log:33.

Round 15's claim

BLOCKER

None open.

NOTE

  • kernel/src/syscall/io.rs:59 — the write side of the reclassification is unmeasured. This patch passes every test: Some(id) => WriteBlock::Pipe(id, WaitClass::Pipe),. process_stats needs a child parked writing to a full connection, and that patch must turn it red.
  • kernel/src/syscall/io.rs:95 — pipe_wait is a matches! with a silent Pipe default. A new pipe-bearing kind is classed by accident. Decide the class in the exhaustive matches that already name every pipe-bearing kind (object/ops.rs:204, :216).
  • tests/toyos-rust-tests/src/bin/process_stats.rs:313–332 — a second copy of process_lifecycle.rs:210–258: the roster decoder, its BLOCKED constant and the bounded poll. One shared decoder would do both.
  • Landing head: main is at ec0a91a (Clipboard copied once into a region the compositor made; copy-once is a type, the compositor forbids unsafe code #557). git merge-tree --write-tree origin/main HEAD exits 0 (tree 3c1eff51), and main's new commits share no kernel, fsd, init or blockd file with the branch. Fast and the nightly are owed at that merge. A blocking_read_window red there goes on --known-red with main's issue, never re-run.

REMOVE

  • PR body l.123, "069722c3 answers it.", and l.99, "The A/B against main at 069722c is what judges the fix." — 0/10 against 2/10 does not separate.
  • PR body row 73, "at 06c6195, which is this arm" — false: 06c6195 has neither the merge of ec06384 nor r15's test edits.
  • PR body rows 72 and 75 and l.125: the "expected EXIT=1" clauses and "run by the orchestrator" with no result — stale now that the runs exist.
  • kernel/src/actuator.rs:260 — "Hold every thread that waits on a watch" is false, more so after this change. PR body l.138, which records leaving it, goes with it.
  • issues/kernel/watch-window-spins-out-a-hold-its-poster-is-queued-behind.md:24 — the orch-runs/ab-brw-536-7.log citation is a scratch path that will not exist.
  • tests/toyos-rust-tests/src/bin/process_stats.rs:260 — "which is what every file call's wait on its server is" is narration.

LAND AFTER NAMED CHANGES

Japabu and others added 4 commits September 28, 2026 18:21
No conflict. The one file both sides change is tests/toyos.rs; main's
hunks add metal_sim_hostile_clipboard (skip row, tier row, carries row,
its function and dispatch arm) and a long-copy check in
metal_sim_client_death, none of which this branch touches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
… is chosen where the kind is

- `ops::pipe_read` and `ops::pipe_write` answer a blocking read's or write's
  pipe together with its wait class, in the exhaustive matches that already
  name every pipe-bearing kind, so a new one cannot compile without choosing.
  `io.rs`'s `pipe_wait`, a `matches!` that fell back to `Pipe`, is deleted;
  `pipe_id_read` and `pipe_id_write` are those two with the class dropped.
- `process_stats` gains `a_wait_to_write_a_full_connection_is_ipc`: a child
  fills a connection, parks writing one byte more, and this process reads to
  make room once the roster has it blocked. Both connection arms share
  `parked_on_a_connection`.
- `tests/toyos-rust-tests/src/roster.rs` is the one roster decoder, its
  `BLOCKED`, the capability it reads with and the bounded poll, included by
  `process_stats` and `process_lifecycle` in place of their two copies.
- Deleted: `watch_window`'s doc in `actuator.rs`, which said it holds every
  watch waiter; the narration in `a_wait_on_a_connection_is_ipc`'s doc; and a
  scratch log path cited in the watch-window issue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
#564's schedule and its 15 deletions meet this branch's storage move.
Where #564 deleted a test or the kernel code only it armed, the deletion
stands; where this branch deleted a test #564 retiered, this branch's
deletion stands; every surviving row takes #564's tier, and the five this
branch adds (fsd_restart, fsd_end_at_mount, fsd_claim_held, fsd_two_data,
blockd_serves_nothing) keep the Fast tier they were added at, as #564 kept
the tier of every row main added.

- kernel/src/actuator.rs: quiesce-last-exit and quiesce-dump go (#564);
  fat-flush-meta-refuse, resize-evict-window and resize-fault-refuse stay
  gone (this branch).
- kernel/src/syscall/machine.rs: the stop serves no dump (#564) and syncs
  no filesystem (this branch): neither block survives.
- toyos-quiesce/src/lib.rs: this branch's FILES_MS, FLUSH_MS and SYNC_MS,
  and #564's LAST_THREAD doc naming quiesce-last-park alone.
- tests/common/power.rs: quiesce_dump_holds_the_stopped goes with its
  actuator; quiesce_wakes_on_the_last_park is #564's inlined body with this
  branch's stop-record ordering in place of "Syncing filesystems...".
- tests/common/storage.rs and tests/toyos-rust-tests/src/bin/so_cache_policy.rs:
  so_cache_refusals and its binary go (#564); this branch's /tmp retarget of
  them serves no surviving test. The so-cache-tiny actuator went with them.
- issues/filesystem/home-budget-refusal-retried-is-red-on-every-nightly.md
  stays deleted: its test and binary are gone on both sides, and the
  reproducer #564's exit names is the kernel's /home over its NVMe driver,
  which this branch deletes. The fsync-budget-spent race it pointed at is
  still tracked by the partition-claim-gives-up issue.
- issues/README.md: boot-media is gated by esp_filesystem, kernel_log_file,
  log_partition_layout and log_partition_identity, plus toybox_cp_volume:
  wall_clock_file went in #564, log_backing_read_error and
  boot_volume_metadata_error on this branch.
- tests/toyos.rs: RUST_SKIP, MACHINE_TESTS, CARRIES and the dispatch lose
  so_cache_refusals, quiesce_wakes_on_the_last_exit and
  quiesce_dump_holds_the_stopped (#564) and keep this branch's deletions
  (cache_eviction, the writeback trio, page_cache_partition_offset,
  quiesce_leaves_the_volume_whole, ftruncate_flush_race,
  log_backing_read_error, boot_volume_metadata_error); the comment of the
  deleted ftruncate_flush_race row goes with it. #564's tier for
  block_duplicate_id, partition_claim_departure, quiesce_refuses_a_second_shutdown,
  quiesce_wakes_on_the_last_park, late_storage_connect, log_partition_layout,
  the root_* rows, log_partition_identity, tls_rebase_window,
  sysret_ss_reload, userdev_residue_is_its_own and
  blockd_lends_within_its_bound, beside this branch's comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
#572 deletes ovmf/ and aavmf/ so every guest boots the host QEMU's own
edk2 firmware; this branch already deleted ftruncate_flush_race and its
kernel actuator/vfs code as part of retargeting the nightly storage
suite onto fsd (193600c). The two touch the same issue file: #572
appended a firmware-comparison row (stock edk2 vs. the committed
ovmf/, both still flaky) to a table whose subject — the test binary
and the VFS lock path it raced on — no longer exists on this branch
(`rg ftruncate_flush_race` finds nothing under kernel/, tests/, or
src/). The evidence adds nothing a surviving issue or test needs, so
the deletion stands; no other file cites the issue by path or bare
name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 17, at d5bbfbb

Gate. CI host COMPLETED SUCCESS at d5bbfbb (run 36458582511, headSha d5bbfbb). The orchestrator's guest runs at d5bbfbb (summary.txt between QUEUE63-DONE and QUEUE64-DONE):

  • Fast EXIT=0, "240 passed, 240 total" (536m-fast.log:937).
  • Nightly EXIT=0, "383 passed, 383 total" (536m-nightly.log:1102).
  • process_stats, --nightly quiesce_refuses_a_second_shutdown and the eleven --weekly rows EXIT=0.
  • The b3 arm EXIT=1: "reading home/overwrite-pinned.bin off the image: NotFound" (536m-b3.log:31).

The T14 reading is still owed. It is a landing condition and outside this review.

Net lines (git diff --shortstat origin/main...HEAD): 279 files, +10551 −10077.

  • Production is +993, unchanged since r15: kernel −4384, userland +4389, other +988.
  • r16 is +171 −128 over 7 files, and its kernel part nets to zero.
  • Tests: tests/ −238, the crate test directories −311. issues/ +56.

Earlier findings (r15)

  • r15 BLOCKERs: none were open.
  • NOTE io.rs:59, the write side unmeasured: CLOSED. a_wait_to_write_a_full_connection_is_ipc is EXIT=0 at 0baa8b3 and d5bbfbb. Under fsd-r16/write-class-pipe.patch it is EXIT=1, "charged 0 ns to ipc and 627150 ns to pipe" (536r16-write-class-pipe.log:46). The read arm passes in the same run (:43), so the red is the write class alone.
  • NOTE pipe_wait's silent Pipe default: CLOSED. pipe_wait is deleted. The class is chosen in ops::pipe_read and ops::pipe_write (kernel/src/object/ops.rs:215–237), which name every kind. The whole-change revert is EXIT=1 at 0baa8b3 (536r16-ipc-class-reverted.log:45).
  • NOTE, the duplicated roster decoder: CLOSED. tests/toyos-rust-tests/src/roster.rs is the one decoder, and both binaries include it.
  • REMOVE items: all CLOSED. The two A/B sentences, row 73's clause, the "expected EXIT=1" clauses, actuator.rs:260 with its body line, the scratch path in the issue and the narration in process_stats are all gone.

Merges

BLOCKER

  • src/redlist.rs (origin/main c089fce, Disable three flaky main-level reds behind their issues #580, row at :64) — git merge-tree --write-tree origin/main HEAD exits 0 (tree 950ac5a5), and yet the composition is red. Disable three flaky main-level reds behind their issues #580 disables quiesce_leaves_the_volume_whole, which this branch deletes. redlist::check (src/redlist.rs:141) then refuses "is disabled and nothing registers it" in check_redlist (tests/toyos.rs:18735, :18796), before any boot and on --list, so the merge queue reds.
    • Merge origin/main.
    • Delete that row, and delete issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md. Its test, quiesce-fsync-refuse, mirror_refuse and Syncing filesystems... are all gone on this branch.
    • Measure cargo test --test toyos-build -- --list EXIT=0 at that merge, and git grep quiesce_leaves_the_volume_whole -- src tests issues must come back empty.

NOTE

REMOVE

SEND BACK

Japabu and others added 2 commits September 28, 2026 19:51
… a test this branch deletes, move the stop-sync's only guest control to Nightly, and rename a misnamed binding

Merging origin/main (c089fce, #580) brought back a redlist row and issue for
`quiesce_leaves_the_volume_whole`, a test this branch already deletes; the
harness refuses a disabled row nothing registers, so both go.
`home_overwrite_reads_back` is the only guest check on init's stop sync and
was priced at Weekly for the old kernel-`/home` test it replaced, so a lost
sync goes unseen by every PR and nightly; it moves to Nightly.
`tests/common/power.rs`'s `synced_at` bound the stop record's line, not
anything synced there. Also deletes two stale doc comments in
`process_stats.rs` that only narrated the code below them, and a roster.rs
clause the reviewer found false for `process_stats`'s own child.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Japabu and others added 2 commits September 29, 2026 08:45
Main's rust pin has not moved since the last merge, so the fork is unchanged.
Every conflict, and how it was resolved:

Modify/delete, main deleted:
- issues/build/a-swaps-redial-races-a-hard-dial-ceiling-against-an-unbounded-guest-gap.md:
  #566 fixed the defect and deleted the issue. This branch had added one sighting to
  it, and a sighting of a fixed defect has no home, so the file stays deleted.
- src/heartbeat.rs: #562 deleted `kernel_heartbeat`'s CPU-mask and gap verdicts with
  the file. This branch had given its done-line table blockd and fsd rows. The table
  goes with the verdict it served.
- tests/doomcase/system.toml: #562 moved the doom audio tests to metal and deleted
  their QEMU config. This branch had added blockd and fsd rows to it. Nothing boots
  it now.

Modify/delete, this branch deleted:
- tests/toyos-rust-tests/src/bin/ftruncate_flush_race.rs,
  tests/toyos-rust-tests/src/bin/quiesce_fsync.rs,
  issues/build/ftruncate-flush-race-reds-intermittently-and-nothing-says-why.md,
  issues/build/quiesce-leaves-the-volume-whole-needs-its-flush-to-close-inside-the-stops-budget.md
  and issues/kernel/a-root-metadata-read-refused-on-budget-is-not-retried.md: main's
  hunks remove timing from them or note its own runs. They are about the kernel FAT
  flush, the stop's kernel sync and the kernel's metadata read, which this branch
  deletes, so they stay deleted.

Content:
- kernel/src/actuator.rs: main's `quiesce_last_teardown` (#549) is kept. The kernel
  FAT actuators `fat_flush_meta_refuse`, `resize_evict_window` and
  `resize_fault_refuse` stay deleted. `process_reopen_selftest` stays where this
  branch has it, with main's doc (#549 also opens every kernel thread's pid).
- src/redlist.rs: both conflicted rows go. `doom_sound_flood` left QEMU with #562,
  and this branch deletes `ftruncate_flush_race`.
- tests/common/gpt.rs: this branch's `device_saying` and decoy `boot` are kept. Main
  drops the `drain_serial` window, so its `qemu` binding is no longer `mut`.
- tests/common/inspect.rs: main's "nothing plays audio" (#562 deleted
  `inspect_plays`) is taken, with this branch's clause on the boot stick.
- tests/common/iommu.rs: main's `panic-reboot-fast` and its wait for the fatal path's
  reset are kept. This branch's `iommu_empty_domain` reads the xHCI's DCBAAP over
  QMP, and QEMU has exited by the time that reset is seen. So `fault_boot` now takes
  a `holding` read, which it runs after the fault line and before it waits for the
  reset, while the fatal path holds its panel. `iommu_context_absent` reads nothing
  there.
- tests/common/origin.rs: main's judgement of `log_ring_keeps_the_owners_slots` is
  taken whole: init says it waited a flush out, or its stop line is missing. That
  drops the millisecond inference between two records, whose record this branch had
  changed from `Syncing filesystems...` to the stop record (#562: no QEMU test
  measures time).
- tests/common/volumes.rs: main's timing edit to `ftruncate_flush_race` goes with
  the test.
- tests/logstallcase/system.toml: main drops `power` and the `shutdown` symlink, since
  the metal row reads `/log` without a stop. This branch's blockd and fsd rows are
  kept, because fsd holds `/log`.
- tests/toyos-rust-tests/src/bin/blockd_io.rs: main's `claim_when_free`, now generic
  and with no deadline, is taken inside this branch's `if let Some(syscap)`. `bench`
  is this branch's blockd-only arm with main's timing removed: no MiB/s, and the
  line says only how many Flushes each run took. The module doc's "timed" goes.
- tests/toyos-rust-tests/src/roster.rs (add/add): both sides wrote one roster
  decoder. Main's is taken whole, because five binaries read it and it has no
  deadline (#562). This branch's copy had a 5 s give-up.
- tests/toyos-rust-tests/src/bin/process_lifecycle.rs: main's is taken whole. This
  branch's only change to it was the move onto its own roster.rs.
- tests/toyos-rust-tests/src/bin/process_stats.rs: main's `refused_calls_are_counted`
  and its roster wait for the held child are kept, and so are this branch's two
  connection arms. The system capability is taken once in `main` and passed to the
  three arms that read the roster, since a second take of the label finds nothing.
  The connection arms now wait on main's `threads_of` for the child's main thread
  to be blocked, with no deadline.
- tests/toyos-rust-tests/src/bin/quiesce_twice.rs: main's `Duration`-only import.
  This branch deletes the owed file, so `File` and `Write` go.
- tests/toyos.rs:
  - RUST_SKIP: main's audio rows are taken. `audio_tone_load` goes, since main
    deleted it. `log_volume_reread` goes, since this branch deletes it.
  - MACHINE_TESTS: `quiesce_leaves_the_volume_whole` stays deleted.
    `quiesce_wakes_on_the_last_teardown` comes from main with main's comment.
    `blockd_serves_nothing` is kept. `hda_tone` and `hda_client_stall` went to metal
    with #562, and `hda_two_live_refused` takes main's comment.
  - CARRIES and dispatch: the same.
  - `nvme_wide_sector`: this branch's blockd arm, which already had no drain window.
- toyos-quiesce/src/lib.rs: this branch's `FILES_MS`, `FLUSH_MS` and `SYNC_MS` are
  kept, with main's `LAST_THREAD` doc, which names both quiesce-last actuators.
- userland/logd/src/policy.rs: this branch deletes the module doc and the
  `LOG_WRITE_BUDGET` paragraphs main edited one line of, so they stay deleted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`issues/build/parallel-tests-red-under-other-suites.md` recorded
`quiesce_leaves_the_volume_whole` red under `quiesce-fsync-refuse` with the
kernel's log volume. The test, the actuator and that volume are all gone on this
branch, so the sighting is deleted and not rewritten.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Japabu and others added 2 commits September 29, 2026 12:52
On the T14 at b607ce7 metal_device_probe failed with
  nvme: no record carries "blockd: no NVMe controller this row names is on this machine"
although the readback's log carries that line:
  {2026-09-29 08:43:12 1.162 blockd} blockd: no NVMe controller this row names is on this machine; serving no partition

The judge handed `unmet` `Readback::kernel()`, which is the log with
every program's line taken out (`bootlog::kernel_records`). On main the
NVMe records were the kernel's; on this branch they are blockd's, a
program's line, so the filter removed the one record the row owed, and
`nvme-served` (NeverSays "blockd: partition ") could never have seen a
served partition either. The unit fixture hid it by writing blockd's
line as a kernel record.

blockd's two records move to their own table, BLOCKD_RECORDS, judged
against `bootlog::lines_of(log, "blockd")`; RECORDS stays judged against
the kernel's records alone, now filtered inside `unmet`, which takes the
whole log. A kernel line or another program saying blockd's words does
not count, nor a program saying "Boot: complete (". A new test holds
BLOCKD_RECORDS to blockd's source. The inventory reads the whole log so
its "blockd: " row reports something.

Oracle: the recorded T14 readback (target/metal/metaldevicecase), run
through both readings as a temporary test: the old one reproduces the
T14's single finding exactly, the new one reports none.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs
Brings #583, #586, #593, #597, #555, #600, #610 and #611. Every conflicted
hunk, and where it went:

- rust: 1471893e39c, which merges main's pin 9c3eea441d8 into this branch's
  62fa74d7a50 with no conflict (main's three bootstrap commits and this
  branch's six std files do not meet). Pushed to ToyOSOrg/rust wt-toyos-fsd.
- kernel/src/actuator.rs: #583 deletes `rtc-zone-east`, and that deletion
  stands. This branch's doc for `leak-rollback-selftest` stands, since the
  FAT reopen control went with the kernel's FAT adapter.
- kernel/src/fat32_adapter.rs (modify/delete): the deletion stands. #583's
  hunk made FAT stamp UTC (`clock::utc_secs`) with the refusal reason at the
  site. FAT stamping is fsd's on this branch, so it is carried there:
  `local_secs`, which recovered a zone through `toyos_wallclock::resolve`
  (deleted by #583) and cited logd's `wall.rs` (deleted by #583), becomes
  `utc_secs`, `clock_epoch()` alone, with main's reason. fsd no longer
  depends on toyos-wallclock; userland/Cargo.lock drops the edge.
- kernel/src/sched/kthread.rs: #586 deleted `lognest` and `log-storm`, and
  this branch deleted `iod`, so klogd is the one kernel thread in every
  build: MAX_KERNEL_TASKS = 1, with no feature split.
- issues/kernel/the-kernel-still-creates-threads.md: #586 met K3 and deleted
  it; this branch meets K5 and deletes it. K6 is blocked on K2 and K4.
- userland/logd/src/main.rs: #583's `boot_secs` rename, beside this branch's
  `Published::new()`, which takes no argument here. The `owed` and
  `retrying_since` fields stay deleted (this branch).
- userland/logd/src/serve.rs: this branch's `serving` flag, with #583's
  `boot_secs`.
- tests/test-durations: #586 deleted `log_conservation_smp1` and this branch
  deleted `log_backing_read_error`; neither row stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 19, at 27b5641 (806da12 and the merge of d1d83f6)

Gate.

  • CI: host COMPLETED SUCCESS at 27b5641 (run 36564193345).
  • The orchestrator's runs at 27b5641:
    • T14, the full metal set: no failure that main's own baseline lacks. metal_device_probe PASS.
    • Fast 242/242 and --nightly 376/376.
    • --weekly: wall_clock_utc, iommu_context_absent, iommu_empty_domain and metal_device_probe each EXIT=0.
  • 806da12: cargo test --lib -- metaldevices EXIT=0, 8/8. Both reader mutations EXIT=101, and the recorded T14 readback serves as oracle (PR body l.139).
  • cargo test --test toyos-build -- --list EXIT=0 at 27b5641 (mine).

Net lines (git diff --shortstat origin/main...HEAD, base d1d83f6): 275 files, +10458 −10041.

  • Production is +973: kernel −4385, userland +4370, other +988. That is −20 since r17.
  • src/ −1, tests/ −235, crate test directories −311, issues/ −9.
  • The merge's own fsd change is +5 −24. Accepted.

Earlier findings (r17)

  • BLOCKER, Disable three flaky main-level reds behind their issues #580's redlist row: CLOSED. git grep quiesce_leaves_the_volume_whole -- src tests issues comes back empty, and --list EXIT=0.
  • NOTE, home_overwrite_reads_back at Weekly: CLOSED. It is Nightly now (tests/toyos.rs:875).
  • NOTE, synced_at: CLOSED. It is stopped_at (tests/common/power.rs:358).
  • REMOVEs: CLOSED. roster.rs's "already decided" clause and process_stats' two narrating docs are gone.

806da12

  • unmet now judges RECORDS against bootlog::kernel_records and BLOCKD_RECORDS against bootlog::lines_of(log, "blockd"). The kernel reader and the whole-log reader are both measured red.
  • The forged-writer and unbooted cases cover a sibling kernel line and a sibling program line.
  • If the tag is renamed in tests/metaldevicecase/system.toml, nvme's Says fails loudly rather than passing vacuously.
  • blockd_writes_the_lines_its_table_reads holds both needles to userland/blockd/src/main.rs:556 and :583.
  • No mutation found that survives.

The merge 27b5641: every conflict hunk is accounted for

BLOCKER

  1. The branch no longer composes with main, and A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587's mtime contract is not met by fsd (userland/fsd/src/volume.rs:146, kernel/src/object/ops.rs:115, :133, :536, :779, kernel/src/revoke_selftest.rs:21).
    • The composition. git merge-tree --write-tree origin/main HEAD (origin/main 8a66b44) exits 1. It conflicts in kernel/src/fat32_adapter.rs (modify/delete again, A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587), kernel/src/vfs.rs, kernel/src/object/ops.rs, kernel/src/leak_selftest.rs, issues/isolation/the-so-caches-refusals-are-narrower-than-its-reach.md and rust. The fork itself merges 90697f1401a cleanly (tree 9479250b).
    • The contract. A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587 is an owner-approved ABI contract: Stat::mtime is UTC nanoseconds since the Unix epoch, DATA keeps the nanosecond, and 0 is undated.
    • What breaks. fsd stamps DATA with clock_epoch()·10⁹, which is whole seconds. Two writes to /home inside one second then carry one mtime, which is the make/ninja failure A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587 exists to prevent. The kernel's /tmp stamps are still nanos_since_boot.
    • Fix:
      • Merge main.
      • Stamp mtime_now at every surviving kernel site.
      • Give fsd one wall-clock reader that meets DATA's row. userland can derive the kernel's anchor exactly: clock_epoch() bracketed by two nanos_since_boot() readings in the same second gives BOOT_SECS, and BOOT_SECS·10⁹ + nanos_since_boot() is utc_nanos. The alternative is that the owner amends A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587's table and the gap is filed.
      • Main's undated-FAT issue now describes fsd's stamp.
    • Must turn red, measured at the merged head:
      • file_mtime extended to /home with two writes that must stamp strictly later. It must turn red under this head's now_nanos.
      • file_mtime_undated extended to /home. It must turn red under now_nanos → clock_epoch().map_or(toyos_abi::clock::nanos_since_boot(), |s| s.saturating_mul(1_000_000_000)). Today it writes /tmp alone, so that fsd mutation passes every test.
      • A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587's site-114, m132-ops and m821-ops must turn file_mtime red.
      • m926-fat-seconds, applied to fsd (drop * 1_000_000_000 at fsd/src/fat.rs:209, :216, :230, :307), must turn file_mtime's /log arm red.
      • file_mtime_survives_a_reboot must be EXIT=0 on fsd's DATA.

Owed before LAND, at the landing head

The controls below are the ones landing needs:

  • The stop's sync. home_overwrite_reads_back (Nightly) must be EXIT=0, and EXIT=1 under b3-no-stop-sync (delete init's self.sync_files(), and put #[allow(dead_code)] on sync_files), failing with "reading home/overwrite-pinned.bin off the image: NotFound".
  • A connection's wait is IPC. process_stats (Fast) must be EXIT=0.
    • Under ipc-class-reverted (kernel/src/object/ops.rs's pipe_read/pipe_write class and kernel/src/syscall/io.rs put back as at 06c6195) it must be EXIT=1 in a_wait_on_a_connection_is_ipc, with "charged 0 ns to ipc".
    • Under write-class-pipe it must be EXIT=1 in a_wait_to_write_a_full_connection_is_ipc, with the read arm's line printed before it.
    • The last reds were at 0baa8b3. Since then the park predicate is main's roster decoder, and sched/payload.rs, the other WaitClass reader's file, has changed.
  • fsd stamps FAT in UTC (the merge's new claim). Under utc-plus-7200 (clock_epoch().map(|s| s + 7200).unwrap_or(0) in fsd/src/main.rs:172), --weekly wall_clock_utc must be EXIT=1, with "this boot's FAT timestamp is 72NN s from the staged instant".
    • This stands for the whole-change revert: the zone recovery it replaced no longer exists on main to revert onto, and the test stages UTC+2, so +7200 is exactly the old stamp.
    • Independent oracle: the host's -rtc base= instant, read back by the host's FAT reader (volumes::root_entries).
  • BLOCKER 1's arms.
  • Each arm's patch goes on the PR as a diff. The fsd-r* scratch patches the table cites are gone.

NOTE

  • userland/fsd/src/main.rs:171: utc_secs is a second wall-clock reader beside fsd::volume::now_nanos, with the same undated rule in another unit. Keep one, and have FAT divide it, as main's stamp(mtime_now()) does.
  • issues/build/two-shared-members-assume-a-bcachefs-home-the-t14-does-not-have.md (Triage main's T14 metal reds: issues only, no redlist rows #615, arrives with the next merge) rests on the kernel's tmpfs /home and NvmeBacking. Both are gone here, and the T14's /home is fsd's bcachefs in RAM (fsd/src/main.rs:362). Account for it at the merge against the landing head's T14 reading of fs_large_file and home_backing_revoked.

REMOVE

  • PR body l.141–149, the "Owed at 27b5641" and "Also owed, from b607ce7" lines. The T14, wall_clock_utc, Fast, the nightly and both iommu weekly rows are now run.
  • PR body l.41, "The T14 at b607ce7 failed … a partition blockd served." This is the investigation story, and 806da12's message carries it.
  • PR body l.42, "which is main's Loader slimming stage 1: the RTC is UTC, KernelArgs refuses another layout by name, IA32_TSC_ADJUST moves to the kernel #583 ruling carried … toyos-wallclock." This is merge provenance, and 27b5641's message carries it.
  • PR body l.56, the provenance paragraph ("Each mutation is a checked patch … none of them is rebuilt there."). It is chronology over nine scratch directories, none of which exists.
  • tests/toyos-rust-tests/src/bin/home_overwrite_zero.rs:59, "fs::metadata is the file_cache::size a read bounds itself by." It is false here: /home is fsd's.

SEND BACK

🤖 Generated with Claude Code

https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs

Japabu and others added 2 commits September 29, 2026 15:19
Brings #587 (a file's mtime is UTC nanoseconds since the Unix epoch),
#603, #614 and #615.

Conflicts, every hunk of both sides:

- kernel/src/fat32_adapter.rs (modify/delete): the deletion stands.
  #587's hunks there are `stamp(mtime)` as whole UTC seconds, `now()` as
  `stamp(mtime_now())`, `create` stamping the VFS's mtime, the flush
  stamping its own instant, and `file_mtime` answering seconds times 10^9.
  FAT stamping is fsd's on this branch: its `lstat`, `list` and
  `node_meta` already answer seconds times 10^9, its level stamps the
  flush's own instant, and the next commit gives it the nanosecond clock
  to divide.
- kernel/src/vfs.rs: this branch deleted `open_backing_identified`'s
  flush of a dirty file, so #587's `mtime_now` there goes with it. #587's
  `file_mtime` doc hunk is taken.
- kernel/src/object/ops.rs: the write and `ftruncate` stamps keep this
  branch's `file_cache::touch` and `set_size`, and stamp `mtime_now`; the
  two open stamps merged clean. #587's comment on dirty state in the cache
  is not taken: this branch's cache keeps no dirt.
- kernel/src/leak_selftest.rs: this branch deleted `fat_reopen_census`
  with the FAT adapter it probed, so #587's `mtime_now` there goes too.
- kernel/src/clock.rs (merged clean): `NANOS_PER_SEC` is private, since
  the FAT adapter it was made public for is gone.
- issues/isolation/the-so-caches-refusals-are-narrower-than-its-reach.md:
  this branch deleted the FAT same-size-rewrite section, because the
  kernel's cache reaches no FAT mount. #587's undated paragraph is still
  true of `/tmp`, the one writable mount a kernel `dlopen` reaches, so the
  section stays for it, with #587's exit condition.
- rust: c4c65e3e87a on ToyOSOrg/rust wt-toyos-fsd merges main's pin
  90697f1401a with no conflict.

Issues arriving with main, against this tree:

- two-shared-members-assume-a-bcachefs-home-the-t14-does-not-have.md
  (#615) is deleted. Its premise is the kernel's tmpfs `/home` and
  `NvmeBacking`, and both are gone: on the T14 `/home` is fsd's bcachefs in
  memory. The T14 run at 27b5641 reads `fs_large_file` exit=0 and
  `home_backing_revoked` exit=0 beside fsd's "this machine has no DATA
  partition; ... are in memory" line.
- the-last-handle-to-close-stamps-the-file-with-its-own-mtime.md is
  deleted. Its mechanism is the kernel write-back's last-close flush, which
  this branch deletes. `/tmp`'s mtime is the file's (`file_cache::touch`),
  and fsd's DATA keeps one mtime per node. What stays true of a handle's
  mtime is filed as a-kernel-files-fstat-answers-its-handles-mtime.md.
- a-fat-files-mtime-reads-finer-through-its-writer-than-after-a-reopen.md
  is deleted. Its exit condition's second arm holds here: the kernel mounts
  no FAT volume a process writes. fsd's FAT reads the entry's stamp through
  a handle, which differs from a reopen's the other way. That is filed as
  a-fat-files-mtime-through-its-handle-is-its-last-level.md.
- std-calls-undated-what-it-did-not-stat.md: the sentence that std reads
  an mtime only off SYS_FSTAT is deleted, since a served file's comes off
  fsd's stat.
- tests/common/wallclock.rs: `mtime_boot`'s tmpfs guard looked for the
  kernel's "are a tmpfs", which nothing prints here. It reads fsd's
  `storage::IN_MEMORY` line instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs
#587's contract: a file's mtime is nanoseconds since the Unix epoch, UTC,
DATA keeps the nanosecond, and 0 is undated. fsd stamped DATA with
`clock_epoch()` times 10^9, whole seconds, so two writes to `/home`
inside one second carried one mtime. It also read the clock a second
time for FAT (`utc_secs`), with the same undated rule in another unit.

- `fsd::volume::now_nanos` is the one reader. It derives the kernel's
  anchor exactly: `SYS_CLOCK_EPOCH` is `BOOT_SECS` plus the whole seconds
  of the counter at the call, and the process's clock page is the same
  counter on the kernel's formula, so a call bracketed by two readings
  inside one second names `BOOT_SECS`. A bracket that straddles a second
  is asked again, and three that straddle panic by name. The anchor is
  taken once: the kernel sets it once, before the first process. A
  refused clock is undated, 0.
- FAT's volume takes the same clock and divides it, as main's
  `stamp(mtime_now())` did. `utc_secs` is deleted.
- `file_mtime` judges `/home` too: two writes, the second strictly
  later, and not both whole seconds. `file_mtime undated` writes `/home`
  beside `/tmp` and requires both undated.
- The host test `the_anchor_is_the_kernels` holds the derivation to the
  kernel's arithmetic, a straddled bracket included.
- userlands-wall-clock-...md: the sentence that a file server can stamp
  only whole seconds is deleted; fsd stamps nanoseconds now.
- home_overwrite_zero.rs: the comment that `fs::metadata` is the kernel's
  `file_cache::size` is deleted; `/home` is fsd's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs
@Japabu

Japabu commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

Mutation patches for round 19, head 5337e791. Each passes git apply --check, the tree builds under it (cargo run -- --build-only EXIT=0; anchor-no-retry is measured by the fsd host suite instead), and git apply -R leaves the tree clean (arm.sh, record arms.out). No guest was booted: every guest arm below is owed to the orchestrator.

home-whole-seconds.patch: DATA stamps whole seconds, as `clock_epoch()·10⁹` did. Expect file_mtime (Fast) EXIT=1

Expected: file_mtime exits 1 with "on /home a write after another is stamped N ns and the one before it N ns" when both writes land in one second, and otherwise with "/home stamps … whole seconds, not the nanosecond DATA keeps".

diff --git a/userland/fsd/src/volume.rs b/userland/fsd/src/volume.rs
index abd8a5215..76aeb6704 100644
--- a/userland/fsd/src/volume.rs
+++ b/userland/fsd/src/volume.rs
@@ -155,7 +155,7 @@ pub fn now_nanos() -> u64 {
     let anchor = ANCHOR.get_or_init(|| {
         anchor(|| (nanos_since_boot(), toyos_abi::syscall::clock_epoch(), nanos_since_boot()))
     });
-    anchor.map_or(0, |secs| secs.saturating_mul(NANOS_PER_SEC).saturating_add(nanos_since_boot()))
+    anchor.map_or(0, |secs| secs.saturating_mul(NANOS_PER_SEC).saturating_add(nanos_since_boot() / NANOS_PER_SEC * NANOS_PER_SEC))
 }
 
 /// The Unix second at the counter's nanosecond zero, the kernel's `BOOT_SECS`,
home-undated-boot-nanos.patch: fsd's undated case answers nanoseconds since boot. Expect --nightly file_mtime_undated EXIT=1

Expected: file_mtime_undated exits 1. The guest panics with "/home/file-mtime-undated was written on a machine whose RTC never answered, and its mtime reads Ok(…)", after /tmp's line has printed.

diff --git a/userland/fsd/src/volume.rs b/userland/fsd/src/volume.rs
index abd8a5215..16ee85b4b 100644
--- a/userland/fsd/src/volume.rs
+++ b/userland/fsd/src/volume.rs
@@ -155,7 +155,7 @@ pub fn now_nanos() -> u64 {
     let anchor = ANCHOR.get_or_init(|| {
         anchor(|| (nanos_since_boot(), toyos_abi::syscall::clock_epoch(), nanos_since_boot()))
     });
-    anchor.map_or(0, |secs| secs.saturating_mul(NANOS_PER_SEC).saturating_add(nanos_since_boot()))
+    anchor.map_or(nanos_since_boot(), |secs| secs.saturating_mul(NANOS_PER_SEC).saturating_add(nanos_since_boot()))
 }
 
 /// The Unix second at the counter's nanosecond zero, the kernel's `BOOT_SECS`,
site-114.patch: create with truncation stamps `nanos_since_boot()` (`ops.rs:115`). Expect file_mtime EXIT=1

Expected: file_mtime exits 1 with "/tmp/file-mtime-truncated is stamped N ns by a create with truncation". The earlier write_judged arms pass, because their write stamps mtime_now after the create.

diff --git a/kernel/src/object/ops.rs b/kernel/src/object/ops.rs
index 991100e4b..f220cf71c 100644
--- a/kernel/src/object/ops.rs
+++ b/kernel/src/object/ops.rs
@@ -112,7 +112,7 @@ pub fn open(table: &mut HandleTable, path: &str, flags: OpenFlags) -> u64 {
         }
 
         let built = if truncate && create {
-            let mtime = crate::clock::mtime_now();
+            let mtime = crate::clock::nanos_since_boot();
             // `NotFound` is not a failure: truncating past a name that was not there is fine.
             // Any `vfs.delete` error other than `NotFound` is propagated, not swallowed: truncating past it could silently create a file over one the mount could not confirm was missing.
             match vfs.delete(target.as_str()) {
m132-ops.patch: create of a missing file stamps `nanos_since_boot()` (`ops.rs:133`). Expect file_mtime EXIT=1

Expected: file_mtime exits 1 with "/tmp/file-mtime-missing is stamped N ns by a create of a missing file".

diff --git a/kernel/src/object/ops.rs b/kernel/src/object/ops.rs
index 991100e4b..c4338d31d 100644
--- a/kernel/src/object/ops.rs
+++ b/kernel/src/object/ops.rs
@@ -130,7 +130,7 @@ pub fn open(table: &mut HandleTable, path: &str, flags: OpenFlags) -> u64 {
                     (file_id, mtime, position)
                 }),
                 Err(SyscallError::NotFound) if create => {
-                    let mtime = crate::clock::mtime_now();
+                    let mtime = crate::clock::nanos_since_boot();
                     vfs.create_file(target.as_str(), mtime).map(|file_id| (file_id, mtime, 0))
                 }
                 Err(e) => Err(e),
m821-ops.patch: `ftruncate` stamps `nanos_since_boot()` (`ops.rs:779`). Expect file_mtime EXIT=1

Expected: file_mtime exits 1 with "/tmp/file-mtime-truncated is stamped N ns by a truncation".

diff --git a/kernel/src/object/ops.rs b/kernel/src/object/ops.rs
index 991100e4b..d179f040d 100644
--- a/kernel/src/object/ops.rs
+++ b/kernel/src/object/ops.rs
@@ -776,7 +776,7 @@ pub fn ftruncate(object: &KObjectRef, size: u64) -> u64 {
     // The seek pointer is not touched (POSIX ftruncate): a shrink leaves it past EOF.
     file.with(|state| {
         file_cache::set_size(state.file_id, size);
-        state.mtime = crate::clock::mtime_now();
+        state.mtime = crate::clock::nanos_since_boot();
         file_cache::touch(state.file_id, state.mtime);
         0
     })
m926-fat-seconds.patch: fsd's FAT answers its whole seconds unconverted (`fat.rs:211`, `:218`, `:232`, `:309`). Expect file_mtime EXIT=1

Expected: file_mtime exits 1 in its /log arm with "/log/file-mtime reads back N ns after a reopen, and the wall clock read … s before its write and … s after", N being about 2·10⁹.

diff --git a/userland/fsd/src/fat.rs b/userland/fsd/src/fat.rs
index 1e39a13a4..e809952dd 100644
--- a/userland/fsd/src/fat.rs
+++ b/userland/fsd/src/fat.rs
@@ -208,14 +208,14 @@ impl<D: Disk> Volume for FatVolume<D> {
         if let Some(open) = self.by_path.get(path).and_then(|n| self.open.get(n)) {
             let len = open.file.len();
             let mtime = self.fs.metadata(path).map(|m| m.modified_unix).unwrap_or(0);
-            return Ok(Meta { kind: Kind::File, size: len, mtime: mtime * 1_000_000_000 });
+            return Ok(Meta { kind: Kind::File, size: len, mtime: mtime });
         }
         let meta = self.fs.metadata(path).map_err(|e| match e {
             Error::NotADirectory => SyscallError::NotFound,
             e => logged("metadata", path, e),
         })?;
         let kind = if meta.is_dir { Kind::Dir } else { Kind::File };
-        Ok(Meta { kind, size: if meta.is_dir { 0 } else { meta.len }, mtime: meta.modified_unix * 1_000_000_000 })
+        Ok(Meta { kind, size: if meta.is_dir { 0 } else { meta.len }, mtime: meta.modified_unix })
     }
 
     fn read_link(&mut self, path: &str) -> Result<String, SyscallError> {
@@ -229,7 +229,7 @@ impl<D: Disk> Volume for FatVolume<D> {
             .into_iter()
             .map(|e| {
                 let kind = if e.is_dir { Kind::Dir } else { Kind::File };
-                (e.name, Meta { kind, size: if e.is_dir { 0 } else { e.len }, mtime: e.modified_unix * 1_000_000_000 })
+                (e.name, Meta { kind, size: if e.is_dir { 0 } else { e.len }, mtime: e.modified_unix })
             })
             .collect())
     }
@@ -306,7 +306,7 @@ impl<D: Disk> Volume for FatVolume<D> {
         let open = self.entry(node)?;
         let (path, size) = (open.path.clone(), open.file.len());
         let mtime = self.fs.metadata(&path).map(|m| m.modified_unix).unwrap_or(0);
-        Ok(Meta { kind: Kind::File, size, mtime: mtime * 1_000_000_000 })
+        Ok(Meta { kind: Kind::File, size, mtime: mtime })
     }
 
     fn read(&mut self, node: Node, offset: u64, out: &mut dyn Out) -> Result<usize, SyscallError> {
utc-plus-7200.patch: fsd's one wall-clock reader two hours east, the zone recovery #583 removed. Expect --weekly wall_clock_utc EXIT=1

Expected: wall_clock_utc exits 1 with "this boot's FAT timestamp is 72NNs from the staged instant". The test stages UTC+2 in the firmware, so +7200 s is exactly the stamp the zone recovery gave. The independent oracle is the host's -rtc base= instant, read back by the host's FAT reader (volumes::root_entries).

diff --git a/userland/fsd/src/volume.rs b/userland/fsd/src/volume.rs
index abd8a5215..c6e5a4264 100644
--- a/userland/fsd/src/volume.rs
+++ b/userland/fsd/src/volume.rs
@@ -153,7 +153,7 @@ pub const NANOS_PER_SEC: u64 = 1_000_000_000;
 pub fn now_nanos() -> u64 {
     static ANCHOR: OnceLock<Option<u64>> = OnceLock::new();
     let anchor = ANCHOR.get_or_init(|| {
-        anchor(|| (nanos_since_boot(), toyos_abi::syscall::clock_epoch(), nanos_since_boot()))
+        anchor(|| (nanos_since_boot(), toyos_abi::syscall::clock_epoch().map(|s| s + 7200), nanos_since_boot()))
     });
     anchor.map_or(0, |secs| secs.saturating_mul(NANOS_PER_SEC).saturating_add(nanos_since_boot()))
 }
b3-no-stop-sync.patch: init's stop syncs no file server. Expect --nightly home_overwrite_reads_back EXIT=1

Expected: home_overwrite_reads_back exits 1 with "reading home/overwrite-pinned.bin off the image: NotFound". sync_files is #[allow(dead_code)] because nothing else reads it.

diff --git a/userland/init/src/main.rs b/userland/init/src/main.rs
index 7fa2469fc..9ebf8152d 100644
--- a/userland/init/src/main.rs
+++ b/userland/init/src/main.rs
@@ -1231,7 +1231,6 @@ impl<'a> Init<'a> {
         };
         say!("{STOPPING} ({how:?})");
         self.log.flush();
-        self.sync_files();
         let refused = match how {
             Stop::Reboot => self.syscap.reboot(),
             Stop::Shutdown => self.syscap.shutdown(),
@@ -1245,6 +1244,7 @@ impl<'a> Init<'a> {
     /// stop waits for no process, so what a server holds and has not written
     /// is lost unless it is asked first. After `logd`'s flush, so the log's
     /// server writes what logd wrote last.
+    #[allow(dead_code)]
     fn sync_files(&self) {
         let roles = self.system.programs.iter().flat_map(|p| p.roles.iter());
         let mut syncing = Vec::new();
ipc-class-reverted.patch: `kernel/src/object/ops.rs` and `kernel/src/syscall/io.rs` as they were at 06c6195, the whole class change (`git diff 0baa8b3 06c6195` over the two files). Expect process_stats (Fast) EXIT=1

Expected: process_stats exits 1 in a_wait_on_a_connection_is_ipc with "a child that parked reading a connection charged 0 ns to ipc and N ns to pipe".

diff --git a/kernel/src/object/ops.rs b/kernel/src/object/ops.rs
index 68d8d6bca..6786ec807 100644
--- a/kernel/src/object/ops.rs
+++ b/kernel/src/object/ops.rs
@@ -9,7 +9,6 @@ use alloc::vec::Vec;
 
 use toyos_abi::handle::{RawHandle, Rights};
 use toyos_abi::syscall::{FileType, OpenFlags, SeekFrom, SyscallError};
-use toyos_sched::task::WaitClass;
 
 use crate::drivers::serial;
 use crate::file_cache;
@@ -203,19 +202,9 @@ pub fn close_all(table: &mut HandleTable) {
 }
 
 pub fn pipe_id_read(object: &KObjectRef) -> Option<PipeId> {
-    pipe_read(object).map(|(id, _)| id)
-}
-
-pub fn pipe_id_write(object: &KObjectRef) -> Option<PipeId> {
-    pipe_write(object).map(|(id, _)| id)
-}
-
-/// The pipe a blocking read of `object` parks on, and the class its wait is
-/// charged to: a connection's is its peer's answer, which is IPC.
-pub fn pipe_read(object: &KObjectRef) -> Option<(PipeId, WaitClass)> {
     match object {
-        KObjectRef::PipeRead(r) => Some((r.id(), WaitClass::Pipe)),
-        KObjectRef::Connection(c) => Some((c.rx(), WaitClass::Ipc)),
+        KObjectRef::PipeRead(r) => Some(r.id()),
+        KObjectRef::Connection(c) => Some(c.rx()),
         KObjectRef::PipeWrite(_) | KObjectRef::File(_) | KObjectRef::Device(_)
         | KObjectRef::Console(_) | KObjectRef::Acceptor(_) | KObjectRef::Inbox(_)
         | KObjectRef::SysCap(_)
@@ -224,11 +213,10 @@ pub fn pipe_read(object: &KObjectRef) -> Option<(PipeId, WaitClass)> {
     }
 }
 
-/// [`pipe_read`]'s answer for a blocking write.
-pub fn pipe_write(object: &KObjectRef) -> Option<(PipeId, WaitClass)> {
+pub fn pipe_id_write(object: &KObjectRef) -> Option<PipeId> {
     match object {
-        KObjectRef::PipeWrite(w) => Some((w.id(), WaitClass::Pipe)),
-        KObjectRef::Connection(c) => Some((c.tx(), WaitClass::Ipc)),
+        KObjectRef::PipeWrite(w) => Some(w.id()),
+        KObjectRef::Connection(c) => Some(c.tx()),
         KObjectRef::PipeRead(_) | KObjectRef::File(_) | KObjectRef::Device(_)
         | KObjectRef::Console(_) | KObjectRef::Acceptor(_) | KObjectRef::Inbox(_)
         | KObjectRef::SysCap(_)
diff --git a/kernel/src/syscall/io.rs b/kernel/src/syscall/io.rs
index ad5c6d25f..20560cc0c 100644
--- a/kernel/src/syscall/io.rs
+++ b/kernel/src/syscall/io.rs
@@ -21,7 +21,7 @@ use super::handles::with_object_ref;
 
 /// What `sys_write` does when the object took nothing.
 enum WriteBlock {
-    Pipe(pipe::PipeId, WaitClass),
+    Pipe(pipe::PipeId),
     Refused(u64),
     /// Carried out of the process's lock: `HandleError::refuse` may take the
     /// process down and cannot run under a guard.
@@ -30,7 +30,7 @@ enum WriteBlock {
 
 /// What `sys_read` parks on when the handle has nothing to give.
 enum ReadBlock {
-    Pipe(alloc::sync::Arc<crate::watch::Watch>, pipe::PipeId, WaitClass),
+    Pipe(alloc::sync::Arc<crate::watch::Watch>, pipe::PipeId),
     VirtioSound,
     Hda,
     /// A claimed keyboard, woken by its own IRQ.
@@ -55,8 +55,8 @@ pub(super) fn sys_write(h: RawHandle, buf: &UserBytes) -> u64 {
             };
             match ops::try_write(object, buf) {
                 Some(n) => Ok((n, ops::pipe_id_write(object))),
-                None => Err(match ops::pipe_write(object) {
-                    Some((id, class)) => WriteBlock::Pipe(id, class),
+                None => Err(match ops::pipe_id_write(object) {
+                    Some(id) => WriteBlock::Pipe(id),
                     None => WriteBlock::Refused(SyscallError::NotFound.to_u64()),
                 }),
             }
@@ -66,14 +66,14 @@ pub(super) fn sys_write(h: RawHandle, buf: &UserBytes) -> u64 {
                 if let Some(id) = pipe_id { process::wake_pipe_readers(id); }
                 return n;
             }
-            Err(WriteBlock::Pipe(id, class)) => match pipe::write_watch(id) {
+            Err(WriteBlock::Pipe(id)) => match pipe::write_watch(id) {
                 Some(end) => {
                     let parkable = crate::scheduler::Parkable::at_entry();
                     if watch::wait_until(
                         &parkable,
                         &end,
                         0,
-                        class,
+                        WaitClass::Pipe,
                         Deadline::never(),
                         || pipe::has_space(id),
                     )
@@ -113,8 +113,8 @@ fn read_block(object: &KObjectRef) -> ReadBlock {
             );
             ReadBlock::Console(Deadline::at(crate::clock::now() + CONSOLE_REPOLL.duration()))
         }
-        _ => match ops::pipe_read(object).and_then(|(id, class)| {
-            pipe::read_watch(id).map(|end| ReadBlock::Pipe(end, id, class))
+        _ => match ops::pipe_id_read(object).and_then(|id| {
+            pipe::read_watch(id).map(|end| ReadBlock::Pipe(end, id))
         }) {
             Some(block) => block,
             None => ReadBlock::Refused(SyscallError::NotFound.to_u64()),
@@ -153,13 +153,13 @@ pub(super) fn sys_read(h: RawHandle, buf: &mut UserBytesMut) -> u64 {
                 if let Some(id) = pipe_id { process::wake_pipe_writers(id); }
                 return n;
             }
-            Err(ReadBlock::Pipe(end, id, class)) => {
+            Err(ReadBlock::Pipe(end, id)) => {
                 let parkable = crate::scheduler::Parkable::at_entry();
                 if watch::wait_until(
                     &parkable,
                     &end,
                     0,
-                    class,
+                    WaitClass::Pipe,
                     Deadline::never(),
                     || pipe::has_data(id),
                 )
write-class-pipe.patch: a blocked write to a connection is charged to `WaitClass::Pipe`, and nothing else changes. Expect process_stats (Fast) EXIT=1

Expected: process_stats exits 1 in a_wait_to_write_a_full_connection_is_ipc with "a child that parked writing a full connection charged 0 ns to ipc and N ns to pipe". The read arm's "a connection's wait: ok" line prints before it.

diff --git a/kernel/src/object/ops.rs b/kernel/src/object/ops.rs
index 991100e4b..cf7c2d822 100644
--- a/kernel/src/object/ops.rs
+++ b/kernel/src/object/ops.rs
@@ -228,7 +228,7 @@ pub fn pipe_read(object: &KObjectRef) -> Option<(PipeId, WaitClass)> {
 pub fn pipe_write(object: &KObjectRef) -> Option<(PipeId, WaitClass)> {
     match object {
         KObjectRef::PipeWrite(w) => Some((w.id(), WaitClass::Pipe)),
-        KObjectRef::Connection(c) => Some((c.tx(), WaitClass::Ipc)),
+        KObjectRef::Connection(c) => Some((c.tx(), WaitClass::Pipe)),
         KObjectRef::PipeRead(_) | KObjectRef::File(_) | KObjectRef::Device(_)
         | KObjectRef::Console(_) | KObjectRef::Acceptor(_) | KObjectRef::Inbox(_)
         | KObjectRef::SysCap(_)
anchor-no-retry.patch: the anchor taken from a bracket that straddled a second. Host: fsd EXIT=101, measured

Measured at 5337e79: cargo test --manifest-path userland/fsd/Cargo.toml --target aarch64-apple-darwin EXIT=101. the_anchor_is_the_kernels fails with left: Some(2000000001) against right: Some(2000000000), and a_clock_that_always_straddles_is_refused_by_name fails with "test did not panic as expected".

diff --git a/userland/fsd/src/volume.rs b/userland/fsd/src/volume.rs
index abd8a5215..12e9c65b5 100644
--- a/userland/fsd/src/volume.rs
+++ b/userland/fsd/src/volume.rs
@@ -167,9 +167,8 @@ fn anchor(mut read: impl FnMut() -> (u64, Option<u64>, u64)) -> Option<u64> {
     for _ in 0..3 {
         let (before, epoch, after) = read();
         let epoch = epoch?;
-        if before / NANOS_PER_SEC == after / NANOS_PER_SEC {
-            return Some(epoch - before / NANOS_PER_SEC);
-        }
+        let _ = after;
+        return Some(epoch - before / NANOS_PER_SEC);
     }
     panic!("fsd: three calls to the wall clock each straddled a second of the counter");
 }
fsd-clock-reverted.patch: the negative control, `userland/fsd/src` as at 27b5641 (`git diff 5337e79 27b5641 -- userland/fsd/src`). Expect file_mtime (Fast) EXIT=1

Expected: file_mtime exits 1 on /home with "on /home a write after another is stamped N ns and the one before it N ns", or "/home stamps … whole seconds". This reverts the whole of 5337e79's fsd change, and its host tests with it.

diff --git a/userland/fsd/src/fat.rs b/userland/fsd/src/fat.rs
index 1e39a13a4..eeff485a9 100644
--- a/userland/fsd/src/fat.rs
+++ b/userland/fsd/src/fat.rs
@@ -25,7 +25,7 @@ use toyos_fat32::{BlockAccess, Error, Fat32, FatTime, IoError};
 
 use crate::cache::Cache;
 use crate::disk::{Disk, BLOCK};
-use crate::volume::{parent, Kind, Meta, Node, OpenHow, Out, Volume, NANOS_PER_SEC};
+use crate::volume::{parent, Kind, Meta, Node, OpenHow, Out, Volume};
 
 /// The most entries one directory listing materialises.
 const MAX_LIST: usize = 16_384;
@@ -93,9 +93,7 @@ pub struct FatVolume<D: Disk> {
     open: BTreeMap<Node, Open>,
     by_path: BTreeMap<String, Node>,
     next: Node,
-    /// What this volume stamps its entries with, [`crate::volume::now_nanos`]'s
-    /// unit. FAT specifies local time; this stamps UTC because the owner ruled
-    /// the hardware clock is UTC.
+    /// Seconds since the epoch this volume stamps its entries with.
     clock: fn() -> u64,
     /// Where a read lands before it goes out: `toyos-fat32` reads into a
     /// slice, and a client's window is never one. Kept, so a read allocates
@@ -154,7 +152,7 @@ impl<D: Disk> FatVolume<D> {
     }
 
     fn time(&self) -> FatTime {
-        FatTime::from_unix_secs((self.clock)() / NANOS_PER_SEC)
+        FatTime::from_unix_secs((self.clock)())
     }
 
     fn entry(&mut self, node: Node) -> Result<&mut Open, SyscallError> {
@@ -513,7 +511,7 @@ mod tests {
         let mut ram = Ram::new((blocks + 2 * CLEAN_LIMIT) as u64);
         ram.write(0, &image).unwrap();
         let refused = Rc::new(Cell::new(u64::MAX));
-        let v = FatVolume::mount(Refusing { ram, refused: Rc::clone(&refused) }, true, || 1_717_245_296 * NANOS_PER_SEC).unwrap();
+        let v = FatVolume::mount(Refusing { ram, refused: Rc::clone(&refused) }, true, || 1_717_245_296).unwrap();
         (v, refused, spec)
     }
 
diff --git a/userland/fsd/src/main.rs b/userland/fsd/src/main.rs
index 703c51215..1cd9cb009 100644
--- a/userland/fsd/src/main.rs
+++ b/userland/fsd/src/main.rs
@@ -166,6 +166,12 @@ struct Stream {
     offset: u64,
 }
 
+/// Seconds since the epoch, UTC. FAT specifies local time; this stamps UTC
+/// because the owner ruled the hardware clock is UTC.
+fn utc_secs() -> u64 {
+    toyos_abi::syscall::clock_epoch().unwrap_or(0)
+}
+
 fn main() {
     let args: Vec<String> = std::env::args().collect();
     let role = args.get(1).and_then(|r| Role::parse(r)).unwrap_or_else(|| {
@@ -362,7 +368,7 @@ fn ram(roots: &[&str], why: &str) -> Box<dyn Volume> {
 }
 
 fn fat_on<D: Disk + 'static>(disk: D, writable: bool) -> Result<Box<dyn Volume>, String> {
-    FatVolume::mount(disk, writable, fsd::volume::now_nanos).map(|v| Box::new(v) as Box<dyn Volume>)
+    FatVolume::mount(disk, writable, utc_secs).map(|v| Box::new(v) as Box<dyn Volume>)
 }
 
 struct Server {
diff --git a/userland/fsd/src/volume.rs b/userland/fsd/src/volume.rs
index abd8a5215..d084a4123 100644
--- a/userland/fsd/src/volume.rs
+++ b/userland/fsd/src/volume.rs
@@ -7,9 +7,6 @@
 //! node whose file was unlinked or renamed over answers `Gone` from then on:
 //! its blocks may be another file's.
 
-use std::sync::OnceLock;
-
-use toyos_abi::clock::nanos_since_boot;
 use toyos_abi::syscall::SyscallError;
 
 #[derive(Clone, Copy, Debug, PartialEq, Eq)]
@@ -145,64 +142,7 @@ pub fn join(dir: &str, name: &str) -> String {
     }
 }
 
-pub const NANOS_PER_SEC: u64 = 1_000_000_000;
-
-/// What a file written now is stamped with (`toyos_abi::syscall::Stat::mtime`):
-/// nanoseconds since the Unix epoch, UTC, as the kernel's `clock::mtime_now`
-/// reckons them, and 0 — undated — on a machine whose RTC never answered.
+/// Nanoseconds since the Unix epoch, now, to the second the clock reads.
 pub fn now_nanos() -> u64 {
-    static ANCHOR: OnceLock<Option<u64>> = OnceLock::new();
-    let anchor = ANCHOR.get_or_init(|| {
-        anchor(|| (nanos_since_boot(), toyos_abi::syscall::clock_epoch(), nanos_since_boot()))
-    });
-    anchor.map_or(0, |secs| secs.saturating_mul(NANOS_PER_SEC).saturating_add(nanos_since_boot()))
-}
-
-/// The Unix second at the counter's nanosecond zero, the kernel's `BOOT_SECS`,
-/// out of `read`'s counter reading, `SYS_CLOCK_EPOCH` and counter reading.
-/// The epoch is the anchor plus the whole seconds of the counter at the call,
-/// so readings on either side within one second name it exactly. `None` is the
-/// clock refused.
-fn anchor(mut read: impl FnMut() -> (u64, Option<u64>, u64)) -> Option<u64> {
-    for _ in 0..3 {
-        let (before, epoch, after) = read();
-        let epoch = epoch?;
-        if before / NANOS_PER_SEC == after / NANOS_PER_SEC {
-            return Some(epoch - before / NANOS_PER_SEC);
-        }
-    }
-    panic!("fsd: three calls to the wall clock each straddled a second of the counter");
-}
-
-#[cfg(test)]
-mod tests {
-    use super::*;
-
-    /// The kernel's arithmetic, `BOOT_SECS + ⌊n/10⁹⌋` at the counter reading
-    /// `n` its call took: the anchor comes back exactly, and a call whose
-    /// readings straddle a second is asked again rather than guessed at.
-    #[test]
-    fn the_anchor_is_the_kernels() {
-        const BOOT_SECS: u64 = 2_000_000_000;
-        let kernel = |n: u64| Some(BOOT_SECS + n / NANOS_PER_SEC);
-
-        assert_eq!(anchor(|| (4_200_000_000, kernel(4_500_000_000), 4_700_000_000)), Some(BOOT_SECS));
-
-        let mut calls = [(4_900_000_000, 5_000_000_100, 5_100_000_000), (5_200_000_000, 5_250_000_000, 5_300_000_000)]
-            .into_iter();
-        let straddled = anchor(|| {
-            let (before, at, after) = calls.next().expect("asked at most twice");
-            (before, kernel(at), after)
-        });
-        assert_eq!(straddled, Some(BOOT_SECS));
-        assert_eq!(calls.next(), None, "the straddled call was asked again");
-
-        assert_eq!(anchor(|| (1, None, 2)), None);
-    }
-
-    #[test]
-    #[should_panic(expected = "each straddled a second")]
-    fn a_clock_that_always_straddles_is_refused_by_name() {
-        anchor(|| (999_999_999, Some(1), 1_000_000_000));
-    }
+    toyos_abi::syscall::clock_epoch().map_or(0, |secs| secs.saturating_mul(1_000_000_000))
 }

🤖 Generated with Claude Code

https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs

@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 20, at 5337e79 (5337e79 and the merge 5d4174f of 8ee3c51)

Gate.

  • CI: host COMPLETED SUCCESS at 5337e79 (run 36575862707).
  • The orchestrator's guest runs at 5337e79 (536r19-*.log): file_mtime 0; file_mtime_undated 0; file_mtime_survives_a_reboot 0 (the device holds 1993799667562105424 ns); wall_clock_utc 0; home_overwrite_reads_back 0; process_stats 0; Fast 0; nightly 380/380.
  • The T14 at this head has not run.

Net lines (git diff --shortstat origin/main...HEAD, base 8ee3c51): 285 files, +10605 −10158.

  • Production is +1023: kernel −4391, userland +4426, other +988. That is +50 since r19: the anchor and its in-crate tests. Accepted.
  • tests/ −212, crate test directories −311, src/ −1, issues/ −52.

Earlier findings (r19)

  • BLOCKER 1, composition with 8ee3c51 and A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587's contract: CLOSED.
    • 5d4174f merges 8ee3c51.
    • fsd-clock-reverted and home-whole-seconds are EXIT=1 with "on /home a write after another is stamped 1790689541000000000 ns and the one before it 1790689541000000000 ns".
    • site-114, m132-ops and m821-ops are each EXIT=1.
    • m926-fat-seconds is EXIT=1 with "/log/file-mtime reads back 1790689672 ns after a reopen".
    • home-undated-boot-nanos is EXIT=1 with "/home/file-mtime-undated … reads Ok(SystemTime(814.792912ms))".
    • anchor-no-retry is EXIT=101 on the host.
    • main has since moved: see the NOTE on issues: the kernel's drivers and threads, and the host binaries the build rests on, each with its owner #607.
  • Owed, the stop's sync: CLOSED. b3-no-stop-sync EXIT=1, "reading home/overwrite-pinned.bin off the image: NotFound".
  • Owed, a connection's wait is IPC: CLOSED. ipc-class-reverted and write-class-pipe are each EXIT=1, "charged 0 ns to ipc".
  • Owed, fsd stamps FAT in UTC: CLOSED. utc-plus-7200 EXIT=1, "this boot's FAT timestamp is 7201s from the staged instant".
  • Owed, the arms' patches as diffs: CLOSED. issuecomment-5891368856 holds all twelve.
  • NOTE, utc_secs a second reader: CLOSED. It is deleted, and FAT divides now_nanos.
  • NOTE, Triage main's T14 metal reds: issues only, no redlist rows #615's issue: CLOSED at the merge. The deletion rests on a T14 reading; see "The anchor, and the T14" below.
  • REMOVEs: CLOSED.

The merge 5d4174f: every conflict hunk is accounted for

  • fat32_adapter.rs. A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587 made four changes there: stamp/now(), file_mtime·10⁹, create stamping the VFS's mtime, and the flush's own instant. Each has its fsd counterpart:
    • fat.rs:157 is time() as now_nanos()/10⁹, and an undated 0 clamps to FatTime::EPOCH as main's did;
    • :211, :218, :232 and :309 answer seconds·10⁹;
    • :257 stamps its own instant at create;
    • level (:170) stamps the flush's instant.
  • Kernel stamps. ops.rs:115, :133, :536 and :779, and revoke_selftest.rs:21, are mtime_now. The vfs and leak_selftest hunks leave with the code this branch deletes. clock::NANOS_PER_SEC is private, and nothing outside clock.rs names it.
  • Issues.
  • rust. c4c65e3e87a contains 90697f1401a and is on origin/wt-toyos-fsd.

The anchor, and the T14

  • The derivation is exact.
    • The kernel answers SYS_CLOCK_EPOCH as ⌊(BOOT_SECS·10⁹ + n)/10⁹⌋ = BOOT_SECS + ⌊n/10⁹⌋.
    • toyos_abi::clock::nanos_between is the kernel's clock::nanos_since_boot formula term for term.
    • So for before ≤ n ≤ after with ⌊before/10⁹⌋ = ⌊after/10⁹⌋, epoch − ⌊before/10⁹⌋ is BOOT_SECS.
  • It cannot panic on a correct machine.
    • Consecutive brackets are disjoint, so three straddles need three distinct second boundaries.
    • A retry begins microseconds past the boundary its predecessor straddled. So brackets two and three must each stay open about a whole second around a counter read and a syscall that does not block. That means fsd is held off-CPU for about a second, twice in a row, inside microsecond windows.
    • The bound is therefore no timing assumption. A machine that does that is not serving anyway.
  • A named panic is the right refusal.
    • Answering 0 would claim "undated" on a dated machine.
    • Answering clock_epoch()·10⁹ is the whole-second path this round deleted.
    • An unbounded loop never fails loudly.
    • init restarts the server.
  • The T14: landing must wait for it.

BLOCKER

  1. userland/fsd/src/fat.rs:210, :308 — self.fs.metadata(path).map(|m| m.modified_unix).unwrap_or(0) answers a refused directory read as mtime 0.
    • Under A file's mtime is wall-clock time: nanoseconds since the Unix epoch, UTC #587's contract, 0 is "undated", so a device error on /log reaches std as Unsupported and not as Io. This is a silent default hiding a failure.
    • main's FAT file_mtime refused the same read (refused(...)), so this regresses the merged tree.
    • Fix: self.fs.metadata(path).map_err(|e| logged("metadata", path, e))?.modified_unix at both sites.
    • Test: an fsd host test in fat.rs's mod tests over Refusing. The entry's block is refused past the cache, and lstat of the open file and node_meta must each answer Err(Io).
    • It must turn red with .unwrap_or(0) restored at each site.
  2. tests/toyos-rust-tests/src/bin/file_mtime.rs:104 — the /home arm judges a write alone. Each of these patches passes every test:
    --- a/userland/fsd/src/data.rs
    @@ fn truncate
             if size >= open.size {
                 open.size = size;
    -            open.mtime = now;
    --- a/userland/fsd/src/data.rs
    @@ fn open
    -                    mapped("create", path, self.fs.create(path, &[], now))?;
    +                    mapped("create", path, self.fs.create(path, &[], 0))?;

NOTE

REMOVE

  • PR body l.46–54, the merge paragraphs ("The merge of origin/main (a7cd327) is 1f96997 …" through 5d4174f). This is merge chronology, and each merge's message carries it.
  • PR body l.178–188, "Size … at b607ce7". The counts are stale and will rot.
  • issues/kernel/every-driver-is-still-in-the-kernel.md:52–55, the widening bullet ("which keeps NVMe …"). It is false once this branch deletes the kernel's NVMe driver.

SEND BACK

🤖 Generated with Claude Code

https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant