Conversation
…inned Stage 1 of the 4 KiB process-memory track. A syscall's byte buffer no longer has to be one physically contiguous run: `window` walks the range page by page into the maximal physical runs its pages sit in, pins each run under the address-space lock, and every copy through a `UserBytes`/`UserBytesMut` is cut at the runs' seams. A buffer whose two demand-paged 2 MiB windows sit in frames the pager handed out in either order was refused with `BadAddress`; it is now copied through both. The decisions are in `toyos-userbound`'s new `segment` module, pure and host-tested: `segments` is the per-4 KiB-page walk that replaces `contiguous`, and `pieces` cuts one copy's range across the runs. Its tests run a fake address space over shuffled 4 KiB frames and check every byte a read returns and every byte a write leaves against the byte a ring 3 access at that address sees. Typed values keep `is_user_object`'s refusal of a 2 MiB straddle, which `abuse_page_straddle` and `abuse_kernel_addr` assert: a 2 MiB page is one frame in order, so a typed value is one run, and anything else panics as the kernel invariant it is. Stage 3 of the track, where leaves become 4 KiB, is where that refusal goes. The guest test `user_copy_spans_windows` faults the upper of two `.bss` windows first, so their frames are out of order, and reads and writes a 64 KiB buffer across the boundary through a pipe and a file. The track issue gains its eight stages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A pipe through `std::io::pipe` with its other end on a scoped thread, in place of the ABI's pipe calls: the ToyOS std passes the caller's buffer to `SYS_READ`/`SYS_WRITE` unchanged, so the kernel copies are the same ones, and no pipe capacity on another system decides the outcome. The same file built for the host with rustc runs green there. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review of #593 at 939783c (high-risk: memory management, user-copy boundary). CI Net: +589/-204. Production is about +66 (user_ptr +18, userbound lib +5, segment +78, span -35). Tests are about +285 (segment +233, guest +100, span -48). The issue gains 34 lines. Negative control: Oracle: POSIX read/write semantics, through the same std-only source built with rustc and run on Darwin (oracle_run.log). That is independent for the byte-delivery claim only. No oracle and no test backs the pin claims (see the BLOCKERs below). user_copy_races_munmap: the PR claims nothing about it, and the branch does not touch Stage: BLOCKER
NOTE
REMOVE
SEND BACK |
…nd the pins are all-or-nothing in one tested place The window walk asks the page tables once per leaf mapping instead of once per 4 KiB page: `AddressSpace::leaf` answers a translation together with the bytes to the end of the leaf that maps it, and `segments` steps by that. A 2 MiB leaf is one lookup and one run. A typed `copy_in`/`copy_out` pins its one run as a `[Segment; 1]` and allocates nothing. A bulk window sizes its run list before the address-space lock is taken, one run per 2 MiB window it touches, so nothing is allocated under the lock; a window holding more runs than that panics as a broken kernel invariant. Pinning every run or none, and unpinning every run on drop, is `toyos_userbound::Pinned` over a `Pins` trait: the PMM in the kernel, a counter in the host tests, which assert every frame's count while pinned, zero after the copy, and zero after a refusal at every run. `munmap_reissues_second_read_window` is the parked-reader staging over a buffer whose second physical run is the one the sibling unmaps and maps again. `user_copy_spans_windows`' seam is no longer 4 KiB-aligned within its buffer. `View::pieces` keeps the bound; `toyos_userbound::pieces` no longer repeats it. The track issue cites the owner's "Start it" and the orchestrator's rulings, and loses stage 1, the optional stage 8 and the sentence restating CLAUDE.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fore it parks A kernel that refuses the buffer answers the read without parking it, and the roster wait would then spin until the harness ceiling. The wait also ends on the read's answer, and an answer before the park is the red. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review round 2 of #593 at f03f81b (high-risk: memory management, the user-copy boundary). CI Guest runs, from the orchestrator's logs at this head:
Net: +946/−227. Production is about +156: segment +121, user_ptr +47, paging +18, lib +5, span −35. Tests are about +530: segment +365, guest +213, span −48. The issue gains 33 lines. Earlier BLOCKERs
BLOCKER
NOTE
REMOVE
SEND BACK |
Stage 1 of
issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md: a user copy no longer needs its window to be one physically contiguous run.What changed and why
AddressSpace::leafreturns a translation together with the number of bytes left to the end of the leaf that maps it.toyos_userbound::segmentssteps by that count and joins leaves whose frames follow each other into maximal runs. A 2 MiB leaf is one lookup and one run. A split window's 4 KiB leaves are one lookup each. A page that is absent, or that refuses the access asked, refuses the whole window.toyos_userbound::Pinned<R, P: Pins>pins every run or none. When a run's pin is refused, it unpins the runs before it and refuses the window. On drop it unpins every run. The kernel'sPinsis the PMM'spin_range/unpin_range, andFramePins<R>isPinned<R, Pmm>, so the kernel has no pin or unpin loop of its own. The pins are taken under the address-space lock over a translation that still names each frame, as before.copy_in/copy_outallocates nothing.object_runwalks the value's one run into a[Segment; 1]and pins that.is_user_objectis unchanged, so a typed value still lies inside one 2 MiB page, andabuse_page_straddlestill asserts that refusal. Stage 3 of the track removes it.windowsizes itsVec<Segment>to one run per 2 MiB window the range touches. That count comes from the prefault loop, before the address-space lock is taken. A window is one frame in order, so it holds at most one run. More runs than windows panic in a#[cold]helper as a broken kernel invariant. With the lock held, the only allocation or free is theVecdropped on a refused pin.UserBytes/UserBytesMutare aViewof runs, an offset and a length.View::piecesholds the one bound check, andtoyos_userbound::piecesno longer repeats it. The empty window borrows an empty run list.munmaprefuses a partial size" to the orchestrator's ruling. Stage 1 is deleted (this PR is its record), and so are the optional stage 8 and the sentence restating CLAUDE.md. Stage 2's exit now requires both guest tests below to go red again under their negative controls, because each test gets two physical runs only while the allocator hands out the lowest free frame first.Tests
toyos-userbound::segment: fake address spaces with leaves of 1, 4 and 16 pages, frames shuffled.a_range_is_looked_up_once_per_leaf: lookups equal leaves touched, and 64 MiB of 2 MiB leaves from an unaligned start is 33 lookups and one run.every_run_is_pinned_until_the_window_is_dropped: every frame's pin count while pinned, including a frame mapped twice, and every count back to zero after the copy.a_run_that_cannot_be_pinned_refuses_the_window_and_leaves_nothing_pinned: refuses the pin at every run in turn. It checks that no run after the refused one is tried and that every count ends at zero. The fake panics on an unpin of a frame nothing pinned, as the PMM does.munmap_reissues_second_read_window(new, shared boot, Fast): two 2 MiB mappings placed side by side withFIXED, the upper one first. A thread parks in a pipereadwhose 4 KiB buffer is split evenly across the seam. A sibling unmaps the upper mapping, maps 2 MiB again, fills it and writes the pipe. The test asserts that the read returned the whole buffer, that the first run reached the lower mapping, and that the sibling's new mapping still holds its own byte.user_copy_spans_windows: unchanged, except that the seam now sits 1000 bytes past a 4 KiB boundary of the buffer, so a page-sized file-cache copy is cut there too.Gates, at f03f81b
cargo run -- --ci hostcargo run -- --build-onlycargo test --test toyos-build -- --list(listsFast munmap_reissues_second_read_windowandFast user_copy_spans_windows)rustc --edition 2021 -D warnings -C debug-assertions=on -C overflow-checks=on user_copy_spans_windows.rson the host, then the binaryGuest runs by the orchestrator at 939783c, the previous head:
user_copy_spans_windows0, whole-change revert with the test kept 1,abuse_page_straddle0,abuse_kernel_addr0, Fast 0. None has been run at f03f81b yet; they are owed, withmunmap_reissues_second_read_windowand the mutation arms below.High-risk checks (memory management)
Negative controls. Each patch is applied with
git apply --check, the image is built withcargo run -- --build-only,cargo test -p toyos-userboundis run, the patch is reversed, andgit diff --quietshows the tree clean (exit 0 each time). Every patch is againsttoyos-userbound/src/segment.rsat f03f81b:Pinned::pinandPinned::dropare the loopswindowandFramePins::droprun, so these are the review's mutations at their new site.take1-pin-and-drop.take(1)in the pin loop and the drop loopmunmap_reissues_second_read_windowreddrop-first-run-onlyevery_run_is_pinned…rollback-plus-one[..refused + 1], thetruncate(refused + 1)of the reviewa_run_that_cannot_be_pinned…rollback-clear[..refused][..0], theclear()of the reviewa_run_that_cannot_be_pinned…rollback-clearwas first written as[..0], but that leavesrefusedunused, and the kernel build denies warnings, so it did not build (exit 101). It was rewritten as above before its measurement.Whole change.
revert-to-main-keep-test.patchrevertskernel/src/user_ptr.rs, bothpaging.rsfiles andtoyos-userbound/to 3d90247 (the merge base withorigin/main) and keeps both guest tests. It applies at f03f81b, and the reverted tree builds (--build-onlyexit 0). The tree was then restored clean. Owed guest result: both tests red.user_copy_spans_windowshas its buffer refused.munmap_reissues_second_read_windowhas its read refused before it parks, and fails on that answer.Independent oracle. POSIX
read/writesemantics, throughuser_copy_spans_windows, which uses onlystd. On ToyOS,std::io::pipeandstd::fspass the caller's buffer straight toSYS_READ/SYS_WRITE. The same source, built for the host with therustcline above, exits 0 on Darwin 27.0.0; the file is the same at f03f81b. That backs the claim that the bytes are delivered. Nothing independent backs the pin claims: they rest on the mutations above, and on the host fake modellingpmm::pin_range's contract ("false, and nothing pinned") andunpin_range's panic.Cost, A/B against main
/private/tmp/claude-502/-Users-jan-Dev-jan-toyos/2280e09e-428b-4b81-bc00-1ede594b7247/scratchpad/mem4k-r2/bench/run.sh 3d902477 954bbce1(toyos-userbound/is the same at f03f81b) buildsbench.rswithrustc -O. The bench takesorigin/main'sspan.rs(contiguous) and this branch'ssegment.rsfrom git, unmodified. It drives both over a four-level table walked asAddressSpace::walkwalks it, with 2 MiB leaves, prefault walks on both sides and no-op pins. It reports the best of five rounds. Exit 0:copy_inof a 16-byte valuecopy_outof a 16-byte valueThe walk counts are exact. The times are this host's and vary from run to run: the same bench built from the working tree earlier gave 174.1 against 296.0 ns for the 64 MiB read, and 35814.2 against 202.4 ns for the 64 MiB write. The extra time on a branch read is the bulk window's
Vec(the host'smalloc, not the kernel'sdlmalloc), taken before the lock. The in-guest cost under the real lock and allocator has not been measured and is owed.What I am unsure of
windowstill faults in 2 MiB steps, because that is the pager's fill size.🤖 Generated with Claude Code