Skip to content

User copies are scatter-gather: a window is its physical runs, each pinned - #593

Open
Japabu wants to merge 5 commits into
mainfrom
wt/toyos-mem4k
Open

Japabu wants to merge 5 commits into
mainfrom
wt/toyos-mem4k

Conversation

@Japabu

@Japabu Japabu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Stage 1 of issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md: a user copy no longer needs its window to be one physically contiguous run.

What changed and why

  • A user window is a list of physical runs, walked once per leaf. AddressSpace::leaf returns a translation together with the number of bytes left to the end of the leaf that maps it. toyos_userbound::segments steps by that count and joins leaves whose frames follow each other into maximal runs. A 2 MiB leaf is one lookup and one run. A split window's 4 KiB leaves are one lookup each. A page that is absent, or that refuses the access asked, refuses the whole window.
  • The pins are all-or-nothing, in one place a host test can refuse. toyos_userbound::Pinned<R, P: Pins> pins every run or none. When a run's pin is refused, it unpins the runs before it and refuses the window. On drop it unpins every run. The kernel's Pins is the PMM's pin_range/unpin_range, and FramePins<R> is Pinned<R, Pmm>, so the kernel has no pin or unpin loop of its own. The pins are taken under the address-space lock over a translation that still names each frame, as before.
  • A typed copy_in/copy_out allocates nothing. object_run walks the value's one run into a [Segment; 1] and pins that. is_user_object is unchanged, so a typed value still lies inside one 2 MiB page, and abuse_page_straddle still asserts that refusal. Stage 3 of the track removes it.
  • A bulk window allocates before the lock. window sizes its Vec<Segment> to one run per 2 MiB window the range touches. That count comes from the prefault loop, before the address-space lock is taken. A window is one frame in order, so it holds at most one run. More runs than windows panic in a #[cold] helper as a broken kernel invariant. With the lock held, the only allocation or free is the Vec dropped on a refused pin.
  • Copies are cut at run seams. UserBytes/UserBytesMut are a View of runs, an offset and a length. View::pieces holds the one bound check, and toyos_userbound::pieces no longer repeats it. The empty window borrows an empty run list.
  • The track issue now says the owner approved starting this work ("Start it") and delegated every kernel-interface decision to the orchestrator. It attributes strict commit and "munmap refuses a partial size" to the orchestrator's ruling. Stage 1 is deleted (this PR is its record), and so are the optional stage 8 and the sentence restating CLAUDE.md. Stage 2's exit now requires both guest tests below to go red again under their negative controls, because each test gets two physical runs only while the allocator hands out the lowest free frame first.

Tests

  • Host, toyos-userbound::segment: fake address spaces with leaves of 1, 4 and 16 pages, frames shuffled.
    • Reads and writes land on the byte a ring 3 access sees, from any offset.
    • The runs are maximal and cover the range exactly.
    • An absent page refuses the window. A read-only page refuses a write and allows a read.
    • An empty range and a kernel-half range are refused before any walk.
    • a_range_is_looked_up_once_per_leaf: lookups equal leaves touched, and 64 MiB of 2 MiB leaves from an unaligned start is 33 lookups and one run.
    • every_run_is_pinned_until_the_window_is_dropped: every frame's pin count while pinned, including a frame mapped twice, and every count back to zero after the copy.
    • a_run_that_cannot_be_pinned_refuses_the_window_and_leaves_nothing_pinned: refuses the pin at every run in turn. It checks that no run after the refused one is tried and that every count ends at zero. The fake panics on an unpin of a frame nothing pinned, as the PMM does.
  • Guest, munmap_reissues_second_read_window (new, shared boot, Fast): two 2 MiB mappings placed side by side with FIXED, the upper one first. A thread parks in a pipe read whose 4 KiB buffer is split evenly across the seam. A sibling unmaps the upper mapping, maps 2 MiB again, fills it and writes the pipe. The test asserts that the read returned the whole buffer, that the first run reached the lower mapping, and that the sibling's new mapping still holds its own byte.
  • Guest, user_copy_spans_windows: unchanged, except that the seam now sits 1000 bytes past a 4 KiB boundary of the buffer, so a page-sized file-cache copy is cut there too.

Gates, at f03f81b

gate exit
cargo run -- --ci host 0
cargo run -- --build-only 0
cargo test --test toyos-build -- --list (lists Fast munmap_reissues_second_read_window and Fast user_copy_spans_windows) 0
rustc --edition 2021 -D warnings -C debug-assertions=on -C overflow-checks=on user_copy_spans_windows.rs on the host, then the binary 0, 0

Guest runs by the orchestrator at 939783c, the previous head: user_copy_spans_windows 0, whole-change revert with the test kept 1, abuse_page_straddle 0, abuse_kernel_addr 0, Fast 0. None has been run at f03f81b yet; they are owed, with munmap_reissues_second_read_window and the mutation arms below.

High-risk checks (memory management)

Negative controls. Each patch is applied with git apply --check, the image is built with cargo run -- --build-only, cargo test -p toyos-userbound is run, the patch is reversed, and git diff --quiet shows the tree clean (exit 0 each time). Every patch is against toyos-userbound/src/segment.rs at f03f81b: Pinned::pin and Pinned::drop are the loops window and FramePins::drop run, so these are the review's mutations at their new site.

patch mutation build host guest (owed)
take1-pin-and-drop .take(1) in the pin loop and the drop loop 0 101: both pin tests munmap_reissues_second_read_window red
drop-first-run-only drop unpins only the first run 0 101: every_run_is_pinned… none: no guest census reads pin counts
rollback-plus-one rollback unpins [..refused + 1], the truncate(refused + 1) of the review 0 101: a_run_that_cannot_be_pinned… none
rollback-clear rollback unpins [..refused][..0], the clear() of the review 0 101: a_run_that_cannot_be_pinned… none

rollback-clear was first written as [..0], but that leaves refused unused, and the kernel build denies warnings, so it did not build (exit 101). It was rewritten as above before its measurement.

Whole change. revert-to-main-keep-test.patch reverts kernel/src/user_ptr.rs, both paging.rs files and toyos-userbound/ to 3d90247 (the merge base with origin/main) and keeps both guest tests. It applies at f03f81b, and the reverted tree builds (--build-only exit 0). The tree was then restored clean. Owed guest result: both tests red. user_copy_spans_windows has its buffer refused. munmap_reissues_second_read_window has its read refused before it parks, and fails on that answer.

Independent oracle. POSIX read/write semantics, through user_copy_spans_windows, which uses only std. On ToyOS, std::io::pipe and std::fs pass the caller's buffer straight to SYS_READ/SYS_WRITE. The same source, built for the host with the rustc line above, exits 0 on Darwin 27.0.0; the file is the same at f03f81b. That backs the claim that the bytes are delivered. Nothing independent backs the pin claims: they rest on the mutations above, and on the host fake modelling pmm::pin_range's contract ("false, and nothing pinned") and unpin_range's panic.

Cost, A/B against main

/private/tmp/claude-502/-Users-jan-Dev-jan-toyos/2280e09e-428b-4b81-bc00-1ede594b7247/scratchpad/mem4k-r2/bench/run.sh 3d902477 954bbce1 (toyos-userbound/ is the same at f03f81b) builds bench.rs with rustc -O. The bench takes origin/main's span.rs (contiguous) and this branch's segment.rs from git, unmodified. It drives both over a four-level table walked as AddressSpace::walk walks it, with 2 MiB leaves, prefault walks on both sides and no-op pins. It reports the best of five rounds. Exit 0:

case main branch
read 64 MiB, frames in order 247.4 ns, 67 walks 310.5 ns, 66 walks
write 64 MiB, frames in order 114431.2 ns, 16419 walks 322.1 ns, 66 walks
read 64 KiB across one 2 MiB boundary 32.8 ns, 5 walks 55.0 ns, 4 walks
write 64 KiB across one 2 MiB boundary 121.2 ns, 19 walks 77.8 ns, 4 walks
read / write 64 MiB, one run per window refused 361.5 / 380.5 ns, 66 walks
copy_in of a 16-byte value 3.8 ns, 3 walks 1.7 ns, 2 walks
copy_out of a 16-byte value 3.5 ns, 4 walks 2.1 ns, 2 walks

The walk counts are exact. The times are this host's and vary from run to run: the same bench built from the working tree earlier gave 174.1 against 296.0 ns for the 64 MiB read, and 35814.2 against 202.4 ns for the 64 MiB write. The extra time on a branch read is the bulk window's Vec (the host's malloc, not the kernel's dlmalloc), taken before the lock. The in-guest cost under the real lock and allocator has not been measured and is owed.

What I am unsure of

  • The host fake is my model of the PMM's pin contract, not the PMM. The PMM counts pins per 2 MiB frame, and the fake counts them per 4 KiB frame.
  • Both guest tests get two physical runs only because the allocator hands out the lowest free frame first. Neither can see a physical address, so stage 2 has to re-run their negative controls.
  • The prefault loop in window still faults in 2 MiB steps, because that is the pager's fill size.

🤖 Generated with Claude Code

Japabu and others added 2 commits September 29, 2026 00:45
…inned

Stage 1 of the 4 KiB process-memory track. A syscall's byte buffer no
longer has to be one physically contiguous run: `window` walks the range
page by page into the maximal physical runs its pages sit in, pins each
run under the address-space lock, and every copy through a
`UserBytes`/`UserBytesMut` is cut at the runs' seams. A buffer whose two
demand-paged 2 MiB windows sit in frames the pager handed out in either
order was refused with `BadAddress`; it is now copied through both.

The decisions are in `toyos-userbound`'s new `segment` module, pure and
host-tested: `segments` is the per-4 KiB-page walk that replaces
`contiguous`, and `pieces` cuts one copy's range across the runs. Its
tests run a fake address space over shuffled 4 KiB frames and check every
byte a read returns and every byte a write leaves against the byte a ring
3 access at that address sees.

Typed values keep `is_user_object`'s refusal of a 2 MiB straddle, which
`abuse_page_straddle` and `abuse_kernel_addr` assert: a 2 MiB page is one
frame in order, so a typed value is one run, and anything else panics as
the kernel invariant it is. Stage 3 of the track, where leaves become
4 KiB, is where that refusal goes.

The guest test `user_copy_spans_windows` faults the upper of two `.bss`
windows first, so their frames are out of order, and reads and writes a
64 KiB buffer across the boundary through a pipe and a file.

The track issue gains its eight stages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A pipe through `std::io::pipe` with its other end on a scoped thread, in
place of the ABI's pipe calls: the ToyOS std passes the caller's buffer
to `SYS_READ`/`SYS_WRITE` unchanged, so the kernel copies are the same
ones, and no pipe capacity on another system decides the outcome. The
same file built for the host with rustc runs green there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu
Japabu marked this pull request as ready for review September 28, 2026 22:52
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review of #593 at 939783c (high-risk: memory management, user-copy boundary). CI host green at this head (run 36494917426). Guest runs are the orchestrator's: user_copy_spans_windows, abuse_page_straddle, abuse_kernel_addr and Fast all EXIT=0. The whole-revert arm is EXIT=1.

Net: +589/-204. Production is about +66 (user_ptr +18, userbound lib +5, segment +78, span -35). Tests are about +285 (segment +233, guest +100, span -48). The issue gains 34 lines.

Negative control: revert-to-main-keep-test.patch is byte-identical, index lines aside, to git diff HEAD e3a1cdc8 over every code file the branch touches. origin/main has not moved those files since the merge base. So the patch reverts the whole change onto the base, and the only thing left is issue prose. Its red is at user_copy_spans_windows.rs:64, Kind(InvalidInput). std maps both BadAddress and InvalidArgument to InvalidInput (rust/library/std/src/sys/pal/toyos/mod.rs:27-28). With a live handle, the only refusal on that SYS_READ path is dispatch's bad_addr when user_bytes_mut is refused. So by elimination it is the right cause, but the log does not show the "BadAddress" the PR predicted.

Oracle: POSIX read/write semantics, through the same std-only source built with rustc and run on Darwin (oracle_run.log). That is independent for the byte-delivery claim only. No oracle and no test backs the pin claims (see the BLOCKERs below).

user_copy_races_munmap: the PR claims nothing about it, and the branch does not touch src/redlist.rs. Its issue's exit condition requires the remap_race::hold redesign plus a one-CPU mutation that reds with a line saying so. This branch changes neither, so one green nightly run is not evidence, and the row must not move.

Stage: origin/main's track has no stages at all. This branch writes the stage list and stage 1's exit, then meets the exit it wrote.

BLOCKER

  • kernel/src/user_ptr.rs:365 — the walk is one page-table lookup per 4 KiB page, under the address-space Lock, which disables preemption while held and makes contending CPUs spin with IF clear. The user chooses the length, so a 1 GiB mapped buffer is 262144 walks where main did 512. The PR says this cost was never measured. Every copy_in/copy_out also now does a heap Vec push and free under that same lock, which the PR body calls "one allocation per bulk copy". Two fixes: have the leaf closure return the leaf's extent, so a 2 MiB leaf costs one lookup now and a 4 KiB leaf costs one at stage 3; and give object no Vec. Then measure both against main, same session: a 64 MiB pipe/file read and a copy_in loop.
  • kernel/src/user_ptr.rs:343,374 — a window of several runs is pinned across a park while a sibling unmaps, and no test covers it. This mutation passes every test: in window, replace runs.iter().position( with runs.iter().take(1).position(; in FramePins::drop, replace for run in &self.0 with for run in self.0.iter().take(1). It must turn red on a munmap_reissues_read_window variant: a parked read whose buffer spans two out-of-order windows, where B unmaps and reissues the second window.
  • kernel/src/user_ptr.rs:343 — a pin leak on the drop path is invisible. Freed frames are still counted free, but alloc_page skips them forever. The mutation for run in self.0.iter().take(1) in FramePins::drop alone passes every test. Add a test that asserts every pin is back to zero after copies of several runs: a PMM pin count in a record the harness checks.
  • kernel/src/user_ptr.rs:374 — the rollback on a refused pin is unreachable in any test. Both runs.truncate(refused + 1) (unpins a run that was never pinned) and runs.clear() (leaks the prefix's pins) pass. Two fixes. Put the all-or-nothing pin of a run list in one place that a test can refuse: either a pure helper in toyos-userbound, host-tested with a pin closure that refuses, or one pmm call under a single BITMAP hold whose rollback sits beside pin_range's own.
  • issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md:41-74 — main holds none of these stages, and nothing cites a source for the rulings the branch writes into them: strict commit, no promotion, no live split, an OWNED bit, a watermark, and munmap refusing a partial size. The last is a syscall behaviour change, and CLAUDE.md requires discussion before one. The branch also deletes "Blocked on: the owner's word" without citing that word. Either take the stage list out of this PR, or cite the owner's ruling for each stage.

NOTE

  • kernel/src/user_ptr.rs:210,221 — View::pieces asserts the bound that toyos_userbound::pieces asserts again. Keep one of them.
  • tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs:5-8 — the test cannot check that the frames really are out of order. That holds only because the PMM is ascending next-fit. Stage 2 rewrites the allocator, and the test would then pass on a one-run kernel. Stage 2 has to re-run this negative control.
  • user_copy_spans_windows.rs:83-92 — the buffer seam falls at offset 32768, which is page-aligned, so no file-cache sub-copy is ever cut at a seam. Only the pipe copies cross it.

REMOVE

  • toyos-userbound/src/segment.rs:5-7 — "two neighbouring pages sit in whichever frames the PMM handed out, in either order" is false today: a 2 MiB leaf is one frame, in order.
  • toyos-userbound/src/segment.rs:92-93 — narrates why 48 pages were chosen.
  • toyos-userbound/src/segment.rs:98 — narrates the shuffle.
  • toyos-userbound/src/segment.rs:250-251 — "as before the walk was per page" describes a past implementation.
  • tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs:5-8 — "never one physically contiguous run" is false in general, and "refuses with BadAddress" describes a past kernel.
  • tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs:83 — "a page at a time" describes the file cache, and will rot.
  • issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md:41 — "Each lands alone and leaves main whole" restates CLAUDE.md.
  • issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md:46-51 — stage 1 becomes history the moment this lands.
  • issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md:73-74 — an "Optional" stage.
  • PR body — "One allocation per bulk copy was added" is false (every typed copy allocates too).
  • PR body — "Expected result: … failing on BadAddress" does not match the log, which says InvalidInput.
  • PR body — "It was not run on Linux; this host has none."
  • PR body — "kernel-loom is not touched."
  • PR body — the user_copy_races_munmap paragraph.

SEND BACK

Japabu and others added 3 commits September 29, 2026 04:13
…nd the pins are all-or-nothing in one tested place

The window walk asks the page tables once per leaf mapping instead of once
per 4 KiB page: `AddressSpace::leaf` answers a translation together with the
bytes to the end of the leaf that maps it, and `segments` steps by that. A
2 MiB leaf is one lookup and one run.

A typed `copy_in`/`copy_out` pins its one run as a `[Segment; 1]` and
allocates nothing. A bulk window sizes its run list before the
address-space lock is taken, one run per 2 MiB window it touches, so nothing
is allocated under the lock; a window holding more runs than that panics as a
broken kernel invariant.

Pinning every run or none, and unpinning every run on drop, is
`toyos_userbound::Pinned` over a `Pins` trait: the PMM in the kernel, a
counter in the host tests, which assert every frame's count while pinned,
zero after the copy, and zero after a refusal at every run.

`munmap_reissues_second_read_window` is the parked-reader staging over a
buffer whose second physical run is the one the sibling unmaps and maps
again.

`user_copy_spans_windows`' seam is no longer 4 KiB-aligned within its buffer.
`View::pieces` keeps the bound; `toyos_userbound::pieces` no longer repeats it.

The track issue cites the owner's "Start it" and the orchestrator's rulings,
and loses stage 1, the optional stage 8 and the sentence restating CLAUDE.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fore it parks

A kernel that refuses the buffer answers the read without parking it, and
the roster wait would then spin until the harness ceiling. The wait also
ends on the read's answer, and an answer before the park is the red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 2 of #593 at f03f81b (high-risk: memory management, the user-copy boundary). CI host passed on run 36514270885, headSha f03f81b. The PR is MERGEABLE/CLEAN.

Guest runs, from the orchestrator's logs at this head: user_copy_spans_windows passed, munmap_reissues_read_window passed, and Fast passed 230 of 230. The single runs of munmap_reissues_second_read_window, abuse_page_straddle and abuse_kernel_addr do not count as reds. Each hit the first-boot budget of 20 s (10 s × width 2, qemu.rs:5071). Two of them printed nothing, so the kernel never ran. The third logged power-on to loader 16501 ms and stopped after init: started logd, before any test ran. All three pass in the Fast run.

revert-to-main-keep-test.patch is byte-identical, index lines aside, to git diff HEAD 3d902477 -- kernel toyos-userbound. origin/main has not moved those files. Under it, user_copy_spans_windows fails at :63 with InvalidInput, and munmap_reissues_second_read_window fails at :86 with "answered -1 before it parked".

Net: +946/−227. Production is about +156: segment +121, user_ptr +47, paging +18, lib +5, span −35. Tests are about +530: segment +365, guest +213, span −48. The issue gains 33 lines.

Earlier BLOCKERs

  • B1 CLOSED — leaf returns the leaf's extent, and a_range_is_looked_up_once_per_leaf holds the count. I re-ran run.sh 3d902477 954bbce1 and run.sh 3d902477 f03f81be, both exit 0. Writing 64 MiB drops from 16419 walks at 26934–27348 ns to 66 walks at 147–179 ns. copy_in drops from 1.6 ns to 0.7 ns. Typed copies go through [Segment; 1].
  • B2 CLOSED — with take1-pin-and-drop applied, munmap_reissues_second_read_window exits 101 at :106: "second run reached a sibling's reissued frame — byte 0 of it is 0xa1". The test passes in Fast at this head.
  • B3 CLOSED by the orchestrator's ruling — with drop-first-run-only applied, every_run_is_pinned_until_the_window_is_dropped fails (101), and the kernel's drop is that same generic Drop.
  • B4 CLOSED — with rollback-plus-one or rollback-clear applied, a_run_that_cannot_be_pinned… fails (101). The rollback-clear log shows a leaked pin on frame 43.
  • B5 CLOSED — the track now attributes "Start it" to the owner and the rulings to the orchestrator.

BLOCKER

  • kernel/src/arch/x86_64/paging.rs:610 — the write-rights check for a split window now depends on this 4096, and no test can fail on it. Before this round, the 4 KiB write step lived in the pure crate. Now segments steps by whatever extent walk reports. Patch: return Some((dm, rights & pte, 4096)); → return Some((dm, rights & pte, PAGE_2M));. With it, a write window is checked only at its first page, up to the end of its 2 MiB window. abuse_readonly_copyout still passes, because every arm there starts on the read-only page. No guest test starts a write on a writable page and runs into a read-only page. Such a layout exists in every binary: a mixed window's pages past the image are mapped present and Read (process.rs:1261). Add the arm and show it red under the patch: find the first page above a .bss static where a 1-byte read answers BadAddress. Assert that page is readable and lies in the same 2 MiB window. Then read(fd, page - 8, 16) must answer BadAddress, and the page's bytes must be unchanged.

NOTE

  • kernel/src/user_ptr.rs:374 — the bulk window's Vec does not block landing. My re-run shows reads at 1.28–1.31× on 64 MiB (+32 to +40 ns) and 1.75–1.83× on 64 KiB (+8 ns). Both are host malloc, paid before the lock, and small against the copy each one gates. The in-guest cost under the kernel heap's lock is owed. Record it in the track, not only in this PR body.
  • kernel/src/user_ptr.rs:365 — impl Pins for Pmm is the one pin path the host tests do not reach. fn unpin(&mut self, _run: Segment) {} leaks every copied frame's pin, and no test is known to go red under it.
  • kernel/src/user_ptr.rs:102,389 — object_run and window repeat the same fault_in → lock → segments → pin sequence and differ only in how they store runs.
  • kernel/src/arch/x86_64/paging.rs:569,577 — leaf restates translate_writable's rights test. translate_now in user_ptr.rs could become leaf(..).map(|(p, _)| p), and one of the two could go.
  • kernel/src/user_ptr.rs:128 — window_split becomes a panic userland can reach once stage 3 maps 4 KiB leaves from arbitrary frames. The stage-3 text does not name it.

REMOVE

  • PR body — "None has been run at f03f81b yet; they are owed", the "guest (owed)" column and "Owed guest result": all false now.
  • PR body — the rollback-clear "was first written as [..0]" paragraph, which is an investigation story.
  • PR body — "The extra time on a branch read is the bulk window's Vec". This attribution was never measured.
  • PR body — "the same bench built from the working tree earlier gave …", which records where a measurement came from.

SEND BACK

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant