From 8bfad2cd2a632e4bd9deebf2a4cfd0e9e3da29db Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 00:45:49 +0200 Subject: [PATCH 1/7] User copies are scatter-gather: a window is its physical runs, each pinned Stage 1 of the 4 KiB process-memory track. A syscall's byte buffer no longer has to be one physically contiguous run: `window` walks the range page by page into the maximal physical runs its pages sit in, pins each run under the address-space lock, and every copy through a `UserBytes`/`UserBytesMut` is cut at the runs' seams. A buffer whose two demand-paged 2 MiB windows sit in frames the pager handed out in either order was refused with `BadAddress`; it is now copied through both. The decisions are in `toyos-userbound`'s new `segment` module, pure and host-tested: `segments` is the per-4 KiB-page walk that replaces `contiguous`, and `pieces` cuts one copy's range across the runs. Its tests run a fake address space over shuffled 4 KiB frames and check every byte a read returns and every byte a write leaves against the byte a ring 3 access at that address sees. Typed values keep `is_user_object`'s refusal of a 2 MiB straddle, which `abuse_page_straddle` and `abuse_kernel_addr` assert: a 2 MiB page is one frame in order, so a typed value is one run, and anything else panics as the kernel invariant it is. Stage 3 of the track, where leaves become 4 KiB, is where that refusal goes. The guest test `user_copy_spans_windows` faults the upper of two `.bss` windows first, so their frames are out of order, and reads and writes a 64 KiB buffer across the boundary through a pipe and a file. The track issue gains its eight stages. Co-Authored-By: Claude Opus 5.5 --- ...b-pages-and-that-caps-the-process-count.md | 40 ++- kernel/src/user_ptr.rs | 244 +++++++------- .../src/bin/user_copy_spans_windows.rs | 96 ++++++ toyos-userbound/src/lib.rs | 13 +- toyos-userbound/src/segment.rs | 311 ++++++++++++++++++ toyos-userbound/src/span.rs | 85 +---- 6 files changed, 585 insertions(+), 204 deletions(-) create mode 100644 tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs create mode 100644 toyos-userbound/src/segment.rs diff --git a/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md b/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md index 037523116b..35af9bbf57 100644 --- a/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md +++ b/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md @@ -18,9 +18,6 @@ IOMMU's superpages, the kernel's own mappings. One paging design serves both, on x86-64 and on the ARM64 the tree is kept portable for (`issues/kernel/toyos-runs-on-arm64.md`). -**Blocked on:** the owner's word to start. Recorded, not started, at the -owner's instruction. - **The constraint that makes this a track, measured on the T14** (run 29 readback, `PMM: 100/16038MB used` at 11 s with four user processes and three kernel threads live; the same record's rows: stack 16 pages held, init-tls 5, @@ -38,3 +35,40 @@ longer shares a 2 MiB page with a neighbour's registers, so the relocation that `kernel/src/pcidev/` does before a hand-over is no longer needed for a window firmware placed apart from its neighbours. Placing a window inside the host bridge's firmware-reported apertures stays correct under either page size. + +## Stages + +Each lands alone and leaves `main` whole. **Commit is strict**: a mapping +reserves its frame count when it is made and is refused by name then, never +killed at a fault. No promotion (the kernel has no thread to do it) and no +live split of a user leaf. + +1. **Scatter-gather user copies.** A user window is the physical runs its + pages sit in, each pinned; a buffer's copy is cut at their seams instead of + refused. Exit: `toyos-userbound`'s segment tests over shuffled 4 KiB frames + and the guest `user_copy_spans_windows` green; `abuse_page_straddle` + unchanged, because a typed value still lies inside one 2 MiB page. +2. **Superframe frame allocator.** 4 KiB frames carved from 2 MiB + superframes; the pin rule held per superframe, so no pinned frame is + reissued and no superframe holding one is handed out whole; a watermark + refuses an allocation before the machine runs dry. Exit: host tests of + carve, free, pin and the refusal; every 2 MiB caller of the PMM unchanged. +3. **4 KiB user leaves.** A software `OWNED` bit in the leaf is the ledger of + the frames a process owns, with per-process counts; stacks are demand-zero + and TLS is 4 KiB. A typed syscall value crossing a page is served in + segments, so `is_user_object` stops refusing a straddle and + `abuse_page_straddle`'s refusals become delivery verdicts. Exit: a + process's floor on the T14 against the 10 to 14 MB above. +4. **Demand-zero anonymous `mmap`.** A fault installs a 2 MiB leaf only for a + whole aligned 2 MiB span of the mapping; unmapping a range no thread touched + sends no IPI; `munmap` refuses a size that is not the whole mapping. Exit: + `mmap` of more than the watermark allows refused at `mmap`, and a guest that + touches one page of a large mapping holds one frame. +5. **File-backed faults at 4 KiB.** ROOT's read-only text is shared + zero-copy. Exit: two processes of one binary hold its text once. +6. **`/apps` image sharing.** The design is the owner's decision, owed when + this stage is reached. +7. **Retire 2 MiB where it no longer pays.** Exit: every remaining 2 MiB user + leaf is a stage-4 whole span or a device or DMA grant. +8. **Optional: PCID-scoped remote flush.** Exit: a measured shootdown cost + that pays for it. diff --git a/kernel/src/user_ptr.rs b/kernel/src/user_ptr.rs index 36b4212f7d..c4d49b089a 100644 --- a/kernel/src/user_ptr.rs +++ b/kernel/src/user_ptr.rs @@ -7,6 +7,8 @@ //! for the copy's life — a typed value's as much as a [`UserBytes`]/ //! [`UserBytesMut`] buffer's — so a sibling's `munmap` cannot reissue the //! backing between the translation and the copy, nor under a copy across a park. +//! A window is the physical runs its pages sit in, and a buffer's copy is cut at +//! their seams; a typed value lies inside one 2 MiB page, so in one run. //! Every single-word user read is `read_volatile`, including the futex word //! and the crash dump's walk, both outside this module. @@ -16,6 +18,7 @@ use alloc::vec::Vec; use core::marker::PhantomData; use toyos_abi::syscall::SyscallError; +use toyos_userbound::Segment; use crate::UserAddr; @@ -90,13 +93,25 @@ pub(crate) fn translate_user(addr: UserAddr, access: Access) -> Option(ptr: UserAddr, access: Access) -> Result<(*mut T, FramePin), SyscallError> { - let size = core::mem::size_of::(); - if !toyos_userbound::is_user_object(ptr.raw(), size as u64, core::mem::align_of::() as u64) { +fn object(ptr: UserAddr, access: Access) -> Result<(*mut T, FramePins), SyscallError> { + let (kptr, pins) = object_run(ptr, core::mem::size_of::(), core::mem::align_of::(), access)?; + Ok((kptr.cast(), pins)) +} + +fn object_run(ptr: UserAddr, size: usize, align: usize, access: Access) -> Result<(*mut u8, FramePins), SyscallError> { + if !toyos_userbound::is_user_object(ptr.raw(), size as u64, align as u64) { return Err(SyscallError::BadAddress); } - let (kptr, pin) = window(ptr, size, access).ok_or(SyscallError::BadAddress)?; - Ok((kptr.cast(), pin)) + let pins = window(ptr, size, access).ok_or(SyscallError::BadAddress)?; + let [run] = pins.0[..] else { object_split(ptr, size, pins.0.len()) }; + Ok((crate::mm::DirectMap::from_phys(run.phys).as_mut_ptr(), pins)) +} + +/// A 2 MiB page is one frame in order, so an object inside one is one run. +#[cold] +#[inline(never)] +fn object_split(ptr: UserAddr, size: usize, runs: usize) -> ! { + panic!("a {size}-byte value at {:#x} inside one 2 MiB page lies in {runs} physical runs", ptr.raw()) } /// Context for a single syscall invocation; the lifetime `'a` keeps validated references from escaping it. @@ -113,24 +128,12 @@ impl<'a> SyscallContext<'a> { /// A bulk buffer the kernel reads out of and never borrows. pub fn user_bytes(&self, ptr: UserAddr, len: u64) -> Option> { - let len = len as usize; - if len == 0 { - let kptr = core::ptr::NonNull::::dangling().as_ptr() as *const u8; - return Some(UserBytes { kptr, len, _pin: None, _scope: PhantomData }); - } - let (kptr, pin) = window(ptr, len, Access::Read)?; - Some(UserBytes { kptr: kptr as *const u8, len, _pin: Some(pin), _scope: PhantomData }) + Some(UserBytes(View::window(ptr, len, Access::Read)?)) } /// A bulk buffer the kernel writes into and never borrows. pub fn user_bytes_mut(&self, ptr: UserAddr, len: u64) -> Option> { - let len = len as usize; - if len == 0 { - let kptr = core::ptr::NonNull::::dangling().as_ptr(); - return Some(UserBytesMut { kptr, len, _pin: None, _scope: PhantomData }); - } - let (kptr, pin) = window(ptr, len, Access::Write)?; - Some(UserBytesMut { kptr, len, _pin: Some(pin), _scope: PhantomData }) + Some(UserBytesMut(View::window(ptr, len, Access::Write)?)) } /// Copy a user string of at most [`MAX_USER_STR`] bytes into kernel memory; an over-long or non-UTF-8 string is `InvalidArgument`, not `BadAddress`. @@ -144,14 +147,14 @@ impl<'a> SyscallContext<'a> { /// Read a typed value out of user memory, copied rather than borrowed. pub fn copy_in(&self, ptr: UserAddr) -> Result { - let (kptr, _pin) = object::(ptr, Access::Read)?; - // SAFETY: `object::` validated size/align inside one page and pinned the frame until `_pin` drops; `T: UserSafe` makes every bit pattern valid; `read_volatile` guards against a concurrent write from another thread of the same process. + let (kptr, _pins) = object::(ptr, Access::Read)?; + // SAFETY: `object::` validated size/align inside one run and pinned its frames until `_pins` drops; `T: UserSafe` makes every bit pattern valid; `read_volatile` guards against a concurrent write from another thread of the same process. Ok(unsafe { kptr.read_volatile() }) } /// Write a typed value into user memory. pub fn copy_out(&self, ptr: UserAddr, value: &T) -> Result<(), SyscallError> { - let (kptr, _pin) = object::(ptr, Access::Write)?; + let (kptr, _pins) = object::(ptr, Access::Write)?; if crate::actuator::copy_meets_a_remap() { remap_race::hold(kptr.cast(), core::mem::size_of::()); } @@ -169,92 +172,115 @@ impl<'a> SyscallContext<'a> { } } -/// A bulk buffer the kernel copies out of and never borrows: no reference exists because another thread of the same process can rewrite the bytes at any time; not volatile per byte, since the missing reference already stops the compiler assuming stability. -pub struct UserBytes<'a> { - kptr: *const u8, +/// Where a window's bytes are: the runs it pinned itself, or its parent's for a +/// [`sub`](UserBytes::sub) view, which borrows the parent's pins through the +/// returned lifetime. +enum Runs<'a> { + Pinned(FramePins), + Borrowed(&'a [Segment]), +} + +/// `len` bytes starting `off` bytes into `runs`: the one shape both window types are. +struct View<'a> { + runs: Runs<'a>, + off: usize, len: usize, - /// Holds the frames covered pinned for this window's life; `None` for the empty window and for a [`sub`](UserBytes::sub) view, which borrows its parent's pin through the returned lifetime. - _pin: Option, - _scope: PhantomData<&'a ()>, } +impl View<'_> { + fn window(ptr: UserAddr, len: u64, access: Access) -> Option { + let len = len as usize; + let pins = match len { + 0 => FramePins(Vec::new()), + _ => window(ptr, len, access)?, + }; + Some(View { runs: Runs::Pinned(pins), off: 0, len }) + } + + fn runs(&self) -> &[Segment] { + match &self.runs { + Runs::Pinned(pins) => &pins.0, + Runs::Borrowed(runs) => runs, + } + } + + /// Hands `copy` each piece of bytes `[off, off + len)` of this view as a + /// direct-map pointer and the piece's offset in that range; panics if out of + /// bounds — the kernel's own arithmetic, never user input. + fn pieces(&self, what: &str, off: usize, len: usize, mut copy: impl FnMut(*mut u8, usize, usize)) { + assert!( + off.checked_add(len).is_some_and(|end| end <= self.len), + "{what} {off}+{len} past a {}-byte window", + self.len + ); + toyos_userbound::pieces(self.runs(), (self.off + off) as u64, len as u64, |phys, at, n| { + copy(crate::mm::DirectMap::from_phys(phys).as_mut_ptr(), at as usize, n as usize) + }); + } + + fn sub(&self, what: &str, off: usize, len: usize) -> View<'_> { + assert!( + off.checked_add(len).is_some_and(|end| end <= self.len), + "{what} {off}+{len} past a {}-byte window", + self.len + ); + View { runs: Runs::Borrowed(self.runs()), off: self.off + off, len } + } +} + +/// A bulk buffer the kernel copies out of and never borrows: no reference exists because another thread of the same process can rewrite the bytes at any time; not volatile per byte, since the missing reference already stops the compiler assuming stability. +pub struct UserBytes<'a>(View<'a>); + impl UserBytes<'_> { pub fn len(&self) -> usize { - self.len + self.0.len } /// Copy `dst.len()` bytes out of the window at `off`; panics if out of bounds — the kernel's own arithmetic, never user input. pub fn read_at(&self, off: usize, dst: &mut [u8]) { - assert!( - off.checked_add(dst.len()).is_some_and(|end| end <= self.len), - "UserBytes::read_at {off}+{} past a {}-byte window", - dst.len(), - self.len - ); - // SAFETY: the assert proves `off + dst.len() <= self.len`, and `window` proved the range is one physically contiguous mapping; `dst` is an owned `&mut`, so the ranges cannot overlap. - unsafe { - core::ptr::copy_nonoverlapping(self.kptr.add(off), dst.as_mut_ptr(), dst.len()); - } + let into = dst.as_mut_ptr(); + self.0.pieces("UserBytes::read_at", off, dst.len(), |run, at, n| { + // SAFETY: `pieces` bounds `[at, at + n)` inside `dst` and `run..+n` inside one run `window` pinned; `dst` is an owned `&mut`, so the ranges cannot overlap. + unsafe { core::ptr::copy_nonoverlapping(run, into.add(at), n) } + }); } /// Copy the window at `off` into one ring run; panics if out of bounds, as /// [`read_at`](Self::read_at) does. `copy` for [`UserBytesMut::write_run`]'s /// reason. pub fn read_run(&self, off: usize, dst: &mut toyos_abi::ring::Dst<'_>) { - assert!( - off.checked_add(dst.len()).is_some_and(|end| end <= self.len), - "UserBytes::read_run {off}+{} past a {}-byte window", - dst.len(), - self.len - ); - // SAFETY: as `UserBytesMut::write_run`, mirrored. - unsafe { - core::ptr::copy(self.kptr.add(off), dst.as_mut_ptr(), dst.len()); - } + let into = dst.as_mut_ptr(); + self.0.pieces("UserBytes::read_run", off, dst.len(), |run, at, n| { + // SAFETY: as `UserBytesMut::write_run`, mirrored. + unsafe { core::ptr::copy(run, into.add(at), n) } + }); } /// The `len`-byte window at `off` inside this one. pub fn sub(&self, off: usize, len: usize) -> UserBytes<'_> { - assert!( - off.checked_add(len).is_some_and(|end| end <= self.len), - "UserBytes::sub {off}+{len} past a {}-byte window", - self.len - ); - // SAFETY: the assert proves `off + len <= self.len`, so the result stays inside the window `window` validated. - UserBytes { kptr: unsafe { self.kptr.add(off) }, len, _pin: None, _scope: PhantomData } + UserBytes(self.0.sub("UserBytes::sub", off, len)) } } /// A bulk buffer the kernel copies into and never reads back, so it cannot act on a value another thread substituted. -pub struct UserBytesMut<'a> { - kptr: *mut u8, - len: usize, - /// As [`UserBytes::_pin`]: the frames stay pinned for this window's life. - _pin: Option, - _scope: PhantomData<&'a mut ()>, -} +pub struct UserBytesMut<'a>(View<'a>); impl UserBytesMut<'_> { pub fn len(&self) -> usize { - self.len + self.0.len } pub fn is_empty(&self) -> bool { - self.len == 0 + self.0.len == 0 } /// Copy `src` into the window at `off`; panics if out of bounds, for the same reason as [`UserBytes::read_at`]. pub fn write_at(&mut self, off: usize, src: &[u8]) { - assert!( - off.checked_add(src.len()).is_some_and(|end| end <= self.len), - "UserBytesMut::write_at {off}+{} past a {}-byte window", - src.len(), - self.len - ); - // SAFETY: the assert proves `off + src.len() <= self.len`, and `window` proved the range is one physically contiguous mapping. - unsafe { - core::ptr::copy_nonoverlapping(src.as_ptr(), self.kptr.add(off), src.len()); - } + let from = src.as_ptr(); + self.0.pieces("UserBytesMut::write_at", off, src.len(), |run, at, n| { + // SAFETY: `pieces` bounds `[at, at + n)` inside `src` and `run..+n` inside one run `window` pinned and proved writable. + unsafe { core::ptr::copy_nonoverlapping(from.add(at), run, n) } + }); } /// Copy one ring run into the window at `off`; panics if out of bounds, as @@ -262,38 +288,24 @@ impl UserBytesMut<'_> { /// `copy_nonoverlapping`**: `SYS_PIPE_MAP` maps the ring's page into the /// caller, which may then name it as this window. pub fn write_run(&mut self, off: usize, src: &toyos_abi::ring::Src<'_>) { - assert!( - off.checked_add(src.len()).is_some_and(|end| end <= self.len), - "UserBytesMut::write_run {off}+{} past a {}-byte window", - src.len(), - self.len - ); - // SAFETY: the assert proves `off + src.len() <= self.len`, `window` proved the range is one physically contiguous mapping, and `Ring::read` proved the run is inside its own data region. Neither side is a reference, so nothing here claims either range is exclusive. - unsafe { - core::ptr::copy(src.as_ptr(), self.kptr.add(off), src.len()); - } + let from = src.as_ptr(); + self.0.pieces("UserBytesMut::write_run", off, src.len(), |run, at, n| { + // SAFETY: `pieces` bounds `[at, at + n)` inside the ring run, which `Ring::read` proved is inside its own data region, and `run..+n` inside one run `window` pinned and proved writable. Neither side is a reference, so nothing here claims either range is exclusive. + unsafe { core::ptr::copy(from.add(at), run, n) } + }); } /// Zero `len` bytes of the window at `off`. pub fn fill_zero(&mut self, off: usize, len: usize) { - assert!( - off.checked_add(len).is_some_and(|end| end <= self.len), - "UserBytesMut::fill_zero {off}+{len} past a {}-byte window", - self.len - ); - // SAFETY: same bound as `write_at`, with a constant zero byte instead of a slice. - unsafe { core::ptr::write_bytes(self.kptr.add(off), 0, len) }; + self.0.pieces("UserBytesMut::fill_zero", off, len, |run, _, n| { + // SAFETY: as `write_at`, with a constant zero byte instead of a slice. + unsafe { core::ptr::write_bytes(run, 0, n) } + }); } /// The `len`-byte window at `off` inside this one. pub fn sub(&mut self, off: usize, len: usize) -> UserBytesMut<'_> { - assert!( - off.checked_add(len).is_some_and(|end| end <= self.len), - "UserBytesMut::sub {off}+{len} past a {}-byte window", - self.len - ); - // SAFETY: [`UserBytes::sub`]'s argument exactly. - UserBytesMut { kptr: unsafe { self.kptr.add(off) }, len, _pin: None, _scope: PhantomData } + UserBytesMut(self.0.sub("UserBytesMut::sub", off, len)) } } @@ -323,20 +335,19 @@ impl ByteSource for UserBytes<'_> { } } -/// A pin on the physical frames a user-copy window covers: the PMM reissues none of them while it lives, so the window's direct-map pointer stays backed even after a sibling unmaps and frees the range across a park. -struct FramePin { - phys: u64, - len: usize, -} +/// A pin on every physical run a user-copy window covers: the PMM reissues none of their frames while it lives, so the window's direct-map pointers stay backed even after a sibling unmaps and frees the range across a park. +struct FramePins(Vec); -impl Drop for FramePin { +impl Drop for FramePins { fn drop(&mut self) { - crate::mm::pmm::unpin_range(self.phys, self.len); + for run in &self.0 { + crate::mm::pmm::unpin_range(run.phys, run.len as usize); + } } } -/// Validate `[ptr, ptr+len)` as one physically contiguous user window, pin every frame it covers, and return its direct-map address with the pin. The pin is taken under the address-space lock over a translation that still names the frame, so a concurrent `munmap` — which needs that same lock to free the range — cannot reissue a frame between the confirmation and the pin. -fn window(ptr: UserAddr, len: usize, access: Access) -> Option<(*mut u8, FramePin)> { +/// Validate `[ptr, ptr+len)` as a user window, walked page by page into its physical runs, and pin every frame they cover. The pins are taken under the address-space lock over a translation that still names each frame, so a concurrent `munmap` — which needs that same lock to free the range — cannot reissue a frame between the confirmation and the pin. +fn window(ptr: UserAddr, len: usize, access: Access) -> Option { if !toyos_userbound::in_user_half(ptr.raw(), len as u64) { return None; } @@ -351,14 +362,21 @@ fn window(ptr: UserAddr, len: usize, access: Access) -> Option<(*mut u8, FramePi } let pt = crate::process::current_address_space(); let guard = pt.lock(); - let phys = toyos_userbound::contiguous(start, len as u64, access, |at| { - translate_now(&guard, UserAddr::new(at), access).map(|dm| dm.phys()) - })?; - if !crate::mm::pmm::pin_range(phys, len) { + let mut runs = Vec::new(); + toyos_userbound::segments( + start, + len as u64, + |at| translate_now(&guard, UserAddr::new(at), access).map(|dm| dm.phys()), + |run| runs.push(run), + )?; + // Cut back to the runs already pinned on a refusal, which the drop then unpins. + if let Some(refused) = runs.iter().position(|run| !crate::mm::pmm::pin_range(run.phys, run.len as usize)) { + runs.truncate(refused); + drop(FramePins(runs)); return None; } drop(guard); - Some((crate::mm::DirectMap::from_phys(phys).as_mut_ptr(), FramePin { phys, len })) + Some(FramePins(runs)) } /// `copy-meets-a-remap`: a sibling's `munmap` and `mmap` staged between a diff --git a/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs b/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs new file mode 100644 index 0000000000..b704c513d7 --- /dev/null +++ b/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs @@ -0,0 +1,96 @@ +//! A syscall's byte buffer that spans two demand-paged windows is copied through +//! both frames, in whichever order the pager handed them out. +//! +//! `.bss` is demand-paged one 2 MiB window at a time, and `SPAN` holds two whole +//! windows nothing else touches. Faulting the upper one first puts it in the +//! frame the allocator hands out first, so the two frames are never one +//! physically contiguous run — the buffer a kernel that copies through one run +//! refuses with `BadAddress`. + +use std::fs::{self, File}; +use std::io::{Read, Write}; + +use toyos_abi::syscall; + +const PAGE_2M: usize = 2 * 1024 * 1024; +/// Three windows: whatever the alignment, two whole ones lie inside. +const SPAN_LEN: usize = 3 * PAGE_2M; +/// Bytes of the buffer on each side of the boundary. +const REACH: usize = 32 * 1024; +const CANARY: u8 = 0xA5; +const PATH: &str = "/tmp/user_copy_spans_windows.bin"; + +static mut SPAN: [u8; SPAN_LEN] = [0; SPAN_LEN]; + +fn pattern(len: usize, salt: u8) -> Vec { + (0..len).map(|i| (i.wrapping_mul(131) ^ (i >> 9)) as u8 ^ salt).collect() +} + +/// The first byte where `got` and `want` differ, and on which side of the boundary. +fn first_difference(got: &[u8], want: &[u8]) -> Option { + let i = got.iter().zip(want).position(|(g, w)| g != w)?; + let side = if i < REACH { "below" } else { "above" }; + Some(format!("byte {i} ({side} the boundary) is {:#04x}, not {:#04x}", got[i], want[i])) +} + +fn main() { + let base = (&raw mut SPAN).cast::(); + let lower = (base as usize).next_multiple_of(PAGE_2M); + let boundary = lower + PAGE_2M; + assert!(boundary + PAGE_2M <= base as usize + SPAN_LEN, "SPAN holds two whole windows"); + let at = |addr: usize| unsafe { base.add(addr - base as usize) }; + + // Upper window first, then lower. + unsafe { + at(boundary).write_volatile(CANARY); + at(boundary - 1).write_volatile(CANARY); + at(boundary - REACH - 1).write_volatile(CANARY); + at(boundary + REACH).write_volatile(CANARY); + } + // SAFETY: `SPAN` is this process's and nothing else names these bytes while the slice lives. + let buf = unsafe { core::slice::from_raw_parts_mut(at(boundary - REACH), 2 * REACH) }; + let outside_intact = || unsafe { + at(boundary - REACH - 1).read_volatile() == CANARY && at(boundary + REACH).read_volatile() == CANARY + }; + + // 1. A pipe read: the kernel writes a ring run into the buffer. + let ends = syscall::pipe().expect("pipe"); + let sent = pattern(buf.len(), 0x5A); + assert_eq!(syscall::write(ends.write, &sent), Ok(sent.len()), "fill the pipe"); + let got = syscall::read(ends.read, buf); + assert!(outside_intact(), "a pipe read into the buffer wrote past its ends"); + assert_eq!(got, Ok(buf.len()), "a pipe read into a buffer across two windows"); + if let Some(diff) = first_difference(buf, &sent) { + panic!("a pipe read across two windows: {diff}"); + } + + // 2. A pipe write: the kernel reads the buffer into a ring run. + let mine = pattern(buf.len(), 0xC3); + buf.copy_from_slice(&mine); + assert_eq!(syscall::write(ends.write, buf), Ok(buf.len()), "a pipe write from a buffer across two windows"); + let mut back = vec![0u8; buf.len()]; + assert_eq!(syscall::read(ends.read, &mut back), Ok(back.len()), "drain the pipe"); + if let Some(diff) = first_difference(&back, &mine) { + panic!("a pipe write from across two windows: {diff}"); + } + syscall::close(ends.read); + syscall::close(ends.write); + + // 3. A file write and read back: the file cache's copies, a page at a time. + let mine = pattern(buf.len(), 0x3C); + buf.copy_from_slice(&mine); + File::create(PATH).and_then(|mut f| f.write_all(buf)).expect("a file write from a buffer across two windows"); + let stored = fs::read(PATH).expect("read the file back"); + if let Some(diff) = first_difference(&stored, &mine) { + panic!("a file write from across two windows: {diff}"); + } + buf.fill(0); + File::open(PATH).and_then(|mut f| f.read_exact(buf)).expect("a file read into a buffer across two windows"); + assert!(outside_intact(), "a file read into the buffer wrote past its ends"); + if let Some(diff) = first_difference(buf, &mine) { + panic!("a file read across two windows: {diff}"); + } + fs::remove_file(PATH).expect("cleanup"); + + println!("a buffer across two demand windows is copied through both frames"); +} diff --git a/toyos-userbound/src/lib.rs b/toyos-userbound/src/lib.rs index 5729263857..31f0472e12 100644 --- a/toyos-userbound/src/lib.rs +++ b/toyos-userbound/src/lib.rs @@ -2,11 +2,14 @@ //! //! **Before a dereference**: is this //! address userland's, is the object at it aligned for the type being read, and -//! does it lie wholly inside one mapping? **Before a placement**: can a length +//! does it lie wholly inside one mapping? **Before a copy**: which physical runs +//! hold a user window, and which of them does each piece of the copy land in? +//! **Before a placement**: can a length //! userland asked for be placed at all, and where does it go? **After a trap**: //! which side did the frame come from? //! -//! [`span`] answers the first, [`place`] the second and [`fault`] the third. +//! [`span`] answers the first, [`segment`] the second, [`place`] the third and +//! [`fault`] the fourth. //! //! Pure. No I/O, no allocation, no `unsafe`, nothing read from a device and //! nothing named outside this crate. The kernel is the only caller — @@ -25,11 +28,13 @@ pub mod fault; pub mod place; +pub mod segment; pub mod span; pub use fault::Ring; pub use place::{PageSpan, Window}; +pub use segment::{pieces, segments, Segment}; pub use span::{ - align_2m_checked, contiguous, in_user_half, is_user_addr, is_user_object, rebase_base, Access, - PAGE_2M, PAGE_4K, USER_TOP, + align_2m_checked, in_user_half, is_user_addr, is_user_object, rebase_base, Access, PAGE_2M, + PAGE_4K, USER_TOP, }; diff --git a/toyos-userbound/src/segment.rs b/toyos-userbound/src/segment.rs new file mode 100644 index 0000000000..af055edb39 --- /dev/null +++ b/toyos-userbound/src/segment.rs @@ -0,0 +1,311 @@ +//! Where a user window's bytes are in physical memory: one [`Segment`] per +//! physically contiguous run, in address order, and each copy cut at their +//! seams. +//! +//! **A window is asked at every 4 KiB page, whatever leaf maps it.** No leaf +//! is smaller, a split window grants writes page by page, and two neighbouring +//! pages sit in whichever frames the PMM handed out, in either order. A page +//! that is absent or does not grant the access refuses the whole window: the +//! kernel copies nothing through a window it could not wholly place. + +use crate::span::{in_user_half, PAGE_4K}; + +/// `len` bytes of a user window, physically contiguous from `phys`. +#[derive(Clone, Copy, PartialEq, Eq, Debug)] +pub struct Segment { + pub phys: u64, + pub len: u64, +} + +/// Walks `[start, start + len)` and hands `emit` each maximal physically +/// contiguous run of it, in order. `leaf` is the page walk: the physical +/// address one user address translates to where it grants the access asked, +/// and `None` where it does not. +/// +/// `None` where the range is empty, leaves the user half, or holds a page +/// `leaf` refuses — and runs before the refused page may already have been +/// emitted, so nothing emitted is acted on until this returns `Some`. +pub fn segments( + start: u64, + len: u64, + mut leaf: impl FnMut(u64) -> Option, + mut emit: impl FnMut(Segment), +) -> Option<()> { + if len == 0 || !in_user_half(start, len) { + return None; + } + let end = start + len; + let mut run = Segment { phys: leaf(start)?, len: 0 }; + let mut at = start; + while at < end { + let phys = if at == start { run.phys } else { leaf(at)? }; + let next = ((at & !(PAGE_4K - 1)) + PAGE_4K).min(end); + if run.phys + run.len != phys { + emit(run); + run = Segment { phys, len: 0 }; + } + run.len += next - at; + at = next; + } + emit(run); + Some(()) +} + +/// Hands `copy` each physically contiguous piece of bytes `[off, off + len)` +/// of the window `segs` lays out, as `(phys, at, n)`: the `n` bytes at `phys` +/// are bytes `at..at + n` of that range. +/// +/// Panics where the range runs past the window's end: the kernel's own +/// arithmetic, bounded before it gets here, never a user length. +pub fn pieces(segs: &[Segment], off: u64, len: u64, mut copy: impl FnMut(u64, u64, u64)) { + let mut skip = off; + let mut done = 0; + for seg in segs { + if done == len { + break; + } + if skip >= seg.len { + skip -= seg.len; + continue; + } + let n = (seg.len - skip).min(len - done); + copy(seg.phys + skip, done, n); + skip = 0; + done += n; + } + assert!(done == len, "pieces: {off}+{len} runs {} bytes past the window", len - done); +} + +#[cfg(test)] +mod tests { + extern crate std; + + use std::vec; + use std::vec::Vec; + + use super::*; + use crate::span::{Access, PAGE_2M, USER_TOP}; + + /// Where the fake user range starts: a 2 MiB boundary, so a range at the + /// top of one fake page and the bottom of the next straddles one too. + const VBASE: u64 = 5 * PAGE_2M; + /// Pages in the fake range: two 2 MiB pages' worth would be 1024, and + /// every property below is about seams, so a few dozen are enough. + const PAGES: u64 = 48; + /// Where fake physical memory starts: nonzero, so a phys of 0 is a bug. + const PBASE: u64 = 0x10_0000_0000; + + /// A deterministic shuffle: xorshift64 driving Fisher-Yates. + fn shuffled(n: usize, mut seed: u64) -> Vec { + let mut order: Vec = (0..n as u64).collect(); + for i in (1..n).rev() { + seed ^= seed << 13; + seed ^= seed >> 7; + seed ^= seed << 17; + order.swap(i, (seed % (i as u64 + 1)) as usize); + } + order + } + + /// A process's pages over fake physical memory: page `i` of the range + /// lives in frame `frame[i]`, and `ro` pages grant no write. + struct Space { + frame: Vec, + ro: Vec, + absent: Vec, + phys: Vec, + } + + impl Space { + /// Every page present and writable, frames in `order`, and each user + /// byte holding a value its virtual address decides. + fn new(order: Vec) -> Self { + let mut space = Space { + frame: order, + ro: Vec::new(), + absent: Vec::new(), + phys: vec![0; (PAGES * PAGE_4K) as usize], + }; + for v in VBASE..VBASE + PAGES * PAGE_4K { + let p = space.leaf(v, Access::Read).expect("every page is present"); + space.phys[(p - PBASE) as usize] = byte_at(v); + } + space + } + + fn leaf(&self, v: u64, access: Access) -> Option { + if !(VBASE..VBASE + PAGES * PAGE_4K).contains(&v) { + return None; + } + let page = (v - VBASE) / PAGE_4K; + if self.absent.contains(&page) || (access == Access::Write && self.ro.contains(&page)) { + return None; + } + Some(PBASE + self.frame[page as usize] * PAGE_4K + v % PAGE_4K) + } + + fn window(&self, start: u64, len: u64, access: Access) -> Option> { + let mut segs = Vec::new(); + segments(start, len, |at| self.leaf(at, access), |s| segs.push(s))?; + Some(segs) + } + + /// The kernel's `read_at`, over fake memory. + fn read(&self, segs: &[Segment], off: u64, len: u64) -> Vec { + let mut out = vec![0u8; len as usize]; + pieces(segs, off, len, |phys, at, n| { + let p = (phys - PBASE) as usize; + out[at as usize..(at + n) as usize].copy_from_slice(&self.phys[p..p + n as usize]); + }); + out + } + + /// The kernel's `write_at`, over fake memory. + fn write(&mut self, segs: &[Segment], off: u64, src: &[u8]) { + let phys = &mut self.phys; + pieces(segs, off, src.len() as u64, |p, at, n| { + let p = (p - PBASE) as usize; + phys[p..p + n as usize].copy_from_slice(&src[at as usize..(at + n) as usize]); + }); + } + + /// What a ring 3 load at `v` sees. + fn load(&self, v: u64) -> u8 { + self.phys[(self.leaf(v, Access::Read).unwrap() - PBASE) as usize] + } + } + + fn byte_at(v: u64) -> u8 { + (v.wrapping_mul(0x9E37_79B9_7F4A_7C15) >> 56) as u8 + } + + /// Ranges that start and end on every kind of seam: at a page start, one + /// byte either side of one, mid-page, and across many pages. + fn ranges() -> Vec<(u64, u64)> { + let mut out = Vec::new(); + for first in [0, 1, 7, 13] { + for head in [0, 1, PAGE_4K / 2, PAGE_4K - 1] { + let start = VBASE + first * PAGE_4K + head; + for len in [1, 2, PAGE_4K - head, PAGE_4K, PAGE_4K + 1, 3 * PAGE_4K - 5, 20 * PAGE_4K + 3] { + if start + len <= VBASE + PAGES * PAGE_4K { + out.push((start, len)); + } + } + } + } + out.push((VBASE, PAGES * PAGE_4K)); + out + } + + #[test] + fn a_read_across_shuffled_frames_gets_the_bytes_at_their_addresses() { + for seed in [1, 0x5eed, 0xdead_beef, 42] { + let space = Space::new(shuffled(PAGES as usize, seed)); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Read).expect("every page is readable"); + let want: Vec = (start..start + len).map(byte_at).collect(); + assert_eq!(space.read(&segs, 0, len), want, "seed {seed:#x}: [{start:#x}, +{len:#x})"); + // A view from any offset inside the window: `UserBytes::sub`. + for off in [1, PAGE_4K - 1, PAGE_4K, len / 2] { + if off < len { + let tail = len - off; + assert_eq!(space.read(&segs, off, tail), want[off as usize..], "seed {seed:#x}: +{off:#x}"); + } + } + } + } + } + + #[test] + fn a_write_across_shuffled_frames_lands_where_a_ring_3_load_reads_it() { + for seed in [3, 0xabcdef, 99] { + for (start, len) in ranges() { + let mut space = Space::new(shuffled(PAGES as usize, seed)); + let segs = space.window(start, len, Access::Write).expect("every page is writable"); + let src: Vec = (0..len).map(|i| !byte_at(i ^ seed)).collect(); + space.write(&segs, 0, &src); + for v in VBASE..VBASE + PAGES * PAGE_4K { + let want = if (start..start + len).contains(&v) { src[(v - start) as usize] } else { byte_at(v) }; + assert_eq!(space.load(v), want, "seed {seed:#x}: [{start:#x}, +{len:#x}) at {v:#x}"); + } + } + } + } + + #[test] + fn the_segments_are_the_maximal_runs_and_cover_the_window_exactly() { + for seed in [7, 0x1234_5678] { + let space = Space::new(shuffled(PAGES as usize, seed)); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Read).unwrap(); + assert_eq!(segs.iter().map(|s| s.len).sum::(), len); + assert_eq!(segs[0].phys, space.leaf(start, Access::Read).unwrap()); + for pair in segs.windows(2) { + assert_ne!(pair[0].phys + pair[0].len, pair[1].phys, "two runs that follow were not joined"); + } + } + } + } + + /// A 2 MiB leaf, or a split window over one frame, is every page in order: + /// one segment, which is one pin and one copy as before the walk was per page. + #[test] + fn frames_in_order_are_one_segment() { + let space = Space::new((0..PAGES).collect()); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Write).unwrap(); + assert_eq!(segs, [Segment { phys: space.leaf(start, Access::Read).unwrap(), len }]); + } + } + + /// Frames in reverse: every page seam is a segment seam. + #[test] + fn frames_in_reverse_are_one_segment_per_page() { + let space = Space::new((0..PAGES).rev().collect()); + let segs = space.window(VBASE + PAGE_4K - 3, PAGE_4K + 6, Access::Read).unwrap(); + assert_eq!( + segs.iter().map(|s| s.len).collect::>(), + [3, PAGE_4K, 3], + "{segs:x?}" + ); + } + + #[test] + fn an_absent_page_anywhere_refuses_the_window() { + for hole in [0, 1, 5, PAGES - 1] { + let mut space = Space::new(shuffled(PAGES as usize, hole + 1)); + space.absent.push(hole); + let h = VBASE + hole * PAGE_4K; + for access in [Access::Read, Access::Write] { + assert_eq!(space.window(VBASE, PAGES * PAGE_4K, access), None, "hole at page {hole}"); + assert_eq!(space.window(h + PAGE_4K - 1, 1, access), None, "the hole's last byte"); + assert_eq!(space.window(h.saturating_sub(1).max(VBASE), 2, access), None, "across the hole's start"); + } + } + } + + #[test] + fn a_read_only_page_between_writable_ones_refuses_a_write_only() { + let mut space = Space::new(shuffled(PAGES as usize, 11)); + space.ro.push(2); + let (start, len) = (VBASE + PAGE_4K, 3 * PAGE_4K); + assert_eq!(space.window(start, len, Access::Write), None); + assert!(space.window(start, len, Access::Read).is_some()); + assert!(space.window(start, PAGE_4K, Access::Write).is_some(), "page 1 alone is writable"); + } + + #[test] + fn an_empty_or_kernel_window_is_refused_before_the_walk() { + let never = |_: u64| -> Option { panic!("walked") }; + assert_eq!(segments(PAGE_2M, 0, never, |_| panic!("emitted")), None); + assert_eq!(segments(USER_TOP - 8, 16, never, |_| panic!("emitted")), None); + assert_eq!(segments(u64::MAX - 4, 8, never, |_| panic!("emitted")), None); + } + + #[test] + #[should_panic(expected = "past the window")] + fn a_piece_past_the_window_is_a_kernel_bug() { + let segs = [Segment { phys: PBASE, len: 8 }, Segment { phys: PBASE + 64, len: 8 }]; + pieces(&segs, 4, 13, |_, _, _| {}); + } +} diff --git a/toyos-userbound/src/span.rs b/toyos-userbound/src/span.rs index c1e30ebdb4..3e7a26379b 100644 --- a/toyos-userbound/src/span.rs +++ b/toyos-userbound/src/span.rs @@ -89,44 +89,9 @@ pub enum Access { Write, } -/// The grain a split window grants rights at. +/// The smallest leaf, and the grain a user window is walked at. pub const PAGE_4K: u64 = 4096; -/// Where `[start, start + len)` begins in physical memory, if the whole range is -/// one physically contiguous run that grants `access`. `leaf` is the page walk: -/// the physical address one user address translates to where it grants -/// `access`, and `None` where it does not. -/// -/// **A write is asked at every 4 KiB page, a read at every 2 MiB page**, both -/// at the first and the last byte. A split window is one 2 MiB frame in order, -/// so contiguity can break only at a 2 MiB boundary and every present page is -/// readable; but it grants writes page by page, so a read-only page between -/// two writable ones is a hole that the ends alone would pass. -pub fn contiguous( - start: u64, - len: u64, - access: Access, - mut leaf: impl FnMut(u64) -> Option, -) -> Option { - if len == 0 || !in_user_half(start, len) { - return None; - } - let phys = leaf(start)?; - let step = match access { - Access::Read => PAGE_2M, - Access::Write => PAGE_4K, - }; - let end = start + len; - let mut at = (start & !(step - 1)) + step; - while at < end { - if leaf(at)? != phys + (at - start) { - return None; - } - at += step; - } - (len == 1 || leaf(end - 1)? == phys + (len - 1)).then_some(phys) -} - #[cfg(test)] mod tests { use super::*; @@ -207,54 +172,6 @@ mod tests { } } - /// A split window at `PAGE_2M` over frame `FRAME`: every page present, - /// and page `ro` readable but not writable. - const FRAME: u64 = 0x4000_0000; - - fn split_with_one_read_only(ro: u64) -> impl Fn(u64, Access) -> Option { - move |at, access| { - if !(PAGE_2M..2 * PAGE_2M).contains(&at) { - return None; - } - let page = (at - PAGE_2M) / PAGE_4K; - (access == Access::Read || page != ro).then_some(FRAME + (at - PAGE_2M)) - } - } - - #[test] - fn a_write_across_a_read_only_page_between_writable_ones_is_refused() { - let walk = split_with_one_read_only(2); - let start = PAGE_2M + PAGE_4K; - let len = 3 * PAGE_4K; - assert_eq!( - contiguous(start, len, Access::Write, |at| walk(at, Access::Write)), - None, - "a write window over pages 1..4 with page 2 read-only was granted" - ); - assert_eq!(contiguous(start, len, Access::Read, |at| walk(at, Access::Read)), Some(FRAME + PAGE_4K)); - assert_eq!( - contiguous(start, PAGE_4K, Access::Write, |at| walk(at, Access::Write)), - Some(FRAME + PAGE_4K), - "page 1 alone is writable" - ); - } - - #[test] - fn a_window_across_two_frames_that_do_not_follow_is_refused() { - let walk = |at: u64| Some(if at < 2 * PAGE_2M { FRAME + at } else { 8 * FRAME + at }); - for access in [Access::Read, Access::Write] { - assert_eq!(contiguous(2 * PAGE_2M - 8, 16, access, walk), None, "{access:?}"); - assert_eq!(contiguous(2 * PAGE_2M - 8, 8, access, walk), Some(FRAME + 2 * PAGE_2M - 8)); - } - } - - #[test] - fn an_empty_or_kernel_window_is_refused_before_the_walk() { - let never = |_: u64| -> Option { panic!("walked") }; - assert_eq!(contiguous(PAGE_2M, 0, Access::Read, never), None); - assert_eq!(contiguous(USER_TOP - 8, 16, Access::Write, never), None); - } - #[test] fn a_zero_sized_object_is_nothing_to_dereference() { assert!(!is_user_object(PAGE_2M, 0, 1)); From 939783c579c66877f8da18dc003f001adfe260d2 Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 00:49:32 +0200 Subject: [PATCH 2/7] user_copy_spans_windows speaks only std, so its source is its own oracle A pipe through `std::io::pipe` with its other end on a scoped thread, in place of the ABI's pipe calls: the ToyOS std passes the caller's buffer to `SYS_READ`/`SYS_WRITE` unchanged, so the kernel copies are the same ones, and no pipe capacity on another system decides the outcome. The same file built for the host with rustc runs green there. Co-Authored-By: Claude Opus 5.5 --- .../src/bin/user_copy_spans_windows.rs | 30 +++++++++++-------- 1 file changed, 17 insertions(+), 13 deletions(-) diff --git a/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs b/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs index b704c513d7..55a6150626 100644 --- a/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs +++ b/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs @@ -6,11 +6,13 @@ //! frame the allocator hands out first, so the two frames are never one //! physically contiguous run — the buffer a kernel that copies through one run //! refuses with `BadAddress`. +//! +//! Plain `std` and nothing of ToyOS's own, so the same source is its own +//! oracle on any other operating system: the bytes it expects are the ones +//! `read` and `write` move everywhere. use std::fs::{self, File}; -use std::io::{Read, Write}; - -use toyos_abi::syscall; +use std::io::{self, Read, Write}; const PAGE_2M: usize = 2 * 1024 * 1024; /// Three windows: whatever the alignment, two whole ones lie inside. @@ -53,28 +55,30 @@ fn main() { at(boundary - REACH - 1).read_volatile() == CANARY && at(boundary + REACH).read_volatile() == CANARY }; - // 1. A pipe read: the kernel writes a ring run into the buffer. - let ends = syscall::pipe().expect("pipe"); + // 1. A pipe read: the kernel writes into the buffer. The other end is a + // thread so that no pipe capacity decides the outcome. + let (mut reader, mut writer) = io::pipe().expect("pipe"); let sent = pattern(buf.len(), 0x5A); - assert_eq!(syscall::write(ends.write, &sent), Ok(sent.len()), "fill the pipe"); - let got = syscall::read(ends.read, buf); + std::thread::scope(|s| { + s.spawn(|| writer.write_all(&sent).expect("fill the pipe")); + reader.read_exact(buf).expect("a pipe read into a buffer across two windows"); + }); assert!(outside_intact(), "a pipe read into the buffer wrote past its ends"); - assert_eq!(got, Ok(buf.len()), "a pipe read into a buffer across two windows"); if let Some(diff) = first_difference(buf, &sent) { panic!("a pipe read across two windows: {diff}"); } - // 2. A pipe write: the kernel reads the buffer into a ring run. + // 2. A pipe write: the kernel reads the buffer. let mine = pattern(buf.len(), 0xC3); buf.copy_from_slice(&mine); - assert_eq!(syscall::write(ends.write, buf), Ok(buf.len()), "a pipe write from a buffer across two windows"); let mut back = vec![0u8; buf.len()]; - assert_eq!(syscall::read(ends.read, &mut back), Ok(back.len()), "drain the pipe"); + std::thread::scope(|s| { + s.spawn(|| reader.read_exact(&mut back).expect("drain the pipe")); + writer.write_all(buf).expect("a pipe write from a buffer across two windows"); + }); if let Some(diff) = first_difference(&back, &mine) { panic!("a pipe write from across two windows: {diff}"); } - syscall::close(ends.read); - syscall::close(ends.write); // 3. A file write and read back: the file cache's copies, a page at a time. let mine = pattern(buf.len(), 0x3C); From 954bbce18b80115c4b3c2d39d6cf94a29dbc095d Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 04:28:28 +0200 Subject: [PATCH 3/7] Round 2 of #593: one walk per leaf, a typed copy allocates nothing, and the pins are all-or-nothing in one tested place The window walk asks the page tables once per leaf mapping instead of once per 4 KiB page: `AddressSpace::leaf` answers a translation together with the bytes to the end of the leaf that maps it, and `segments` steps by that. A 2 MiB leaf is one lookup and one run. A typed `copy_in`/`copy_out` pins its one run as a `[Segment; 1]` and allocates nothing. A bulk window sizes its run list before the address-space lock is taken, one run per 2 MiB window it touches, so nothing is allocated under the lock; a window holding more runs than that panics as a broken kernel invariant. Pinning every run or none, and unpinning every run on drop, is `toyos_userbound::Pinned` over a `Pins` trait: the PMM in the kernel, a counter in the host tests, which assert every frame's count while pinned, zero after the copy, and zero after a refusal at every run. `munmap_reissues_second_read_window` is the parked-reader staging over a buffer whose second physical run is the one the sibling unmaps and maps again. `user_copy_spans_windows`' seam is no longer 4 KiB-aligned within its buffer. `View::pieces` keeps the bound; `toyos_userbound::pieces` no longer repeats it. The track issue cites the owner's "Start it" and the orchestrator's rulings, and loses stage 1, the optional stage 8 and the sentence restating CLAUDE.md. Co-Authored-By: Claude Opus 5.5 --- ...b-pages-and-that-caps-the-process-count.md | 23 +- kernel/src/arch/aarch64/paging.rs | 4 + kernel/src/arch/x86_64/paging.rs | 28 +- kernel/src/user_ptr.rs | 119 +++--- .../bin/munmap_reissues_second_read_window.rs | 110 ++++++ .../src/bin/user_copy_spans_windows.rs | 23 +- toyos-userbound/src/lib.rs | 10 +- toyos-userbound/src/segment.rs | 351 +++++++++++++----- toyos-userbound/src/span.rs | 2 +- 9 files changed, 500 insertions(+), 170 deletions(-) create mode 100644 tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs diff --git a/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md b/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md index 35af9bbf57..2553125972 100644 --- a/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md +++ b/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md @@ -38,21 +38,21 @@ host bridge's firmware-reported apertures stays correct under either page size. ## Stages -Each lands alone and leaves `main` whole. **Commit is strict**: a mapping +The owner approved starting this work ("Start it") and delegated every +kernel-interface decision to the orchestrator, whose stages and rulings +follow. **Commit is strict**, by the orchestrator's ruling: a mapping reserves its frame count when it is made and is refused by name then, never killed at a fault. No promotion (the kernel has no thread to do it) and no live split of a user leaf. -1. **Scatter-gather user copies.** A user window is the physical runs its - pages sit in, each pinned; a buffer's copy is cut at their seams instead of - refused. Exit: `toyos-userbound`'s segment tests over shuffled 4 KiB frames - and the guest `user_copy_spans_windows` green; `abuse_page_straddle` - unchanged, because a typed value still lies inside one 2 MiB page. 2. **Superframe frame allocator.** 4 KiB frames carved from 2 MiB superframes; the pin rule held per superframe, so no pinned frame is reissued and no superframe holding one is handed out whole; a watermark refuses an allocation before the machine runs dry. Exit: host tests of - carve, free, pin and the refusal; every 2 MiB caller of the PMM unchanged. + carve, free, pin and the refusal; every 2 MiB caller of the PMM unchanged; + `user_copy_spans_windows` and `munmap_reissues_second_read_window` red + again under their negative controls, because each is two physical runs + only while the allocator hands out the lowest free frame first. 3. **4 KiB user leaves.** A software `OWNED` bit in the leaf is the ledger of the frames a process owns, with per-process counts; stacks are demand-zero and TLS is 4 KiB. A typed syscall value crossing a page is served in @@ -61,14 +61,13 @@ live split of a user leaf. process's floor on the T14 against the 10 to 14 MB above. 4. **Demand-zero anonymous `mmap`.** A fault installs a 2 MiB leaf only for a whole aligned 2 MiB span of the mapping; unmapping a range no thread touched - sends no IPI; `munmap` refuses a size that is not the whole mapping. Exit: - `mmap` of more than the watermark allows refused at `mmap`, and a guest that - touches one page of a large mapping holds one frame. + sends no IPI; `munmap` refuses a size that is not the whole mapping, by the + orchestrator's ruling. Exit: `mmap` of more than the watermark allows + refused at `mmap`, and a guest that touches one page of a large mapping + holds one frame. 5. **File-backed faults at 4 KiB.** ROOT's read-only text is shared zero-copy. Exit: two processes of one binary hold its text once. 6. **`/apps` image sharing.** The design is the owner's decision, owed when this stage is reached. 7. **Retire 2 MiB where it no longer pays.** Exit: every remaining 2 MiB user leaf is a stage-4 whole span or a device or DMA grant. -8. **Optional: PCID-scoped remote flush.** Exit: a measured shootdown cost - that pays for it. diff --git a/kernel/src/arch/aarch64/paging.rs b/kernel/src/arch/aarch64/paging.rs index 5f532fa241..eb969396e1 100644 --- a/kernel/src/arch/aarch64/paging.rs +++ b/kernel/src/arch/aarch64/paging.rs @@ -72,6 +72,10 @@ impl AddressSpace { match self.never {} } + pub fn leaf(&self, _vaddr: UserAddr, _access: toyos_userbound::Access) -> Option<(u64, u64)> { + match self.never {} + } + pub fn alloc_region(&mut self, _size: u64, _kind: RegionKind) -> Option { match self.never {} } diff --git a/kernel/src/arch/x86_64/paging.rs b/kernel/src/arch/x86_64/paging.rs index 43991575b4..1c65c277a3 100644 --- a/kernel/src/arch/x86_64/paging.rs +++ b/kernel/src/arch/x86_64/paging.rs @@ -42,6 +42,8 @@ const ADDR_MASK_2M: u64 = 0x000F_FFFF_FFE0_0000; /// Every upper-level table entry's flags: present, writable, user. const TABLE_FLAGS: u64 = PAGE_PRESENT | PAGE_WRITE | PAGE_USER; +/// The rights a ring 3 store needs at every level of its walk. +const STORE: u64 = PAGE_USER | PAGE_WRITE; impl Prot { /// The permission bits a leaf entry carries; address and cache policy stay the caller's. @@ -556,7 +558,7 @@ impl AddressSpace { /// Checked here, not at the callers: a user space shallow-copies the /// kernel PML4 half, so a kernel address would otherwise walk to a writable kernel page. pub fn translate(&self, vaddr: UserAddr) -> Option { - self.walk(vaddr).map(|(dm, _)| dm) + self.walk(vaddr).map(|(dm, _, _)| dm) } /// As [`translate`](Self::translate), but only where a user store would @@ -565,12 +567,24 @@ impl AddressSpace { /// memory goes through this, so a syscall cannot write a page the process /// itself may not — the clock page, a shared library's `.text`. pub fn translate_writable(&self, vaddr: UserAddr) -> Option { - const STORE: u64 = PAGE_USER | PAGE_WRITE; - self.walk(vaddr).and_then(|(dm, rights)| (rights & STORE == STORE).then_some(dm)) + self.walk(vaddr).and_then(|(dm, rights, _)| (rights & STORE == STORE).then_some(dm)) + } + + /// `vaddr`'s physical address as [`translate`](Self::translate) or + /// [`translate_writable`](Self::translate_writable) answers it for + /// `access`, and the bytes from it to the end of the leaf that maps it, + /// which that one walk answers for too. + pub fn leaf(&self, vaddr: UserAddr, access: toyos_userbound::Access) -> Option<(u64, u64)> { + let (dm, rights, size) = self.walk(vaddr)?; + let granted = match access { + toyos_userbound::Access::Read => true, + toyos_userbound::Access::Write => rights & STORE == STORE, + }; + granted.then(|| (dm.phys(), size - (vaddr.raw() & (size - 1)))) } - /// The direct-map address of `vaddr` and the rights every level of its walk grants in common. - fn walk(&self, vaddr: UserAddr) -> Option<(crate::mm::DirectMap, u64)> { + /// The direct-map address of `vaddr`, the rights every level of its walk grants in common, and the size of the leaf that maps it. + fn walk(&self, vaddr: UserAddr) -> Option<(crate::mm::DirectMap, u64, u64)> { let va = vaddr.raw(); if !toyos_userbound::is_user_addr(va) { return None; @@ -593,11 +607,11 @@ impl AddressSpace { return None; } let dm = crate::mm::DirectMap::from_phys((pte & ADDR_MASK) + (va & 0xFFF)); - return Some((dm, rights & pte)); + return Some((dm, rights & pte, 4096)); } let page_phys = pde & ADDR_MASK_2M; let offset = va & (PAGE_2M - 1); - Some((crate::mm::DirectMap::from_phys(page_phys + offset), rights)) + Some((crate::mm::DirectMap::from_phys(page_phys + offset), rights, PAGE_2M)) } /// Where `span` goes, top-down and never below the floor: a region the diff --git a/kernel/src/user_ptr.rs b/kernel/src/user_ptr.rs index c4d49b089a..6665ae32bb 100644 --- a/kernel/src/user_ptr.rs +++ b/kernel/src/user_ptr.rs @@ -2,11 +2,11 @@ //! user memory lands only where a ring 3 store from the process itself could. //! //! SMAP stays enabled; access goes through the direct map, never stac/clac. -//! Values are copied out, never referenced. **Every copy goes through -//! [`window`]**, which pins every frame it covers under the address-space lock -//! for the copy's life — a typed value's as much as a [`UserBytes`]/ -//! [`UserBytesMut`] buffer's — so a sibling's `munmap` cannot reissue the -//! backing between the translation and the copy, nor under a copy across a park. +//! Values are copied out, never referenced. **Every copy pins every frame it +//! covers** under the address-space lock for the copy's life — [`object_run`] +//! a typed value's, [`window`] a [`UserBytes`]/[`UserBytesMut`] buffer's — so a +//! sibling's `munmap` cannot reissue the backing between the translation and +//! the copy, nor under a copy across a park. //! A window is the physical runs its pages sit in, and a buffer's copy is cut at //! their seams; a typed value lies inside one 2 MiB page, so in one run. //! Every single-word user read is `read_volatile`, including the futex word @@ -18,7 +18,7 @@ use alloc::vec::Vec; use core::marker::PhantomData; use toyos_abi::syscall::SyscallError; -use toyos_userbound::Segment; +use toyos_userbound::{Pinned, Pins, Segment}; use crate::UserAddr; @@ -93,25 +93,42 @@ pub(crate) fn translate_user(addr: UserAddr, access: Access) -> Option(ptr: UserAddr, access: Access) -> Result<(*mut T, FramePins), SyscallError> { +fn object(ptr: UserAddr, access: Access) -> Result<(*mut T, FramePins<[Segment; 1]>), SyscallError> { let (kptr, pins) = object_run(ptr, core::mem::size_of::(), core::mem::align_of::(), access)?; Ok((kptr.cast(), pins)) } -fn object_run(ptr: UserAddr, size: usize, align: usize, access: Access) -> Result<(*mut u8, FramePins), SyscallError> { +/// The one run a typed value lies in, pinned, with nothing allocated. +fn object_run(ptr: UserAddr, size: usize, align: usize, access: Access) -> Result<(*mut u8, FramePins<[Segment; 1]>), SyscallError> { if !toyos_userbound::is_user_object(ptr.raw(), size as u64, align as u64) { return Err(SyscallError::BadAddress); } - let pins = window(ptr, size, access).ok_or(SyscallError::BadAddress)?; - let [run] = pins.0[..] else { object_split(ptr, size, pins.0.len()) }; + fault_in(ptr, size, access).ok_or(SyscallError::BadAddress)?; + let pt = crate::process::current_address_space(); + let guard = pt.lock(); + let mut run = None; + toyos_userbound::segments( + ptr.raw(), + size as u64, + |at| guard.leaf(UserAddr::new(at), access), + |next| { + if run.replace(next).is_some() { + window_split(ptr, size) + } + }, + ) + .ok_or(SyscallError::BadAddress)?; + let run = run.expect("`segments` emits a run whenever it places the window"); + let pins = FramePins::pin([run], Pmm).ok_or(SyscallError::BadAddress)?; + drop(guard); Ok((crate::mm::DirectMap::from_phys(run.phys).as_mut_ptr(), pins)) } -/// A 2 MiB page is one frame in order, so an object inside one is one run. +/// A 2 MiB window is one frame in order, so it holds at most one run. #[cold] #[inline(never)] -fn object_split(ptr: UserAddr, size: usize, runs: usize) -> ! { - panic!("a {size}-byte value at {:#x} inside one 2 MiB page lies in {runs} physical runs", ptr.raw()) +fn window_split(ptr: UserAddr, len: usize) -> ! { + panic!("[{:#x}, +{len:#x}) lies in more physical runs than it touches 2 MiB windows", ptr.raw()) } /// Context for a single syscall invocation; the lifetime `'a` keeps validated references from escaping it. @@ -176,7 +193,7 @@ impl<'a> SyscallContext<'a> { /// [`sub`](UserBytes::sub) view, which borrows the parent's pins through the /// returned lifetime. enum Runs<'a> { - Pinned(FramePins), + Pinned(FramePins>), Borrowed(&'a [Segment]), } @@ -190,16 +207,16 @@ struct View<'a> { impl View<'_> { fn window(ptr: UserAddr, len: u64, access: Access) -> Option { let len = len as usize; - let pins = match len { - 0 => FramePins(Vec::new()), - _ => window(ptr, len, access)?, + let runs = match len { + 0 => Runs::Borrowed(&[]), + _ => Runs::Pinned(window(ptr, len, access)?), }; - Some(View { runs: Runs::Pinned(pins), off: 0, len }) + Some(View { runs, off: 0, len }) } fn runs(&self) -> &[Segment] { match &self.runs { - Runs::Pinned(pins) => &pins.0, + Runs::Pinned(pins) => pins.runs(), Runs::Borrowed(runs) => runs, } } @@ -336,47 +353,59 @@ impl ByteSource for UserBytes<'_> { } /// A pin on every physical run a user-copy window covers: the PMM reissues none of their frames while it lives, so the window's direct-map pointers stay backed even after a sibling unmaps and frees the range across a park. -struct FramePins(Vec); +type FramePins = Pinned; -impl Drop for FramePins { - fn drop(&mut self) { - for run in &self.0 { - crate::mm::pmm::unpin_range(run.phys, run.len as usize); - } +/// The frame allocator's pins. +struct Pmm; + +impl Pins for Pmm { + fn pin(&mut self, run: Segment) -> bool { + crate::mm::pmm::pin_range(run.phys, run.len as usize) } -} -/// Validate `[ptr, ptr+len)` as a user window, walked page by page into its physical runs, and pin every frame they cover. The pins are taken under the address-space lock over a translation that still names each frame, so a concurrent `munmap` — which needs that same lock to free the range — cannot reissue a frame between the confirmation and the pin. -fn window(ptr: UserAddr, len: usize, access: Access) -> Option { - if !toyos_userbound::in_user_half(ptr.raw(), len as u64) { - return None; + fn unpin(&mut self, run: Segment) { + crate::mm::pmm::unpin_range(run.phys, run.len as usize); } - let start = ptr.raw(); - let end = start + len as u64; - // Fault every page of the range in; what it maps is confirmed under the lock below. +} + +/// Faults every 2 MiB window `[ptr, ptr + len)` touches in, and answers how +/// many that is; what they map is confirmed under the lock afterwards. +fn fault_in(ptr: UserAddr, len: usize, access: Access) -> Option { translate_user(ptr, access)?; - let mut boundary = (start & !(crate::mm::PAGE_2M - 1)) + crate::mm::PAGE_2M; + let end = ptr.raw() + len as u64; + let mut windows = 1; + let mut boundary = (ptr.raw() & !(crate::mm::PAGE_2M - 1)) + crate::mm::PAGE_2M; while boundary < end { translate_user(UserAddr::new(boundary), access)?; boundary += crate::mm::PAGE_2M; + windows += 1; } + Some(windows) +} + +/// Validate `[ptr, ptr+len)` as a user window, walked leaf by leaf into its physical runs, and pin every frame they cover. The pins are taken under the address-space lock over a translation that still names each frame, so a concurrent `munmap` — which needs that same lock to free the range — cannot reissue a frame between the confirmation and the pin. Nothing is allocated or freed under the lock but on a refused pin. +fn window(ptr: UserAddr, len: usize, access: Access) -> Option>> { + if !toyos_userbound::in_user_half(ptr.raw(), len as u64) { + return None; + } + let windows = fault_in(ptr, len, access)?; + let mut runs = Vec::with_capacity(windows); let pt = crate::process::current_address_space(); let guard = pt.lock(); - let mut runs = Vec::new(); toyos_userbound::segments( - start, + ptr.raw(), len as u64, - |at| translate_now(&guard, UserAddr::new(at), access).map(|dm| dm.phys()), - |run| runs.push(run), + |at| guard.leaf(UserAddr::new(at), access), + |run| { + if runs.len() == windows { + window_split(ptr, len) + } + runs.push(run) + }, )?; - // Cut back to the runs already pinned on a refusal, which the drop then unpins. - if let Some(refused) = runs.iter().position(|run| !crate::mm::pmm::pin_range(run.phys, run.len as usize)) { - runs.truncate(refused); - drop(FramePins(runs)); - return None; - } + let pins = FramePins::pin(runs, Pmm)?; drop(guard); - Some(FramePins(runs)) + Some(pins) } /// `copy-meets-a-remap`: a sibling's `munmap` and `mmap` staged between a diff --git a/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs b/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs new file mode 100644 index 0000000000..b7e5291598 --- /dev/null +++ b/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs @@ -0,0 +1,110 @@ +//! A parked reader whose buffer spans two mappings holds both frames: a +//! sibling that unmaps the second one and maps again is never handed the frame +//! the parked copy is about to land in. +//! +//! `munmap_reissues_read_window`'s staging, over a buffer that ends in a second +//! mapping. The two mappings are placed side by side with `FIXED`, the upper +//! one first: the frame allocator hands out the lowest free frame first, so +//! the upper mapping's frame lies below the lower one's and the buffer is two +//! physical runs, of which the upper is the second. Nothing here can see a +//! physical address, so that premise is the allocator's and is not checked. +//! +//! The assertion is that the sibling's new mapping still holds its own byte. + +use std::sync::atomic::{AtomicI64, AtomicU32, Ordering}; +use std::sync::Arc; +use std::thread; + +use toyos::endow::{Endowments, SYSCAP_LABEL}; +use toyos::syscap::SysCap; +use toyos_abi::syscall::{close, mmap, munmap, pipe, read, write, MmapFlags, MmapProt}; + +const PAGE_2M: usize = 2 * 1024 * 1024; + +/// Bytes of the buffer in the lower mapping and in the upper one; the whole +/// buffer is far under the pipe ring, so one write delivers it. +const BELOW: usize = 2048; +const ABOVE: usize = 2048; + +/// What the parked reader copies out of the pipe, and the sibling never writes. +const PATTERN_A: u8 = 0xA1; +/// What the sibling writes into its new mapping, and a safe kernel leaves there. +const PATTERN_B: u8 = 0xB2; + +fn map(at: *mut u8, len: usize, prot: MmapProt, flags: MmapFlags) -> *mut u8 { + let p = unsafe { mmap(at, len, prot, MmapFlags::ANONYMOUS | MmapFlags::PRIVATE | flags) }; + assert!(!p.is_null(), "mmap of {len:#x} bytes at {at:?} failed"); + p +} + +#[path = "../roster.rs"] +mod roster; + +fn main() { + let cap: SysCap = Endowments::get() + .take(SYSCAP_LABEL) + .expect("test-runner endows every binary it spawns a system capability"); + let ends = pipe().expect("a pipe"); + let read_end = ends.read; + let write_end = ends.write; + + // Two 2 MiB slots side by side: reserved, released, then mapped upper first. + let rw = MmapProt::READ | MmapProt::WRITE; + let lower = map(core::ptr::null_mut(), 2 * PAGE_2M, MmapProt::NONE, MmapFlags(0)); + unsafe { munmap(lower, 2 * PAGE_2M) }.expect("release the reservation"); + let upper = unsafe { lower.add(PAGE_2M) }; + assert_eq!(map(upper, PAGE_2M, rw, MmapFlags::FIXED), upper); + assert_eq!(map(lower, PAGE_2M, rw, MmapFlags::FIXED), lower); + let buf_addr = upper as usize - BELOW; + + let ready = Arc::new(AtomicU32::new(0)); + let result = Arc::new(AtomicI64::new(i64::MIN)); + + let a = { + let ready = Arc::clone(&ready); + let result = Arc::clone(&result); + thread::spawn(move || { + // Valid when formed and when the read begins; the kernel owns the + // pointer once it parks, and nothing in this thread dereferences it. + let buf = unsafe { core::slice::from_raw_parts_mut(buf_addr as *mut u8, BELOW + ABOVE) }; + ready.store(1, Ordering::SeqCst); + let n = match read(read_end, buf) { + Ok(n) => n as i64, + Err(_) => -1, + }; + result.store(n, Ordering::SeqCst); + }) + }; + + roster::await_true(|| { + roster::my_threads(&cap).iter().any(|&(is_thread, state)| is_thread && state == roster::BLOCKED) + && ready.load(Ordering::SeqCst) == 1 + }); + + unsafe { munmap(upper, PAGE_2M) }.expect("munmap the buffer's second mapping"); + + let sibling = map(core::ptr::null_mut(), PAGE_2M, rw, MmapFlags(0)); + unsafe { core::ptr::write_bytes(sibling, PATTERN_B, ABOVE) }; + + write(write_end, &[PATTERN_A; BELOW + ABOVE]).expect("write to wake the reader"); + + a.join().expect("the reader thread panicked"); + close(read_end); + close(write_end); + + let n = result.load(Ordering::SeqCst); + assert_eq!(n, (BELOW + ABOVE) as i64, "the parked reader's read answered {n}, so the copy the test is about never ran"); + let first = unsafe { core::slice::from_raw_parts(buf_addr as *const u8, BELOW) }; + assert!(first.iter().all(|&b| b == PATTERN_A), "the copy's first run never reached the mapping still in place"); + + let got = unsafe { core::slice::from_raw_parts(sibling, ABOVE) }; + if let Some(bad) = got.iter().position(|&b| b != PATTERN_B) { + panic!( + "a parked reader's second run reached a sibling's reissued frame — byte {bad} of it is {:#x}, not {PATTERN_B:#x}", + got[bad] + ); + } + unsafe { munmap(sibling, PAGE_2M) }.expect("munmap the sibling's mapping"); + unsafe { munmap(lower, PAGE_2M) }.expect("munmap the buffer's first mapping"); + println!("munmap_reissues_second_read_window: both frames held under the copy"); +} diff --git a/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs b/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs index 55a6150626..ce0d015f75 100644 --- a/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs +++ b/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs @@ -2,10 +2,7 @@ //! both frames, in whichever order the pager handed them out. //! //! `.bss` is demand-paged one 2 MiB window at a time, and `SPAN` holds two whole -//! windows nothing else touches. Faulting the upper one first puts it in the -//! frame the allocator hands out first, so the two frames are never one -//! physically contiguous run — the buffer a kernel that copies through one run -//! refuses with `BadAddress`. +//! windows nothing else touches. //! //! Plain `std` and nothing of ToyOS's own, so the same source is its own //! oracle on any other operating system: the bytes it expects are the ones @@ -17,8 +14,10 @@ use std::io::{self, Read, Write}; const PAGE_2M: usize = 2 * 1024 * 1024; /// Three windows: whatever the alignment, two whole ones lie inside. const SPAN_LEN: usize = 3 * PAGE_2M; -/// Bytes of the buffer on each side of the boundary. -const REACH: usize = 32 * 1024; +/// Bytes of the buffer below the boundary and above it: the boundary is not +/// 4 KiB-aligned within the buffer. +const BELOW: usize = 32 * 1024 + 1000; +const ABOVE: usize = 32 * 1024; const CANARY: u8 = 0xA5; const PATH: &str = "/tmp/user_copy_spans_windows.bin"; @@ -31,7 +30,7 @@ fn pattern(len: usize, salt: u8) -> Vec { /// The first byte where `got` and `want` differ, and on which side of the boundary. fn first_difference(got: &[u8], want: &[u8]) -> Option { let i = got.iter().zip(want).position(|(g, w)| g != w)?; - let side = if i < REACH { "below" } else { "above" }; + let side = if i < BELOW { "below" } else { "above" }; Some(format!("byte {i} ({side} the boundary) is {:#04x}, not {:#04x}", got[i], want[i])) } @@ -46,13 +45,13 @@ fn main() { unsafe { at(boundary).write_volatile(CANARY); at(boundary - 1).write_volatile(CANARY); - at(boundary - REACH - 1).write_volatile(CANARY); - at(boundary + REACH).write_volatile(CANARY); + at(boundary - BELOW - 1).write_volatile(CANARY); + at(boundary + ABOVE).write_volatile(CANARY); } // SAFETY: `SPAN` is this process's and nothing else names these bytes while the slice lives. - let buf = unsafe { core::slice::from_raw_parts_mut(at(boundary - REACH), 2 * REACH) }; + let buf = unsafe { core::slice::from_raw_parts_mut(at(boundary - BELOW), BELOW + ABOVE) }; let outside_intact = || unsafe { - at(boundary - REACH - 1).read_volatile() == CANARY && at(boundary + REACH).read_volatile() == CANARY + at(boundary - BELOW - 1).read_volatile() == CANARY && at(boundary + ABOVE).read_volatile() == CANARY }; // 1. A pipe read: the kernel writes into the buffer. The other end is a @@ -80,7 +79,7 @@ fn main() { panic!("a pipe write from across two windows: {diff}"); } - // 3. A file write and read back: the file cache's copies, a page at a time. + // 3. A file write and read back: the file cache's copies. let mine = pattern(buf.len(), 0x3C); buf.copy_from_slice(&mine); File::create(PATH).and_then(|mut f| f.write_all(buf)).expect("a file write from a buffer across two windows"); diff --git a/toyos-userbound/src/lib.rs b/toyos-userbound/src/lib.rs index 31f0472e12..c7fdc63f4f 100644 --- a/toyos-userbound/src/lib.rs +++ b/toyos-userbound/src/lib.rs @@ -3,10 +3,10 @@ //! **Before a dereference**: is this //! address userland's, is the object at it aligned for the type being read, and //! does it lie wholly inside one mapping? **Before a copy**: which physical runs -//! hold a user window, and which of them does each piece of the copy land in? -//! **Before a placement**: can a length -//! userland asked for be placed at all, and where does it go? **After a trap**: -//! which side did the frame come from? +//! hold a user window, which of them does each piece of the copy land in, and +//! is every one of them pinned? **Before a placement**: can a length userland +//! asked for be placed at all, and where does it go? **After a trap**: which +//! side did the frame come from? //! //! [`span`] answers the first, [`segment`] the second, [`place`] the third and //! [`fault`] the fourth. @@ -33,7 +33,7 @@ pub mod span; pub use fault::Ring; pub use place::{PageSpan, Window}; -pub use segment::{pieces, segments, Segment}; +pub use segment::{pieces, segments, Pinned, Pins, Segment}; pub use span::{ align_2m_checked, in_user_half, is_user_addr, is_user_object, rebase_base, Access, PAGE_2M, PAGE_4K, USER_TOP, diff --git a/toyos-userbound/src/segment.rs b/toyos-userbound/src/segment.rs index af055edb39..ba44f2b363 100644 --- a/toyos-userbound/src/segment.rs +++ b/toyos-userbound/src/segment.rs @@ -1,14 +1,13 @@ //! Where a user window's bytes are in physical memory: one [`Segment`] per -//! physically contiguous run, in address order, and each copy cut at their -//! seams. +//! physically contiguous run, in address order, each copy cut at their seams, +//! and every run pinned for the window's life. //! -//! **A window is asked at every 4 KiB page, whatever leaf maps it.** No leaf -//! is smaller, a split window grants writes page by page, and two neighbouring -//! pages sit in whichever frames the PMM handed out, in either order. A page -//! that is absent or does not grant the access refuses the whole window: the -//! kernel copies nothing through a window it could not wholly place. +//! **A window is asked once per leaf that maps it.** A page that is absent or +//! does not grant the access refuses the whole window, and so does a run that +//! cannot be pinned: the kernel copies nothing through a window it could not +//! wholly place and hold. -use crate::span::{in_user_half, PAGE_4K}; +use crate::span::in_user_half; /// `len` bytes of a user window, physically contiguous from `phys`. #[derive(Clone, Copy, PartialEq, Eq, Debug)] @@ -18,9 +17,10 @@ pub struct Segment { } /// Walks `[start, start + len)` and hands `emit` each maximal physically -/// contiguous run of it, in order. `leaf` is the page walk: the physical -/// address one user address translates to where it grants the access asked, -/// and `None` where it does not. +/// contiguous run of it, in order. `leaf` is the page walk: for one user +/// address, where it grants the access asked, the physical address it +/// translates to and the bytes from it to the end of the leaf that maps it, +/// at least one; `None` where it does not. /// /// `None` where the range is empty, leaves the user half, or holds a page /// `leaf` refuses — and runs before the refused page may already have been @@ -28,24 +28,28 @@ pub struct Segment { pub fn segments( start: u64, len: u64, - mut leaf: impl FnMut(u64) -> Option, + mut leaf: impl FnMut(u64) -> Option<(u64, u64)>, mut emit: impl FnMut(Segment), ) -> Option<()> { if len == 0 || !in_user_half(start, len) { return None; } let end = start + len; - let mut run = Segment { phys: leaf(start)?, len: 0 }; let mut at = start; - while at < end { - let phys = if at == start { run.phys } else { leaf(at)? }; - let next = ((at & !(PAGE_4K - 1)) + PAGE_4K).min(end); + let (mut phys, mut extent) = leaf(at)?; + let mut run = Segment { phys, len: 0 }; + loop { if run.phys + run.len != phys { emit(run); run = Segment { phys, len: 0 }; } - run.len += next - at; - at = next; + let n = extent.min(end - at); + run.len += n; + at += n; + if at == end { + break; + } + (phys, extent) = leaf(at)?; } emit(run); Some(()) @@ -53,10 +57,8 @@ pub fn segments( /// Hands `copy` each physically contiguous piece of bytes `[off, off + len)` /// of the window `segs` lays out, as `(phys, at, n)`: the `n` bytes at `phys` -/// are bytes `at..at + n` of that range. -/// -/// Panics where the range runs past the window's end: the kernel's own -/// arithmetic, bounded before it gets here, never a user length. +/// are bytes `at..at + n` of that range. The caller bounds the range inside +/// the window. pub fn pieces(segs: &[Segment], off: u64, len: u64, mut copy: impl FnMut(u64, u64, u64)) { let mut skip = off; let mut done = 0; @@ -73,29 +75,68 @@ pub fn pieces(segs: &[Segment], off: u64, len: u64, mut copy: impl FnMut(u64, u6 skip = 0; done += n; } - assert!(done == len, "pieces: {off}+{len} runs {} bytes past the window", len - done); +} + +/// What holds a frame against reissue: the kernel's frame allocator, or a +/// test's counter. +pub trait Pins { + /// Pins every frame `run` touches; `false`, with nothing pinned, when it + /// cannot. + fn pin(&mut self, run: Segment) -> bool; + /// Releases a pin [`pin`](Self::pin) took over the same run. + fn unpin(&mut self, run: Segment); +} + +/// A window's runs, every one of them pinned until this drops. +pub struct Pinned, P: Pins> { + runs: R, + pins: P, +} + +impl, P: Pins> Pinned { + /// Pins every run, or none: a run that cannot be pinned unpins the runs + /// before it, and the window is refused. + pub fn pin(runs: R, mut pins: P) -> Option { + if let Some(refused) = runs.as_ref().iter().position(|&run| !pins.pin(run)) { + for &run in &runs.as_ref()[..refused] { + pins.unpin(run); + } + return None; + } + Some(Pinned { runs, pins }) + } + + pub fn runs(&self) -> &[Segment] { + self.runs.as_ref() + } +} + +impl, P: Pins> Drop for Pinned { + fn drop(&mut self) { + for &run in self.runs.as_ref() { + self.pins.unpin(run); + } + } } #[cfg(test)] mod tests { extern crate std; + use std::cell::{Cell, RefCell}; use std::vec; use std::vec::Vec; use super::*; - use crate::span::{Access, PAGE_2M, USER_TOP}; + use crate::span::{Access, PAGE_2M, PAGE_4K, USER_TOP}; - /// Where the fake user range starts: a 2 MiB boundary, so a range at the - /// top of one fake page and the bottom of the next straddles one too. + /// Where the fake user range starts: a 2 MiB boundary. const VBASE: u64 = 5 * PAGE_2M; - /// Pages in the fake range: two 2 MiB pages' worth would be 1024, and - /// every property below is about seams, so a few dozen are enough. + /// Pages in the fake range. const PAGES: u64 = 48; /// Where fake physical memory starts: nonzero, so a phys of 0 is a bug. const PBASE: u64 = 0x10_0000_0000; - /// A deterministic shuffle: xorshift64 driving Fisher-Yates. fn shuffled(n: usize, mut seed: u64) -> Vec { let mut order: Vec = (0..n as u64).collect(); for i in (1..n).rev() { @@ -107,41 +148,58 @@ mod tests { order } - /// A process's pages over fake physical memory: page `i` of the range - /// lives in frame `frame[i]`, and `ro` pages grant no write. + /// A process's pages over fake physical memory, mapped by leaves of + /// `grain` pages: leaf `i` of the range lives in frame `frame[i]`, and `ro` + /// and `absent` pages (with a `grain` of one) grant no write and nothing. struct Space { + grain: u64, frame: Vec, ro: Vec, absent: Vec, phys: Vec, + lookups: Cell, } impl Space { - /// Every page present and writable, frames in `order`, and each user - /// byte holding a value its virtual address decides. - fn new(order: Vec) -> Self { + /// Every page present and writable, leaves in frames `order`, and each + /// user byte holding a value its virtual address decides. + fn new(order: Vec, grain: u64) -> Self { + assert_eq!(order.len() as u64 * grain, PAGES); let mut space = Space { + grain, frame: order, ro: Vec::new(), absent: Vec::new(), phys: vec![0; (PAGES * PAGE_4K) as usize], + lookups: Cell::new(0), }; for v in VBASE..VBASE + PAGES * PAGE_4K { - let p = space.leaf(v, Access::Read).expect("every page is present"); + let p = space.phys_of(v).expect("every page is present"); space.phys[(p - PBASE) as usize] = byte_at(v); } space } - fn leaf(&self, v: u64, access: Access) -> Option { + fn leaf_bytes(&self) -> u64 { + self.grain * PAGE_4K + } + + fn phys_of(&self, v: u64) -> Option { if !(VBASE..VBASE + PAGES * PAGE_4K).contains(&v) { return None; } - let page = (v - VBASE) / PAGE_4K; + let off = v - VBASE; + Some(PBASE + self.frame[(off / self.leaf_bytes()) as usize] * self.leaf_bytes() + off % self.leaf_bytes()) + } + + /// The kernel's `AddressSpace::leaf`, over fake memory. + fn leaf(&self, v: u64, access: Access) -> Option<(u64, u64)> { + self.lookups.set(self.lookups.get() + 1); + let page = v.checked_sub(VBASE)? / PAGE_4K; if self.absent.contains(&page) || (access == Access::Write && self.ro.contains(&page)) { return None; } - Some(PBASE + self.frame[page as usize] * PAGE_4K + v % PAGE_4K) + Some((self.phys_of(v)?, self.leaf_bytes() - (v - VBASE) % self.leaf_bytes())) } fn window(&self, start: u64, len: u64, access: Access) -> Option> { @@ -171,7 +229,7 @@ mod tests { /// What a ring 3 load at `v` sees. fn load(&self, v: u64) -> u8 { - self.phys[(self.leaf(v, Access::Read).unwrap() - PBASE) as usize] + self.phys[(self.phys_of(v).unwrap() - PBASE) as usize] } } @@ -197,19 +255,24 @@ mod tests { out } + /// Each leaf size the fake maps with, in pages. + const GRAINS: [u64; 3] = [1, 4, 16]; + #[test] fn a_read_across_shuffled_frames_gets_the_bytes_at_their_addresses() { - for seed in [1, 0x5eed, 0xdead_beef, 42] { - let space = Space::new(shuffled(PAGES as usize, seed)); - for (start, len) in ranges() { - let segs = space.window(start, len, Access::Read).expect("every page is readable"); - let want: Vec = (start..start + len).map(byte_at).collect(); - assert_eq!(space.read(&segs, 0, len), want, "seed {seed:#x}: [{start:#x}, +{len:#x})"); - // A view from any offset inside the window: `UserBytes::sub`. - for off in [1, PAGE_4K - 1, PAGE_4K, len / 2] { - if off < len { - let tail = len - off; - assert_eq!(space.read(&segs, off, tail), want[off as usize..], "seed {seed:#x}: +{off:#x}"); + for grain in GRAINS { + for seed in [1, 0x5eed, 0xdead_beef, 42] { + let space = Space::new(shuffled((PAGES / grain) as usize, seed), grain); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Read).expect("every page is readable"); + let want: Vec = (start..start + len).map(byte_at).collect(); + assert_eq!(space.read(&segs, 0, len), want, "grain {grain}, seed {seed:#x}: [{start:#x}, +{len:#x})"); + // A view from any offset inside the window: `UserBytes::sub`. + for off in [1, PAGE_4K - 1, PAGE_4K, len / 2] { + if off < len { + let tail = len - off; + assert_eq!(space.read(&segs, off, tail), want[off as usize..], "grain {grain}, seed {seed:#x}: +{off:#x}"); + } } } } @@ -218,15 +281,17 @@ mod tests { #[test] fn a_write_across_shuffled_frames_lands_where_a_ring_3_load_reads_it() { - for seed in [3, 0xabcdef, 99] { - for (start, len) in ranges() { - let mut space = Space::new(shuffled(PAGES as usize, seed)); - let segs = space.window(start, len, Access::Write).expect("every page is writable"); - let src: Vec = (0..len).map(|i| !byte_at(i ^ seed)).collect(); - space.write(&segs, 0, &src); - for v in VBASE..VBASE + PAGES * PAGE_4K { - let want = if (start..start + len).contains(&v) { src[(v - start) as usize] } else { byte_at(v) }; - assert_eq!(space.load(v), want, "seed {seed:#x}: [{start:#x}, +{len:#x}) at {v:#x}"); + for grain in GRAINS { + for seed in [3, 0xabcdef, 99] { + for (start, len) in ranges() { + let mut space = Space::new(shuffled((PAGES / grain) as usize, seed), grain); + let segs = space.window(start, len, Access::Write).expect("every page is writable"); + let src: Vec = (0..len).map(|i| !byte_at(i ^ seed)).collect(); + space.write(&segs, 0, &src); + for v in VBASE..VBASE + PAGES * PAGE_4K { + let want = if (start..start + len).contains(&v) { src[(v - start) as usize] } else { byte_at(v) }; + assert_eq!(space.load(v), want, "grain {grain}, seed {seed:#x}: [{start:#x}, +{len:#x}) at {v:#x}"); + } } } } @@ -234,46 +299,77 @@ mod tests { #[test] fn the_segments_are_the_maximal_runs_and_cover_the_window_exactly() { - for seed in [7, 0x1234_5678] { - let space = Space::new(shuffled(PAGES as usize, seed)); - for (start, len) in ranges() { - let segs = space.window(start, len, Access::Read).unwrap(); - assert_eq!(segs.iter().map(|s| s.len).sum::(), len); - assert_eq!(segs[0].phys, space.leaf(start, Access::Read).unwrap()); - for pair in segs.windows(2) { - assert_ne!(pair[0].phys + pair[0].len, pair[1].phys, "two runs that follow were not joined"); + for grain in GRAINS { + for seed in [7, 0x1234_5678] { + let space = Space::new(shuffled((PAGES / grain) as usize, seed), grain); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Read).unwrap(); + assert_eq!(segs.iter().map(|s| s.len).sum::(), len); + assert_eq!(segs[0].phys, space.phys_of(start).unwrap()); + for pair in segs.windows(2) { + assert_ne!(pair[0].phys + pair[0].len, pair[1].phys, "two runs that follow were not joined"); + } } } } } - /// A 2 MiB leaf, or a split window over one frame, is every page in order: - /// one segment, which is one pin and one copy as before the walk was per page. + /// The cost the walk is held to: one lookup per leaf the range touches, + /// whatever the leaf's size. + #[test] + fn a_range_is_looked_up_once_per_leaf() { + for grain in GRAINS { + let space = Space::new(shuffled((PAGES / grain) as usize, 5), grain); + for (start, len) in ranges() { + space.lookups.set(0); + space.window(start, len, Access::Read).unwrap(); + let leaves = (start + len - 1 - VBASE) / space.leaf_bytes() - (start - VBASE) / space.leaf_bytes() + 1; + assert_eq!(space.lookups.get(), leaves, "grain {grain}: [{start:#x}, +{len:#x})"); + } + } + // 64 MiB from a byte past a 2 MiB boundary: 33 leaves of 2 MiB, frames + // in order, so 33 lookups and one run. + let (start, len) = (VBASE + 1, 32 * PAGE_2M); + let mut lookups = 0; + let mut segs = Vec::new(); + segments( + start, + len, + |at| { + lookups += 1; + Some((PBASE + at, PAGE_2M - at % PAGE_2M)) + }, + |s| segs.push(s), + ) + .unwrap(); + assert_eq!(lookups, 33); + assert_eq!(segs, [Segment { phys: PBASE + start, len }]); + } + + /// A 2 MiB leaf, or a split window over one frame, is every page in order. #[test] fn frames_in_order_are_one_segment() { - let space = Space::new((0..PAGES).collect()); - for (start, len) in ranges() { - let segs = space.window(start, len, Access::Write).unwrap(); - assert_eq!(segs, [Segment { phys: space.leaf(start, Access::Read).unwrap(), len }]); + for grain in GRAINS { + let space = Space::new((0..PAGES / grain).collect(), grain); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Write).unwrap(); + assert_eq!(segs, [Segment { phys: space.phys_of(start).unwrap(), len }]); + } } } - /// Frames in reverse: every page seam is a segment seam. + /// Frames in reverse: every leaf seam is a segment seam. #[test] - fn frames_in_reverse_are_one_segment_per_page() { - let space = Space::new((0..PAGES).rev().collect()); + fn frames_in_reverse_are_one_segment_per_leaf() { + let space = Space::new((0..PAGES).rev().collect(), 1); let segs = space.window(VBASE + PAGE_4K - 3, PAGE_4K + 6, Access::Read).unwrap(); - assert_eq!( - segs.iter().map(|s| s.len).collect::>(), - [3, PAGE_4K, 3], - "{segs:x?}" - ); + assert_eq!(segs.iter().map(|s| s.len).collect::>(), [3, PAGE_4K, 3], "{segs:x?}"); } #[test] fn an_absent_page_anywhere_refuses_the_window() { for hole in [0, 1, 5, PAGES - 1] { - let mut space = Space::new(shuffled(PAGES as usize, hole + 1)); + let mut space = Space::new(shuffled(PAGES as usize, hole + 1), 1); space.absent.push(hole); let h = VBASE + hole * PAGE_4K; for access in [Access::Read, Access::Write] { @@ -286,7 +382,7 @@ mod tests { #[test] fn a_read_only_page_between_writable_ones_refuses_a_write_only() { - let mut space = Space::new(shuffled(PAGES as usize, 11)); + let mut space = Space::new(shuffled(PAGES as usize, 11), 1); space.ro.push(2); let (start, len) = (VBASE + PAGE_4K, 3 * PAGE_4K); assert_eq!(space.window(start, len, Access::Write), None); @@ -296,16 +392,95 @@ mod tests { #[test] fn an_empty_or_kernel_window_is_refused_before_the_walk() { - let never = |_: u64| -> Option { panic!("walked") }; + let never = |_: u64| -> Option<(u64, u64)> { panic!("walked") }; assert_eq!(segments(PAGE_2M, 0, never, |_| panic!("emitted")), None); assert_eq!(segments(USER_TOP - 8, 16, never, |_| panic!("emitted")), None); assert_eq!(segments(u64::MAX - 4, 8, never, |_| panic!("emitted")), None); } + /// Pin counts per fake 4 KiB frame, which refuse the pin call numbered + /// `refuse` and panic on an unpin of a frame nothing pinned, as the PMM does. + struct Frames { + count: RefCell>, + calls: Cell, + refuse: Option, + } + + impl Frames { + fn new(refuse: Option) -> Self { + Frames { count: RefCell::new(vec![0; PAGES as usize]), calls: Cell::new(0), refuse } + } + + fn of(run: Segment) -> core::ops::RangeInclusive { + ((run.phys - PBASE) / PAGE_4K) as usize..=((run.phys + run.len - 1 - PBASE) / PAGE_4K) as usize + } + + fn pinned(&self) -> Vec { + self.count.borrow().clone() + } + } + + impl Pins for &Frames { + fn pin(&mut self, run: Segment) -> bool { + let call = self.calls.get(); + self.calls.set(call + 1); + if self.refuse == Some(call) { + return false; + } + for f in Frames::of(run) { + self.count.borrow_mut()[f] += 1; + } + true + } + + fn unpin(&mut self, run: Segment) { + for f in Frames::of(run) { + let mut count = self.count.borrow_mut(); + count[f] = count[f].checked_sub(1).expect("unpin of a frame that was not pinned"); + } + } + } + + /// Frames in a shuffled order with one frame mapped twice, as a shared + /// mapping is: every run holds its own pin on it. + fn twice_mapped() -> Space { + let mut order = shuffled(PAGES as usize, 0x77); + order[9] = order[3]; + Space::new(order, 1) + } + + #[test] + fn every_run_is_pinned_until_the_window_is_dropped() { + let space = twice_mapped(); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Read).unwrap(); + let frames = Frames::new(None); + let mut want = vec![0; PAGES as usize]; + for &run in &segs { + for f in Frames::of(run) { + want[f] += 1; + } + } + { + let pinned = Pinned::pin(&segs[..], &frames).expect("nothing refuses"); + assert_eq!(frames.pinned(), want, "[{start:#x}, +{len:#x}) while pinned"); + let got = space.read(pinned.runs(), 0, len); + assert_eq!(got, (start..start + len).map(|v| space.load(v)).collect::>()); + } + assert_eq!(frames.pinned(), vec![0; PAGES as usize], "[{start:#x}, +{len:#x}) after the copy"); + } + } + #[test] - #[should_panic(expected = "past the window")] - fn a_piece_past_the_window_is_a_kernel_bug() { - let segs = [Segment { phys: PBASE, len: 8 }, Segment { phys: PBASE + 64, len: 8 }]; - pieces(&segs, 4, 13, |_, _, _| {}); + fn a_run_that_cannot_be_pinned_refuses_the_window_and_leaves_nothing_pinned() { + let space = twice_mapped(); + let segs = space.window(VBASE + PAGE_4K / 2, 20 * PAGE_4K, Access::Read).unwrap(); + assert!(segs.len() > 3, "{segs:x?}"); + for refused in 0..segs.len() { + let frames = Frames::new(Some(refused)); + assert!(Pinned::pin(&segs[..], &frames).is_none(), "run {refused} refused"); + assert_eq!(frames.calls.get(), refused + 1, "no run after the refused one is tried"); + assert_eq!(frames.pinned(), vec![0; PAGES as usize], "run {refused} refused"); + } } } diff --git a/toyos-userbound/src/span.rs b/toyos-userbound/src/span.rs index 3e7a26379b..0fa7f759cb 100644 --- a/toyos-userbound/src/span.rs +++ b/toyos-userbound/src/span.rs @@ -89,7 +89,7 @@ pub enum Access { Write, } -/// The smallest leaf, and the grain a user window is walked at. +/// The grain a split window grants rights at. pub const PAGE_4K: u64 = 4096; #[cfg(test)] From f03f81be800fff841e4dafbbbf7faea37bd96796 Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 04:39:53 +0200 Subject: [PATCH 4/7] munmap_reissues_second_read_window fails at once on a read refused before it parks A kernel that refuses the buffer answers the read without parking it, and the roster wait would then spin until the harness ceiling. The wait also ends on the read's answer, and an answer before the park is the red. Co-Authored-By: Claude Opus 5.5 --- .../src/bin/munmap_reissues_second_read_window.rs | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs b/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs index b7e5291598..765a30bc24 100644 --- a/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs +++ b/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs @@ -76,10 +76,14 @@ fn main() { }) }; + // A read the kernel refuses returns without parking. roster::await_true(|| { - roster::my_threads(&cap).iter().any(|&(is_thread, state)| is_thread && state == roster::BLOCKED) - && ready.load(Ordering::SeqCst) == 1 + result.load(Ordering::SeqCst) != i64::MIN + || roster::my_threads(&cap).iter().any(|&(is_thread, state)| is_thread && state == roster::BLOCKED) + && ready.load(Ordering::SeqCst) == 1 }); + let early = result.load(Ordering::SeqCst); + assert_eq!(early, i64::MIN, "the reader's read answered {early} before it parked"); unsafe { munmap(upper, PAGE_2M) }.expect("munmap the buffer's second mapping"); From fa99614b74dbe48ded12a2ce48c81afa696f96e3 Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 08:45:53 +0200 Subject: [PATCH 5/7] Round 3 of #593: a write across a split window's pages is tested per page, and one pin path serves both copies The review found that the write-rights check for a split window now rests on the 4096 `walk` reports for a 4 KiB leaf, and that no test fails if it says 2 MiB: `segments` then checks a write window only at its first page, to the end of its 2 MiB window. `abuse_readonly_copyout` gains the arm that starts a write on a writable page and runs into a read-only one. It finds the first page above a `.bss` static that refuses a 1-byte `read`, without changing a byte (each probed page is offered its own byte back), asserts the kernel can read that page, and then asserts `read(fd, page - 8, 16)` answers `BadAddress` with the page's bytes unchanged. In this build the last RW segment ends at 0x9260c and the static sits at 0x91398, so the arm straddles 0x93000 inside the image's first window. The AArch64 `leaf` is `match self.never {}`: no leaf size, so nothing to mutate there. Also from the review: - `translate_writable` is gone; `translate_user` asks `leaf`, which now carries the ring 3 store contract. - `object_run` and `window` share `pinned`: fault in, walk under the lock, place each run, pin. The one-run-per-window bound and its panic live there. - The track records the in-guest cost of the bulk window's `Vec` as stage 1's owed measurement, and stage 3 names `window_split` as the bound it breaks. - Filed: a user-copy pin that is never released reds no test. Co-Authored-By: Claude Opus 5.5 --- ...pin-that-is-never-released-reds-no-test.md | 26 ++++++++ ...b-pages-and-that-caps-the-process-count.md | 14 +++- kernel/src/arch/aarch64/paging.rs | 4 -- kernel/src/arch/x86_64/paging.rs | 23 +++---- kernel/src/user_ptr.rs | 49 +++++++------- .../src/bin/abuse_readonly_copyout.rs | 65 ++++++++++++++++++- 6 files changed, 133 insertions(+), 48 deletions(-) create mode 100644 issues/kernel/a-user-copy-pin-that-is-never-released-reds-no-test.md diff --git a/issues/kernel/a-user-copy-pin-that-is-never-released-reds-no-test.md b/issues/kernel/a-user-copy-pin-that-is-never-released-reds-no-test.md new file mode 100644 index 0000000000..2f89426942 --- /dev/null +++ b/issues/kernel/a-user-copy-pin-that-is-never-released-reds-no-test.md @@ -0,0 +1,26 @@ +--- +status: open +kind: defect +opened: 2026-09-29 +--- + +# A user-copy pin that is never released reds no test + +`kernel/src/user_ptr.rs`'s `impl Pins for Pmm` is the one pin path no host +test reaches: the host tests of `toyos_userbound::Pinned` drive a counting +fake. With its `unpin` emptied (`fn unpin(&mut self, _run: Segment) {}`), +every frame a syscall ever copied through keeps a pin. `pmm::free_page` still +counts such a frame free, and `alloc_page` and `alloc_contiguous` skip it for +the rest of the boot, so the machine loses memory that no free count shows and +no known guest test reads. + +## Exit condition + +A test that goes red under that mutation: for example, a guest job that copies +into and unmaps more frames than the guest has, or a census that counts free +frames still holding a pin and a test that asserts it returns to zero once no +copy is in flight. Then this file is deleted. + +## Owner + +`kernel/src/user_ptr.rs`, `kernel/src/mm/pmm.rs`. Nobody holds it. diff --git a/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md b/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md index 2553125972..4dea86b1b3 100644 --- a/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md +++ b/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md @@ -45,6 +45,15 @@ reserves its frame count when it is made and is refused by name then, never killed at a fault. No promotion (the kernel has no thread to do it) and no live split of a user leaf. +1. **Scatter-gather user copies** (#593) owe one measurement: what the bulk + window's `Vec` costs under the kernel heap's lock. It is one + heap allocation per bulk copy, taken before the address-space lock; the + host A/B in #593's review put a 64 MiB `read` at 1.28–1.31× main's, 32 + to 40 ns more. No job measures it in-guest yet: the stage owes + `copy_cost`, a job beside `syscall_cost` that times `read` into a 64 KiB + and a 64 MiB window, and its metal row, so that + `cargo test -- --metal copy_cost` at the stage's merge and at its base is + the A/B. 2. **Superframe frame allocator.** 4 KiB frames carved from 2 MiB superframes; the pin rule held per superframe, so no pinned frame is reissued and no superframe holding one is handed out whole; a watermark @@ -57,7 +66,10 @@ live split of a user leaf. the frames a process owns, with per-process counts; stacks are demand-zero and TLS is 4 KiB. A typed syscall value crossing a page is served in segments, so `is_user_object` stops refusing a straddle and - `abuse_page_straddle`'s refusals become delivery verdicts. Exit: a + `abuse_page_straddle`'s refusals become delivery verdicts. A 2 MiB window + of 4 KiB leaves from arbitrary frames is up to 512 runs, so the one-run + bound `user_ptr::window_split` panics on becomes reachable from userland + and goes with the stage. Exit: a process's floor on the T14 against the 10 to 14 MB above. 4. **Demand-zero anonymous `mmap`.** A fault installs a 2 MiB leaf only for a whole aligned 2 MiB span of the mapping; unmapping a range no thread touched diff --git a/kernel/src/arch/aarch64/paging.rs b/kernel/src/arch/aarch64/paging.rs index eb969396e1..0fd3db1943 100644 --- a/kernel/src/arch/aarch64/paging.rs +++ b/kernel/src/arch/aarch64/paging.rs @@ -68,10 +68,6 @@ impl AddressSpace { match self.never {} } - pub fn translate_writable(&self, _vaddr: UserAddr) -> Option { - match self.never {} - } - pub fn leaf(&self, _vaddr: UserAddr, _access: toyos_userbound::Access) -> Option<(u64, u64)> { match self.never {} } diff --git a/kernel/src/arch/x86_64/paging.rs b/kernel/src/arch/x86_64/paging.rs index 1c65c277a3..466d1ecf4b 100644 --- a/kernel/src/arch/x86_64/paging.rs +++ b/kernel/src/arch/x86_64/paging.rs @@ -42,8 +42,6 @@ const ADDR_MASK_2M: u64 = 0x000F_FFFF_FFE0_0000; /// Every upper-level table entry's flags: present, writable, user. const TABLE_FLAGS: u64 = PAGE_PRESENT | PAGE_WRITE | PAGE_USER; -/// The rights a ring 3 store needs at every level of its walk. -const STORE: u64 = PAGE_USER | PAGE_WRITE; impl Prot { /// The permission bits a leaf entry carries; address and cache policy stay the caller's. @@ -561,20 +559,15 @@ impl AddressSpace { self.walk(vaddr).map(|(dm, _, _)| dm) } - /// As [`translate`](Self::translate), but only where a user store would - /// land: every level of the walk grants `USER` and `WRITE`, as the MMU - /// demands of a ring 3 store under `CR0.WP`. A kernel copy into user - /// memory goes through this, so a syscall cannot write a page the process - /// itself may not — the clock page, a shared library's `.text`. - pub fn translate_writable(&self, vaddr: UserAddr) -> Option { - self.walk(vaddr).and_then(|(dm, rights, _)| (rights & STORE == STORE).then_some(dm)) - } - - /// `vaddr`'s physical address as [`translate`](Self::translate) or - /// [`translate_writable`](Self::translate_writable) answers it for - /// `access`, and the bytes from it to the end of the leaf that maps it, - /// which that one walk answers for too. + /// `vaddr`'s physical address and the bytes from it to the end of the leaf + /// that maps it, which that one walk answers for. A `Write` is answered + /// only where a user store would land: every level of the walk grants + /// `USER` and `WRITE`, as the MMU demands of a ring 3 store under + /// `CR0.WP`. A kernel copy into user memory goes through this, so a syscall + /// cannot write a page the process itself may not — the clock page, a + /// shared library's `.text`. pub fn leaf(&self, vaddr: UserAddr, access: toyos_userbound::Access) -> Option<(u64, u64)> { + const STORE: u64 = PAGE_USER | PAGE_WRITE; let (dm, rights, size) = self.walk(vaddr)?; let granted = match access { toyos_userbound::Access::Read => true, diff --git a/kernel/src/user_ptr.rs b/kernel/src/user_ptr.rs index 6665ae32bb..2f2ea160f5 100644 --- a/kernel/src/user_ptr.rs +++ b/kernel/src/user_ptr.rs @@ -72,10 +72,7 @@ fn translate_now( addr: UserAddr, access: Access, ) -> Option { - match access { - Access::Read => space.translate(addr), - Access::Write => space.translate_writable(addr), - } + space.leaf(addr, access).map(|(phys, _)| crate::mm::DirectMap::from_phys(phys)) } /// Translate a user virtual address to its direct-map address, demand-paging it in if needed; `pub(crate)` because the futex word outlives its syscall. @@ -103,25 +100,10 @@ fn object_run(ptr: UserAddr, size: usize, align: usize, access: Access) -> Resul if !toyos_userbound::is_user_object(ptr.raw(), size as u64, align as u64) { return Err(SyscallError::BadAddress); } - fault_in(ptr, size, access).ok_or(SyscallError::BadAddress)?; - let pt = crate::process::current_address_space(); - let guard = pt.lock(); - let mut run = None; - toyos_userbound::segments( - ptr.raw(), - size as u64, - |at| guard.leaf(UserAddr::new(at), access), - |next| { - if run.replace(next).is_some() { - window_split(ptr, size) - } - }, - ) - .ok_or(SyscallError::BadAddress)?; - let run = run.expect("`segments` emits a run whenever it places the window"); - let pins = FramePins::pin([run], Pmm).ok_or(SyscallError::BadAddress)?; - drop(guard); - Ok((crate::mm::DirectMap::from_phys(run.phys).as_mut_ptr(), pins)) + let pins = pinned(ptr, size, access, |_| [Segment { phys: 0, len: 0 }], |one, i, run| one[i] = run) + .ok_or(SyscallError::BadAddress)?; + let phys = pins.runs()[0].phys; + Ok((crate::mm::DirectMap::from_phys(phys).as_mut_ptr(), pins)) } /// A 2 MiB window is one frame in order, so it holds at most one run. @@ -383,13 +365,25 @@ fn fault_in(ptr: UserAddr, len: usize, access: Access) -> Option { Some(windows) } -/// Validate `[ptr, ptr+len)` as a user window, walked leaf by leaf into its physical runs, and pin every frame they cover. The pins are taken under the address-space lock over a translation that still names each frame, so a concurrent `munmap` — which needs that same lock to free the range — cannot reissue a frame between the confirmation and the pin. Nothing is allocated or freed under the lock but on a refused pin. +/// A bulk buffer's runs, one per 2 MiB window at most, allocated before the lock. fn window(ptr: UserAddr, len: usize, access: Access) -> Option>> { if !toyos_userbound::in_user_half(ptr.raw(), len as u64) { return None; } + pinned(ptr, len, access, Vec::with_capacity, |runs, _, run| runs.push(run)) +} + +/// Faults `[ptr, ptr+len)` in, walks it leaf by leaf into its physical runs, and pins every frame they cover. `runs` makes the store for as many runs as the range touches 2 MiB windows, and `place` puts run `i` in it. The pins are taken under the address-space lock over a translation that still names each frame, so a concurrent `munmap` — which needs that same lock to free the range — cannot reissue a frame between the confirmation and the pin. Nothing is allocated or freed under the lock but on a refused pin. +fn pinned>( + ptr: UserAddr, + len: usize, + access: Access, + runs: impl FnOnce(usize) -> R, + mut place: impl FnMut(&mut R, usize, Segment), +) -> Option> { let windows = fault_in(ptr, len, access)?; - let mut runs = Vec::with_capacity(windows); + let mut runs = runs(windows); + let mut placed = 0; let pt = crate::process::current_address_space(); let guard = pt.lock(); toyos_userbound::segments( @@ -397,10 +391,11 @@ fn window(ptr: UserAddr, len: usize, access: Access) -> Option u64 { @@ -66,6 +79,55 @@ fn refused(what: &str, addr: u64) { syscall::close(fd); } +/// The first page at or above [`BSS`], inside its 2 MiB window, a `read` may +/// not write; each page below it took a 1-byte `read` of its own first byte. +fn first_unwritable_page(probe: RawHandle) -> u64 { + let pipe = syscall::pipe().expect("pipe"); + let bss = &raw const BSS as u64; + let window_end = (bss & !(PAGE_2M as u64 - 1)) + PAGE_2M as u64; + let mut page = bss & !(PAGE_4K - 1); + while page < window_end { + let ret = raw(SYS_WRITE, pipe.write.0 as u64, page, 1); + assert_eq!(ret, 1, "write from {page:#x}, in .bss's window: {ret:#x}"); + syscall::read(pipe.read, &mut [0u8; 1]).expect("drain the pipe"); + let own = unsafe { (page as *const u8).read_volatile() }; + syscall::seek(probe, SeekFrom::Start(own as u64)).expect("seek the probe"); + let ret = raw(SYS_READ, probe.0 as u64, page, 1); + if SyscallError::from_u64(ret) == Some(SyscallError::BadAddress) { + syscall::close(pipe.read); + syscall::close(pipe.write); + return page; + } + assert_eq!(ret, 1, "a 1-byte read into {page:#x}: {ret:#x}"); + page += PAGE_4K; + } + panic!("no page above .bss at {bss:#x} refuses a write below {window_end:#x}: the image ends on its window's edge"); +} + +/// A `read` that starts on a writable page and runs into the read-only page +/// after it, in one 2 MiB window, is refused whole. +fn refused_across_pages() { + let flags = OpenFlags::READ | OpenFlags::WRITE | OpenFlags::CREATE | OpenFlags::TRUNCATE; + let probe = syscall::open(PROBE_PATH, flags).expect("create the probe file"); + let every: Vec = (0..=255).collect(); + syscall::write(probe, &every).expect("fill the probe file"); + let page = first_unwritable_page(probe); + assert!(page > &raw const BSS as u64, "BSS's own page refuses a write"); + let at = page - (STRADDLE / 2) as u64; + let before = snapshot(at); + // The writable half is offered its own bytes and the read-only half their + // complement, so a write to the read-only page shows and one below harms nothing. + let offered: [u8; STRADDLE] = core::array::from_fn(|i| if i < STRADDLE / 2 { before[i] } else { !before[i] }); + syscall::seek(probe, SeekFrom::Start(STRADDLE_AT)).expect("seek the probe"); + syscall::write(probe, &offered).expect("write the straddle's bytes"); + syscall::seek(probe, SeekFrom::Start(STRADDLE_AT)).expect("seek the probe"); + let ret = raw(SYS_READ, probe.0 as u64, at, STRADDLE as u64); + assert_eq!(before, snapshot(at), "read wrote across a writable page into the read-only one at {page:#x}"); + assert_eq!(SyscallError::from_u64(ret), Some(SyscallError::BadAddress), "read across into {page:#x}: {ret:#x}"); + syscall::close(probe); + syscall::delete(PROBE_PATH).expect("delete the probe file"); +} + fn main() { if std::env::args().nth(1).as_deref() == Some(CHILD_ARG) { let page = unsafe { core::ptr::read_volatile(CLOCK_PAGE as *const ClockPage) }; @@ -95,6 +157,7 @@ fn main() { unsafe { syscall::munmap(ro, PAGE_2M) }.expect("munmap"); refused("this program's own text", main as *const () as u64); + refused_across_pages(); // Last: a written clock page asserts in every stamp this process and every // other one takes, the verdict's own path included, so the arms whose harm From 12a452000e01b3ab1d60b929de171ff90fa876d3 Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 12:37:47 +0200 Subject: [PATCH 6/7] Delete TECHNICAL_DEBT section 2 and the translate_writable issue This branch makes both false. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs --- ...writes-through-a-read-only-user-mapping.md | 34 ------------------- 1 file changed, 34 deletions(-) delete mode 100644 issues/isolation/a-syscall-writes-through-a-read-only-user-mapping.md diff --git a/issues/isolation/a-syscall-writes-through-a-read-only-user-mapping.md b/issues/isolation/a-syscall-writes-through-a-read-only-user-mapping.md deleted file mode 100644 index 995d7dbd8e..0000000000 --- a/issues/isolation/a-syscall-writes-through-a-read-only-user-mapping.md +++ /dev/null @@ -1,34 +0,0 @@ ---- -status: assigned -kind: defect -opened: 2026-09-26 ---- - -# A syscall writes through a read-only user mapping - -Found in the review of the logging track (PR #527) and fixed on that branch; -held by it, and deleted by whichever change first lands after it on `main`. - -**On `main`, a kernel copy into user memory checks that the page is present, -never that the process may store to it.** `user_ptr::window` and -`user_ptr::copy_out` translate through `AddressSpace::translate`, which walks to -any present leaf, and then write through the direct map, where the leaf's -`WRITE` bit does not apply. So any syscall that fills a caller's buffer — -`read`, `fstat`, `sched_info`, `process_stats` — writes a page mapped -`Prot::Read` or `Prot::ReadExec`: - -- an `mmap(PROT_READ)` region, which the process then reads back rewritten; -- its own `.text`; -- a shared library's `.text`, which `LibMemory::Shared` maps from one cached - image into every process that loads it — so the write is **cross-process**; -- on the branch, the clock page, one frame every address space maps. - -`tests/toyos-rust-tests/src/bin/abuse_readonly_copyout.rs` is the gate. With -the branch's `translate_writable` and the `Access` it threads through -`user_ptr` reverted as one patch, it reds: `read wrote into a read-only mmap`, -exit 1. Put back, it is green, exit 0. The code that patch reverts to is -`main`'s byte for byte. - -The oracle is the MMU's own rule (Intel SDM vol. 3A §4.6.1): a ring 3 store -needs `R/W` and `U/S` set at every level of the walk under `CR0.WP`, which is -what `translate_writable` checks. From c671f22d3abc0113773f6050fc26439af4ff7401 Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 12:50:51 +0200 Subject: [PATCH 7/7] Delete section 2 of kernel/TECHNICAL_DEBT.md This branch makes it false. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01U6SVYFkdvV2t38KzNrESxs --- kernel/TECHNICAL_DEBT.md | 16 ---------------- 1 file changed, 16 deletions(-) diff --git a/kernel/TECHNICAL_DEBT.md b/kernel/TECHNICAL_DEBT.md index 823b02167d..da503538fc 100644 --- a/kernel/TECHNICAL_DEBT.md +++ b/kernel/TECHNICAL_DEBT.md @@ -8,22 +8,6 @@ Removed. Kernel PML4 has no PML4[0..255] entries. Process page tables start empty and build user mappings on demand. SMP trampoline uses the bootloader's PML4 for AP transition, then switches to kernel PML4 in ap_entry. -## 2. user_ptr assumes physically contiguous user buffers - -`user_ptr::window` walks every 2 MiB boundary a buffer crosses and refuses one -whose pages are not physically adjacent, so a non-contiguous buffer is a -`BadAddress` and not a silent misread. What remains is that it *only* refuses: -ToyOS allocates user memory in contiguous 2 MiB blocks, so nothing produces one -today, and swap, COW fork or non-contiguous VMAs would turn ordinary reads and -writes into refusals rather than into corruption. - -**Fix, when one of those arrives:** `UserBytes`/`UserBytesMut` already hand out -no reference, so the change is confined to how a window addresses its pages — -a run list instead of one base pointer, with `read_at`/`write_at` splitting at -the boundaries. No caller's signature moves. - -**Not urgent while all user allocations are 2MB-aligned contiguous blocks.** - ## ~~3. No DmaAddr newtype~~ DONE Added `DmaAddr` newtype in addr.rs. `DmaPool::page_phys()` returns `DmaAddr`.