diff --git a/issues/isolation/a-syscall-writes-through-a-read-only-user-mapping.md b/issues/isolation/a-syscall-writes-through-a-read-only-user-mapping.md deleted file mode 100644 index 995d7dbd8e..0000000000 --- a/issues/isolation/a-syscall-writes-through-a-read-only-user-mapping.md +++ /dev/null @@ -1,34 +0,0 @@ ---- -status: assigned -kind: defect -opened: 2026-09-26 ---- - -# A syscall writes through a read-only user mapping - -Found in the review of the logging track (PR #527) and fixed on that branch; -held by it, and deleted by whichever change first lands after it on `main`. - -**On `main`, a kernel copy into user memory checks that the page is present, -never that the process may store to it.** `user_ptr::window` and -`user_ptr::copy_out` translate through `AddressSpace::translate`, which walks to -any present leaf, and then write through the direct map, where the leaf's -`WRITE` bit does not apply. So any syscall that fills a caller's buffer — -`read`, `fstat`, `sched_info`, `process_stats` — writes a page mapped -`Prot::Read` or `Prot::ReadExec`: - -- an `mmap(PROT_READ)` region, which the process then reads back rewritten; -- its own `.text`; -- a shared library's `.text`, which `LibMemory::Shared` maps from one cached - image into every process that loads it — so the write is **cross-process**; -- on the branch, the clock page, one frame every address space maps. - -`tests/toyos-rust-tests/src/bin/abuse_readonly_copyout.rs` is the gate. With -the branch's `translate_writable` and the `Access` it threads through -`user_ptr` reverted as one patch, it reds: `read wrote into a read-only mmap`, -exit 1. Put back, it is green, exit 0. The code that patch reverts to is -`main`'s byte for byte. - -The oracle is the MMU's own rule (Intel SDM vol. 3A §4.6.1): a ring 3 store -needs `R/W` and `U/S` set at every level of the walk under `CR0.WP`, which is -what `translate_writable` checks. diff --git a/issues/kernel/a-user-copy-pin-that-is-never-released-reds-no-test.md b/issues/kernel/a-user-copy-pin-that-is-never-released-reds-no-test.md new file mode 100644 index 0000000000..2f89426942 --- /dev/null +++ b/issues/kernel/a-user-copy-pin-that-is-never-released-reds-no-test.md @@ -0,0 +1,26 @@ +--- +status: open +kind: defect +opened: 2026-09-29 +--- + +# A user-copy pin that is never released reds no test + +`kernel/src/user_ptr.rs`'s `impl Pins for Pmm` is the one pin path no host +test reaches: the host tests of `toyos_userbound::Pinned` drive a counting +fake. With its `unpin` emptied (`fn unpin(&mut self, _run: Segment) {}`), +every frame a syscall ever copied through keeps a pin. `pmm::free_page` still +counts such a frame free, and `alloc_page` and `alloc_contiguous` skip it for +the rest of the boot, so the machine loses memory that no free count shows and +no known guest test reads. + +## Exit condition + +A test that goes red under that mutation: for example, a guest job that copies +into and unmaps more frames than the guest has, or a census that counts free +frames still holding a pin and a test that asserts it returns to zero once no +copy is in flight. Then this file is deleted. + +## Owner + +`kernel/src/user_ptr.rs`, `kernel/src/mm/pmm.rs`. Nobody holds it. diff --git a/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md b/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md index 037523116b..4dea86b1b3 100644 --- a/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md +++ b/issues/kernel/process-memory-is-2-mib-pages-and-that-caps-the-process-count.md @@ -18,9 +18,6 @@ IOMMU's superpages, the kernel's own mappings. One paging design serves both, on x86-64 and on the ARM64 the tree is kept portable for (`issues/kernel/toyos-runs-on-arm64.md`). -**Blocked on:** the owner's word to start. Recorded, not started, at the -owner's instruction. - **The constraint that makes this a track, measured on the T14** (run 29 readback, `PMM: 100/16038MB used` at 11 s with four user processes and three kernel threads live; the same record's rows: stack 16 pages held, init-tls 5, @@ -38,3 +35,51 @@ longer shares a 2 MiB page with a neighbour's registers, so the relocation that `kernel/src/pcidev/` does before a hand-over is no longer needed for a window firmware placed apart from its neighbours. Placing a window inside the host bridge's firmware-reported apertures stays correct under either page size. + +## Stages + +The owner approved starting this work ("Start it") and delegated every +kernel-interface decision to the orchestrator, whose stages and rulings +follow. **Commit is strict**, by the orchestrator's ruling: a mapping +reserves its frame count when it is made and is refused by name then, never +killed at a fault. No promotion (the kernel has no thread to do it) and no +live split of a user leaf. + +1. **Scatter-gather user copies** (#593) owe one measurement: what the bulk + window's `Vec` costs under the kernel heap's lock. It is one + heap allocation per bulk copy, taken before the address-space lock; the + host A/B in #593's review put a 64 MiB `read` at 1.28–1.31× main's, 32 + to 40 ns more. No job measures it in-guest yet: the stage owes + `copy_cost`, a job beside `syscall_cost` that times `read` into a 64 KiB + and a 64 MiB window, and its metal row, so that + `cargo test -- --metal copy_cost` at the stage's merge and at its base is + the A/B. +2. **Superframe frame allocator.** 4 KiB frames carved from 2 MiB + superframes; the pin rule held per superframe, so no pinned frame is + reissued and no superframe holding one is handed out whole; a watermark + refuses an allocation before the machine runs dry. Exit: host tests of + carve, free, pin and the refusal; every 2 MiB caller of the PMM unchanged; + `user_copy_spans_windows` and `munmap_reissues_second_read_window` red + again under their negative controls, because each is two physical runs + only while the allocator hands out the lowest free frame first. +3. **4 KiB user leaves.** A software `OWNED` bit in the leaf is the ledger of + the frames a process owns, with per-process counts; stacks are demand-zero + and TLS is 4 KiB. A typed syscall value crossing a page is served in + segments, so `is_user_object` stops refusing a straddle and + `abuse_page_straddle`'s refusals become delivery verdicts. A 2 MiB window + of 4 KiB leaves from arbitrary frames is up to 512 runs, so the one-run + bound `user_ptr::window_split` panics on becomes reachable from userland + and goes with the stage. Exit: a + process's floor on the T14 against the 10 to 14 MB above. +4. **Demand-zero anonymous `mmap`.** A fault installs a 2 MiB leaf only for a + whole aligned 2 MiB span of the mapping; unmapping a range no thread touched + sends no IPI; `munmap` refuses a size that is not the whole mapping, by the + orchestrator's ruling. Exit: `mmap` of more than the watermark allows + refused at `mmap`, and a guest that touches one page of a large mapping + holds one frame. +5. **File-backed faults at 4 KiB.** ROOT's read-only text is shared + zero-copy. Exit: two processes of one binary hold its text once. +6. **`/apps` image sharing.** The design is the owner's decision, owed when + this stage is reached. +7. **Retire 2 MiB where it no longer pays.** Exit: every remaining 2 MiB user + leaf is a stage-4 whole span or a device or DMA grant. diff --git a/kernel/TECHNICAL_DEBT.md b/kernel/TECHNICAL_DEBT.md index 823b02167d..da503538fc 100644 --- a/kernel/TECHNICAL_DEBT.md +++ b/kernel/TECHNICAL_DEBT.md @@ -8,22 +8,6 @@ Removed. Kernel PML4 has no PML4[0..255] entries. Process page tables start empty and build user mappings on demand. SMP trampoline uses the bootloader's PML4 for AP transition, then switches to kernel PML4 in ap_entry. -## 2. user_ptr assumes physically contiguous user buffers - -`user_ptr::window` walks every 2 MiB boundary a buffer crosses and refuses one -whose pages are not physically adjacent, so a non-contiguous buffer is a -`BadAddress` and not a silent misread. What remains is that it *only* refuses: -ToyOS allocates user memory in contiguous 2 MiB blocks, so nothing produces one -today, and swap, COW fork or non-contiguous VMAs would turn ordinary reads and -writes into refusals rather than into corruption. - -**Fix, when one of those arrives:** `UserBytes`/`UserBytesMut` already hand out -no reference, so the change is confined to how a window addresses its pages — -a run list instead of one base pointer, with `read_at`/`write_at` splitting at -the boundaries. No caller's signature moves. - -**Not urgent while all user allocations are 2MB-aligned contiguous blocks.** - ## ~~3. No DmaAddr newtype~~ DONE Added `DmaAddr` newtype in addr.rs. `DmaPool::page_phys()` returns `DmaAddr`. diff --git a/kernel/src/arch/aarch64/paging.rs b/kernel/src/arch/aarch64/paging.rs index 5f532fa241..0fd3db1943 100644 --- a/kernel/src/arch/aarch64/paging.rs +++ b/kernel/src/arch/aarch64/paging.rs @@ -68,7 +68,7 @@ impl AddressSpace { match self.never {} } - pub fn translate_writable(&self, _vaddr: UserAddr) -> Option { + pub fn leaf(&self, _vaddr: UserAddr, _access: toyos_userbound::Access) -> Option<(u64, u64)> { match self.never {} } diff --git a/kernel/src/arch/x86_64/paging.rs b/kernel/src/arch/x86_64/paging.rs index 43991575b4..466d1ecf4b 100644 --- a/kernel/src/arch/x86_64/paging.rs +++ b/kernel/src/arch/x86_64/paging.rs @@ -556,21 +556,28 @@ impl AddressSpace { /// Checked here, not at the callers: a user space shallow-copies the /// kernel PML4 half, so a kernel address would otherwise walk to a writable kernel page. pub fn translate(&self, vaddr: UserAddr) -> Option { - self.walk(vaddr).map(|(dm, _)| dm) + self.walk(vaddr).map(|(dm, _, _)| dm) } - /// As [`translate`](Self::translate), but only where a user store would - /// land: every level of the walk grants `USER` and `WRITE`, as the MMU - /// demands of a ring 3 store under `CR0.WP`. A kernel copy into user - /// memory goes through this, so a syscall cannot write a page the process - /// itself may not — the clock page, a shared library's `.text`. - pub fn translate_writable(&self, vaddr: UserAddr) -> Option { + /// `vaddr`'s physical address and the bytes from it to the end of the leaf + /// that maps it, which that one walk answers for. A `Write` is answered + /// only where a user store would land: every level of the walk grants + /// `USER` and `WRITE`, as the MMU demands of a ring 3 store under + /// `CR0.WP`. A kernel copy into user memory goes through this, so a syscall + /// cannot write a page the process itself may not — the clock page, a + /// shared library's `.text`. + pub fn leaf(&self, vaddr: UserAddr, access: toyos_userbound::Access) -> Option<(u64, u64)> { const STORE: u64 = PAGE_USER | PAGE_WRITE; - self.walk(vaddr).and_then(|(dm, rights)| (rights & STORE == STORE).then_some(dm)) + let (dm, rights, size) = self.walk(vaddr)?; + let granted = match access { + toyos_userbound::Access::Read => true, + toyos_userbound::Access::Write => rights & STORE == STORE, + }; + granted.then(|| (dm.phys(), size - (vaddr.raw() & (size - 1)))) } - /// The direct-map address of `vaddr` and the rights every level of its walk grants in common. - fn walk(&self, vaddr: UserAddr) -> Option<(crate::mm::DirectMap, u64)> { + /// The direct-map address of `vaddr`, the rights every level of its walk grants in common, and the size of the leaf that maps it. + fn walk(&self, vaddr: UserAddr) -> Option<(crate::mm::DirectMap, u64, u64)> { let va = vaddr.raw(); if !toyos_userbound::is_user_addr(va) { return None; @@ -593,11 +600,11 @@ impl AddressSpace { return None; } let dm = crate::mm::DirectMap::from_phys((pte & ADDR_MASK) + (va & 0xFFF)); - return Some((dm, rights & pte)); + return Some((dm, rights & pte, 4096)); } let page_phys = pde & ADDR_MASK_2M; let offset = va & (PAGE_2M - 1); - Some((crate::mm::DirectMap::from_phys(page_phys + offset), rights)) + Some((crate::mm::DirectMap::from_phys(page_phys + offset), rights, PAGE_2M)) } /// Where `span` goes, top-down and never below the floor: a region the diff --git a/kernel/src/user_ptr.rs b/kernel/src/user_ptr.rs index 36b4212f7d..2f2ea160f5 100644 --- a/kernel/src/user_ptr.rs +++ b/kernel/src/user_ptr.rs @@ -2,11 +2,13 @@ //! user memory lands only where a ring 3 store from the process itself could. //! //! SMAP stays enabled; access goes through the direct map, never stac/clac. -//! Values are copied out, never referenced. **Every copy goes through -//! [`window`]**, which pins every frame it covers under the address-space lock -//! for the copy's life — a typed value's as much as a [`UserBytes`]/ -//! [`UserBytesMut`] buffer's — so a sibling's `munmap` cannot reissue the -//! backing between the translation and the copy, nor under a copy across a park. +//! Values are copied out, never referenced. **Every copy pins every frame it +//! covers** under the address-space lock for the copy's life — [`object_run`] +//! a typed value's, [`window`] a [`UserBytes`]/[`UserBytesMut`] buffer's — so a +//! sibling's `munmap` cannot reissue the backing between the translation and +//! the copy, nor under a copy across a park. +//! A window is the physical runs its pages sit in, and a buffer's copy is cut at +//! their seams; a typed value lies inside one 2 MiB page, so in one run. //! Every single-word user read is `read_volatile`, including the futex word //! and the crash dump's walk, both outside this module. @@ -16,6 +18,7 @@ use alloc::vec::Vec; use core::marker::PhantomData; use toyos_abi::syscall::SyscallError; +use toyos_userbound::{Pinned, Pins, Segment}; use crate::UserAddr; @@ -69,10 +72,7 @@ fn translate_now( addr: UserAddr, access: Access, ) -> Option { - match access { - Access::Read => space.translate(addr), - Access::Write => space.translate_writable(addr), - } + space.leaf(addr, access).map(|(phys, _)| crate::mm::DirectMap::from_phys(phys)) } /// Translate a user virtual address to its direct-map address, demand-paging it in if needed; `pub(crate)` because the futex word outlives its syscall. @@ -90,13 +90,27 @@ pub(crate) fn translate_user(addr: UserAddr, access: Access) -> Option(ptr: UserAddr, access: Access) -> Result<(*mut T, FramePin), SyscallError> { - let size = core::mem::size_of::(); - if !toyos_userbound::is_user_object(ptr.raw(), size as u64, core::mem::align_of::() as u64) { +fn object(ptr: UserAddr, access: Access) -> Result<(*mut T, FramePins<[Segment; 1]>), SyscallError> { + let (kptr, pins) = object_run(ptr, core::mem::size_of::(), core::mem::align_of::(), access)?; + Ok((kptr.cast(), pins)) +} + +/// The one run a typed value lies in, pinned, with nothing allocated. +fn object_run(ptr: UserAddr, size: usize, align: usize, access: Access) -> Result<(*mut u8, FramePins<[Segment; 1]>), SyscallError> { + if !toyos_userbound::is_user_object(ptr.raw(), size as u64, align as u64) { return Err(SyscallError::BadAddress); } - let (kptr, pin) = window(ptr, size, access).ok_or(SyscallError::BadAddress)?; - Ok((kptr.cast(), pin)) + let pins = pinned(ptr, size, access, |_| [Segment { phys: 0, len: 0 }], |one, i, run| one[i] = run) + .ok_or(SyscallError::BadAddress)?; + let phys = pins.runs()[0].phys; + Ok((crate::mm::DirectMap::from_phys(phys).as_mut_ptr(), pins)) +} + +/// A 2 MiB window is one frame in order, so it holds at most one run. +#[cold] +#[inline(never)] +fn window_split(ptr: UserAddr, len: usize) -> ! { + panic!("[{:#x}, +{len:#x}) lies in more physical runs than it touches 2 MiB windows", ptr.raw()) } /// Context for a single syscall invocation; the lifetime `'a` keeps validated references from escaping it. @@ -113,24 +127,12 @@ impl<'a> SyscallContext<'a> { /// A bulk buffer the kernel reads out of and never borrows. pub fn user_bytes(&self, ptr: UserAddr, len: u64) -> Option> { - let len = len as usize; - if len == 0 { - let kptr = core::ptr::NonNull::::dangling().as_ptr() as *const u8; - return Some(UserBytes { kptr, len, _pin: None, _scope: PhantomData }); - } - let (kptr, pin) = window(ptr, len, Access::Read)?; - Some(UserBytes { kptr: kptr as *const u8, len, _pin: Some(pin), _scope: PhantomData }) + Some(UserBytes(View::window(ptr, len, Access::Read)?)) } /// A bulk buffer the kernel writes into and never borrows. pub fn user_bytes_mut(&self, ptr: UserAddr, len: u64) -> Option> { - let len = len as usize; - if len == 0 { - let kptr = core::ptr::NonNull::::dangling().as_ptr(); - return Some(UserBytesMut { kptr, len, _pin: None, _scope: PhantomData }); - } - let (kptr, pin) = window(ptr, len, Access::Write)?; - Some(UserBytesMut { kptr, len, _pin: Some(pin), _scope: PhantomData }) + Some(UserBytesMut(View::window(ptr, len, Access::Write)?)) } /// Copy a user string of at most [`MAX_USER_STR`] bytes into kernel memory; an over-long or non-UTF-8 string is `InvalidArgument`, not `BadAddress`. @@ -144,14 +146,14 @@ impl<'a> SyscallContext<'a> { /// Read a typed value out of user memory, copied rather than borrowed. pub fn copy_in(&self, ptr: UserAddr) -> Result { - let (kptr, _pin) = object::(ptr, Access::Read)?; - // SAFETY: `object::` validated size/align inside one page and pinned the frame until `_pin` drops; `T: UserSafe` makes every bit pattern valid; `read_volatile` guards against a concurrent write from another thread of the same process. + let (kptr, _pins) = object::(ptr, Access::Read)?; + // SAFETY: `object::` validated size/align inside one run and pinned its frames until `_pins` drops; `T: UserSafe` makes every bit pattern valid; `read_volatile` guards against a concurrent write from another thread of the same process. Ok(unsafe { kptr.read_volatile() }) } /// Write a typed value into user memory. pub fn copy_out(&self, ptr: UserAddr, value: &T) -> Result<(), SyscallError> { - let (kptr, _pin) = object::(ptr, Access::Write)?; + let (kptr, _pins) = object::(ptr, Access::Write)?; if crate::actuator::copy_meets_a_remap() { remap_race::hold(kptr.cast(), core::mem::size_of::()); } @@ -169,92 +171,115 @@ impl<'a> SyscallContext<'a> { } } -/// A bulk buffer the kernel copies out of and never borrows: no reference exists because another thread of the same process can rewrite the bytes at any time; not volatile per byte, since the missing reference already stops the compiler assuming stability. -pub struct UserBytes<'a> { - kptr: *const u8, +/// Where a window's bytes are: the runs it pinned itself, or its parent's for a +/// [`sub`](UserBytes::sub) view, which borrows the parent's pins through the +/// returned lifetime. +enum Runs<'a> { + Pinned(FramePins>), + Borrowed(&'a [Segment]), +} + +/// `len` bytes starting `off` bytes into `runs`: the one shape both window types are. +struct View<'a> { + runs: Runs<'a>, + off: usize, len: usize, - /// Holds the frames covered pinned for this window's life; `None` for the empty window and for a [`sub`](UserBytes::sub) view, which borrows its parent's pin through the returned lifetime. - _pin: Option, - _scope: PhantomData<&'a ()>, } +impl View<'_> { + fn window(ptr: UserAddr, len: u64, access: Access) -> Option { + let len = len as usize; + let runs = match len { + 0 => Runs::Borrowed(&[]), + _ => Runs::Pinned(window(ptr, len, access)?), + }; + Some(View { runs, off: 0, len }) + } + + fn runs(&self) -> &[Segment] { + match &self.runs { + Runs::Pinned(pins) => pins.runs(), + Runs::Borrowed(runs) => runs, + } + } + + /// Hands `copy` each piece of bytes `[off, off + len)` of this view as a + /// direct-map pointer and the piece's offset in that range; panics if out of + /// bounds — the kernel's own arithmetic, never user input. + fn pieces(&self, what: &str, off: usize, len: usize, mut copy: impl FnMut(*mut u8, usize, usize)) { + assert!( + off.checked_add(len).is_some_and(|end| end <= self.len), + "{what} {off}+{len} past a {}-byte window", + self.len + ); + toyos_userbound::pieces(self.runs(), (self.off + off) as u64, len as u64, |phys, at, n| { + copy(crate::mm::DirectMap::from_phys(phys).as_mut_ptr(), at as usize, n as usize) + }); + } + + fn sub(&self, what: &str, off: usize, len: usize) -> View<'_> { + assert!( + off.checked_add(len).is_some_and(|end| end <= self.len), + "{what} {off}+{len} past a {}-byte window", + self.len + ); + View { runs: Runs::Borrowed(self.runs()), off: self.off + off, len } + } +} + +/// A bulk buffer the kernel copies out of and never borrows: no reference exists because another thread of the same process can rewrite the bytes at any time; not volatile per byte, since the missing reference already stops the compiler assuming stability. +pub struct UserBytes<'a>(View<'a>); + impl UserBytes<'_> { pub fn len(&self) -> usize { - self.len + self.0.len } /// Copy `dst.len()` bytes out of the window at `off`; panics if out of bounds — the kernel's own arithmetic, never user input. pub fn read_at(&self, off: usize, dst: &mut [u8]) { - assert!( - off.checked_add(dst.len()).is_some_and(|end| end <= self.len), - "UserBytes::read_at {off}+{} past a {}-byte window", - dst.len(), - self.len - ); - // SAFETY: the assert proves `off + dst.len() <= self.len`, and `window` proved the range is one physically contiguous mapping; `dst` is an owned `&mut`, so the ranges cannot overlap. - unsafe { - core::ptr::copy_nonoverlapping(self.kptr.add(off), dst.as_mut_ptr(), dst.len()); - } + let into = dst.as_mut_ptr(); + self.0.pieces("UserBytes::read_at", off, dst.len(), |run, at, n| { + // SAFETY: `pieces` bounds `[at, at + n)` inside `dst` and `run..+n` inside one run `window` pinned; `dst` is an owned `&mut`, so the ranges cannot overlap. + unsafe { core::ptr::copy_nonoverlapping(run, into.add(at), n) } + }); } /// Copy the window at `off` into one ring run; panics if out of bounds, as /// [`read_at`](Self::read_at) does. `copy` for [`UserBytesMut::write_run`]'s /// reason. pub fn read_run(&self, off: usize, dst: &mut toyos_abi::ring::Dst<'_>) { - assert!( - off.checked_add(dst.len()).is_some_and(|end| end <= self.len), - "UserBytes::read_run {off}+{} past a {}-byte window", - dst.len(), - self.len - ); - // SAFETY: as `UserBytesMut::write_run`, mirrored. - unsafe { - core::ptr::copy(self.kptr.add(off), dst.as_mut_ptr(), dst.len()); - } + let into = dst.as_mut_ptr(); + self.0.pieces("UserBytes::read_run", off, dst.len(), |run, at, n| { + // SAFETY: as `UserBytesMut::write_run`, mirrored. + unsafe { core::ptr::copy(run, into.add(at), n) } + }); } /// The `len`-byte window at `off` inside this one. pub fn sub(&self, off: usize, len: usize) -> UserBytes<'_> { - assert!( - off.checked_add(len).is_some_and(|end| end <= self.len), - "UserBytes::sub {off}+{len} past a {}-byte window", - self.len - ); - // SAFETY: the assert proves `off + len <= self.len`, so the result stays inside the window `window` validated. - UserBytes { kptr: unsafe { self.kptr.add(off) }, len, _pin: None, _scope: PhantomData } + UserBytes(self.0.sub("UserBytes::sub", off, len)) } } /// A bulk buffer the kernel copies into and never reads back, so it cannot act on a value another thread substituted. -pub struct UserBytesMut<'a> { - kptr: *mut u8, - len: usize, - /// As [`UserBytes::_pin`]: the frames stay pinned for this window's life. - _pin: Option, - _scope: PhantomData<&'a mut ()>, -} +pub struct UserBytesMut<'a>(View<'a>); impl UserBytesMut<'_> { pub fn len(&self) -> usize { - self.len + self.0.len } pub fn is_empty(&self) -> bool { - self.len == 0 + self.0.len == 0 } /// Copy `src` into the window at `off`; panics if out of bounds, for the same reason as [`UserBytes::read_at`]. pub fn write_at(&mut self, off: usize, src: &[u8]) { - assert!( - off.checked_add(src.len()).is_some_and(|end| end <= self.len), - "UserBytesMut::write_at {off}+{} past a {}-byte window", - src.len(), - self.len - ); - // SAFETY: the assert proves `off + src.len() <= self.len`, and `window` proved the range is one physically contiguous mapping. - unsafe { - core::ptr::copy_nonoverlapping(src.as_ptr(), self.kptr.add(off), src.len()); - } + let from = src.as_ptr(); + self.0.pieces("UserBytesMut::write_at", off, src.len(), |run, at, n| { + // SAFETY: `pieces` bounds `[at, at + n)` inside `src` and `run..+n` inside one run `window` pinned and proved writable. + unsafe { core::ptr::copy_nonoverlapping(from.add(at), run, n) } + }); } /// Copy one ring run into the window at `off`; panics if out of bounds, as @@ -262,38 +287,24 @@ impl UserBytesMut<'_> { /// `copy_nonoverlapping`**: `SYS_PIPE_MAP` maps the ring's page into the /// caller, which may then name it as this window. pub fn write_run(&mut self, off: usize, src: &toyos_abi::ring::Src<'_>) { - assert!( - off.checked_add(src.len()).is_some_and(|end| end <= self.len), - "UserBytesMut::write_run {off}+{} past a {}-byte window", - src.len(), - self.len - ); - // SAFETY: the assert proves `off + src.len() <= self.len`, `window` proved the range is one physically contiguous mapping, and `Ring::read` proved the run is inside its own data region. Neither side is a reference, so nothing here claims either range is exclusive. - unsafe { - core::ptr::copy(src.as_ptr(), self.kptr.add(off), src.len()); - } + let from = src.as_ptr(); + self.0.pieces("UserBytesMut::write_run", off, src.len(), |run, at, n| { + // SAFETY: `pieces` bounds `[at, at + n)` inside the ring run, which `Ring::read` proved is inside its own data region, and `run..+n` inside one run `window` pinned and proved writable. Neither side is a reference, so nothing here claims either range is exclusive. + unsafe { core::ptr::copy(from.add(at), run, n) } + }); } /// Zero `len` bytes of the window at `off`. pub fn fill_zero(&mut self, off: usize, len: usize) { - assert!( - off.checked_add(len).is_some_and(|end| end <= self.len), - "UserBytesMut::fill_zero {off}+{len} past a {}-byte window", - self.len - ); - // SAFETY: same bound as `write_at`, with a constant zero byte instead of a slice. - unsafe { core::ptr::write_bytes(self.kptr.add(off), 0, len) }; + self.0.pieces("UserBytesMut::fill_zero", off, len, |run, _, n| { + // SAFETY: as `write_at`, with a constant zero byte instead of a slice. + unsafe { core::ptr::write_bytes(run, 0, n) } + }); } /// The `len`-byte window at `off` inside this one. pub fn sub(&mut self, off: usize, len: usize) -> UserBytesMut<'_> { - assert!( - off.checked_add(len).is_some_and(|end| end <= self.len), - "UserBytesMut::sub {off}+{len} past a {}-byte window", - self.len - ); - // SAFETY: [`UserBytes::sub`]'s argument exactly. - UserBytesMut { kptr: unsafe { self.kptr.add(off) }, len, _pin: None, _scope: PhantomData } + UserBytesMut(self.0.sub("UserBytesMut::sub", off, len)) } } @@ -323,42 +334,73 @@ impl ByteSource for UserBytes<'_> { } } -/// A pin on the physical frames a user-copy window covers: the PMM reissues none of them while it lives, so the window's direct-map pointer stays backed even after a sibling unmaps and frees the range across a park. -struct FramePin { - phys: u64, - len: usize, -} +/// A pin on every physical run a user-copy window covers: the PMM reissues none of their frames while it lives, so the window's direct-map pointers stay backed even after a sibling unmaps and frees the range across a park. +type FramePins = Pinned; -impl Drop for FramePin { - fn drop(&mut self) { - crate::mm::pmm::unpin_range(self.phys, self.len); +/// The frame allocator's pins. +struct Pmm; + +impl Pins for Pmm { + fn pin(&mut self, run: Segment) -> bool { + crate::mm::pmm::pin_range(run.phys, run.len as usize) } -} -/// Validate `[ptr, ptr+len)` as one physically contiguous user window, pin every frame it covers, and return its direct-map address with the pin. The pin is taken under the address-space lock over a translation that still names the frame, so a concurrent `munmap` — which needs that same lock to free the range — cannot reissue a frame between the confirmation and the pin. -fn window(ptr: UserAddr, len: usize, access: Access) -> Option<(*mut u8, FramePin)> { - if !toyos_userbound::in_user_half(ptr.raw(), len as u64) { - return None; + fn unpin(&mut self, run: Segment) { + crate::mm::pmm::unpin_range(run.phys, run.len as usize); } - let start = ptr.raw(); - let end = start + len as u64; - // Fault every page of the range in; what it maps is confirmed under the lock below. +} + +/// Faults every 2 MiB window `[ptr, ptr + len)` touches in, and answers how +/// many that is; what they map is confirmed under the lock afterwards. +fn fault_in(ptr: UserAddr, len: usize, access: Access) -> Option { translate_user(ptr, access)?; - let mut boundary = (start & !(crate::mm::PAGE_2M - 1)) + crate::mm::PAGE_2M; + let end = ptr.raw() + len as u64; + let mut windows = 1; + let mut boundary = (ptr.raw() & !(crate::mm::PAGE_2M - 1)) + crate::mm::PAGE_2M; while boundary < end { translate_user(UserAddr::new(boundary), access)?; boundary += crate::mm::PAGE_2M; + windows += 1; } - let pt = crate::process::current_address_space(); - let guard = pt.lock(); - let phys = toyos_userbound::contiguous(start, len as u64, access, |at| { - translate_now(&guard, UserAddr::new(at), access).map(|dm| dm.phys()) - })?; - if !crate::mm::pmm::pin_range(phys, len) { + Some(windows) +} + +/// A bulk buffer's runs, one per 2 MiB window at most, allocated before the lock. +fn window(ptr: UserAddr, len: usize, access: Access) -> Option>> { + if !toyos_userbound::in_user_half(ptr.raw(), len as u64) { return None; } + pinned(ptr, len, access, Vec::with_capacity, |runs, _, run| runs.push(run)) +} + +/// Faults `[ptr, ptr+len)` in, walks it leaf by leaf into its physical runs, and pins every frame they cover. `runs` makes the store for as many runs as the range touches 2 MiB windows, and `place` puts run `i` in it. The pins are taken under the address-space lock over a translation that still names each frame, so a concurrent `munmap` — which needs that same lock to free the range — cannot reissue a frame between the confirmation and the pin. Nothing is allocated or freed under the lock but on a refused pin. +fn pinned>( + ptr: UserAddr, + len: usize, + access: Access, + runs: impl FnOnce(usize) -> R, + mut place: impl FnMut(&mut R, usize, Segment), +) -> Option> { + let windows = fault_in(ptr, len, access)?; + let mut runs = runs(windows); + let mut placed = 0; + let pt = crate::process::current_address_space(); + let guard = pt.lock(); + toyos_userbound::segments( + ptr.raw(), + len as u64, + |at| guard.leaf(UserAddr::new(at), access), + |run| { + if placed == windows { + window_split(ptr, len) + } + place(&mut runs, placed, run); + placed += 1; + }, + )?; + let pins = FramePins::pin(runs, Pmm)?; drop(guard); - Some((crate::mm::DirectMap::from_phys(phys).as_mut_ptr(), FramePin { phys, len })) + Some(pins) } /// `copy-meets-a-remap`: a sibling's `munmap` and `mmap` staged between a diff --git a/tests/toyos-rust-tests/src/bin/abuse_readonly_copyout.rs b/tests/toyos-rust-tests/src/bin/abuse_readonly_copyout.rs index 45702cf264..cf87aa185d 100644 --- a/tests/toyos-rust-tests/src/bin/abuse_readonly_copyout.rs +++ b/tests/toyos-rust-tests/src/bin/abuse_readonly_copyout.rs @@ -12,14 +12,27 @@ use std::process::Command; use toyos_abi::clock::{ClockPage, CLOCK_MAGIC, CLOCK_PAGE}; -use toyos_abi::syscall::{self, MmapFlags, MmapProt, OpenFlags, SyscallError, SYS_FSTAT, SYS_READ}; +use toyos_abi::syscall::{ + self, MmapFlags, MmapProt, OpenFlags, SeekFrom, SyscallError, SYS_FSTAT, SYS_READ, SYS_WRITE, +}; use toyos_abi::RawHandle; const SELF_PATH: &str = "/system/bin/test_rs_abuse_readonly_copyout"; const CHILD_ARG: &str = "reads-the-clock"; const PAGE_2M: usize = 2 * 1024 * 1024; +const PAGE_4K: u64 = 4096; /// Longer than `Stat`, and ends inside the file this reads. const LEN: usize = 64; +/// Bytes `0..=255`, so a probe can offer any page its own byte back, then the +/// bytes the straddling `read` offers, at [`STRADDLE_AT`]. +const PROBE_PATH: &[u8] = b"/tmp/abuse_readonly_copyout.bin"; +const STRADDLE_AT: u64 = 256; +/// The straddling `read`: half on the last writable page, half on the next. +const STRADDLE: usize = 16; + +/// In `.bss`, so its 2 MiB window also holds the pages past the image's end, +/// which the pager maps read-only. +static mut BSS: u8 = 0; /// The typed wrappers take a `&mut [u8]`, which a read-only page cannot be. fn raw(num: u64, a1: u64, a2: u64, a3: u64) -> u64 { @@ -66,6 +79,55 @@ fn refused(what: &str, addr: u64) { syscall::close(fd); } +/// The first page at or above [`BSS`], inside its 2 MiB window, a `read` may +/// not write; each page below it took a 1-byte `read` of its own first byte. +fn first_unwritable_page(probe: RawHandle) -> u64 { + let pipe = syscall::pipe().expect("pipe"); + let bss = &raw const BSS as u64; + let window_end = (bss & !(PAGE_2M as u64 - 1)) + PAGE_2M as u64; + let mut page = bss & !(PAGE_4K - 1); + while page < window_end { + let ret = raw(SYS_WRITE, pipe.write.0 as u64, page, 1); + assert_eq!(ret, 1, "write from {page:#x}, in .bss's window: {ret:#x}"); + syscall::read(pipe.read, &mut [0u8; 1]).expect("drain the pipe"); + let own = unsafe { (page as *const u8).read_volatile() }; + syscall::seek(probe, SeekFrom::Start(own as u64)).expect("seek the probe"); + let ret = raw(SYS_READ, probe.0 as u64, page, 1); + if SyscallError::from_u64(ret) == Some(SyscallError::BadAddress) { + syscall::close(pipe.read); + syscall::close(pipe.write); + return page; + } + assert_eq!(ret, 1, "a 1-byte read into {page:#x}: {ret:#x}"); + page += PAGE_4K; + } + panic!("no page above .bss at {bss:#x} refuses a write below {window_end:#x}: the image ends on its window's edge"); +} + +/// A `read` that starts on a writable page and runs into the read-only page +/// after it, in one 2 MiB window, is refused whole. +fn refused_across_pages() { + let flags = OpenFlags::READ | OpenFlags::WRITE | OpenFlags::CREATE | OpenFlags::TRUNCATE; + let probe = syscall::open(PROBE_PATH, flags).expect("create the probe file"); + let every: Vec = (0..=255).collect(); + syscall::write(probe, &every).expect("fill the probe file"); + let page = first_unwritable_page(probe); + assert!(page > &raw const BSS as u64, "BSS's own page refuses a write"); + let at = page - (STRADDLE / 2) as u64; + let before = snapshot(at); + // The writable half is offered its own bytes and the read-only half their + // complement, so a write to the read-only page shows and one below harms nothing. + let offered: [u8; STRADDLE] = core::array::from_fn(|i| if i < STRADDLE / 2 { before[i] } else { !before[i] }); + syscall::seek(probe, SeekFrom::Start(STRADDLE_AT)).expect("seek the probe"); + syscall::write(probe, &offered).expect("write the straddle's bytes"); + syscall::seek(probe, SeekFrom::Start(STRADDLE_AT)).expect("seek the probe"); + let ret = raw(SYS_READ, probe.0 as u64, at, STRADDLE as u64); + assert_eq!(before, snapshot(at), "read wrote across a writable page into the read-only one at {page:#x}"); + assert_eq!(SyscallError::from_u64(ret), Some(SyscallError::BadAddress), "read across into {page:#x}: {ret:#x}"); + syscall::close(probe); + syscall::delete(PROBE_PATH).expect("delete the probe file"); +} + fn main() { if std::env::args().nth(1).as_deref() == Some(CHILD_ARG) { let page = unsafe { core::ptr::read_volatile(CLOCK_PAGE as *const ClockPage) }; @@ -95,6 +157,7 @@ fn main() { unsafe { syscall::munmap(ro, PAGE_2M) }.expect("munmap"); refused("this program's own text", main as *const () as u64); + refused_across_pages(); // Last: a written clock page asserts in every stamp this process and every // other one takes, the verdict's own path included, so the arms whose harm diff --git a/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs b/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs new file mode 100644 index 0000000000..765a30bc24 --- /dev/null +++ b/tests/toyos-rust-tests/src/bin/munmap_reissues_second_read_window.rs @@ -0,0 +1,114 @@ +//! A parked reader whose buffer spans two mappings holds both frames: a +//! sibling that unmaps the second one and maps again is never handed the frame +//! the parked copy is about to land in. +//! +//! `munmap_reissues_read_window`'s staging, over a buffer that ends in a second +//! mapping. The two mappings are placed side by side with `FIXED`, the upper +//! one first: the frame allocator hands out the lowest free frame first, so +//! the upper mapping's frame lies below the lower one's and the buffer is two +//! physical runs, of which the upper is the second. Nothing here can see a +//! physical address, so that premise is the allocator's and is not checked. +//! +//! The assertion is that the sibling's new mapping still holds its own byte. + +use std::sync::atomic::{AtomicI64, AtomicU32, Ordering}; +use std::sync::Arc; +use std::thread; + +use toyos::endow::{Endowments, SYSCAP_LABEL}; +use toyos::syscap::SysCap; +use toyos_abi::syscall::{close, mmap, munmap, pipe, read, write, MmapFlags, MmapProt}; + +const PAGE_2M: usize = 2 * 1024 * 1024; + +/// Bytes of the buffer in the lower mapping and in the upper one; the whole +/// buffer is far under the pipe ring, so one write delivers it. +const BELOW: usize = 2048; +const ABOVE: usize = 2048; + +/// What the parked reader copies out of the pipe, and the sibling never writes. +const PATTERN_A: u8 = 0xA1; +/// What the sibling writes into its new mapping, and a safe kernel leaves there. +const PATTERN_B: u8 = 0xB2; + +fn map(at: *mut u8, len: usize, prot: MmapProt, flags: MmapFlags) -> *mut u8 { + let p = unsafe { mmap(at, len, prot, MmapFlags::ANONYMOUS | MmapFlags::PRIVATE | flags) }; + assert!(!p.is_null(), "mmap of {len:#x} bytes at {at:?} failed"); + p +} + +#[path = "../roster.rs"] +mod roster; + +fn main() { + let cap: SysCap = Endowments::get() + .take(SYSCAP_LABEL) + .expect("test-runner endows every binary it spawns a system capability"); + let ends = pipe().expect("a pipe"); + let read_end = ends.read; + let write_end = ends.write; + + // Two 2 MiB slots side by side: reserved, released, then mapped upper first. + let rw = MmapProt::READ | MmapProt::WRITE; + let lower = map(core::ptr::null_mut(), 2 * PAGE_2M, MmapProt::NONE, MmapFlags(0)); + unsafe { munmap(lower, 2 * PAGE_2M) }.expect("release the reservation"); + let upper = unsafe { lower.add(PAGE_2M) }; + assert_eq!(map(upper, PAGE_2M, rw, MmapFlags::FIXED), upper); + assert_eq!(map(lower, PAGE_2M, rw, MmapFlags::FIXED), lower); + let buf_addr = upper as usize - BELOW; + + let ready = Arc::new(AtomicU32::new(0)); + let result = Arc::new(AtomicI64::new(i64::MIN)); + + let a = { + let ready = Arc::clone(&ready); + let result = Arc::clone(&result); + thread::spawn(move || { + // Valid when formed and when the read begins; the kernel owns the + // pointer once it parks, and nothing in this thread dereferences it. + let buf = unsafe { core::slice::from_raw_parts_mut(buf_addr as *mut u8, BELOW + ABOVE) }; + ready.store(1, Ordering::SeqCst); + let n = match read(read_end, buf) { + Ok(n) => n as i64, + Err(_) => -1, + }; + result.store(n, Ordering::SeqCst); + }) + }; + + // A read the kernel refuses returns without parking. + roster::await_true(|| { + result.load(Ordering::SeqCst) != i64::MIN + || roster::my_threads(&cap).iter().any(|&(is_thread, state)| is_thread && state == roster::BLOCKED) + && ready.load(Ordering::SeqCst) == 1 + }); + let early = result.load(Ordering::SeqCst); + assert_eq!(early, i64::MIN, "the reader's read answered {early} before it parked"); + + unsafe { munmap(upper, PAGE_2M) }.expect("munmap the buffer's second mapping"); + + let sibling = map(core::ptr::null_mut(), PAGE_2M, rw, MmapFlags(0)); + unsafe { core::ptr::write_bytes(sibling, PATTERN_B, ABOVE) }; + + write(write_end, &[PATTERN_A; BELOW + ABOVE]).expect("write to wake the reader"); + + a.join().expect("the reader thread panicked"); + close(read_end); + close(write_end); + + let n = result.load(Ordering::SeqCst); + assert_eq!(n, (BELOW + ABOVE) as i64, "the parked reader's read answered {n}, so the copy the test is about never ran"); + let first = unsafe { core::slice::from_raw_parts(buf_addr as *const u8, BELOW) }; + assert!(first.iter().all(|&b| b == PATTERN_A), "the copy's first run never reached the mapping still in place"); + + let got = unsafe { core::slice::from_raw_parts(sibling, ABOVE) }; + if let Some(bad) = got.iter().position(|&b| b != PATTERN_B) { + panic!( + "a parked reader's second run reached a sibling's reissued frame — byte {bad} of it is {:#x}, not {PATTERN_B:#x}", + got[bad] + ); + } + unsafe { munmap(sibling, PAGE_2M) }.expect("munmap the sibling's mapping"); + unsafe { munmap(lower, PAGE_2M) }.expect("munmap the buffer's first mapping"); + println!("munmap_reissues_second_read_window: both frames held under the copy"); +} diff --git a/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs b/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs new file mode 100644 index 0000000000..ce0d015f75 --- /dev/null +++ b/tests/toyos-rust-tests/src/bin/user_copy_spans_windows.rs @@ -0,0 +1,99 @@ +//! A syscall's byte buffer that spans two demand-paged windows is copied through +//! both frames, in whichever order the pager handed them out. +//! +//! `.bss` is demand-paged one 2 MiB window at a time, and `SPAN` holds two whole +//! windows nothing else touches. +//! +//! Plain `std` and nothing of ToyOS's own, so the same source is its own +//! oracle on any other operating system: the bytes it expects are the ones +//! `read` and `write` move everywhere. + +use std::fs::{self, File}; +use std::io::{self, Read, Write}; + +const PAGE_2M: usize = 2 * 1024 * 1024; +/// Three windows: whatever the alignment, two whole ones lie inside. +const SPAN_LEN: usize = 3 * PAGE_2M; +/// Bytes of the buffer below the boundary and above it: the boundary is not +/// 4 KiB-aligned within the buffer. +const BELOW: usize = 32 * 1024 + 1000; +const ABOVE: usize = 32 * 1024; +const CANARY: u8 = 0xA5; +const PATH: &str = "/tmp/user_copy_spans_windows.bin"; + +static mut SPAN: [u8; SPAN_LEN] = [0; SPAN_LEN]; + +fn pattern(len: usize, salt: u8) -> Vec { + (0..len).map(|i| (i.wrapping_mul(131) ^ (i >> 9)) as u8 ^ salt).collect() +} + +/// The first byte where `got` and `want` differ, and on which side of the boundary. +fn first_difference(got: &[u8], want: &[u8]) -> Option { + let i = got.iter().zip(want).position(|(g, w)| g != w)?; + let side = if i < BELOW { "below" } else { "above" }; + Some(format!("byte {i} ({side} the boundary) is {:#04x}, not {:#04x}", got[i], want[i])) +} + +fn main() { + let base = (&raw mut SPAN).cast::(); + let lower = (base as usize).next_multiple_of(PAGE_2M); + let boundary = lower + PAGE_2M; + assert!(boundary + PAGE_2M <= base as usize + SPAN_LEN, "SPAN holds two whole windows"); + let at = |addr: usize| unsafe { base.add(addr - base as usize) }; + + // Upper window first, then lower. + unsafe { + at(boundary).write_volatile(CANARY); + at(boundary - 1).write_volatile(CANARY); + at(boundary - BELOW - 1).write_volatile(CANARY); + at(boundary + ABOVE).write_volatile(CANARY); + } + // SAFETY: `SPAN` is this process's and nothing else names these bytes while the slice lives. + let buf = unsafe { core::slice::from_raw_parts_mut(at(boundary - BELOW), BELOW + ABOVE) }; + let outside_intact = || unsafe { + at(boundary - BELOW - 1).read_volatile() == CANARY && at(boundary + ABOVE).read_volatile() == CANARY + }; + + // 1. A pipe read: the kernel writes into the buffer. The other end is a + // thread so that no pipe capacity decides the outcome. + let (mut reader, mut writer) = io::pipe().expect("pipe"); + let sent = pattern(buf.len(), 0x5A); + std::thread::scope(|s| { + s.spawn(|| writer.write_all(&sent).expect("fill the pipe")); + reader.read_exact(buf).expect("a pipe read into a buffer across two windows"); + }); + assert!(outside_intact(), "a pipe read into the buffer wrote past its ends"); + if let Some(diff) = first_difference(buf, &sent) { + panic!("a pipe read across two windows: {diff}"); + } + + // 2. A pipe write: the kernel reads the buffer. + let mine = pattern(buf.len(), 0xC3); + buf.copy_from_slice(&mine); + let mut back = vec![0u8; buf.len()]; + std::thread::scope(|s| { + s.spawn(|| reader.read_exact(&mut back).expect("drain the pipe")); + writer.write_all(buf).expect("a pipe write from a buffer across two windows"); + }); + if let Some(diff) = first_difference(&back, &mine) { + panic!("a pipe write from across two windows: {diff}"); + } + + // 3. A file write and read back: the file cache's copies. + let mine = pattern(buf.len(), 0x3C); + buf.copy_from_slice(&mine); + File::create(PATH).and_then(|mut f| f.write_all(buf)).expect("a file write from a buffer across two windows"); + let stored = fs::read(PATH).expect("read the file back"); + if let Some(diff) = first_difference(&stored, &mine) { + panic!("a file write from across two windows: {diff}"); + } + buf.fill(0); + File::open(PATH).and_then(|mut f| f.read_exact(buf)).expect("a file read into a buffer across two windows"); + assert!(outside_intact(), "a file read into the buffer wrote past its ends"); + if let Some(diff) = first_difference(buf, &mine) { + panic!("a file read across two windows: {diff}"); + } + fs::remove_file(PATH).expect("cleanup"); + + println!("a buffer across two demand windows is copied through both frames"); +} diff --git a/toyos-userbound/src/lib.rs b/toyos-userbound/src/lib.rs index 5729263857..c7fdc63f4f 100644 --- a/toyos-userbound/src/lib.rs +++ b/toyos-userbound/src/lib.rs @@ -2,11 +2,14 @@ //! //! **Before a dereference**: is this //! address userland's, is the object at it aligned for the type being read, and -//! does it lie wholly inside one mapping? **Before a placement**: can a length -//! userland asked for be placed at all, and where does it go? **After a trap**: -//! which side did the frame come from? +//! does it lie wholly inside one mapping? **Before a copy**: which physical runs +//! hold a user window, which of them does each piece of the copy land in, and +//! is every one of them pinned? **Before a placement**: can a length userland +//! asked for be placed at all, and where does it go? **After a trap**: which +//! side did the frame come from? //! -//! [`span`] answers the first, [`place`] the second and [`fault`] the third. +//! [`span`] answers the first, [`segment`] the second, [`place`] the third and +//! [`fault`] the fourth. //! //! Pure. No I/O, no allocation, no `unsafe`, nothing read from a device and //! nothing named outside this crate. The kernel is the only caller — @@ -25,11 +28,13 @@ pub mod fault; pub mod place; +pub mod segment; pub mod span; pub use fault::Ring; pub use place::{PageSpan, Window}; +pub use segment::{pieces, segments, Pinned, Pins, Segment}; pub use span::{ - align_2m_checked, contiguous, in_user_half, is_user_addr, is_user_object, rebase_base, Access, - PAGE_2M, PAGE_4K, USER_TOP, + align_2m_checked, in_user_half, is_user_addr, is_user_object, rebase_base, Access, PAGE_2M, + PAGE_4K, USER_TOP, }; diff --git a/toyos-userbound/src/segment.rs b/toyos-userbound/src/segment.rs new file mode 100644 index 0000000000..ba44f2b363 --- /dev/null +++ b/toyos-userbound/src/segment.rs @@ -0,0 +1,486 @@ +//! Where a user window's bytes are in physical memory: one [`Segment`] per +//! physically contiguous run, in address order, each copy cut at their seams, +//! and every run pinned for the window's life. +//! +//! **A window is asked once per leaf that maps it.** A page that is absent or +//! does not grant the access refuses the whole window, and so does a run that +//! cannot be pinned: the kernel copies nothing through a window it could not +//! wholly place and hold. + +use crate::span::in_user_half; + +/// `len` bytes of a user window, physically contiguous from `phys`. +#[derive(Clone, Copy, PartialEq, Eq, Debug)] +pub struct Segment { + pub phys: u64, + pub len: u64, +} + +/// Walks `[start, start + len)` and hands `emit` each maximal physically +/// contiguous run of it, in order. `leaf` is the page walk: for one user +/// address, where it grants the access asked, the physical address it +/// translates to and the bytes from it to the end of the leaf that maps it, +/// at least one; `None` where it does not. +/// +/// `None` where the range is empty, leaves the user half, or holds a page +/// `leaf` refuses — and runs before the refused page may already have been +/// emitted, so nothing emitted is acted on until this returns `Some`. +pub fn segments( + start: u64, + len: u64, + mut leaf: impl FnMut(u64) -> Option<(u64, u64)>, + mut emit: impl FnMut(Segment), +) -> Option<()> { + if len == 0 || !in_user_half(start, len) { + return None; + } + let end = start + len; + let mut at = start; + let (mut phys, mut extent) = leaf(at)?; + let mut run = Segment { phys, len: 0 }; + loop { + if run.phys + run.len != phys { + emit(run); + run = Segment { phys, len: 0 }; + } + let n = extent.min(end - at); + run.len += n; + at += n; + if at == end { + break; + } + (phys, extent) = leaf(at)?; + } + emit(run); + Some(()) +} + +/// Hands `copy` each physically contiguous piece of bytes `[off, off + len)` +/// of the window `segs` lays out, as `(phys, at, n)`: the `n` bytes at `phys` +/// are bytes `at..at + n` of that range. The caller bounds the range inside +/// the window. +pub fn pieces(segs: &[Segment], off: u64, len: u64, mut copy: impl FnMut(u64, u64, u64)) { + let mut skip = off; + let mut done = 0; + for seg in segs { + if done == len { + break; + } + if skip >= seg.len { + skip -= seg.len; + continue; + } + let n = (seg.len - skip).min(len - done); + copy(seg.phys + skip, done, n); + skip = 0; + done += n; + } +} + +/// What holds a frame against reissue: the kernel's frame allocator, or a +/// test's counter. +pub trait Pins { + /// Pins every frame `run` touches; `false`, with nothing pinned, when it + /// cannot. + fn pin(&mut self, run: Segment) -> bool; + /// Releases a pin [`pin`](Self::pin) took over the same run. + fn unpin(&mut self, run: Segment); +} + +/// A window's runs, every one of them pinned until this drops. +pub struct Pinned, P: Pins> { + runs: R, + pins: P, +} + +impl, P: Pins> Pinned { + /// Pins every run, or none: a run that cannot be pinned unpins the runs + /// before it, and the window is refused. + pub fn pin(runs: R, mut pins: P) -> Option { + if let Some(refused) = runs.as_ref().iter().position(|&run| !pins.pin(run)) { + for &run in &runs.as_ref()[..refused] { + pins.unpin(run); + } + return None; + } + Some(Pinned { runs, pins }) + } + + pub fn runs(&self) -> &[Segment] { + self.runs.as_ref() + } +} + +impl, P: Pins> Drop for Pinned { + fn drop(&mut self) { + for &run in self.runs.as_ref() { + self.pins.unpin(run); + } + } +} + +#[cfg(test)] +mod tests { + extern crate std; + + use std::cell::{Cell, RefCell}; + use std::vec; + use std::vec::Vec; + + use super::*; + use crate::span::{Access, PAGE_2M, PAGE_4K, USER_TOP}; + + /// Where the fake user range starts: a 2 MiB boundary. + const VBASE: u64 = 5 * PAGE_2M; + /// Pages in the fake range. + const PAGES: u64 = 48; + /// Where fake physical memory starts: nonzero, so a phys of 0 is a bug. + const PBASE: u64 = 0x10_0000_0000; + + fn shuffled(n: usize, mut seed: u64) -> Vec { + let mut order: Vec = (0..n as u64).collect(); + for i in (1..n).rev() { + seed ^= seed << 13; + seed ^= seed >> 7; + seed ^= seed << 17; + order.swap(i, (seed % (i as u64 + 1)) as usize); + } + order + } + + /// A process's pages over fake physical memory, mapped by leaves of + /// `grain` pages: leaf `i` of the range lives in frame `frame[i]`, and `ro` + /// and `absent` pages (with a `grain` of one) grant no write and nothing. + struct Space { + grain: u64, + frame: Vec, + ro: Vec, + absent: Vec, + phys: Vec, + lookups: Cell, + } + + impl Space { + /// Every page present and writable, leaves in frames `order`, and each + /// user byte holding a value its virtual address decides. + fn new(order: Vec, grain: u64) -> Self { + assert_eq!(order.len() as u64 * grain, PAGES); + let mut space = Space { + grain, + frame: order, + ro: Vec::new(), + absent: Vec::new(), + phys: vec![0; (PAGES * PAGE_4K) as usize], + lookups: Cell::new(0), + }; + for v in VBASE..VBASE + PAGES * PAGE_4K { + let p = space.phys_of(v).expect("every page is present"); + space.phys[(p - PBASE) as usize] = byte_at(v); + } + space + } + + fn leaf_bytes(&self) -> u64 { + self.grain * PAGE_4K + } + + fn phys_of(&self, v: u64) -> Option { + if !(VBASE..VBASE + PAGES * PAGE_4K).contains(&v) { + return None; + } + let off = v - VBASE; + Some(PBASE + self.frame[(off / self.leaf_bytes()) as usize] * self.leaf_bytes() + off % self.leaf_bytes()) + } + + /// The kernel's `AddressSpace::leaf`, over fake memory. + fn leaf(&self, v: u64, access: Access) -> Option<(u64, u64)> { + self.lookups.set(self.lookups.get() + 1); + let page = v.checked_sub(VBASE)? / PAGE_4K; + if self.absent.contains(&page) || (access == Access::Write && self.ro.contains(&page)) { + return None; + } + Some((self.phys_of(v)?, self.leaf_bytes() - (v - VBASE) % self.leaf_bytes())) + } + + fn window(&self, start: u64, len: u64, access: Access) -> Option> { + let mut segs = Vec::new(); + segments(start, len, |at| self.leaf(at, access), |s| segs.push(s))?; + Some(segs) + } + + /// The kernel's `read_at`, over fake memory. + fn read(&self, segs: &[Segment], off: u64, len: u64) -> Vec { + let mut out = vec![0u8; len as usize]; + pieces(segs, off, len, |phys, at, n| { + let p = (phys - PBASE) as usize; + out[at as usize..(at + n) as usize].copy_from_slice(&self.phys[p..p + n as usize]); + }); + out + } + + /// The kernel's `write_at`, over fake memory. + fn write(&mut self, segs: &[Segment], off: u64, src: &[u8]) { + let phys = &mut self.phys; + pieces(segs, off, src.len() as u64, |p, at, n| { + let p = (p - PBASE) as usize; + phys[p..p + n as usize].copy_from_slice(&src[at as usize..(at + n) as usize]); + }); + } + + /// What a ring 3 load at `v` sees. + fn load(&self, v: u64) -> u8 { + self.phys[(self.phys_of(v).unwrap() - PBASE) as usize] + } + } + + fn byte_at(v: u64) -> u8 { + (v.wrapping_mul(0x9E37_79B9_7F4A_7C15) >> 56) as u8 + } + + /// Ranges that start and end on every kind of seam: at a page start, one + /// byte either side of one, mid-page, and across many pages. + fn ranges() -> Vec<(u64, u64)> { + let mut out = Vec::new(); + for first in [0, 1, 7, 13] { + for head in [0, 1, PAGE_4K / 2, PAGE_4K - 1] { + let start = VBASE + first * PAGE_4K + head; + for len in [1, 2, PAGE_4K - head, PAGE_4K, PAGE_4K + 1, 3 * PAGE_4K - 5, 20 * PAGE_4K + 3] { + if start + len <= VBASE + PAGES * PAGE_4K { + out.push((start, len)); + } + } + } + } + out.push((VBASE, PAGES * PAGE_4K)); + out + } + + /// Each leaf size the fake maps with, in pages. + const GRAINS: [u64; 3] = [1, 4, 16]; + + #[test] + fn a_read_across_shuffled_frames_gets_the_bytes_at_their_addresses() { + for grain in GRAINS { + for seed in [1, 0x5eed, 0xdead_beef, 42] { + let space = Space::new(shuffled((PAGES / grain) as usize, seed), grain); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Read).expect("every page is readable"); + let want: Vec = (start..start + len).map(byte_at).collect(); + assert_eq!(space.read(&segs, 0, len), want, "grain {grain}, seed {seed:#x}: [{start:#x}, +{len:#x})"); + // A view from any offset inside the window: `UserBytes::sub`. + for off in [1, PAGE_4K - 1, PAGE_4K, len / 2] { + if off < len { + let tail = len - off; + assert_eq!(space.read(&segs, off, tail), want[off as usize..], "grain {grain}, seed {seed:#x}: +{off:#x}"); + } + } + } + } + } + } + + #[test] + fn a_write_across_shuffled_frames_lands_where_a_ring_3_load_reads_it() { + for grain in GRAINS { + for seed in [3, 0xabcdef, 99] { + for (start, len) in ranges() { + let mut space = Space::new(shuffled((PAGES / grain) as usize, seed), grain); + let segs = space.window(start, len, Access::Write).expect("every page is writable"); + let src: Vec = (0..len).map(|i| !byte_at(i ^ seed)).collect(); + space.write(&segs, 0, &src); + for v in VBASE..VBASE + PAGES * PAGE_4K { + let want = if (start..start + len).contains(&v) { src[(v - start) as usize] } else { byte_at(v) }; + assert_eq!(space.load(v), want, "grain {grain}, seed {seed:#x}: [{start:#x}, +{len:#x}) at {v:#x}"); + } + } + } + } + } + + #[test] + fn the_segments_are_the_maximal_runs_and_cover_the_window_exactly() { + for grain in GRAINS { + for seed in [7, 0x1234_5678] { + let space = Space::new(shuffled((PAGES / grain) as usize, seed), grain); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Read).unwrap(); + assert_eq!(segs.iter().map(|s| s.len).sum::(), len); + assert_eq!(segs[0].phys, space.phys_of(start).unwrap()); + for pair in segs.windows(2) { + assert_ne!(pair[0].phys + pair[0].len, pair[1].phys, "two runs that follow were not joined"); + } + } + } + } + } + + /// The cost the walk is held to: one lookup per leaf the range touches, + /// whatever the leaf's size. + #[test] + fn a_range_is_looked_up_once_per_leaf() { + for grain in GRAINS { + let space = Space::new(shuffled((PAGES / grain) as usize, 5), grain); + for (start, len) in ranges() { + space.lookups.set(0); + space.window(start, len, Access::Read).unwrap(); + let leaves = (start + len - 1 - VBASE) / space.leaf_bytes() - (start - VBASE) / space.leaf_bytes() + 1; + assert_eq!(space.lookups.get(), leaves, "grain {grain}: [{start:#x}, +{len:#x})"); + } + } + // 64 MiB from a byte past a 2 MiB boundary: 33 leaves of 2 MiB, frames + // in order, so 33 lookups and one run. + let (start, len) = (VBASE + 1, 32 * PAGE_2M); + let mut lookups = 0; + let mut segs = Vec::new(); + segments( + start, + len, + |at| { + lookups += 1; + Some((PBASE + at, PAGE_2M - at % PAGE_2M)) + }, + |s| segs.push(s), + ) + .unwrap(); + assert_eq!(lookups, 33); + assert_eq!(segs, [Segment { phys: PBASE + start, len }]); + } + + /// A 2 MiB leaf, or a split window over one frame, is every page in order. + #[test] + fn frames_in_order_are_one_segment() { + for grain in GRAINS { + let space = Space::new((0..PAGES / grain).collect(), grain); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Write).unwrap(); + assert_eq!(segs, [Segment { phys: space.phys_of(start).unwrap(), len }]); + } + } + } + + /// Frames in reverse: every leaf seam is a segment seam. + #[test] + fn frames_in_reverse_are_one_segment_per_leaf() { + let space = Space::new((0..PAGES).rev().collect(), 1); + let segs = space.window(VBASE + PAGE_4K - 3, PAGE_4K + 6, Access::Read).unwrap(); + assert_eq!(segs.iter().map(|s| s.len).collect::>(), [3, PAGE_4K, 3], "{segs:x?}"); + } + + #[test] + fn an_absent_page_anywhere_refuses_the_window() { + for hole in [0, 1, 5, PAGES - 1] { + let mut space = Space::new(shuffled(PAGES as usize, hole + 1), 1); + space.absent.push(hole); + let h = VBASE + hole * PAGE_4K; + for access in [Access::Read, Access::Write] { + assert_eq!(space.window(VBASE, PAGES * PAGE_4K, access), None, "hole at page {hole}"); + assert_eq!(space.window(h + PAGE_4K - 1, 1, access), None, "the hole's last byte"); + assert_eq!(space.window(h.saturating_sub(1).max(VBASE), 2, access), None, "across the hole's start"); + } + } + } + + #[test] + fn a_read_only_page_between_writable_ones_refuses_a_write_only() { + let mut space = Space::new(shuffled(PAGES as usize, 11), 1); + space.ro.push(2); + let (start, len) = (VBASE + PAGE_4K, 3 * PAGE_4K); + assert_eq!(space.window(start, len, Access::Write), None); + assert!(space.window(start, len, Access::Read).is_some()); + assert!(space.window(start, PAGE_4K, Access::Write).is_some(), "page 1 alone is writable"); + } + + #[test] + fn an_empty_or_kernel_window_is_refused_before_the_walk() { + let never = |_: u64| -> Option<(u64, u64)> { panic!("walked") }; + assert_eq!(segments(PAGE_2M, 0, never, |_| panic!("emitted")), None); + assert_eq!(segments(USER_TOP - 8, 16, never, |_| panic!("emitted")), None); + assert_eq!(segments(u64::MAX - 4, 8, never, |_| panic!("emitted")), None); + } + + /// Pin counts per fake 4 KiB frame, which refuse the pin call numbered + /// `refuse` and panic on an unpin of a frame nothing pinned, as the PMM does. + struct Frames { + count: RefCell>, + calls: Cell, + refuse: Option, + } + + impl Frames { + fn new(refuse: Option) -> Self { + Frames { count: RefCell::new(vec![0; PAGES as usize]), calls: Cell::new(0), refuse } + } + + fn of(run: Segment) -> core::ops::RangeInclusive { + ((run.phys - PBASE) / PAGE_4K) as usize..=((run.phys + run.len - 1 - PBASE) / PAGE_4K) as usize + } + + fn pinned(&self) -> Vec { + self.count.borrow().clone() + } + } + + impl Pins for &Frames { + fn pin(&mut self, run: Segment) -> bool { + let call = self.calls.get(); + self.calls.set(call + 1); + if self.refuse == Some(call) { + return false; + } + for f in Frames::of(run) { + self.count.borrow_mut()[f] += 1; + } + true + } + + fn unpin(&mut self, run: Segment) { + for f in Frames::of(run) { + let mut count = self.count.borrow_mut(); + count[f] = count[f].checked_sub(1).expect("unpin of a frame that was not pinned"); + } + } + } + + /// Frames in a shuffled order with one frame mapped twice, as a shared + /// mapping is: every run holds its own pin on it. + fn twice_mapped() -> Space { + let mut order = shuffled(PAGES as usize, 0x77); + order[9] = order[3]; + Space::new(order, 1) + } + + #[test] + fn every_run_is_pinned_until_the_window_is_dropped() { + let space = twice_mapped(); + for (start, len) in ranges() { + let segs = space.window(start, len, Access::Read).unwrap(); + let frames = Frames::new(None); + let mut want = vec![0; PAGES as usize]; + for &run in &segs { + for f in Frames::of(run) { + want[f] += 1; + } + } + { + let pinned = Pinned::pin(&segs[..], &frames).expect("nothing refuses"); + assert_eq!(frames.pinned(), want, "[{start:#x}, +{len:#x}) while pinned"); + let got = space.read(pinned.runs(), 0, len); + assert_eq!(got, (start..start + len).map(|v| space.load(v)).collect::>()); + } + assert_eq!(frames.pinned(), vec![0; PAGES as usize], "[{start:#x}, +{len:#x}) after the copy"); + } + } + + #[test] + fn a_run_that_cannot_be_pinned_refuses_the_window_and_leaves_nothing_pinned() { + let space = twice_mapped(); + let segs = space.window(VBASE + PAGE_4K / 2, 20 * PAGE_4K, Access::Read).unwrap(); + assert!(segs.len() > 3, "{segs:x?}"); + for refused in 0..segs.len() { + let frames = Frames::new(Some(refused)); + assert!(Pinned::pin(&segs[..], &frames).is_none(), "run {refused} refused"); + assert_eq!(frames.calls.get(), refused + 1, "no run after the refused one is tried"); + assert_eq!(frames.pinned(), vec![0; PAGES as usize], "run {refused} refused"); + } + } +} diff --git a/toyos-userbound/src/span.rs b/toyos-userbound/src/span.rs index c1e30ebdb4..0fa7f759cb 100644 --- a/toyos-userbound/src/span.rs +++ b/toyos-userbound/src/span.rs @@ -92,41 +92,6 @@ pub enum Access { /// The grain a split window grants rights at. pub const PAGE_4K: u64 = 4096; -/// Where `[start, start + len)` begins in physical memory, if the whole range is -/// one physically contiguous run that grants `access`. `leaf` is the page walk: -/// the physical address one user address translates to where it grants -/// `access`, and `None` where it does not. -/// -/// **A write is asked at every 4 KiB page, a read at every 2 MiB page**, both -/// at the first and the last byte. A split window is one 2 MiB frame in order, -/// so contiguity can break only at a 2 MiB boundary and every present page is -/// readable; but it grants writes page by page, so a read-only page between -/// two writable ones is a hole that the ends alone would pass. -pub fn contiguous( - start: u64, - len: u64, - access: Access, - mut leaf: impl FnMut(u64) -> Option, -) -> Option { - if len == 0 || !in_user_half(start, len) { - return None; - } - let phys = leaf(start)?; - let step = match access { - Access::Read => PAGE_2M, - Access::Write => PAGE_4K, - }; - let end = start + len; - let mut at = (start & !(step - 1)) + step; - while at < end { - if leaf(at)? != phys + (at - start) { - return None; - } - at += step; - } - (len == 1 || leaf(end - 1)? == phys + (len - 1)).then_some(phys) -} - #[cfg(test)] mod tests { use super::*; @@ -207,54 +172,6 @@ mod tests { } } - /// A split window at `PAGE_2M` over frame `FRAME`: every page present, - /// and page `ro` readable but not writable. - const FRAME: u64 = 0x4000_0000; - - fn split_with_one_read_only(ro: u64) -> impl Fn(u64, Access) -> Option { - move |at, access| { - if !(PAGE_2M..2 * PAGE_2M).contains(&at) { - return None; - } - let page = (at - PAGE_2M) / PAGE_4K; - (access == Access::Read || page != ro).then_some(FRAME + (at - PAGE_2M)) - } - } - - #[test] - fn a_write_across_a_read_only_page_between_writable_ones_is_refused() { - let walk = split_with_one_read_only(2); - let start = PAGE_2M + PAGE_4K; - let len = 3 * PAGE_4K; - assert_eq!( - contiguous(start, len, Access::Write, |at| walk(at, Access::Write)), - None, - "a write window over pages 1..4 with page 2 read-only was granted" - ); - assert_eq!(contiguous(start, len, Access::Read, |at| walk(at, Access::Read)), Some(FRAME + PAGE_4K)); - assert_eq!( - contiguous(start, PAGE_4K, Access::Write, |at| walk(at, Access::Write)), - Some(FRAME + PAGE_4K), - "page 1 alone is writable" - ); - } - - #[test] - fn a_window_across_two_frames_that_do_not_follow_is_refused() { - let walk = |at: u64| Some(if at < 2 * PAGE_2M { FRAME + at } else { 8 * FRAME + at }); - for access in [Access::Read, Access::Write] { - assert_eq!(contiguous(2 * PAGE_2M - 8, 16, access, walk), None, "{access:?}"); - assert_eq!(contiguous(2 * PAGE_2M - 8, 8, access, walk), Some(FRAME + 2 * PAGE_2M - 8)); - } - } - - #[test] - fn an_empty_or_kernel_window_is_refused_before_the_walk() { - let never = |_: u64| -> Option { panic!("walked") }; - assert_eq!(contiguous(PAGE_2M, 0, Access::Read, never), None); - assert_eq!(contiguous(USER_TOP - 8, 16, Access::Write, never), None); - } - #[test] fn a_zero_sized_object_is_nothing_to_dereference() { assert!(!is_user_object(PAGE_2M, 0, 1));