Kernel: stop reserving the stack's image offset as a physical range - #584
Conversation
`KernelArgs::kernel_stack_addr` is an offset into the kernel image, not an address. The loader writes `StackedImage::place`'s `stack` there (bootloader/src/main.rs, `stack_offset: placed.stack`), and both entries read it that way: x86-64 `_start` adds it to `kernel_memory_addr`, and the AArch64 `_start` loads it under the name `stack_offset`. `kernel_main` alone read it as physical and reserved `[kernel_stack_addr, +8 MiB)`. The kernel is a static PIE at base 0, so the offset is the image's `vaddr` extent rounded up to a page: 0x42c000 for main's x86-64 kernel, measured with llvm-readobj -l on target/kernel-x86_64-4be7ccf4859a9aa6 (the last PT_LOAD ends at 0x200240 + 0x22b940 = 0x42bb80). The bogus region [0x42c000, 0xc2c000) touches the 2 MiB frames at 0x400000, 0x600000, 0x800000, 0xa00000 and 0xc00000, so the PMM withheld each of those five (up to 10 MiB) that lies wholly inside a usable firmware entry, and nothing used them. On QEMU's AArch64 `virt` machine RAM starts at 0x4000_0000, so the range reached no RAM there. The real stack was never at risk. The loader allocates the image and the stack as one block: `mem_size = placed.size = stack + stack_size`, and it passes that block as `kernel_memory_size`. So the stack is the image's tail, [kernel_memory_addr + stack, kernel_memory_addr + kernel_memory_size), and the image's own reservation already keeps it. The line is deleted rather than corrected, because a corrected version would duplicate the image reservation. Two checks replace it: - an assert that `kernel_stack_addr + kernel_stack_size <= kernel_memory_size`, the loader fact the deletion relies on; - every reservation except the architecture's own page must lie inside one `LoaderData` descriptor of the firmware map. Each region in that list is an allocation made by the loader: the image, the ELF, the black-box page and ROOT. A region not held that way withholds memory nothing uses. This is the check that refuses the deleted line if it is ever reinstated. The architecture's page moves to the end of the array so the check can exclude it by position. The boot record printed the offset in an address's slot. It now reads `stack image+<offset>+<size>`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
…s controls can fail CLAUDE.md's Firmware paragraph now reads as the orchestrator ruled: the kernel calls no UEFI service, and every UEFI call is the loader's, before ExitBootServices. The loader's GetVariable, SetVariable, GetTime and ResetSystem come before the handover, so the old wording was false of it. The track: - Stage 1 matches #583 at 8565cc5: KernelArgs::layout (0x5459_0001), kernel_args_layout_refused with the loader-writes-no-layout actuator, the probe's realtime=, and the kernel's IA32_TSC_ADJUST line. No test fails without that line, and the stage says so. - Stage 3 drops "why nothing is weaker". A security version admits an older build the same key signed at that version. The floor issue records that as the owner's accepted cost. Raising the version is a reviewed PR that edits one constant and names the security fix. The loader deletes the build-time ToyOSImageFloor- variables instead of leaving them behind. - Stage 4's panic handler writes loader.log and powers off, never resets. Its test panics on a floor planted in 9 bytes, a failure the machine causes. KernelArgs' layout word rises to 0x5459_0002, and kernel_args_last_layout_refused fails if it does not. - Stage 5 refuses a signed kernel the loader cannot load inside verify, so the other slot boots instead of the pass bricking the machine. An install's priority rises above the kept slot's. update --good, run by init at the health gate under the slots claim, writes the good flag, and an image without that claim is never good. Each rule gets a named guest control: update_floor_waits_for_good, update_readonly_stick_boots_nothing (red under `let persisted = true;`) and update_unloadable_kernel_boots_the_other_slot. The slot-table oracle is decoded without production code. - Every #539 piece the review listed is placed or deleted. The #539-only issue names and the stack-offset closure (#584's) are gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review, round 1, at b0c5640Ready: CI Negative control: red for the named reason. The assert runs after BLOCKER
NOTE
REMOVE
SEND BACK |
…not a copy The inline predicate in kernel_main's reservation loop duplicated toyos_rootimage::handoff::held without its checked_add, and no test covered the copy. Both crates were already kernel dependencies, so the loop now builds a Descriptor iterator and calls held(..., block=1) — block 1 because the ELF region is not page-aligned — the same call kernel/src/rootfs.rs already makes for ROOT's image. The architecture's own page is now named (a `loader` array of the four loader allocations, checked, with arch::boot::reserved() appended after) so a later region can't land inside the exempted slot by position. Files the misleading kernel_stack_addr name (toyos-abi/src/boot.rs:8) as an issue for after #583, and removes the tracker entry this branch's exit condition already closes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review, round 2, at 4126414Ready: CI Round-1 blockers
Negative control and oracle
BLOCKER
NOTE
REMOVE
SEND BACK |
`reserved`'s construction copied `loader` by position, so a fifth region
added to `loader` compiled and passed the containment check but was silently
dropped from what `mm::init` withholds — exactly what mutation-r2.patch had
done to `root_image`. Destructuring `loader` by name (`image`, `elf`,
`black_box`, `root`) makes a fifth element a compile error instead.
The issue's owner line named a role ("whoever lands #583"), not a concrete
owner; it now names PR #583 (wt/toyos-loader1) itself, and drops the line
number citation that rots with the next edit to toyos-abi/src/boot.rs. The
same rotting citation is dropped from main.rs's comment on the
architecture's own reserved page.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review, round 3, at 65591b9Ready: CI Round-2 blockers
Round-2 NOTEs, now measured
BLOCKERNone. NOTE
REMOVE
LAND AFTER NAMED CHANGES |
KernelArgs::kernel_stack_addris an offset into the kernel image, not an address.kernel_mainread it as a physical address and reserved[kernel_stack_addr, +8 MiB), so the PMM held back low physical frames that nothing uses. The deleted line is replaced by two boot-time checks.Changes
kernel_stack_addr + kernel_stack_size <= kernel_memory_size, which is exactly the loader fact the deletion depends on.LoaderDatadescriptor of the firmware map. The image, the ELF, the black-box page and ROOT are named in aloaderarray and checked against it. The check callstoyos_rootimage::handoff::held— the functionkernel/src/rootfs.rsalready calls for ROOT's image — instead of an inline copy; block size 1, because the ELF region is not page-aligned.reserveddestructuresloaderby name, not by index.let [image, elf, black_box, root] = loader; let reserved = [image, elf, black_box, root, arch::boot::reserved()];— indexing (loader[0], …) let a region added toloadercompile, pass the containment check, and still be silently dropped from whatmm::initwithholds. Destructuring by name turns that mutation into a compile error. Proof: appending a fifth element toloader(a checked patch, applied and reverted, not part of this diff) makescargo run -- --build-onlyfail witherror[E0527]: pattern requires 4 elements but array has 5atsrc/main.rs:392:9, EXIT=101; the unmutated tree builds at EXIT=0.stack image+<offset>+<size>. Before, it printed the offset in the slot where an address goes. No test parses that record.The architecture's own reserved page is not a loader allocation and is appended to
reservedafter the containment check runs, rather than exempted by position.No
KernelArgslayout change. The misleading namekernel_stack_addris filed asissues/kernel/kernel_stack_addr-names-an-offset-not-an-address.md, owned by the author of PR #583 (wt/toyos-loader1), which is already changing that struct.Gates, at 65591b9
cargo run -- --ci hostcargo run -- --build-onlyHigh-risk checks (memory management)
reservedarray, just ahead of the architecture's page, and revert the two asserts, theheldcall and theloader/reservedsplit with it — the whole change, back onto 762e4ba. Not yet re-run at this head; round 1's run at b0c5640 panicked beforemm::initfor the named reason.ExitBootServicesis real and independent of the kernel's own computation — the loader copiesentry.ty.0verbatim (bootloader/src/main.rs:771), and the containment assert's refusal is OVMF's descriptors disagreeing with a reserved region. No such oracle exists for the T14 (no reading is on record at any head) or for AArch64 (no guest run reaches the check), and it says nothing about where the stack is — that premise rests on the stack-in-image assert alone, which has never been seen red on any run.What I'm unsure of
Measured by the orchestrator: pmm_accounting's withheld total is 29360128 bytes at 762e4ba and 20971520 at 65591b9; the 8388608-byte difference is the stack's size.
MAP_MARGIN(tracked inissues/panic-path/the-loaders-truncated-map-refusal-is-executed-by-nothing.md), would now panic at boot instead of booting; the image and ELF pool allocations lying in one descriptor each is unmeasured on the T14, and its first boot of this head is that measurement.virt_*tests stop attest-early-panic, before the reservations.🤖 Generated with Claude Code