Skip to content

Track: the kernel mitigates what Linux mitigates on the T14, and matches its hardening defaults - #594

Merged
Japabu merged 12 commits into
mainfrom
wt/toyos-secparity
Sep 29, 2026
Merged

Japabu merged 12 commits into
mainfrom
wt/toyos-secparity

Conversation

@Japabu

@Japabu Japabu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

This PR files a track and nine defects. It changes no production code.

  • issues/kernel/the-kernel-mitigates-what-linux-mitigates-on-the-t14.md is a kind: track. The T14 is an i5-1135G7 (Tiger Lake, 06_8C). The track makes ToyOS mitigate there what the T14's own Linux mitigates, and match every hardening default that Linux's config sets.
  • Nine defects, each a row of the track's hardening table:
    • issues/kernel/every-syscall-runs-at-one-kernel-stack-offset.md
    • issues/kernel/kernel-functions-return-with-their-used-registers-intact.md
    • issues/kernel/a-threads-kernel-stack-has-no-guard-page.md
    • issues/kernel/kernel-text-is-writable-and-every-kernel-page-executable.md
    • issues/kernel/the-kernel-heap-has-none-of-slubs-hardening.md
    • issues/kernel/tsx-stays-as-firmware-left-it.md
    • issues/kernel/a-device-without-a-domain-of-its-own-reaches-all-memory.md
    • issues/kernel/user-programs-run-without-a-shadow-stack.md
    • issues/boot-media/the-loader-never-sets-the-firmwares-memory-overwrite-request.md

The oracle

The oracle is the kernel the T14 runs, Ubuntu's 6.8.0-142-generic.

  • Source. Tag Ubuntu-6.8.0-142.142 of git.launchpad.net/~ubuntu-kernel/ubuntu/+source/linux/+git/noble, commit 53e5d07aac028a1523ab0b115f079d6d1bc831ef (git ls-remote).
  • How it was read. From a depth-1 clone of the tag, whose HEAD is 53e5d07a, and from the source package linux_6.8.0-142.142 (linux_6.8.0.orig.tar.gz plus linux_6.8.0-142.142.diff.gz, both matching the md5 sums in the archive's .dsc). One diff hunk did not apply, in net/netfilter/xt_RATEEST.c, because macOS folds that name's case; no citation is in net/.
  • LLVM. rust/src/llvm-project at a79bc52c1d5e, the gitlink rust/ pins, read with git show a79bc52c:<path>.
  • Config. /boot/config-6.8.0-142-generic, extracted from linux-modules-6.8.0-142-generic_6.8.0-142.142_amd64.deb on the Ubuntu archive. Its sha256 is 3b8533dd9d235ca634ac58f82c5ce1ee35f12ef620693e17033184d2c9ca5890. The hardening table's numbers are that file's lines.
  • QEMU. QEMU citations are qemu/qemu at v11.1.1. qemu64 is CPUID_VENDOR_AMD, family 15 (target/i386/cpu.c:3545-3549).

Decisions

  • Parity with the pinned config, and no more. The track states one rule: it matches every hardening default the pinned config sets, and no more.
    • A hardening default is an option the config sets that, under the default command line, makes an attack on the kernel or a process harder and that no vulnerabilities line reports.
    • Access policy (the LSMs, lockdown, *_RESTRICT, STRICT_DEVMEM) belongs to the capability model and is not in the table.
    • Two options are split out by the tree:
      • IOMMU has a row because ToyOS binds every function to one identity domain over all memory (vtd/mod.rs:148,456-490), where Linux's default translates every device's DMA.
      • SHUFFLE_PAGE_ALLOCATOR is not applicable, because at this tag only page_alloc.shuffle=1 enables it (mm/shuffle.c:12-30). The Kconfig help's memory-side-cache detection has no caller.
    • The IBT line is gone: the rule covers it, since the config leaves CONFIG_X86_KERNEL_IBT unset.
  • SLS is int3 in the kernel's own assembly.
    • With thunk-extern and the external retpoline thunks, compiled code keeps no raw ret or indirect jmp, so no compiler option acts here.
    • S5 owes int3 after every ret and indirect jmp in the thunks and the entry code.
    • The mutation is deleting the int3 after one thunk's ret.
  • S0 is a one-time capture.
    • The kernel image, s0.cpio and busybox are never committed, and no build or test boots them. Only the captured text outputs are committed.
    • S0 builds the ToyOS boot-actuators facts line.
    • It also captures CPUID 0x80000000, 0x80000008 and 0x80000021, and MSR 0x10F where RTM_ALWAYS_ABORT and TSX_FORCE_ABORT are present.
  • S1's inputs include the AMD leaves. The TCG model is AuthenticAMD. Linux reads CPUID 0x80000008 EBX and 0x80000021 EAX (common.c:1072-1084) and folds their bits into IBRS, IBPB, STIBP and SSBD (common.c:994-1011).
  • S2's base is S1's whole x86_spec_ctrl_base, RRSBA_DIS_S included (bugs.c:1736,1800,1934,2224).
  • S4 names its -cpu override. Arch::cpu is one string per accelerator. The no-SMAP test drops +smap under TCG and uses host,-smap under KVM.
  • The MOR defect's Exit matches the oracle's own condition, and its test plants the variable. libstub/tpm.c:36-40 writes MemoryOverwriteRequestControl only where the firmware already defines it, and :42-45 ignores the write's result. The guest test plants it as 0 under e20939be-32d4-41be-a150-897f85d49829 (tpm.c:20-21) with vars::plant and asserts vars::live reads 1 after boot; deleting the loader's SetVariable reds it. A second boot plants nothing and asserts nothing exists under that GUID afterward; making the write unconditional reds it. vars::plant and vars::live gain the vendor as an argument.
  • S9 carries an LLVM change and a rustc change; no ToyOS-owned mechanism exists.
    • On x86_64-unknown-none, X86 lowers the guard to the segment slot only when hasStackGuardSlotTLS holds (X86ISelLoweringCall.cpp:548-551,564); otherwise it falls through (:604, :638-640) to the global __stack_chk_guard.
    • Measured with rust/build/aarch64-apple-darwin/ci-llvm/bin/llc (LLVM 22.1.8-rust-1.99.0-nightly, the CI LLVM, not the fork's build) on one sspstrong function with the module flags tls, gs, 40 under -code-model=kernel: x86_64-unknown-none-elf emits movq __stack_chk_guard(%rip), and x86_64-unknown-linux-gnu emits movq %gs:40. Both exit 0.
    • The other candidates are refused: LOAD_STACK_GUARD is 64-bit Mach-O only (X86ISelLowering.cpp:2770-2772), and a glibc or musl llvm-target for the kernel is a false environment claim.
    • Step 1 makes X86 honour an explicit stack-protector-guard=tls on any triple, as RISC-V does (RISCVISelLowering.cpp:25703-25707). Clang accepts the flag on every x86 triple (Clang.cpp:3479-3491,3561-3570), so the change is an upstream fix. Its test is RUN lines in stack-protector-3.ll.
    • Step 2 adds rustc's three -Zstack-protector-guard* options, set as module flags the way CodeGenModule.cpp:1543-1553 sets them. Its test is an assembly test for x86_64-unknown-none.
    • The orchestrator's ruling: both steps proceed. Root CLAUDE.md's dependency rule governs generally, not just at src/forkcheck.rs's existing target-arm dispatch site — "a fork carries a change written to upstream quality and goes when upstream has it" — so a general cross-platform option written to upstream quality is admitted when ToyOS needs it and upstream lacks it. Steps 1 and 2 are named steps of S9 with no more waiting; the stage that lands them amends src/forkcheck.rs's module header to admit exactly this class of change, in that stage's own PR.
  • S9's test reads the guard a protected frame compares.
    • One boot with smp: 1. Two threads each read %gs:N twice from inside a protected SYS_DEBUG frame, handing off to each other between the reads. The test asserts that each thread's reads agree and that the two threads' reads differ.
    • Overwriting the frame's own guard with the thread's own read returns, which proves the frame compares against %gs:N. Overwriting it with the other thread's read panics as the boot's last event.
    • Deleting context_switch's write of the incoming thread's guard makes both threads' reads agree and lets the overflow with the other thread's guard return, so both assertions red.
    • Linux's copy is __switch_to_asm (entry_64.S:193-196).

Gates

  • cargo run -- --ci host: EXIT=0 ("Host: 54 step(s), all green"), on the tree committed as d73f123.

Unsure

  • The MOR test's first boot assumes OVMF's variable driver accepts the loader's write. Nothing has measured that; the first run shows it.
  • The TCG capture recipe has not been run. It uses a static busybox /init in an initramfs, with direct kernel boot on q35 and -nodefaults -serial stdio. The first S0 run shows whether it works.
  • The "not applicable" rows for Rust are claims about safe Rust. INIT_STACK_ALL_ZERO and INIT_ON_ALLOC (slab) rest on Rust reading no local or allocation before writing it. copy_out's UserSafe no-padding contract is checked by hand (user_ptr.rs:30). An unsafe read of uninitialised memory is undefined behaviour and is outside what those rows cover.

🤖 Generated with Claude Code

… the T14

Owner ruling: ToyOS needs to be at least as secure as Linux on the same
hardware. Today the kernel carries no Spectre/MDS/GDS/ITS mitigation, no
KASLR or user ASLR, and no compiler stack protector, verified by grep against
main. Stages S0-S9 the orchestrator commissioned, with S7 (microcode loading)
parked for the owner because it is the one case CLAUDE.md's vendor-firmware
rule does not carve out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu
Japabu marked this pull request as ready for review September 28, 2026 23:04
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 1, at 4f790a4

CI: host success on 4f790a4 (run 36496028306). Growth: +158 −0, all tracker prose; production 0, tests 0. Linux references are torvalds/linux at v6.8 unless marked master; QEMU is qemu/qemu master target/i386/cpu.c. M = issues/kernel/the-kernel-mitigates-what-linux-mitigates-on-the-t14.md.

BLOCKER

  • M:44-47,67-72,78-80,85-88,112-114 — the QEMU exits for S1, S2, S3's IBPB and S6 cannot run on the fast tier — src/arch.rs:169 gives an x86 TCG guest qemu64,…, and system-mode TCG sets CPUID_7_0_EDX_KERNEL_FEATURES to 0 (SPEC_CTRL, ARCH_CAPABILITIES and SSBD exist only under CONFIG_USER_ONLY), so every one of those exits tests only the "absent" branch; -cpu host needs KVM/HVF on an x86 host. Each stage must name the accelerator its exit runs on, or say T14-only.
  • M:61-63 — CPUID.(7,0):EDX bit 29 enumerates IA32_ARCH_CAPABILITIES and bit 31 is SSBD (cpufeatures.h X86_FEATURE_ARCH_CAPABILITIES 1832+29, SPEC_CTRL_SSBD 1832+31). The stage calls 29 SSBD, invents "a different leaf" for 31, and gives no gate before its rdmsr 0x10A, which faults with #GP when bit 29 is clear.
  • M:64-72,155-156 — Linux does not classify from CPUID alone. "Affected" and "Not affected" come from model tables (common.c cpu_vuln_whitelist TIGERLAKE_L NO_MMIO; master cpu_vuln_blacklist TIGERLAKE_L GDS|ITS|ITS_NATIVE_ONLY). The bit list also omits SSB_NO, BHI_NO, PBRSB_NO, RRSBA, RFDS_NO, ITS_NO and CPUID.(7,2):EDX BHI_CTRL. S1 cannot reproduce Linux's lines, so the exit's "same CPUID-derived reason" cannot be checked. The negative control is also wrong: qemu64 is AuthenticAMD family 15, and Linux reports meltdown/ssb/l1tf/mds/mmio/itlb_multihit "Not affected" for it (common.c VULNWL_AMD(0x0f,…)). An S1 faithful to Linux fails "every line unmitigated-and-vulnerable".
  • M:75-77 — S2 sets the GDS lock bit and never clears GDS_MITG_DIS. On firmware that left DIS set, this locks the mitigation off until reset. Linux clears GDS_MITG_DIS and never writes GDS_MITG_LOCKED (bugs.c update_gds_msr, gds_select_mitigation). The only T14 reading, on unmerged 59c096b, is "Mitigation: Microcode", not "(locked)".
  • M:84,143-146 — Linux's default is conditional IBPB. SPECTRE_V2_USER_CMD_AUTO enables switch_mm_cond_ibpb (bugs.c spectre_v2_user_select_mitigation), and tlb.c cond_mitigation issues IBPB on an mm switch only when one side has TIF_SPEC_IB, which prints "IBPB: conditional". An IBPB on every cross-process switch is spectre_v2_user=on, which prints "IBPB: always-on". So the Decision's rationale is false, it contradicts the recorded orchestrator decision (per-process opt-in), and the track's own exit reds on it.
  • M:85-88 — S3's IBPB exit cannot fail on the idle path. kernel/src/arch/x86_64/hw.rs:444,450 activates a root on every switch, including idle's kernel root, and Cr3::activate(self) (paging.rs:300) has no prev. Mutation: issue IBPB only when neither the previous nor the next root is the kernel root. A→idle→B then gets no IBPB, and both stated arms stay green. The exit must switch A→idle→B (IBPB) and A→idle→A (none), which is Linux's per-CPU last_user_mm_spec.
  • M:81-82 — the BHB-clear condition is wrong. Linux clears the BHB in software on syscall entry when the CPU lacks BHI_DIS_S (X86_FEATURE_BHI_CTRL; master bugs.c bhi_apply_mitigation), and BHI is precisely the eIBRS-era attack. "When the CPU lacks eIBRS's hardware clearing" gates on the wrong bit, and Tiger Lake has eIBRS without BHI_DIS_S.
  • M:91-99 — S4 misses the sites Linux's spectre_v1 line names. Those are the usercopy barrier (lib/usercopy.c:21 barrier_nospec after access_ok), __user pointer sanitization and swapgs barriers. S4 does not cover kernel/src/user_ptr.rs's range check. ToyOS executes no swapgs (kernel/src/arch/x86_64/syscall.rs:99; FSGSBASE is CR4_FORBIDDEN at control_regs.rs:66), which by Linux's own rule (spectre_v1_select_mitigation: FSGSBASE || !smap_works_speculatively) makes the swapgs fences moot only while SMAP is present, and SMAP is optional (control_regs.rs:60). The "existing syscall fuzzing suite" does not exist (rg -i fuzz finds only the sched-sim's --fuzz-sweep at src/flags.rs:586), and the lint's detection rule is undefined, so the exit cannot be checked.
  • M:100-108 — ITS is mis-defined, and its oracle cannot fail. ITS hits indirect branches and RETs whose last byte is in the lower half of a cacheline, and Linux's fix is thunks whose branch sits in the upper half (master Documentation/admin-guide/hw-vuln/indirect-target-selection.rst, "Mitigation"). Mutation: place __x86_indirect_thunk_* and __x86_return_thunk with their jmp/ret at a cacheline start. The scan stays green. The oracle must assert that each thunk branch's last byte has addr & 63 >= 32.
  • M:126-133 — S8's exit cannot fail on a deterministic base. Mutation: base = USER_VM_BASE + spawn_count * PAGE_2M makes two spawns differ, both inside the user half. The scope is also below Linux's: the stack sits directly under the image by construction (kernel/src/vma.rs:9-12, ALLOC_CEILING = STACK_BASE), the mmap region is not randomized, and 2 MiB alignment caps entropy at 26 bits over the whole user half. The exit must tie the base to an arch::entropy::draw and state its entropy against the T14's vm.mmap_rnd_bits, which S0 does not capture.
  • M:134-139 — S9's exit passes with a constant canary. The kernel target is x86_64-unknown-none (kernel/.cargo/config.toml:2), so LLVM reads a global __stack_chk_guard that the kernel must define and seed before any protected frame runs; the stage names neither. One global canary is also weaker than Linux's per-task one. Mutation: static __stack_chk_guard: u64 = 0, and the planted-overflow test stays green.
  • M:152-158 — the exit cites "the LLVM toolchain bar (issues/build/toyos-builds-itself.md)", but at this head that file has no bar and no sysfs lines; both exist only on unmerged wt/toyos-llvmbar (59c096b). The scope is also short of the owner ruling. The issue says ToyOS has no KASLR (M:32), and Linux defaults KASLR on (arch/x86/Kconfig RANDOMIZE_BASE default y), as it does X86_KERNEL_IBT (def_bool y, and Tiger Lake has CET). No stage, decision or filed issue covers either, and the exit checks only sysfs lines, so S8 and S9 are outside it too.
  • M:85-87,109-114,131 — the exits add a "test syscall" and a "probe syscall", and S6's opt-in is new ABI, against root CLAUDE.md's "Never add or change a syscall without discussion". The tree already has both mechanisms: the boot-actuators feature-gated kernel and the build-time byte gate on syscall_entry (src/build.rs:1555-1600). Checking that the BHB sequence is present belongs to that gate, not to a counter on the hot path.

NOTE

  • M:51-52 — 0x123 is IA32_MCU_OPT_CTRL (the GDS bits), and IA32_TSX_CTRL is 0x122 (msr-index.h:236,240). S0 records the register S2 depends on under the wrong name and never reads TSX_CTRL.
  • M:52-53 — reading 0x48 on a running Linux shows the value Linux wrote, not the reset value. The reset value is ToyOS's read before its first write.
  • M:53-54 — reading 0x8B without first writing it returns a stale revision. The sequence is write 0, CPUID(1), then read (master arch/x86/include/asm/microcode.h intel_get_microcode_revision).
  • M:73-80,109-114 — S2 declares SPEC_CTRL once and asserts it like CR0/CR4/EFER, while S6 makes it per-thread state switched on every context switch. The stages contradict each other, and S6 has to rewrite S2's self_check invariant.
  • M:102-104,135-136 — -Zretpoline-external-thunk is a target modifier (rustc_session/src/options.rs:2769), so core/alloc come from the toolchain's x86_64-unknown-none sysroot, and src/toolchain.rs has to rebuild them. Neither S5 nor S9 names that site.
  • M:73-90 — the T14's spectre_v2 line includes "PBRSB-eIBRS: SW sequence" and "KVM: SW loop", both VM-exit mitigations with no ToyOS counterpart. The exit must say how those parts of the line are classified.
  • M:67-72 — S1 needs a different -cpu per test, but src/arch.rs:166 is one declaration for every guest. The stage must say how a control CPU model reaches the harness.

REMOVE

  • M:11-12 — "Linux 6.8's": no evidence in the tree; S0 captures the version.
  • M:21-24 — the parenthetical listing rg's "unrelated matches": it is false (rg swapgs hits syscall.rs:99) and will rot.
  • M:21 — "and no swapgs timing mitigation": misleading, since ToyOS executes no swapgs.
  • M:33-40 — the ToyOS-target stack-protector citations: the kernel builds x86_64-unknown-none, and the rust line numbers will rot.
  • M:44-47 — the QEMU/T14-only definition: false under TCG.
  • M:62-63 — "(SSBD enumerated via a different leaf on some parts — read the SDM table S0 pulled)": false.
  • M:100-101 — the ITS parenthetical definition: false.
  • M:110-111 — "seccomp": 6.8's default mode is prctl-only (bugs.c "…disabled via prctl").
  • M:128 — "(SYS_RANDOM, already implemented)": the kernel draws from arch/x86_64/entropy.rs, not a syscall.
  • M:131 — "(e.g. through SYS_SYSINFO or a probe syscall)".
  • M:136 — "(or the strategy S0/S1's threat model picks)": neither stage has a threat model.
  • M:147-150 — the second Decision bullet: it restates S7.
  • PR body — "each claim checked against main before writing" and "418 passed, 3 ignored": narration, and a count with no log for an issue-only change.

SEND BACK

Japabu and others added 2 commits September 29, 2026 01:23
…logic

The review's blockers, answered against torvalds/linux v6.16 and QEMU
v11.1.1 source:

- The decision is a pure host-tested function carrying v6.16's
  cpu_vuln_whitelist/blacklist, intel-ucode-defs.h and bugs.c selection, fed
  S0's T14 facts and the TCG model's; its expected output is the captured
  Linux lines. QEMU exercises only the absent branches and the software
  sequences the TCG model selects (CPUID_7_0_EDX_KERNEL_FEATURES is 0 in
  system mode).
- CPUID.(7,0):EDX bit 29 is ARCH_CAPABILITIES and 31 SSBD; 0x10A is read
  only when bit 29 is set.
- GDS clears GDS_MITG_DIS and never writes GDS_MITG_LOCKED
  (update_gds_msr).
- IBPB is conditional per-process opt-in (switch_mm_cond_ibpb), decided
  from the per-CPU last-user state, host-tested on A->idle->B and
  A->idle->A.
- BHB clearing is gated on BHI_CTRL (CPUID.(7,2):EDX[4]), not eIBRS.
- Spectre v1 names the usercopy barrier and pointer masking in user_ptr's
  window/object; SMAP becomes required; the lint and the nonexistent fuzzing
  suite are gone.
- ITS thunks are asserted to end their branch at addr & 63 >= 32.
- ASLR covers image, stack and mmap bases with entropy stated against
  mmap_rnd_bits and a per-bit frequency test that reds on BASE + n*2MiB.
- The stack protector defines and seeds __stack_chk_guard; a zero guard
  reds kernel_stack_canary; the per-task canary is a recorded gap.
- The exit cites only main; kernel KASLR (S10) and kernel IBT (S11) are
  added, and S8-S11 sit inside the exit.
- Probes use boot-actuators and the syscall_entry byte gate; the one ABI
  change is a flag in SpawnArgs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 2, at c0cce22

CI: host success on c0cce22 (run 36498933149). Growth: +204 −0, all tracker prose. Production 0, tests 0. Linux citations are torvalds/linux at v6.16. M = issues/kernel/the-kernel-mitigates-what-linux-mitigates-on-the-t14.md.

Round-1 blockers

  • B1 OPEN — the file assumes QEMU is TCG qemu64, but the nightly's 12 guest shards and its audio shards run with --device=/dev/kvm (.github/workflows/nightly.yml:92-94). There, Arch::accel returns Kvm and Arch::cpu gives host,… (src/arch.rs:148-153,168), so the present branches (SPEC_CTRL, SSBD, the 0x10A read, IBT on an Intel runner) run on whatever CPU GitHub hands out, and no S0 fixture covers it. S3's QEMU RSB count also depends on the runner: on an eIBRS or AutoIBRS host, v6.16 fills no RSB on switch (bugs.c:2027-2034).
  • B2 CLOSED — M:68-72.
  • B3 OPEN — the whitelist half of S1's negative control cannot go red. v6.16's TIGERLAKE_L whitelist row carries only NO_MMIO (common.c:1148), and nothing in v6.16 reads NO_MMIO: MMIO is classified by arch_cap_mmio_immune plus the blacklist (common.c:1485-1488). Deleting that row changes no line. The AMD family 0xf part is closed, since the TCG facts are now fed.
  • B4 CLOSED as text. A new mutation below.
  • B5 CLOSED — M:192-196.
  • B6 CLOSED — M:100-104.
  • B7 CLOSED — M:90-94, matching v6.16 bugs.c:2114-2143.
  • B8 OPEN — see the S4 line below.
  • B9 CLOSED — M:130-131.
  • B10 OPEN — see the S8 line below.
  • B11 OPEN — see the S9 line below.
  • B12 CLOSED — S10 and S11 are added, and the exit is at M:202-204.
  • B13 CLOSED — M:197-198. The orchestrator approved the SYS_SPAWN flag.

BLOCKER

  • M:13,50-57 — S0 says "Under Linux v6.16 on the T14", but the T14 runs Ubuntu 24.04.4 with 6.8.0-142-generic (59c096b on wt/toyos-llvmbar). The file never says:

    • where a v6.16 kernel comes from;
    • which config it uses, although mmap_rnd_bits and every CONFIG_MITIGATION_* default depend on the config;
    • how it boots without displacing the Ubuntu entry src/metal.rs returns to;
    • which QEMU runs the TCG capture, although Ubuntu 24.04 ships 8.2.2 and .github/qemu-version pins 11.1.1 because versions disagree.

    A capture from 6.8.0-142 fails a v6.16 fixture on the file set alone: v6.16 has old_microcode and tsa. S0 must either name the build, config, boot path and QEMU, and record /proc/version in the fixture, or re-pin the oracle to the Ubuntu source the T14 runs.

  • M:202-203 — the Exit accepts "prints the line S1 computes". Mutation: gather arch_cap: 0 instead of rdmsr(0x10A). The T14 then prints a different computed line and the Exit still holds. The T14's printed lines must equal S0's Linux strings.

  • M:84-89 — delete the GDS_MITG_DIS clear, and S2's read-back stays green on any T14 whose firmware left DIS clear. Linux's "Mitigation: Microcode" reading on the T14 allows that. A boot-actuators arm must set DIS before init and assert it cleared.

  • M:90-98 — the syscall_entry byte gate sees that the sequence is present, not that it runs. Invert the runtime condition (if decision.bhi_ctrl { clear_bhb() }) and the gate stays green while the T14 clears nothing. A T14 exit must show the loop ran on syscall entry.

  • M:105 — "The T14 counts the IA32_PRED_CMD writes" names no scenario and no expected count. Drop the IBPB write from the switch path, keep the pure function, and nothing goes red.

  • M:116-118 — the S4 gadget goes through window only and has no pass threshold. Delete the lfence after is_user_object in object (kernel/src/user_ptr.rs:95), or delete the user-half mask, and no exit goes red.

  • M:119-131 — nothing tests that the thunk body is "the one S1 selects". Mutation: fn thunk_body(_: &Decision) -> Body { Body::ItsAligned }. The retpoline-mode TCG model then runs a bare jmp *%reg, and the alignment gate stays green. The QEMU boot must read the live thunk back and compare it with S1's selection.

  • M:136-140 — S6 claims inheritance, but the exit tests only a flag set directly. Drop || parent.flagged at spawn and it stays green. The exit must spawn a child of a flagged process with the bit clear in SpawnArgs and read SSBD set.

  • M:152-163 — B10. Replace arch::entropy::draw() on the spawn path with a constant-seeded splitmix64. 256 spawns in one boot then pass every bit test, yet every boot repeats the same sequence. Two boots' sequences must differ, or the first spawn's bases must be bit-tested across S10's 64 boots.

  • M:164-173 — B11. static __stack_chk_guard: u64 = 0x5a5a_5a5a_5a5a_5a5a still panics on a zero overflow, so the exit stays green. S10's 64 boots must print the guard and bit-test it.

NOTE

  • M:98-99 — (b) S3's retpoline claim holds in the v6.16 source. VULNWL_AMD(0x0f, …) lacks NO_SPECTRE_V2 (common.c:1187,1414), and AUTO without IBRS_ENHANCED picks RETPOLINE (bugs.c:2158-2164,1970-1978). S0's TCG capture also measures it.
  • M:174-182 — (b) S10 names no Tier (src/tiers.rs) and prices its 64 boots nowhere. It must place itself in a tier and state the wall time from a measured boot.
  • M:65-67 — v6.16 reads NO_MMIO nowhere, so the port must not carry that whitelist bit as dead data.
  • M:152-182 — entropy::draw returns None (kernel/src/arch/x86_64/entropy.rs:18). S8, S9 and S10 must refuse the spawn or the boot on None, never fall back to a fixed value.
  • M:168-170 — the per-task canary gap is recorded only as a sentence inside a track. It needs its own issue with an owner and an exit condition.
  • M:174-179 — S10 does not say whether the loader or the kernel draws the bases. It also does not say how PHYS_OFFSET, declared twice (kernel/src/mm/mod.rs:33, bootloader/src/main.rs:310), becomes one handed-over value.
  • M:116 — the SMAP refusal is checked by a mutation "run once", so no standing test goes red if SMAP moves back to CR4_OPTIONAL.
  • M:138-140 — userland cannot rdmsr. The exit must name the boot-actuators probe that reads SSBD back.

REMOVE

  • M:45-48 — "QEMU exercises the absent branches …": false on the KVM guest lanes.
  • M:77-78 — "The QEMU boot prints the TCG model's lines …": it asserts nothing.
  • M:198 — "the one ABI change is S6's bit": it restates S6.
  • PR body — "## Round 1, per finding": review chronology does not belong in main's merge record.
  • PR body — the S3 "Unsure" bullet: the v6.16 source settles it.
  • PR body — "## Gates": a host-lib run says nothing about a tracker-only change.

SEND BACK

… names what breaks it

The oracle moves from torvalds/linux v6.16 to Ubuntu's 6.8.0-142-generic,
the kernel the T14 boots (59c096b on wt/toyos-llvmbar). Its source is tag
Ubuntu-6.8.0-142.142 of the noble kernel tree: `git ls-remote` lists it as
the only tag with ABI 142 (tag object e230fb4a, commit 53e5d07a), and its
debian.master/changelog heads with 6.8.0-142.142. Its config is
/boot/config-6.8.0-142-generic from
linux-modules-6.8.0-142-generic_6.8.0-142.142_amd64.deb (package sha256
ee4feb47..., config sha256 3b8533dd...). Every Linux path:line was re-read
at that tag.

What the re-pin changed:
- The config leaves CONFIG_X86_KERNEL_IBT unset, so kernel IBT is not on
  the scoreboard and S11 is deleted.
- CONFIG_SLS=y is the one CPU_MITIGATIONS menu entry no vulnerabilities
  line reports; S5 now owes int3 after every ret and indirect jmp. The rustc
  fork has no option for it.
- CONFIG_ARCH_MMAP_RND_BITS=32 is S8's figure.
- NO_MMIO is read at this tag (common.c:1495-1499).
- The microcode table S1 carries is intel.c's spectre_bad_microcodes;
  intel-ucode-defs.h does not exist at this tag.
- X86_USER_SHADOW_STACK=y. S0 measures whether Ubuntu's processes use it.

S0 names its exact commands, verifies the kernel, package and config
hashes, and ends only once the capture is committed, before Ubuntu is
wiped. The TCG capture runs on the development host's QEMU, the version
.github/qemu-version pins.

KVM: the nightly's guest shards run -cpu host, so a QEMU test asserts
wiring against S1 over the facts the guest reports. Independence comes
from the T14 alone.

S1's negative control deletes GDS from the TIGERLAKE_L blacklist row.
Each other exit names the mutation that breaks it and the value it
breaks:
- the T14 prints S0's strings, with a per-line not-worse rule;
- the GDS_MITG_DIS arm;
- the BHB counter at 1000 per 1000 syscalls;
- RSB fills computed from the reported facts;
- 2 IBPB writes over the scripted sequence B, C, idle, B, A, idle, A, B;
- S4's gadget at <= 16 of 1000 trials per site (false red 7.6e-7), plus a
  standing no-SMAP boot;
- the thunk-body read-back;
- S6 inheritance across spawn;
- the first spawn's bases and the canary, bit-tested across S10's 64
  weekly boots.
Binomial tails were computed with awk: 1.01e-7 at n=64 [12,52], 3.31e-7 at
n=256 [88,168], and 7.61e-7 for P(X>16) with X ~ Bin(1000, 1/256).

S10 is Tier::Weekly and times one boot first. The loader draws both
kernel bases, and PHYS_OFFSET becomes one handed-over value. Every
entropy draw refuses on None.

Filed issues/kernel/the-kernel-stack-canary-is-one-global-not-per-task.md.

Removed: "QEMU exercises the absent branches ...", "The QEMU boot prints the
TCG model's lines ...", and "the one ABI change is S6's bit".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 3, at 8df7bc1

CI: host success on 8df7bc1 (run 36502859985). Growth: +300 −0, all tracker prose; production 0, tests 0. M = issues/kernel/the-kernel-mitigates-what-linux-mitigates-on-the-t14.md. Linux is read at Ubuntu-6.8.0-142.142. git ls-remote gives the tag's peeled commit as 53e5d07a, which matches M:14. The config is read from the tag's debian.master/config/annotations.

Round-2 blockers

  • B1 CLOSED — M:49-57. S2, S3 and S6 compare KVM against S1 over the reported facts.
  • B3 CLOSED — X86_BUG_GDS is set only at common.c:1521-1523, from the TIGERLAKE_L row at 1312. cpu_show_common answers "Not affected" at bugs.c:3308-3309.
  • B8 CLOSED — M:162-166. Each of the three sites has a threshold, and each has its own mutation.
  • B10 CLOSED — M:223-225.
  • B11 CLOSED — M:236-238.
  • S0 against v6.16: CLOSED — the oracle is re-pinned to the tag and config at M:13-19, and the capture is refused on a mismatch (M:86-88).
  • Exit reads a computed line: CLOSED — M:266-281. The claim about ARCH_CAPABILITIES read as 0 holds at bugs.c:1878-1895 and at bugs.c:831-842 (UCODE_NEEDED).
  • GDS_MITG_DIS arm: CLOSED — M:120-123.
  • BHB clear runs: CLOSED as asked. The fix opens the S3 blocker below.
  • IBPB count: CLOSED as asked. See the S3 blocker below.
  • S4 per site: CLOSED.
  • S5 thunk_body: CLOSED — M:183-187.
  • S6 inheritance: CLOSED — M:195-198.

BLOCKER

  • M:20-25 — the track keeps five of the pinned config's hardening defaults and names no rule for leaving out the rest.

    • The config also sets these (annotations line numbers):
      • RANDOMIZE_KSTACK_OFFSET_DEFAULT=y (10676)
      • ZERO_CALL_USED_REGS=y (15523)
      • INIT_ON_ALLOC_DEFAULT_ON=y (6453)
      • VMAP_STACK=y (15139)
      • STRICT_KERNEL_RWX=y (13454)
      • HARDENED_USERCOPY=y (5569)
      • SLAB_FREELIST_HARDENED and SLAB_FREELIST_RANDOM (12177-12178)
      • SHUFFLE_PAGE_ALLOCATOR=y (12144)
    • The IBT deletion follows the config (annotations:834, 'amd64': 'n'), so it is parity and it is accepted. The omissions above do not follow the config. The Exit can therefore be declared at "at least as secure as Linux" while those gaps stand.
    • Each item must become one of four things:
      • a stage;
      • a filed defect;
      • a named ToyOS equivalent at a path;
      • "not applicable", with the reason. For example, INIT_STACK_ALL_ZERO has nothing to do in safe Rust.
  • M:176-183 — S5's SLS mutation cannot go red.

    • -Zharden-sls does not exist upstream. -Zharden-sls flag (target modifier) added to enable mitigation against straight line speculation (SLS) rust-lang/rust#136597 is OPEN and unmerged, and neither upstream/main nor the fork pin 1b236638 carries harden_sls.
    • More to the point, S5's own gate allows no raw ret or indirect jmp outside hand-written thunks and entry asm. -Zfunction-return=thunk-extern and -Zretpoline-external-thunk leave none in compiled code. A compiler SLS option, or LLVM's harden-sls-ret/harden-sls-ijmp, would therefore change no byte.
    • What is missing is int3 in the kernel's own asm, as Linux's RET macro does (arch/x86/include/asm/linkage.h:46,58).
    • Replace "building without the SLS hardening" with: delete the int3 after one thunk's ret, and the gate goes red.
  • M:132-149 — S3's counters prove that a sequence runs, not that it is the right sequence. The syscall_entry content check was dropped in this round. Each of these mutations keeps every count and passes:

    • movl $5,%ecx → movl $1,%ecx in the BHB loop, or deleting its trailing lfence (compare entry_64.S:1534-1569);
    • an RSB fill of 1 entry instead of RSB_CLEAR_LOOPS 32;
    • wrmsr(IA32_PRED_CMD, 0) instead of PRED_CMD_IBPB.

    S3 needs two fixes:

    • a kernel.elf gate that compares the BHB and RSB sequences with the tag's clear_bhb_loop and __FILL_RETURN_BUFFER;
    • an IBPB probe that counts only writes of value 1.

NOTE

  • M:59-90 — the file never says the TCG capture is one-time. It also never says that vmlinuz-6.8.0-142-generic, s0.cpio and busybox are left out of the commit and are booted by no gate or test. "those outputs committed" (M:89) reads as including both binaries. S0 must say that only the text output is committed, and that nothing in build or test boots the Linux guest.
  • M:83-86 — S0 needs a ToyOS boot-actuators facts line, and no stage builds it.
  • M:91-94 — S1's inputs omit CPUID 0x80000008 EBX and 0x80000021, which Linux reads on AMD. The TCG model is AuthenticAMD family 0xf. S0 captures neither leaf.
  • M:111-115 — S2's base leaves out SPEC_CTRL_RRSBA_DIS_S (bugs.c:1736). The base must be S1's full selection.
  • M:159-161 — the harness has one -cpu string (src/arch.rs Arch::cpu). The standing no-SMAP test needs an override that the stage does not name, and on KVM lanes that override must be host,-smap.
  • Round-2 NOTE at M:65-67 withdrawn — at the tag, NO_MMIO is read at common.c:1498. The implementer's refutation holds.

REMOVE

  • M:108-110 — "(59c096b on wt/toyos-llvmbar)": it cites a commit on an unmerged worktree branch, and that will rot.
  • M:177-178 — "the rustc fork has no option for it …": it misleads, because no compiler option would act here.
  • PR body — the "Every exit names the mutation" table restates the track in main's record.
  • PR body — the "Unsure" bullet that cites 59c096b.

SEND BACK

Japabu and others added 2 commits September 29, 2026 02:56
…isposition

The track now states its rule: it matches every hardening default the
pinned Ubuntu config sets, and no more. A table gives each of 27 options
exactly one disposition, verified at Ubuntu-6.8.0-142.142 and at this tree:
6 to a stage, 10 to one of seven defects filed here, 4 to a named ToyOS
mechanism with file:line, and 7 not applicable with the reason.

Filed:
- every-syscall-runs-at-one-kernel-stack-offset (RANDOMIZE_KSTACK_OFFSET_DEFAULT)
- kernel-functions-return-with-their-used-registers-intact (ZERO_CALL_USED_REGS)
- a-threads-kernel-stack-has-no-guard-page (VMAP_STACK)
- kernel-text-is-writable-and-every-kernel-page-executable (STRICT_KERNEL_RWX, DEBUG_WX)
- the-kernel-heap-has-none-of-slubs-hardening (SLAB_FREELIST_RANDOM,
  SLAB_FREELIST_HARDENED, RANDOM_KMALLOC_CACHES)
- tsx-stays-as-firmware-left-it (X86_INTEL_TSX_MODE_OFF)
- a-device-without-a-domain-of-its-own-reaches-all-memory (INTEL_IOMMU_DEFAULT_ON)

S5: no compiler option acts on SLS, because thunk-extern and the external
retpoline thunks leave compiled code no raw ret or indirect jmp. The int3 is
owed in the kernel's own assembly, as Linux's RET does
(linkage.h:46-47,58-59), and the mutation is deleting the int3 after one
thunk's ret.

S3: a kernel.elf gate compares the BHB clear and the RSB fill with
clear_bhb_loop (entry_64.S:1534-1569) and __FILL_RETURN_BUFFER with
RSB_CLEAR_LOOPS (nospec-branch.h:132,137-162); the IBPB probe counts only
writes of PRED_CMD_IBPB. The three mutations the review named are the
ones that must go red.

S0 builds the ToyOS facts line, captures CPUID 0x80000000, 0x80000008 and
0x80000021 and MSR 0x10F, and says the capture is one-time: only its text
outputs are committed. S1 takes the two AMD leaves the TCG model's
AuthenticAMD path reads (common.c:1072-1084). S2's base is S1's whole
x86_spec_ctrl_base, RRSBA_DIS_S included. S4 names the per-test -cpu
override. The citation of an unmerged branch's commit is removed.

Linux was read from the source package linux_6.8.0-142.142 (orig tarball
plus diff, both matching the .dsc's md5), whose lines agree with the
review's reads of the tag.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu Japabu changed the title Track: the kernel mitigates every CPU vulnerability Linux mitigates on the T14 Track: the kernel mitigates what Linux mitigates on the T14, and matches its hardening defaults Sep 29, 2026
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 4, at 53da60c

CI: host success on 53da60c (run 36508042871). Growth: +534 −0, all tracker prose; production 0, tests 0. M = issues/kernel/the-kernel-mitigates-what-linux-mitigates-on-the-t14.md.

How this round was checked:

  • Config. /boot/config-6.8.0-142-generic, extracted from the archive's linux-modules-6.8.0-142-generic_6.8.0-142.142_amd64.deb, has sha256 3b8533dd…5890, as M:19 says. All 31 lines the table cites read what the table says.
  • Hardening section. The config's Kernel hardening options section (11462-11488) sets three options: INIT_STACK_ALL_ZERO, INIT_ON_ALLOC_DEFAULT_ON and ZERO_CALL_USED_REGS. All three have rows.
  • Linux source. I read Linux at the tag itself: a blobless fetch of Ubuntu-6.8.0-142.142 resolves to 53e5d07a. At the tag, these cited lines say what the track and the defects say:
    • entry_64.S:1534-1569
    • nospec-branch.h:132-162
    • msr-index.h:61-62
    • mm/shuffle.c:12-30
    • common.c:360-370,1072-1084
    • sched/core.c:5957-5958
    • bugs.c:1736,1800,1934,2224
    • intel/iommu.c:227
    • iommu.c:197-210
    • fork.c:314-318,1161
    • entry-common.h:76-85
  • The pin. The source package is equivalent to the tag for the cited text files. The file names the tag, which is the right pin. That the package was what got read is provenance, and it belongs only in the PR body, where it is.

Round-3 blockers

  • Hardening defaults without a disposition: CLOSED. Every option the round-3 review named now has exactly one row, and M:25-32 states the parity rule.
  • The mechanism rows hold:
    • pmm.rs:280-282,318-320: every PMM handout is zeroed.
    • driver.rs:521,592,925-935: these are both SchedPass::begin sites. The word is at the allocation's low end, which is where the stack ends. Every kernel stack is painted in alloc_kernel_stack.
    • user_ptr.rs: every copy is bounded by the type's extent.
    • control_regs.rs:60,212: CPUID.(7,0):ECX bit 2 sets UMIP.
  • The not-applicable rows are true of ToyOS:
    • DMA, pipe and user memory all come from zeroed PMM pages.
    • kernel/src maps no vsyscall page and loads no kernel code.
    • At the tag, only page_alloc.shuffle enables the shuffle key.
  • SLS: CLOSED (M:234-243). The int3 goes in the kernel's own assembly, a byte gate checks it, and the mutation deletes one thunk's int3.
  • S3 IBPB probe: CLOSED (M:199-205). It counts only writes of PRED_CMD_IBPB (1), and writing 0 reds it.
  • S3 sequence gate: OPEN (M:183-187). The gate compares loop counts, the fill's call count and the final lfences. It does not compare the sequences with Linux's, which is what the round-4 ruling requires. See the first BLOCKER below.

BLOCKER

  • M:183-187 — S3's gate checks features of the sequences, not the sequences themselves. It compares the two loop counts, the call count and the trailing lfence. Each of these mutations keeps all of those and stays green:

    • (a) delete 3: jmp 4f / nop from clear_bhb_loop's inner loop (entry_64.S:1558-1560 at the tag). This halves the taken branches per inner pass, while both movl $5 and the lfence stay.
    • (b) delete the int3 after each fill call. That is __FILL_RETURN_SLOT's speculation trap (nospec-branch.h:137-141), which M:182 itself claims.
    • (c) replace the call 1f/call 2f/RET nesting with direct jmps.

    The fix: the gate compares each decoded instruction stream, with opcodes, immediates and normalised relative targets, against clear_bhb_loop and __FILL_RETURN_BUFFER(RSB_CLEAR_LOOPS). Mutations (a) and (b) must red it.

  • M:34-62 — RESET_ATTACK_MITIGATION=y (config 2457) has no row.

    • At the tag, the x86 EFI stub runs efi_enable_reset_attack_mitigation() on every boot (drivers/firmware/efi/libstub/x86-stub.c:971). That sets MemoryOverwriteRequestControl=1 (libstub/tpm.c:29-46), so firmware wipes RAM after an unclean reset.
    • bootloader/src never writes MOR. The M:25 parity sentence can therefore be declared met while this default is missing.
    • It needs a row, and the disposition is a filed defect: the write belongs to the loader, before ExitBootServices.
  • M:41 with M:128-129 and M:340-341 — the X86_USER_SHADOW_STACK row's disposition, "S0", closes nothing.

    • If S0 sees shstk, the owed work is "filed as their own track".
    • The Exit waits only on "every defect the hardening table cites", so it never waits on that track.
    • The row must name what builds user shadow stacks, with the Exit waiting on it, or give a one-clause not-applicable reason.

NOTE

  • M:26-28 — the definition leaves out two of its own rows. "makes an attack on the kernel harder" excludes ARCH_MMAP_RND_BITS and X86_USER_SHADOW_STACK, which protect processes. The definition decides which options belong in the table, so it must cover processes too.
  • M:34-62 — three options have no row. BPF_JIT_ALWAYS_ON (124), MODULE_SIG (982) and KEXEC_SIG (318) each need a one-clause not-applicable row: ToyOS has no BPF, no modules and no kexec.
  • M:57 — the reason names only copy_out. Three sites publish struct bytes outside UserSafe, through from_raw_parts: kernel/src/object/ops.rs:444 and kernel/src/syscall/device.rs:279,307. They are covered only by the size assertions at toyos-abi/src/pci.rs:86,103,124. The reason must name that path.
  • issues/kernel/the-kernel-stack-canary-is-one-global-not-per-task.md — this kind: defect is not "real, reproducible" (issues/README.md). Its own line 15 says the gap opens only when S9 lands. It belongs in S9's Exit, or it is filed when S9 lands.
  • M:1-341 — the track's length breaks issues/README.md. The README says a track "does not carry a design, a stage table, a rationale… A track that has grown past a screen is a plan again, and is cut back." This file is 341 lines. The orchestrator should put the owner's rulings against that rule.

REMOVE

  • M:236-238 — "as Linux's RET and ASM_RET are under CONFIG_SLS (linkage.h:46-47,58-59)" is false under the pinned config. The config sets MITIGATION_RETHUNK (557). linkage.h:43-44 then defines RET as jmp __x86_return_thunk. The ret; int3 Linux runs is written by patch_return (alternative.c:985-1000). The PR body's SLS bullet makes the same claim.
  • issues/kernel/a-device-without-a-domain-of-its-own-reaches-all-memory.md:18-19 — "The refusal … tracks is this defect's fix" is false. That track's refusal is the EIM rule and the isolation-scope rule. Neither moves a function off the identity domain.
  • PR body — the disposition-count table (6/10/4/7) and "27 rows". These counts restate the track and will rot.
  • PR body — "the review's :46,58 are the #ifdef lines". This is review history.

SEND BACK

Japabu and others added 2 commits September 29, 2026 04:03
S3's gate compared loop counts, the fill's call count and the trailing
lfence with Linux's clear_bhb_loop and __FILL_RETURN_BUFFER; that let a
gutted sequence keep every count and stay green. The Exit now decodes
each instruction, opcode, immediate and normalised relative target,
against both functions at the pinned tag, and names three mutations
that must red it: dropping the inner loop's jmp/nop, dropping the
int3 __FILL_RETURN_SLOT places after each fill call, and flattening
the call/RET nesting into direct jmps.

RESET_ATTACK_MITIGATION and X86_USER_SHADOW_STACK now each carry a
filed defect instead of a disposition ("S0") that closed nothing: S0
was only ever going to file a track, and the Exit waits on defects,
not a track filed later. The MOR write is the loader's, before
ExitBootServices, and points at the loader-slimming track it belongs
with.

Fixed on the way: the hardening definition now covers process
protections, not just the kernel; BPF_JIT_ALWAYS_ON, MODULE_SIG and
KEXEC_SIG get not-applicable rows; the per-task-canary defect is
deleted because its own body said the gap it describes does not exist
until S9 lands, and S9 now says the gap is filed then, not before.

Removed, not rewritten: the false claim that Linux's RET/ASM_RET are
int3'd under CONFIG_SLS (the pinned config selects the return-thunk
instead), and the IOMMU defect's claim that another track's refusal is
this defect's fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

Review, round 5, at 956301f

CI: host success on 956301f. Growth: +562 −0, all tracker prose; production 0, tests 0. M = issues/kernel/the-kernel-mitigates-what-linux-mitigates-on-the-t14.md. I read Linux from a fresh fetch of refs/tags/Ubuntu-6.8.0-142.142, which resolves to 53e5d07aac028a1523ab0b115f079d6d1bc831ef. Config values come from debian.master/config/annotations at that tag.

Round-4 blockers

  • S3 sequence gate (M:187-194): CLOSED.
    • The Exit now compares decoded instructions (opcode, immediate, normalised relative target) against clear_bhb_loop and __FILL_RETURN_BUFFER(RSB_CLEAR_LOOPS). Mutations (a), (b) and (c) are each named as reds.
    • At the tag, the cited lines say what M says:
      • entry_64.S:1534-1569 is SYM_FUNC_START(clear_bhb_loop) through EXPORT_SYMBOL_GPL.
      • :1558-1559 is 3: jmp 4f / nop.
      • nospec-branch.h:132 is RSB_CLEAR_LOOPS 32.
      • :137-141 is __FILL_RETURN_SLOT, with call 772f; int3.
      • :151-162 is the loop: 16 passes of 2 slots, then lfence.
  • RESET_ATTACK_MITIGATION row (M:43): CLOSED.
    • At the tag, x86-stub.c:971 calls efi_enable_reset_attack_mitigation().
    • bootloader/src writes no MOR: its only set_variable calls are bootnext.rs:55 and floor.rs:104.
    • The defect's Exit is wrong. See the first BLOCKER.
  • X86_USER_SHADOW_STACK (M:42): CLOSED. The table now cites a filed defect, so the track's Exit waits on it. kernel/src has no CET, U_CET or shadow-stack mapping.

BLOCKER

  • issues/boot-media/the-loader-never-sets-the-firmwares-memory-overwrite-request.md:9-13,21-22 — the Exit asks for "sets MemoryOverwriteRequestControl=1 … on every boot", which is not what the oracle does.
    • At the tag, libstub/tpm.c:36-40 calls get_efi_var first and returns without writing when the answer is EFI_NOT_FOUND. It writes only a variable the firmware already defines, and it ignores the status of set_efi_var (:42-45).
    • Built as written, the loader creates a non-volatile variable under the TCG GUID e20939be-… on firmware that has no MOR, such as the harness's OVMF. A test there then passes without measuring anything.
    • Fix: the body and the Exit state the condition, citing tpm.c:36-40.
  • M:298-300 — the per-task canary is left owned by nothing that the Exit waits on.
    • "filed as a defect when this stage lands" puts the per-task defect in no hardening-table row and no Exit.
    • M:348-349 waits only on S1-S10 and on "every defect the hardening table cites". S9's Exit (M:301-306) is met by one global guard.
    • The track therefore closes at STACKPROTECTOR_STRONG parity while Linux's canary stays per task (kernel/fork.c:1161). This is the hole round 4 sent back for shadow stacks. The round-4 NOTE's "or it is filed when S9 lands" was wrong.
    • Fix: S9's Exit requires a per-thread canary, or S9 files the defect and the STACKPROTECTOR_STRONG row (M:40) cites it.

NOTE

  • M:190 — the fill reference contains an instruction ToyOS has no counterpart for. "compares each decoded instruction … against … nospec-branch.h:132-162" includes line 161, ASM_CREDIT_CALL_DEPTH. Under the pinned CALL_DEPTH_TRACKING=y (annotations; nospec-branch.h:76,83-84) that is movq $-1, %gs:pcpu_hot+X86_call_depth. clear_bhb_loop's .align 64, 0xcc runs (entry_64.S:1540,1549) depend on where the function is placed. The Exit should name both as outside the comparison and keep the int3 after each fill call inside it.
  • M:113 — nothing reads grep x86_Thread_features /proc/self/status any more. Its conditional was deleted this round. Delete the capture too.
  • issues/kernel/user-programs-run-without-a-shadow-stack.md:16 — the Exit prescribes the ABI channel ("ask for a shadow stack at spawn"). Linux's opt-in is the process's own arch_prctl. An Exit should state the behaviour and leave the syscall shape to the discussion that CLAUDE.md requires.
  • M:59 — the copy_out reason still names only UserSafe. Three struct publishes do not go through it: ops.rs:444 and device.rs:279,307. Their layout assertions carry them. This NOTE is unchanged from round 4.
  • M:1-349 — the track is still past a screen, against issues/README.md's track rule. This is for the orchestrator to rule on, as in round 4.
  • The three new not-applicable rows (M:65-67) are true. kernel/src has no bpf, kexec or module loader.

REMOVE

  • PR body — the "## Round 4" section and "Carried from round 2". Both are review history in what becomes main's merge record.

SEND BACK

…ion, S9's canary is per thread, and a narrower S3 comparison

- issues/boot-media/the-loader-never-sets-the-firmwares-memory-overwrite-request.md: the Exit now matches libstub/tpm.c:36-40 — the loader writes MemoryOverwriteRequestControl only where firmware already defines it and ignores the write's result — with a guest test on OVMF (which defines none) asserting no variable is created.
- the-kernel-mitigates-what-linux-mitigates-on-the-t14.md: S9's Exit now requires a per-thread canary (a per-CPU %gs slot the scheduler overwrites on every switch, matching switch_to's fixed_percpu_data.stack_canary), with a cross-thread distinctness assertion that a shared global canary must red. S3's instruction comparison excludes clear_bhb_loop's placement-dependent .align padding and CALL_DEPTH_TRACKING's ASM_CREDIT_CALL_DEPTH, each with its reason. The dead x86_Thread_features capture is deleted.
- user-programs-run-without-a-shadow-stack.md: the Exit states the behaviour (a process can have a shadow stack of its own) instead of prescribing the spawn-time syscall shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 6, at 18af9f8

CI: host success on 18af9f8. The body records cargo run -- --ci host EXIT=0. Growth: +583 −0, all tracker prose; production 0, tests 0. M = issues/kernel/the-kernel-mitigates-what-linux-mitigates-on-the-t14.md, MOR = issues/boot-media/the-loader-never-sets-the-firmwares-memory-overwrite-request.md. Linux was read from a fresh fetch of refs/tags/Ubuntu-6.8.0-142.142 (53e5d07a). LLVM was read from rust/src/llvm-project (ToyOSOrg/llvm-project, 52ed14fc).

Round-5 blockers

  • MOR condition: CLOSED. At 53e5d07a, tpm.c:36-40 returns on EFI_NOT_FOUND, and :42-45 discards set_efi_var's status. MOR:11-17 and :26-29 state exactly that.
  • S9 per-thread canary: OPEN. See the first two BLOCKERs.

BLOCKER

  • M:310-316 — the distinctness assertion reads the threads' stored canaries, not the value a protected frame compares.
    • This patch still passes: keep the per-thread draw, and delete the scheduler's write of the incoming thread's canary to %gs:N on switch.
    • Under it, both zeroed overflows are still caught, because zero differs from whatever %gs:N holds. The two stored canaries still differ, so the test stays green while every thread on a CPU shares one guard.
    • Fix: each thread reports the guard its own protected frame loaded while it ran, read before and after a yield to the other thread. That removed switch-time write must red the test.
  • M:297-303 — S9's mechanism does not work on the kernel's target in this tree's LLVM.
    • X86ISelLoweringCall.cpp:564 honours stack-protector-guard-reg/-offset only when hasStackGuardSlotTLS (:548-551: glibc, musl, Fuchsia, Android) holds.
    • The kernel builds for x86_64-unknown-none (kernel/.cargo/config.toml:2). There it falls through to TargetLowering::getIRStackGuard/insertSSPDeclarations (:604, :638-640), which is one global __stack_chk_guard.
    • So "this rust fork's codegen backend gains them" yields exactly the global the Exit claims to exclude. The stage owes a change to the forked LLVM's X86 lowering, and it must name that change.
    • The claim was one file read away from being checked.
  • MOR:29-31 — the one test the Exit names is green on origin/main today.
    • bootloader/src writes no variable under that GUID, so the test cannot fail on the claim "sets it to 1 where defined".
    • Its premise, "OVMF, which defines no MOR variable", is not measured anywhere.
    • The harness already writes and reads OVMF's VARS (tests/common/update.rs:631, vars::plant; :711, vars::live). Fix: plant MemoryOverwriteRequestControl=0 under e20939be-…, boot, and assert it reads 1. Deleting the loader's write must red that test. Keep the no-MOR arm.

NOTE

  • M:310-313 — the test cannot run in one boot as written. "overflows … on each of two threads and asserts each panic …, then asserts" puts a stack-protector panic, which ends the kernel, before further assertions. Name the boots.
  • M:316-318 — "Each of S10's 64 boots prints the guard" no longer names a value, because there is a guard per thread. Name whose guard is printed.
  • S3 exclusions (M:190-194): CLOSED. nospec-branch.h:161 is ASM_CREDIT_CALL_DEPTH. The .align 64, 0xcc lines are there, and the fill int3 stays inside the comparison.
  • x86_Thread_features capture deleted: CLOSED. No reference is left in issues/.
  • Shadow-stack Exit states the property: CLOSED.
  • Carried from round 5 and unchanged: M:59 copy_out, and the length of the track file.

REMOVE

  • M:304-306 and the PR body's S9 decision — "arch/x86/kernel/process_64.c" is wrong. That file has no canary line at 53e5d07a. The copy is in __switch_to_asm, at entry_64.S:193-196.
  • M:300-302 — "the flags Clang's -mstack-protector-guard=sysreg sets for Linux's own build (arch/x86/Makefile)" is false. sysreg is arm64 (arch/arm64/Makefile:59), and x86's guard flags are in the 32-bit branch (arch/x86/Makefile:115-121).
  • PR body, "How it was read" and "Config" — "This round…", "agree with the review's reads" and "re-measured this round". These are review chronology in the merge record.
  • PR body, MOR decision — "OVMF, which defines no such variable". It is unmeasured.

SEND BACK

Japabu and others added 3 commits September 29, 2026 05:30
…st reads the guard a frame compares, and the MOR test plants the variable

S9. The kernel builds for x86_64-unknown-none. On that triple, X86's
getIRStackGuard takes the segment slot only for glibc, musl, Fuchsia and
Android (X86ISelLoweringCall.cpp:548-551,564 at a79bc52c). Every other
triple falls through to the global __stack_chk_guard, and the guard-reg and
guard-offset flags do nothing there. The CI LLVM's llc confirms it: with
tls/gs/40 module flags under the kernel code model, x86_64-unknown-none-elf
emits __stack_chk_guard(%rip) and x86_64-unknown-linux-gnu emits %gs:40.

No ToyOS-owned mechanism reaches the slot. S9 is therefore four steps:
- an LLVM change that honours an explicit tls guard on any x86 triple, as
  RISC-V does;
- rustc options that set the module flags;
- the kernel's slot, which is written only at entry and on context switch;
- the guest test.

The guest test reads %gs:N from inside a protected frame on two threads.
Overwriting a frame's guard with its own thread's value returns. Overwriting
it with the other thread's value panics. Deleting the switch-time write reds
both.

MOR. The guest test plants MemoryOverwriteRequestControl=0 under
e20939be-32d4-41be-a150-897f85d49829 and asserts it reads 1 after boot.
Deleting the loader's write reds that. The arm that plants nothing stays.

Removed: the process_64.c citation, the sysreg/arch/x86/Makefile claim, and
the unmeasured claim that OVMF defines no MOR variable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Root CLAUDE.md's dependency rule governs a fork's changes generally, not
just target arms: a general cross-platform option written to upstream
quality is admitted when ToyOS needs it and upstream lacks it. S9's LLVM
guard-slot fix and its rustc stack-protector-guard options are such changes,
so they proceed as named steps rather than waiting on a ruling; the stage
that lands them amends src/forkcheck.rs's module header to admit exactly
this class of change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 8, at d73f123

CI: host COMPLETED/SUCCESS on d73f123. The body records cargo run -- --ci host EXIT=0. Growth: +652 −0 against origin/main, all under issues/; production 0, tests 0. The branch contains origin/main. LLVM was read with git show a79bc52c:<path>; rust/ at d73f123 is 1b236638, which pins src/llvm-project at a79bc52c1d5e. M = issues/kernel/the-kernel-mitigates-what-linux-mitigates-on-the-t14.md, MOR = issues/boot-media/the-loader-never-sets-the-firmwares-memory-overwrite-request.md.

Round-6 blockers

  • S9 test reads the stored canaries, not the compared guard: CLOSED.
    • M:363-378 is one smp: 1 boot. Threads A and B each read %gs:N twice from inside a protected SYS_DEBUG frame, with a hand-off between the reads.
    • Overwriting the frame's guard with A's read returns; overwriting it with B's read panics. That proves which value the frame compares.
    • Deleting context_switch's write leaves one value in the slot, so A == B, and the B-value overflow returns: both assertions go red.
    • A per-process or boot-wide draw also gives A == B, because A and B are threads of one program. A write of the outgoing thread's guard makes A's two reads differ.
  • S9 mechanism: CLOSED. Every citation holds at a79bc52c.
    • X86ISelLoweringCall.cpp:548-551: hasStackGuardSlotTLS is glibc, musl, Fuchsia or Android. :564 gates the segment slot on it, :604 falls through, and :637-640 declares the global.
    • TargetLoweringBase.cpp:2388-2411 is the __stack_chk_guard global.
    • RISCVISelLowering.cpp:25703-25707 is the "tls" branch that step 1 mirrors.
    • X86ISelLowering.cpp:2770-2772 makes LOAD_STACK_GUARD Mach-O-64 only.
    • Clang.cpp:3479-3491,3561-3570 accepts tls/gs on every x86 triple. CodeGenModule.cpp:1543-1553 sets the module flags.
    • StackProtector.cpp:530-533 loads the IR guard under tls once getIRStackGuard returns the slot, so step 1 is sufficient in the IR pass.
    • The rustc fork at 1b236638 has no stack-protector-guard option (rustc_session/src/options.rs:2843 is stack_protector alone).
    • The PR body records the llc measurement: x86_64-unknown-none-elf gives __stack_chk_guard(%rip) and x86_64-unknown-linux-gnu gives %gs:40, both EXIT 0.
    • Each step's test reds on the unchanged lowering or on the flags dropped from context.rs. The fork-rule amendment is scoped to S9's own PR, per the ruling.
  • MOR test green on main: CLOSED. MOR:30-38 plants MemoryOverwriteRequestControl=0 under e20939be-… with vars::plant, and asserts vars::live reads 1. Deleting the loader's SetVariable leaves 0, which reds it. The no-MOR arm is kept, and an unconditional write reds it.

BLOCKER

None.

NOTE

  • MOR:35-36 — the second arm's premise is unmeasured.
    • The premise: the host's firmware (src/firmware.rs, whatever the host's QEMU declares, Debian's edk2 on CI) creates no variable under that GUID when nothing is planted.
    • A TPM2-enabled edk2 build carries TcgMor, which may create MOR=0 at boot. That would red the arm on a correct loader.
    • The failure is loud, not silent. Add it to "Unsure" beside the first arm's assumption, or measure it on the first run.
  • M:363-373 — the overwrite action has to locate the guard slot. LLVM's protected-object layout does not promise one-past-the-array. A miss makes the B-value arm return and fail loudly, so this is only an implementation hazard.
  • Carried and unchanged: M:59 copy_out, and the track's length.

REMOVE

  • PR body, LLVM bullet — "Every cited file is identical at 52ed14fc, its ancestor." It is review chronology, and the pinned commit is what the record needs.

LAND

@Japabu
Japabu added this pull request to the merge queue Sep 29, 2026
Merged via the queue into main with commit 0368861 Sep 29, 2026
1 check passed
@Japabu
Japabu deleted the wt/toyos-secparity branch September 29, 2026 04:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant