Skip to content

The kernel declares the CPU's HWP request, and a perf-state claim reads the power envelope back per CPU - #590

Open
Japabu wants to merge 8 commits into
mainfrom
wt/toyos-perfstate
Open

Japabu wants to merge 8 commits into
mainfrom
wt/toyos-perfstate

Conversation

@Japabu

@Japabu Japabu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Stage 1 of the new track issues/kernel/the-kernel-owns-cpu-performance-state.md. The kernel now declares the CPU's HWP request from the one CPU-state declaration, and a capability-gated perf-state claim reads the power envelope's registers back per CPU. There is no new syscall.

The only open landing condition is the T14 metal row: perf_request on the T14, plus mutation-hwp-check-no-asserts.patch on the perfdiverge boot. Both are owed by the orchestrator (see "The metal row, owed" below).

What changed, per decision

The declaration lives in kernel/src/arch/x86_64/control_regs.rs. CPUID alone decides whether a CPU gets a request (hwp_declared). Then hwp_init runs on the BSP and on every AP, after self_check:

  1. IA32_HWP_INTERRUPT (0x773) = 0, only where CPUID.06H:EAX[8] enumerates it.
  2. IA32_PM_ENABLE (0x770) = 1.
  3. Only now are IA32_HWP_CAPABILITIES (0x771) and MSR_PLATFORM_INFO (0xCE) read, and the request computed from them.
  4. IA32_HWP_REQUEST (0x774), IA32_HWP_REQUEST_PKG (0x772) = 0x8000ff01 and IA32_ENERGY_PERF_BIAS (0x1B0) = 6.

Each register is written whole, then read back and asserted on each CPU. Each CPU logs one line: control_regs: cpuN pm_enable= hwp_request= hwp_request_pkg= epb= hwp_interrupt=<Some(v)|None> hwp_capabilities= platform_info=. No read-modify-write decides any of them.

The order is the SDM's. Intel SDM Vol. 3B, §17.4.2 "Enabling HWP", p. 17-7, Order Number 253669-093US (September 2026; the PDF I read has SHA-256 a8335f8ccb8b2e82ac51a8ff31360699fc33f334e66133e822038702f52cd276):

Additional MSRs associated with HWP may only be accessed after HWP is enabled, with the exception of IA32_HWP_INTERRUPT and MSR_PPERF. Accessing the IA32_HWP_INTERRUPT MSR requires only HWP is present as enumerated by CPUID but does not require enabling HWP.

intel_pstate does the same thing independently: intel_pstate_hwp_enable clears the interrupt and enables HWP, and only then does intel_pstate_get_hwp_cap read the capabilities.

The values.

  • The request's own fields: desired 0, EPP 128, activity window 0, package control off.
  • min = MSR_PLATFORM_INFO[47:40], the maximum-efficiency ratio. max = IA32_HWP_CAPABILITIES[7:0], the highest level, turbo included. This is how intel_pstate derives min and max.
  • The kernel derives the two ratios per CPU rather than writing a literal, because a literal would cap or misstate another CPU's range.

All or nothing, refused by name. toyos_perfstate::refusal requires every one of these:

  • HWP, EPP, the package request, EPB and package thermal status, each by its CPUID bit.
  • An Intel vendor. The request's minimum comes from MSR_PLATFORM_INFO, which is Intel's.
  • DisplayFamily 06H. No CPUID bit enumerates MSR_PLATFORM_INFO, and SDM Vol. 4 documents it only in family 06H's model tables.
  • Not hybrid.

A CPU that fails any of them gets no register written, and the BSP logs the reason once. Every AP must reach the BSP's verdict, or it is named. The turbo bit (IA32_MISC_ENABLE bit 38) is read back and not written. It is stage 3.

The proof is minted only after every CPU applied the declaration. control_regs::HwpDeclared::ask asserts that control_regs::report has run. report runs after SMP bring-up, and every committed AP runs control_regs::init before it echoes. So "HWP is enabled on every CPU" is now enforced, not just true because of boot order.

The arithmetic lives in toyos-perfstate, a pure no_std host crate shared by the kernel, the guest tests and the program.

The read-back is a device class. The ABI is DeviceType::PerfState = 9 => "perf-state", plus toyos_abi::perf's two all-u64 records:

  • PackageRegisters: 0x772, 0xCE, 0x1B1.
  • CpuRegisters: 0x770, 0x771, 0x774, 0x1B0, 0x1A0.

Every one of those registers is enumerated by CPUID, is architectural (0x1A0), or is read at boot under HwpDeclared (0xCE). A read can therefore reach no register that boot has not already touched. The /system/bin/perfstate row holds devices = ["perf-state"]. On a machine with no declaration, the claim is NotFound.

A read owns its ask. A read issues a generation on a second shootdown::Shootdown, answers for its own CPU inline, and kicks the others. Each other CPU answers from drain_irqs. The ask is a perf_state::Ask held on sys_read's stack: the read's first look makes it, and it ends when the read ends. The claim holds no ask. So a read that ends answered, refused Io, cancelled, or as a nonblocking WouldBlock leaves nothing behind that a later read could be answered from or refused by. This is the contract in toyos-abi/src/syscall.rs's PerfState doc: "each CPU's taken on that CPU after the read asked".

  • Reader::cancel and DeviceClaim::cancel_perf_state are gone, because nothing is left to take back.
  • ABI doc change: "reads of one claim that overlap share one ask" is deleted. Every read now makes its own ask. One answer from a CPU still covers every ask issued before it (fetch_max).
  • Poll: a read asks when it runs, so nothing is ready before one. has_data is false and read_watch is None for the class, so a poll on a claim is refused NotSupported. It no longer reports a readiness that the next nonblocking read would not honour. A nonblocking read asks and answers only if every CPU has already answered, which on SMP means WouldBlock.
  • The bound: past ANSWER (250 ms) the read is refused Io, and each silent CPU is logged as perf_state: cpuN did not answer Generation(G) within 250ms.

Memory ordering, target to initiator is proven by the existing kernel-loom/tests/tlb_shootdown.rs, not a new model: the target's serve closure writes its slot, and the initiator reads it once served answers — the same edge tlb_shootdown already checks for its own tlb store. mutation-served-relaxed.patch (below) turns it red.

The witness. A claim holds an Option<HwpDeclared>, and every read function takes one. On AArch64, arch::perf_state::Declared is uninhabited, so the claim is refused by name there.

Two test actuators, compiled out of a shipping kernel.

  • perf-state-deaf-cpu grants the claim where nothing is declared, answering zeros. No CPU answers a kick, so every CPU except a read's asker stays silent.
  • perf-request-diverges makes cpu1 move its request one ratio off the declaration when it answers a read, then run hwp_check exactly as boot does. It is ruled flashable in src/metal.rs's FLASHABLE.

Tests

  • perf_request (Fast, QEMU): the kernel refuses by name, no CPU logs hwp_request=, and test_rs_perf_state's claim is refused NotFound.
  • perf_state_silent_cpu (Fast, QEMU): boots 2 CPUs with perf-state-deaf-cpu. test_rs_perf_state_silent first makes a nonblocking read, which is the boot's first ask and must return WouldBlock. It then makes two blocking reads, each refused Io. The judge requires exactly two refusal lines, naming Generation(2) and Generation(3) once each, and each naming one CPU. The asker answers inline and no CPU answers a kick, so which CPU names itself does not depend on any vCPU being scheduled in time. The only time bounds left are the harness's liveness bounds. A read that is answered or refused from an ask not its own names Generation(1) and reds.
  • perf_request metal row: two boots, owed (below).

Gates, at 148c2078 (after merging origin/main 89dd9d2b)

gate exit
cargo run -- --ci host 0 (Host: 54 step(s), all green)
cargo run -- --build-only 0 (Build finished.)
cargo test --test toyos-build -- --list (builds the guest binaries, boots nothing) 0; lists Fast perf_request and Fast perf_state_silent_cpu
kernel cargo check, with and without --features boot-actuators 0 and 0, no warnings

I ran no QEMU guest and no T14 run. The orchestrator's measured QEMU runs at this head:

run exit
perf_request 0
perf_state_silent_cpu 0
perf_state_silent_cpu under mutation-claim-held-ask.patch 1
perf_state_silent_cpu under mutation-read-unbounded.patch 1
Fast tier 0

High-risk: CPU state, a device claim, a cross-CPU protocol

Negative controls. Each patch is applied with git apply --check and then git apply, shown to build, run, reversed, and the file compared byte for byte afterwards.

  • mutation-served-relaxed.patch (Shootdown::served's Acquire → Relaxed). Builds (--no-run EXIT=0). --test tlb_shootdown: EXIT=101, with an_acknowledged_flush_postdates_the_page_table_write and one_serve_answers_two_concurrent_shootdowns FAILED — the loom oracle for the perf-state read's target-to-initiator ordering, since the flush's own postdating check reads through the same served edge a perf-state answer's registers are read through.
  • mutation-claim-held-ask.patch reverts the ask's ownership to the round-2 mechanism: an ask held by the claim, cleared only when a read is answered or refused. That is ee20f53d with claim.cancel_perf_state() deleted, and it also carries ee20f53d's nonblocking leak. The deleted line has no site at this head: the ask lives on the read's stack, so this patch is that deletion's equivalent. The mutated kernel builds (cargo check, with and without actuators: EXIT=0, 0). Under perf_state_silent_cpu: EXIT=1. The blocking read joins the nonblocking read's ask, so the refusals name Generation(1) and Generation(2).
  • mutation-read-unbounded.patch removes the deadline arm (if !answered && !ask.deadline.reached(now) → if !answered). It is regenerated for this head and builds (EXIT=0). Under perf_state_silent_cpu: EXIT=1, at the harness's 300 s ceiling.
  • mutation-hwp-check-no-asserts.patch replaces hwp_check's assert loop with let _ = want;. It is regenerated for this head and builds (EXIT=0). It is metal only: no QEMU CPU reaches hwp_check, because TCG and KVM both reduce leaf 6 to ARAT. Expected: perf_request's perfdiverge boot reds. This is part of the owed metal row. It is not a blocker I can close.
  • Host, from round 1 and still in the tree: refusal without the family check, hwp_request's max taken from the guaranteed level, and hwp_notification reading bit 9. Each makes cargo test -p toyos-perfstate exit 101.

The 250 ms bound is proven twice: under QEMU, mutation-read-unbounded.patch exits 1 at this head (above); on metal, the T14 row. No QEMU verdict here waits on guest time.

Independent oracles.

  • The Intel SDM sentence quoted above, for the order of enable and read.
  • intel_pstate's source, for the same order and for how min and max are derived.
  • The T14's Linux readings under intel_pstate: 0x774 = 0x80002a04 on every CPU, 0x771 = 0x010d182a or 0x010e182a, 0x772 = 0x8000ff01, 0x1B0 = 6, 0x770 = 1. These live in unmerged Define the T14 LLVM bar as a libc++ stage-3 recipe with an in-tree judge; the bar value is owed #568's samples. Neither this PR nor main carries a command or log for them. The host tests hold toyos_perfstate to those values, and the metal row re-measures them on ToyOS.

What I am unsure of

  • MSR_PLATFORM_INFO was not read on the T14. Its ratio 4 is inferred from Linux's cpuinfo_min_freq of 400000 kHz. The per-CPU line prints the register, so the first T14 boot settles it. If it is not 4, the metal row reds on hwp_request.
  • Whether the T14 enumerates HWP notification (CPUID.06H:EAX[8]) is unmeasured. The line says which: hwp_interrupt=Some(0) or hwp_interrupt=None.
  • The perfdiverge boot is the first metal boot to end in a panic raised from a job. Whether the panic record on the page after the reset carries the assert's text is unmeasured.
  • No test launches /system/bin/perfstate. Deleting its row's devices stays green. A launch test needs a launcher-driven boot. This is recorded in the track's stage 1 as present-state debt.
  • Multi-package machines. The package registers are those of the package holding the reading CPU. The T14 has one package.
  • The generation check relies on the test being the boot's only asker. That holds in its boot: the claim is exclusive, and nothing else claims it there.

The metal row, owed

perf_request on the T14. Neither assertion has run yet:

  1. testcases boot. Every CPU cpuN that the SMP records bring up logs control_regs: cpuN pm_enable=1 hwp_request=0x80002a04 hwp_request_pkg=0x8000ff01 epb=6 , and test_rs_perf_state exits 0. That binary reads every CPU back holding its declaration, twice.
  2. perfdiverge boot (perf-request-diverges, job test_rs_perf_state). The loader pass after the reset carries, after Previous boot's panic:, control_regs: cpu1 holds hwp_request=0x80002a05, the declaration is 0x80002a04. This is also the only measurement of hwp_check's assert loop.

🤖 Generated with Claude Code

…s it back per CPU

The self-hosting bar is measured under a power envelope that must be read back
for a whole build span, and ToyOS left every register of it where firmware put
it. This is the first stage of the track that makes the kernel own it
(issues/kernel/the-kernel-owns-cpu-performance-state.md).

The declaration. control_regs.rs, the one CPU-state declaration, now also
writes IA32_PM_ENABLE, IA32_HWP_REQUEST, IA32_HWP_REQUEST_PKG and
IA32_ENERGY_PERF_BIAS whole on the BSP and every AP, and asserts each on
each. The values are the bar's: min is the package's maximum-efficiency
ratio (MSR_PLATFORM_INFO[47:40]), max the CPU's HWP highest performance,
desired 0, EPP 128, window 0, no package control, which on the T14's inputs
is 0x80002a04, what Linux's intel_pstate held there; the package request is
0x8000ff01 and EPB 6. The declaration is all or nothing. A CPU missing any
register it names (HWP, EPP, package request, EPB, package thermal status,
or not Intel, or hybrid) gets no request, and the BSP says why once:
"control_regs: no performance request is declared: no HWP ...". Whether the
machine declared one is the BSP's verdict and every AP must reach it; the
request itself is per CPU, since a CPU's highest performance is its own.
IA32_MISC_ENABLE's turbo bit is read back and not written, because its other
bits are model-specific and firmware's.

The arithmetic is toyos-perfstate, a pure host-tested crate, because no QEMU
CPU has HWP. Its oracle is the T14's Linux MSR readings.

The read-back needs no new syscall. It is a device class, perf-state (class 9,
the manifest's `devices = ["perf-state"]`), whose read answers
toyos_abi::perf's records: the package's registers (HWP package request,
PLATFORM_INFO, RAPL unit, PKG_POWER_LIMIT, PKG_ENERGY_STATUS, package thermal
status, TEMPERATURE_TARGET), then each CPU's (PM_ENABLE, HWP capabilities,
HWP request, EPB, MISC_ENABLE). A per-CPU MSR is readable only on its CPU,
so a read asks every CPU. It issues a generation on a second
shootdown::Shootdown (the loom-modelled TLB ack protocol), answers for its own
CPU, and kicks the rest. Each answers from its next scheduler pass in
drain_irqs, which costs two relaxed loads when nothing is owed, and posts
perf_state::WATCH. The read parks there, bounded by 250 ms, past which it is
refused Io and the silent CPUs are named. A claim exists only with the
control_regs::HwpDeclared proof, so no rdmsr of these registers is reachable
on a CPU without them. On AArch64 that proof is an uninhabited type and the
claim is refused by name.

/system/bin/perfstate holds the row and prints one read, checked against the
declaration. The guest binary perf_state runs the same code, and perf_request
drives it: in QEMU it asserts the named refusal, no request line, and the
claim refused NotFound; on the T14 it asserts every CPU logged the bar's
literal values and the binary read them all back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu
Japabu marked this pull request as ready for review September 28, 2026 19:51
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 1, at 545e9cd

CI host success at 545e9cd; git merge-tree against origin/main (e3a1cdc) is clean. Orchestrator runs at this head: perf_request EXIT=0, Fast EXIT=0 (233 passed), control guest-claim-granted.patch EXIT=1. No T14 reading exists. Net +1141/−17: tests about 228 lines, the issue 51, lockfiles 22, production about 840.

BLOCKER

  • (gate) The metal row is missing. Under reviewer.md, a change that targets hardware and has no reading from that hardware is NOT READY FOR REVIEW, so the proposal to land it with the row owed is refused. QEMU runs none of the declaration's writes and none of hwp_check's asserts. It runs none of Reader's cross-CPU path either, except under the mutated kernel. All of that runs at every T14 boot, and one wrong assumption there (the next finding) panics the boot of every metal row for every branch that lands after this one.
  • kernel/src/arch/x86_64/control_regs.rs:171 — The BSP reads IA32_HWP_CAPABILITIES before IA32_PM_ENABLE is written. Linux, the branch's own oracle, never does this. turbostat's print_hwp returns before reading HWP_CAPABILITIES when PM_ENABLE bit 0 is clear ("HWP is enabled and MSRs visible"), and intel_pstate enables HWP before intel_pstate_get_hwp_cap. On firmware that leaves PM_ENABLE=0, that is a #GP at boot. Fix: let CPUID decide, write PM_ENABLE, then read the capabilities and compute the request. Or cite the SDM sentence that allows the read before enable.
  • kernel/src/arch/x86_64/perf_state.rs:32-36 — 0x606, 0x610, 0x611 and 0x1A2 are not enumerated by any CPUID bit and are never read at boot. hwp_check reads only 0x770, 0x771, 0x772, 0x774, 0x1B0 and 0xCE. So the first read of those four comes from a userland SYS_READ on the claim. The kernel has no rdmsr fault fixup (git grep finds none), so on a CPU with HWP and without one of those registers, userland triggers a kernel #GP panic. Fix: drop the four fields until stage 2 or 4 lands a consumer; nothing here reads them except a println. Dropping them also takes MSR_PKG_ENERGY_STATUS off a row any session can launch; Linux restricted that register to root for CVE-2020-8694. The alternative is to read them in hwp_check so that HwpDeclared covers them.
  • kernel/src/perf_state.rs:90 — The 250 ms refusal has no test anywhere. With if !answered && !ask.deadline.reached(now) patched to if !answered, QEMU stays green (claim NotFound) and so does metal (every CPU answers). An arm is needed that silences one CPU (the dump's deaf-window actuator is how the tree does that) and asserts Io plus perf_state: cpuN did not answer.
  • kernel/src/arch/x86_64/control_regs.rs:398 — Deleting hwp_check's assert loop passes every test. The QEMU arm never reaches it, and the metal judge reads the line logged before the asserts run. control_regs_negative's actuator cannot stand in for a test, because self_check fires first. "Asserted on each CPU" needs a test that turns red on that deletion.

NOTE

  • The negative control's red is valid for its claim. The harness requires perf-state: refused NotFound, so any granted claim reds whatever rdmsr returns. The log shows the program's package check firing first, on TCG's zeros. TCG answering 0 for an unimplemented MSR means QEMU cannot model the #GP, which is why the third blocker is visible only by reading the code. The control did measure one thing: the cross-CPU read answering on two QEMU CPUs.
  • kernel/src/perf_state.rs:88 — A cancelled read (the PerfState arm's return cancelled() in syscall/io.rs) leaves pending set. The next read then either returns that old ask's slots, sampled before it asked (which contradicts toyos-abi/src/syscall.rs:1268), or a spurious Io once the old deadline has passed.
  • toyos-perfstate/src/lib.rs:201 — HwpCapabilities has one production reader, of one field (.highest), and capabilities as u8 replaces it. Its other three fields are tested for nothing the code uses. Delete it.
  • kernel/src/arch/x86_64/control_regs.rs:171 — MSR_PLATFORM_INFO (0xCE) is enumerated by no CPUID bit, so a CPU without it takes an unnamed #GP at boot. That is loud, but it is not a refusal by name.
  • control_regs.rs — IA32_HWP_INTERRUPT (0x773, CPUID.06H:EAX[8]) becomes live once PM_ENABLE is set, and nothing declares it. Declare it 0, as intel_pstate does, or record it with the turbo bit.
  • Rulings asked for. The 250 ms bound is allowed: it waits on the event (WATCH and answered()), under the tree's Budget, and names the CPUs that stayed silent. The CPU-state rule holds for the four registers: written whole, with no read-modify-write, by the BSP and every AP, and asserted on each. Leaving the turbo bit to stage 3 is acceptable, because the track records it with an exit. The device class, as class 9 plus two records with no syscall, is minimal except for the fields in the third blocker.

REMOVE

  • PR body, "Expected result: … rdmsr of 0x770 is #GP under TCG" — measured false.
  • PR body, "A virtual CPU that advertises HWP without them would #GP at boot. That is loud, not silent." — false: four of those registers are read only from the claim.
  • PR body, "The first --clippy run was red on AArch64…" — history.
  • PR body and kernel/src/perf_state.rs:135, "two relaxed loads" — a cost claim with no measurement behind it.
  • kernel/src/perf_state.rs:33, "the blocked-task dump's for the same question" — a cross-reference that rots.
  • kernel/src/shootdown.rs:1-2, "— a TLB shootdown, or a performance-state read —" — a list of callers that will grow.
  • toyos-perfstate/Cargo.toml:1-4 — narration; description and src/hostws.rs already say it.
  • toyos-perfstate/src/lib.rs:25-26, "every Intel core since Sandy Bridge has them…" — unmeasured, and the third blocker makes it unnecessary.
  • issues/kernel/the-kernel-owns-cpu-performance-state.md:11, "Until this track, ToyOS left…" — chronology, and false once this lands.
  • issues/kernel/the-kernel-owns-cpu-performance-state.md:14, "The values the kernel programs are the bar's, not its own." — false: the PR body says deriving min and max is the kernel's choice.

SEND BACK

Japabu and others added 3 commits September 28, 2026 23:33
…no claim reaches an unproven register

The review sent #590 back with five blockers. Each is answered here, except
the reading from the T14, which only that machine can give.

- control_regs: CPUID alone now decides whether a CPU gets a request.
  IA32_HWP_INTERRUPT is set to 0 where CPUID.06H:EAX[8] enumerates it, and
  IA32_PM_ENABLE is written next. Only after that are IA32_HWP_CAPABILITIES
  and MSR_PLATFORM_INFO read, and the request computed and written. This is
  intel_pstate's order (intel_pstate_hwp_enable, then
  intel_pstate_get_hwp_cap). The HWP writes moved out of the
  no-ap-control-regs skip, into hwp_init after self_check.
- toyos-perfstate: no CPUID bit enumerates MSR_PLATFORM_INFO (0xCE). SDM
  Vol. 4 documents it in the model tables of DisplayFamily 06H, so a CPU of
  any other family is refused by name (Refusal::NotFamily6) before the
  register is read.
- The read-back drops MSR_RAPL_POWER_UNIT, MSR_PKG_POWER_LIMIT,
  MSR_PKG_ENERGY_STATUS and MSR_TEMPERATURE_TARGET. No CPUID bit enumerates
  them, and nothing at boot read them, so a userland read was the first to
  touch them. The track's stages 2 and 4 now require each one to be proven
  present at boot first, and keep the energy counter (CVE-2020-8694) off any
  row a session can launch.
- A cancelled perf-state read takes its ask back (Reader::cancel, reached from
  sys_read's cancelled wait), so no later read is answered from registers
  sampled before it began. The ABI doc now says that overlapping reads of one
  claim share one ask.
- HwpCapabilities is deleted. The request's maximum is `capabilities as u8`.
- perf-state-deaf-cpu grants the claim on a machine with no declaration and
  answers zeros. The last CPU answers no ask. perf_state_silent_cpu (Fast)
  asserts two reads refused Io, each naming that CPU alone.
- perf-request-diverges has cpu1 move its request one ratio off the
  declaration when it answers a read, then run hwp_check. perf_request's
  metal row gains a second boot, perfdiverge, which must carry that panic on
  the page after the reset. It is priced in the metal profile and ruled
  flashable.
- TESTCASES now runs test_rs_perf_state, before null_sink_client_exits. The
  metal row's job_passed("test_rs_perf_state") named a job that no boot ran.
- Deleted per the review's REMOVEs: the narration in toyos-perfstate's
  Cargo.toml; "since Sandy Bridge"; "two relaxed loads"; the cross-reference
  to the dump's bound; shootdown.rs's list of callers; the track's
  chronology and the claim that its values are the bar's. The stale "four"
  in io.rs's device-class count goes too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A second Io read cannot tell a cleared ask from a stale one, since a stale
ask past its deadline is refused Io too. So the second read shows that the
claim still answers after a refusal, and nothing more. The comments now say
only that.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 2, at ee20f53

CI host success at ee20f53. Orchestrator runs at this head: perf_request EXIT=0, perf_state_silent_cpu EXIT=0, Fast EXIT=0, and perf_state_silent_cpu under mutation-read-unbounded.patch EXIT=1. There is no T14 reading. git merge-tree HEAD origin/main (8a8fe27) exits 1, with conflicts in src/metal.rs and tests/toyos.rs. Net +1453/−21: tests about 420, the issue 68, lockfiles 22, production about 930.

Round-1 blockers

  • Metal row missing — OPEN. No T14 reading exists. The orchestrator holds this as a landing condition.
  • IA32_HWP_CAPABILITIES read before IA32_PM_ENABLE — CLOSED in code. hwp_declared (control_regs.rs:152) decides from CPUID alone. hwp_init (:216) then writes HWP_INTERRUPT, then PM_ENABLE, and only then reads 0x771 and 0xCE. No run exercises this path; its first measurement is the metal row's first boot.
  • Registers first touched by a userland read — CLOSED. The records hold only 0x770, 0x771, 0x774, 0x1B0 and 0x1A0 per CPU, and 0x772, 0xCE and 0x1B1 per package. Each of these is enumerated by CPUID (06H:EAX[6], EAX[7], EAX[11], ECX[3]), is architectural (0x1A0), or is read at boot behind the family gate (0xCE).
  • The 250 ms refusal had no test — CLOSED. perf_state_silent_cpu passes (EXIT=0). Under mutation-read-unbounded.patch it goes red (EXIT=1) with timed out after 300s … did not finish and no did not answer line. The patch applies cleanly at ee20f53. It removes exactly the deadline arm and nothing else, so it reverts what it claims. It is a mutation of the bound. It is not a revert of the whole change.
  • Deleting hwp_check's assert loop passed every test — OPEN. The perfdiverge boot is sound in design: its needle , the declaration is cannot match userland's , and the declaration is. But neither its green arm nor mutation-hwp-check-no-asserts.patch has run. Both are metal only.

BLOCKER

  • kernel/src/perf_state.rs:100-106 — A SYS_READ_NONBLOCK that returns WouldBlock (io.rs:280) leaves its ask in pending. Any later read of the claim, however much later, is then answered from that old ask's slots. Where a CPU never answered, it is instead refused Io without asking again. This contradicts toyos-abi/src/syscall.rs:1268 ("taken on that CPU after the read asked") and the PR's "never answered with registers sampled before it began". The cancel half has no test either: deleting claim.cancel_perf_state(); (io.rs:234) stays green. Fix the behaviour, and add a test that turns red under that deletion and under the nonblock-then-read sequence.
  • tests/toyos.rs:19806 — "each naming cpu1 alone" depends on the kernel's clock. If the test thread lands on cpu1, the deaf CPU, then cpu0 must answer a kick within the kernel's 250 ms. Otherwise cpu0 is named too. A starved TCG vCPU misses that window. This is the timing verdict that main's No QEMU test measures time, and audio is judged on metal only #562 moved out of QEMU; METAL_ONLY gives dump_nmi_probe the same reason. Make the verdict independent of time. For example, deafen every CPU but the one asking, and have the judge count exactly one named CPU per read.
  • kernel/src/shootdown.rs:73 — perf_state reads SLOTS through served's Acquire, in the target→initiator direction. kernel-loom/tests/tlb_shootdown.rs never models that direction. Patching self.flushed[cpu].load(Ordering::Acquire) to Ordering::Relaxed stays green everywhere: x86 is TSO, and loom models only initiator→target. Add a kernel-loom test in which serve's closure writes a loom cell and the initiator reads it after served(), red under that patch.

NOTE

  • Merge: conflicts with origin/main 8a8fe27 in src/metal.rs and tests/toyos.rs. The merged perf_state_silent_cpu must meet No QEMU test measures time, and audio is judged on metal only #562's rule (second blocker).
  • Oracles. intel_pstate is independent and correctly named: intel_pstate_hwp_enable, then intel_pstate_get_hwp_cap. The SDM is cited by volume only, with no section and no revision, and the sentence on which HWP MSRs are accessible before HWP_ENABLE is not quoted, so the ordering rests on intel_pstate alone. The T14 Linux readings (0x774 = 0x80002a04, 0x771, 0x772) have no command or log in this PR or in the tree. They live in unmerged Define the T14 LLVM bar as a libc++ stage-3 recipe with an in-tree judge; the bar value is owed #568's samples.
  • CPU-state rule. It holds for the five HWP registers: each is written whole, from inputs in read-only registers (0x771, 0xCE, as CR4's optional bits come from CPUID), by the BSP and every AP, and asserted on each. It is not yet true of the whole envelope: IA32_MISC_ENABLE bit 38 stays firmware's (stage 3). That is known, tracked and still true.
  • control_regs.rs:127 — HWP becomes DECLARED at the BSP's decision, before any AP has run hwp_init. HwpDeclared's doc still says HWP "is enabled there" on every CPU. It is true only because userland starts after SMP bring-up, and nothing enforces that.
  • system.toml:112 and userland/perfstate/src/main.rs — no test launches /system/bin/perfstate. Deleting devices = ["perf-state"] stays green on QEMU and would stay green on metal.
  • toyos-perfstate/src/lib.rs:88 — the NotIntel reason names RAPL, which this stage no longer reads.
  • control_regs.rs:440 — Enumerated is 12 lines for one format; inline it.

REMOVE

  • toyos-perfstate/src/lib.rs:8-10 — cites issues/build/toyos-builds-itself.md for the envelope, but neither this branch nor main has it there.
  • kernel/src/perf_state.rs:32-33 "A kicked CPU reaches a pass within one timer interrupt" — unmeasured.
  • kernel/src/shootdown.rs:73 "nothing reads through this edge yet" — false since this branch.
  • src/metal.rs:807-810 "the reset that ends the boot clears IA32_PM_ENABLE, which only a reset does" — unmeasured, and the flash ruling does not rest on it.
  • system.toml:109-110 — restates the row.
  • PR body "the harness's 30 s ceiling ends it" — measured at 300 s.
  • PR body "The 250 ms bound is the kernel's choice, and the reviewer allowed it." — history.

SEND BACK

Japabu and others added 2 commits September 29, 2026 03:27
Two conflicts, each resolved by keeping both sides:

- src/metal.rs, FLASHABLE: main's `dump-deaf-cpu` and `watch-window`
  rows, then this branch's `perf-request-diverges`. Its comment loses
  "the reset that ends the boot clears IA32_PM_ENABLE, which only a
  reset does", which the round-2 review removed as unmeasured.
- tests/toyos.rs, TESTCASES: main's job list, with `test_rs_perf_state`
  after `test_rs_abuse_short_sleep` and before `test_rs_syscall_cost`,
  so `test_rs_null_sink_client_exits` stays the last job before
  `log-close`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dict rests on time

A read's ask now lives on `sys_read`'s stack (`perf_state::Ask`), made
by the read's first look and gone with the read. The claim holds no ask,
so a read that ends answered, refused, cancelled or as a nonblocking
`WouldBlock` leaves nothing a later read is answered from or refused by,
which is what toyos-abi/src/syscall.rs's `PerfState` doc promises.
`Reader::cancel` and `DeviceClaim::cancel_perf_state` are deleted: there
is nothing left to take back. The doc's "reads of one claim that overlap
share one ask" is deleted, since every read now asks; a CPU's one answer
still serves every ask issued before it.

With no standing ask there is no readiness before a read: `has_data` is
false and `read_watch` is `None` for the class, so a poll on a claim is
refused `NotSupported` rather than reporting a readiness the next
nonblocking read would not honour.

`perf-state-deaf-cpu` now has no CPU answer a kick: the asker answers
itself inline in `Ask::issue`, so every read on two CPUs names exactly
the one other CPU whichever CPU the test runs on, and no verdict waits
on a kicked vCPU being scheduled inside 250 ms. The refusal names the
ask (`did not answer Generation(N)`). `test_rs_perf_state_silent` makes
a nonblocking read first (the boot's first ask, `WouldBlock`), then two
blocking reads, and `perf_state_silent_cpu` requires the refusals to
name Generation(2) and Generation(3), once each, one CPU each: a read
answered or refused from an ask not its own reds.

kernel-loom/tests/shootdown_answer.rs models the target-to-initiator
edge the slots use: `serve`'s closure writes a loom cell and the
initiator reads it once `served` answers.

`HwpDeclared::ask` asserts `control_regs::report` has run, which is after
every committed CPU ran `init`, so the proof's "enabled on every CPU" is
enforced rather than true by boot order alone.

Also: `Enumerated` inlined as `{:x?}`; the `NotIntel` reason names
MSR_PLATFORM_INFO rather than RAPL; the removed comments (toyos-perfstate
lib.rs's envelope citation, perf_state.rs's "within one timer
interrupt", shootdown.rs's "nothing reads through this edge yet",
system.toml's row comment) are deleted; the track issue drops the same
false citation, says the Linux readings are #568's unmerged samples, and
records that no test launches /system/bin/perfstate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 3, at 148c207

The gate holds. gh pr view 590 --json mergeable returns MERGEABLE. CI run 36509434491 (host) is success on headSha 148c207. The orchestrator's runs at this head, read from the 590r3-*.log files:

  • perf_request: EXIT=0.
  • perf_state_silent_cpu: EXIT=0.
  • Fast: 230 of 230 passed.
  • perf_state_silent_cpu under mutation-claim-held-ask.patch: EXIT=1. The refusals name Generation(1) and Generation(2).
  • perf_state_silent_cpu under mutation-read-unbounded.patch: EXIT=1, at the 300 s ceiling.

Net diff is +1490/−18. Tests are about 327 lines, the issue 71, and lockfiles and manifests about 36. Production is about 1056 lines, including toyos-perfstate's 352. This round adds a net +45.

Earlier blockers

  • Round 2, a nonblocking or cancelled read left its ask on the claim: CLOSED. The ask now lives on sys_read's stack, and Reader holds only declared. mutation-claim-held-ask.patch reverts to the claim-held ask with no cancel, and it turns perf_state_silent_cpu red (EXIT=1, Generation(1) named).
  • Round 2, the verdict depended on time: CLOSED. Under perf-state-deaf-cpu, serve_if_owed never answers, and Ask::issue always serves the asker under the process lock, with preemption off. So each refusal names exactly the one CPU that is not the asker, whatever the schedule. The only clocks left are the harness's liveness bounds.
  • Round 2, loom did not model the target→initiator edge: CLOSED. I extracted the head into a scratch directory and ran cargo test --manifest-path kernel-loom/Cargo.toml --test shootdown_answer: EXIT=0 unpatched. I then patched Shootdown::served from Acquire to Relaxed: EXIT=101, Causality violation: Concurrent read and write accesses.
  • Round 1, T14 metal row, and mutation-hwp-check-no-asserts.patch on the perfdiverge boot: OPEN. Neither has run. The orchestrator holds both as landing conditions.
  • Round 2's HwpDeclared NOTE was also fixed. HwpDeclared::ask asserts APPLIED, which report sets after boot_aps. Every committed AP runs control_regs::init (in percpu::init_ap) before it echoes AP_STARTED.

BLOCKER

None new.

NOTE

  • kernel-loom/tests/shootdown_answer.rs — delete it: it duplicates tlb_shootdown.rs. Under the same served → Relaxed patch, --test tlb_shootdown also exits 101, with an_acknowledged_flush_postdates_the_page_table_write and one_serve_answers_two_concurrent_shootdowns FAILED. That model already reads, through served, a Relaxed atomic that the target's serve closure wrote, and that is the same shape as SLOTS. Round 2's premise that this direction was unmodelled was wrong.
  • kernel/src/object/ops.rs:396 — on SMP, a nonblocking read always issues an ask, sends an IPI to every other CPU, and answers WouldBlock. No retry can ever succeed, and poll is refused. Refusing it NotSupported by name, like poll, would delete this arm's &mut None path. The silent test would then need another way to make its first ask.
  • No test covers the claim that a poll on a perf-state claim is refused NotSupported.

REMOVE

  • kernel/src/syscall/io.rs:95 — "these four device classes": the match lists three.
  • issues/kernel/the-kernel-owns-cpu-performance-state.md — "(Define the T14 LLVM bar as a libc++ stage-3 recipe with an in-tree judge; the bar value is owed #568's samples, not merged; the metal row re-measures them)": this rots.
  • PR body — "Not ready to land: the T14 run is owed": this rots in main's merge commit.
  • PR body — "The orchestrator's runs at ee20f53d" paragraph: history.
  • PR body — every "Expected … Queued." clause, and the "Runs to queue" section.
  • PR body — "On the host: mutation-read-unbounded.patch measured EXIT=1": that run was QEMU, not the host.
  • PR body — "Run alone, --test tlb_shootdown also goes red … The new model states the perf-state direction with a cell rather than an atomic.": goes with the loom file's deletion.

The only things between this PR and landing are the two open round-1 blockers, the T14 perf_request row and its perfdiverge assert mutation, which the orchestrator holds. No change by the implementer is required beyond the NOTE and REMOVE lines.

SEND BACK

Japabu and others added 2 commits September 29, 2026 05:35
`shootdown_answer.rs` duplicated `tlb_shootdown.rs`, which already goes red
under the same `served` Acquire-to-Relaxed mutation, so its own model was
dead weight. `kernel/src/syscall/io.rs`'s device-class count was already
wrong against its match arm, and the issue's citation of an unmerged PR's
unlogged numbers rotted the moment it was written. The nonblocking-read NOTE
is filed rather than fixed, since it stays a NOTE this round.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant