Conversation
…s it back per CPU The self-hosting bar is measured under a power envelope that must be read back for a whole build span, and ToyOS left every register of it where firmware put it. This is the first stage of the track that makes the kernel own it (issues/kernel/the-kernel-owns-cpu-performance-state.md). The declaration. control_regs.rs, the one CPU-state declaration, now also writes IA32_PM_ENABLE, IA32_HWP_REQUEST, IA32_HWP_REQUEST_PKG and IA32_ENERGY_PERF_BIAS whole on the BSP and every AP, and asserts each on each. The values are the bar's: min is the package's maximum-efficiency ratio (MSR_PLATFORM_INFO[47:40]), max the CPU's HWP highest performance, desired 0, EPP 128, window 0, no package control, which on the T14's inputs is 0x80002a04, what Linux's intel_pstate held there; the package request is 0x8000ff01 and EPB 6. The declaration is all or nothing. A CPU missing any register it names (HWP, EPP, package request, EPB, package thermal status, or not Intel, or hybrid) gets no request, and the BSP says why once: "control_regs: no performance request is declared: no HWP ...". Whether the machine declared one is the BSP's verdict and every AP must reach it; the request itself is per CPU, since a CPU's highest performance is its own. IA32_MISC_ENABLE's turbo bit is read back and not written, because its other bits are model-specific and firmware's. The arithmetic is toyos-perfstate, a pure host-tested crate, because no QEMU CPU has HWP. Its oracle is the T14's Linux MSR readings. The read-back needs no new syscall. It is a device class, perf-state (class 9, the manifest's `devices = ["perf-state"]`), whose read answers toyos_abi::perf's records: the package's registers (HWP package request, PLATFORM_INFO, RAPL unit, PKG_POWER_LIMIT, PKG_ENERGY_STATUS, package thermal status, TEMPERATURE_TARGET), then each CPU's (PM_ENABLE, HWP capabilities, HWP request, EPB, MISC_ENABLE). A per-CPU MSR is readable only on its CPU, so a read asks every CPU. It issues a generation on a second shootdown::Shootdown (the loom-modelled TLB ack protocol), answers for its own CPU, and kicks the rest. Each answers from its next scheduler pass in drain_irqs, which costs two relaxed loads when nothing is owed, and posts perf_state::WATCH. The read parks there, bounded by 250 ms, past which it is refused Io and the silent CPUs are named. A claim exists only with the control_regs::HwpDeclared proof, so no rdmsr of these registers is reachable on a CPU without them. On AArch64 that proof is an uninhabited type and the claim is refused by name. /system/bin/perfstate holds the row and prints one read, checked against the declaration. The guest binary perf_state runs the same code, and perf_request drives it: in QEMU it asserts the named refusal, no request line, and the claim refused NotFound; on the T14 it asserts every CPU logged the bar's literal values and the binary read them all back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
Review, round 1, at 545e9cdCI BLOCKER
NOTE
REMOVE
SEND BACK |
…no claim reaches an unproven register The review sent #590 back with five blockers. Each is answered here, except the reading from the T14, which only that machine can give. - control_regs: CPUID alone now decides whether a CPU gets a request. IA32_HWP_INTERRUPT is set to 0 where CPUID.06H:EAX[8] enumerates it, and IA32_PM_ENABLE is written next. Only after that are IA32_HWP_CAPABILITIES and MSR_PLATFORM_INFO read, and the request computed and written. This is intel_pstate's order (intel_pstate_hwp_enable, then intel_pstate_get_hwp_cap). The HWP writes moved out of the no-ap-control-regs skip, into hwp_init after self_check. - toyos-perfstate: no CPUID bit enumerates MSR_PLATFORM_INFO (0xCE). SDM Vol. 4 documents it in the model tables of DisplayFamily 06H, so a CPU of any other family is refused by name (Refusal::NotFamily6) before the register is read. - The read-back drops MSR_RAPL_POWER_UNIT, MSR_PKG_POWER_LIMIT, MSR_PKG_ENERGY_STATUS and MSR_TEMPERATURE_TARGET. No CPUID bit enumerates them, and nothing at boot read them, so a userland read was the first to touch them. The track's stages 2 and 4 now require each one to be proven present at boot first, and keep the energy counter (CVE-2020-8694) off any row a session can launch. - A cancelled perf-state read takes its ask back (Reader::cancel, reached from sys_read's cancelled wait), so no later read is answered from registers sampled before it began. The ABI doc now says that overlapping reads of one claim share one ask. - HwpCapabilities is deleted. The request's maximum is `capabilities as u8`. - perf-state-deaf-cpu grants the claim on a machine with no declaration and answers zeros. The last CPU answers no ask. perf_state_silent_cpu (Fast) asserts two reads refused Io, each naming that CPU alone. - perf-request-diverges has cpu1 move its request one ratio off the declaration when it answers a read, then run hwp_check. perf_request's metal row gains a second boot, perfdiverge, which must carry that panic on the page after the reset. It is priced in the metal profile and ruled flashable. - TESTCASES now runs test_rs_perf_state, before null_sink_client_exits. The metal row's job_passed("test_rs_perf_state") named a job that no boot ran. - Deleted per the review's REMOVEs: the narration in toyos-perfstate's Cargo.toml; "since Sandy Bridge"; "two relaxed loads"; the cross-reference to the dump's bound; shootdown.rs's list of callers; the track's chronology and the claim that its values are the bar's. The stale "four" in io.rs's device-class count goes too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A second Io read cannot tell a cleared ask from a stale one, since a stale ask past its deadline is refused Io too. So the second read shows that the claim still answers after a refusal, and nothing more. The comments now say only that. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review, round 2, at ee20f53CI Round-1 blockers
BLOCKER
NOTE
REMOVE
SEND BACK |
Two conflicts, each resolved by keeping both sides: - src/metal.rs, FLASHABLE: main's `dump-deaf-cpu` and `watch-window` rows, then this branch's `perf-request-diverges`. Its comment loses "the reset that ends the boot clears IA32_PM_ENABLE, which only a reset does", which the round-2 review removed as unmeasured. - tests/toyos.rs, TESTCASES: main's job list, with `test_rs_perf_state` after `test_rs_abuse_short_sleep` and before `test_rs_syscall_cost`, so `test_rs_null_sink_client_exits` stays the last job before `log-close`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dict rests on time
A read's ask now lives on `sys_read`'s stack (`perf_state::Ask`), made
by the read's first look and gone with the read. The claim holds no ask,
so a read that ends answered, refused, cancelled or as a nonblocking
`WouldBlock` leaves nothing a later read is answered from or refused by,
which is what toyos-abi/src/syscall.rs's `PerfState` doc promises.
`Reader::cancel` and `DeviceClaim::cancel_perf_state` are deleted: there
is nothing left to take back. The doc's "reads of one claim that overlap
share one ask" is deleted, since every read now asks; a CPU's one answer
still serves every ask issued before it.
With no standing ask there is no readiness before a read: `has_data` is
false and `read_watch` is `None` for the class, so a poll on a claim is
refused `NotSupported` rather than reporting a readiness the next
nonblocking read would not honour.
`perf-state-deaf-cpu` now has no CPU answer a kick: the asker answers
itself inline in `Ask::issue`, so every read on two CPUs names exactly
the one other CPU whichever CPU the test runs on, and no verdict waits
on a kicked vCPU being scheduled inside 250 ms. The refusal names the
ask (`did not answer Generation(N)`). `test_rs_perf_state_silent` makes
a nonblocking read first (the boot's first ask, `WouldBlock`), then two
blocking reads, and `perf_state_silent_cpu` requires the refusals to
name Generation(2) and Generation(3), once each, one CPU each: a read
answered or refused from an ask not its own reds.
kernel-loom/tests/shootdown_answer.rs models the target-to-initiator
edge the slots use: `serve`'s closure writes a loom cell and the
initiator reads it once `served` answers.
`HwpDeclared::ask` asserts `control_regs::report` has run, which is after
every committed CPU ran `init`, so the proof's "enabled on every CPU" is
enforced rather than true by boot order alone.
Also: `Enumerated` inlined as `{:x?}`; the `NotIntel` reason names
MSR_PLATFORM_INFO rather than RAPL; the removed comments (toyos-perfstate
lib.rs's envelope citation, perf_state.rs's "within one timer
interrupt", shootdown.rs's "nothing reads through this edge yet",
system.toml's row comment) are deleted; the track issue drops the same
false citation, says the Linux readings are #568's unmerged samples, and
records that no test launches /system/bin/perfstate.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review, round 3, at 148c207The gate holds.
Net diff is Earlier blockers
BLOCKERNone new. NOTE
REMOVE
The only things between this PR and landing are the two open round-1 blockers, the T14 SEND BACK |
`shootdown_answer.rs` duplicated `tlb_shootdown.rs`, which already goes red under the same `served` Acquire-to-Relaxed mutation, so its own model was dead weight. `kernel/src/syscall/io.rs`'s device-class count was already wrong against its match arm, and the issue's citation of an unmerged PR's unlogged numbers rotted the moment it was written. The nonblocking-read NOTE is filed rather than fixed, since it stays a NOTE this round. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stage 1 of the new track
issues/kernel/the-kernel-owns-cpu-performance-state.md. The kernel now declares the CPU's HWP request from the one CPU-state declaration, and a capability-gatedperf-stateclaim reads the power envelope's registers back per CPU. There is no new syscall.The only open landing condition is the T14 metal row:
perf_requeston the T14, plusmutation-hwp-check-no-asserts.patchon theperfdivergeboot. Both are owed by the orchestrator (see "The metal row, owed" below).What changed, per decision
The declaration lives in
kernel/src/arch/x86_64/control_regs.rs. CPUID alone decides whether a CPU gets a request (hwp_declared). Thenhwp_initruns on the BSP and on every AP, afterself_check:IA32_HWP_INTERRUPT(0x773) = 0, only where CPUID.06H:EAX[8] enumerates it.IA32_PM_ENABLE(0x770) = 1.IA32_HWP_CAPABILITIES(0x771) andMSR_PLATFORM_INFO(0xCE) read, and the request computed from them.IA32_HWP_REQUEST(0x774),IA32_HWP_REQUEST_PKG(0x772) =0x8000ff01andIA32_ENERGY_PERF_BIAS(0x1B0) = 6.Each register is written whole, then read back and asserted on each CPU. Each CPU logs one line:
control_regs: cpuN pm_enable= hwp_request= hwp_request_pkg= epb= hwp_interrupt=<Some(v)|None> hwp_capabilities= platform_info=. No read-modify-write decides any of them.The order is the SDM's. Intel SDM Vol. 3B, §17.4.2 "Enabling HWP", p. 17-7, Order Number 253669-093US (September 2026; the PDF I read has SHA-256
a8335f8ccb8b2e82ac51a8ff31360699fc33f334e66133e822038702f52cd276):intel_pstate does the same thing independently:
intel_pstate_hwp_enableclears the interrupt and enables HWP, and only then doesintel_pstate_get_hwp_capread the capabilities.The values.
MSR_PLATFORM_INFO[47:40], the maximum-efficiency ratio. max =IA32_HWP_CAPABILITIES[7:0], the highest level, turbo included. This is how intel_pstate derives min and max.All or nothing, refused by name.
toyos_perfstate::refusalrequires every one of these:MSR_PLATFORM_INFO, which is Intel's.MSR_PLATFORM_INFO, and SDM Vol. 4 documents it only in family 06H's model tables.A CPU that fails any of them gets no register written, and the BSP logs the reason once. Every AP must reach the BSP's verdict, or it is named. The turbo bit (
IA32_MISC_ENABLEbit 38) is read back and not written. It is stage 3.The proof is minted only after every CPU applied the declaration.
control_regs::HwpDeclared::askasserts thatcontrol_regs::reporthas run.reportruns after SMP bring-up, and every committed AP runscontrol_regs::initbefore it echoes. So "HWP is enabled on every CPU" is now enforced, not just true because of boot order.The arithmetic lives in
toyos-perfstate, a pure no_std host crate shared by the kernel, the guest tests and the program.The read-back is a device class. The ABI is
DeviceType::PerfState = 9 => "perf-state", plustoyos_abi::perf's two all-u64records:PackageRegisters: 0x772, 0xCE, 0x1B1.CpuRegisters: 0x770, 0x771, 0x774, 0x1B0, 0x1A0.Every one of those registers is enumerated by CPUID, is architectural (0x1A0), or is read at boot under
HwpDeclared(0xCE). A read can therefore reach no register that boot has not already touched. The/system/bin/perfstaterow holdsdevices = ["perf-state"]. On a machine with no declaration, the claim isNotFound.A read owns its ask. A read issues a generation on a second
shootdown::Shootdown, answers for its own CPU inline, and kicks the others. Each other CPU answers fromdrain_irqs. The ask is aperf_state::Askheld onsys_read's stack: the read's first look makes it, and it ends when the read ends. The claim holds no ask. So a read that ends answered, refusedIo, cancelled, or as a nonblockingWouldBlockleaves nothing behind that a later read could be answered from or refused by. This is the contract intoyos-abi/src/syscall.rs'sPerfStatedoc: "each CPU's taken on that CPU after the read asked".Reader::cancelandDeviceClaim::cancel_perf_stateare gone, because nothing is left to take back.fetch_max).has_dataisfalseandread_watchisNonefor the class, so a poll on a claim is refusedNotSupported. It no longer reports a readiness that the next nonblocking read would not honour. A nonblocking read asks and answers only if every CPU has already answered, which on SMP meansWouldBlock.ANSWER(250 ms) the read is refusedIo, and each silent CPU is logged asperf_state: cpuN did not answer Generation(G) within 250ms.Memory ordering, target to initiator is proven by the existing
kernel-loom/tests/tlb_shootdown.rs, not a new model: the target'sserveclosure writes its slot, and the initiator reads it onceservedanswers — the same edgetlb_shootdownalready checks for its owntlbstore.mutation-served-relaxed.patch(below) turns it red.The witness. A claim holds an
Option<HwpDeclared>, and every read function takes one. On AArch64,arch::perf_state::Declaredis uninhabited, so the claim is refused by name there.Two test actuators, compiled out of a shipping kernel.
perf-state-deaf-cpugrants the claim where nothing is declared, answering zeros. No CPU answers a kick, so every CPU except a read's asker stays silent.perf-request-divergesmakes cpu1 move its request one ratio off the declaration when it answers a read, then runhwp_checkexactly as boot does. It is ruled flashable insrc/metal.rs'sFLASHABLE.Tests
perf_request(Fast, QEMU): the kernel refuses by name, no CPU logshwp_request=, andtest_rs_perf_state's claim is refusedNotFound.perf_state_silent_cpu(Fast, QEMU): boots 2 CPUs withperf-state-deaf-cpu.test_rs_perf_state_silentfirst makes a nonblocking read, which is the boot's first ask and must returnWouldBlock. It then makes two blocking reads, each refusedIo. The judge requires exactly two refusal lines, namingGeneration(2)andGeneration(3)once each, and each naming one CPU. The asker answers inline and no CPU answers a kick, so which CPU names itself does not depend on any vCPU being scheduled in time. The only time bounds left are the harness's liveness bounds. A read that is answered or refused from an ask not its own namesGeneration(1)and reds.perf_requestmetal row: two boots, owed (below).Gates, at
148c2078(after merging origin/main89dd9d2b)cargo run -- --ci hostHost: 54 step(s), all green)cargo run -- --build-onlyBuild finished.)cargo test --test toyos-build -- --list(builds the guest binaries, boots nothing)Fast perf_requestandFast perf_state_silent_cpucargo check, with and without--features boot-actuatorsI ran no QEMU guest and no T14 run. The orchestrator's measured QEMU runs at this head:
perf_requestperf_state_silent_cpuperf_state_silent_cpuundermutation-claim-held-ask.patchperf_state_silent_cpuundermutation-read-unbounded.patchHigh-risk: CPU state, a device claim, a cross-CPU protocol
Negative controls. Each patch is applied with
git apply --checkand thengit apply, shown to build, run, reversed, and the file compared byte for byte afterwards.mutation-served-relaxed.patch(Shootdown::served'sAcquire→Relaxed). Builds (--no-runEXIT=0).--test tlb_shootdown: EXIT=101, withan_acknowledged_flush_postdates_the_page_table_writeandone_serve_answers_two_concurrent_shootdownsFAILED — the loom oracle for the perf-state read's target-to-initiator ordering, since the flush's own postdating check reads through the sameservededge a perf-state answer's registers are read through.mutation-claim-held-ask.patchreverts the ask's ownership to the round-2 mechanism: an ask held by the claim, cleared only when a read is answered or refused. That isee20f53dwithclaim.cancel_perf_state()deleted, and it also carriesee20f53d's nonblocking leak. The deleted line has no site at this head: the ask lives on the read's stack, so this patch is that deletion's equivalent. The mutated kernel builds (cargo check, with and without actuators: EXIT=0, 0). Underperf_state_silent_cpu: EXIT=1. The blocking read joins the nonblocking read's ask, so the refusals nameGeneration(1)andGeneration(2).mutation-read-unbounded.patchremoves the deadline arm (if !answered && !ask.deadline.reached(now)→if !answered). It is regenerated for this head and builds (EXIT=0). Underperf_state_silent_cpu: EXIT=1, at the harness's 300 s ceiling.mutation-hwp-check-no-asserts.patchreplaceshwp_check's assert loop withlet _ = want;. It is regenerated for this head and builds (EXIT=0). It is metal only: no QEMU CPU reacheshwp_check, because TCG and KVM both reduce leaf 6 toARAT. Expected:perf_request'sperfdivergeboot reds. This is part of the owed metal row. It is not a blocker I can close.refusalwithout the family check,hwp_request's max taken from the guaranteed level, andhwp_notificationreading bit 9. Each makescargo test -p toyos-perfstateexit 101.The 250 ms bound is proven twice: under QEMU,
mutation-read-unbounded.patchexits 1 at this head (above); on metal, the T14 row. No QEMU verdict here waits on guest time.Independent oracles.
0x80002a04on every CPU, 0x771 =0x010d182aor0x010e182a, 0x772 =0x8000ff01, 0x1B0 = 6, 0x770 = 1. These live in unmerged Define the T14 LLVM bar as a libc++ stage-3 recipe with an in-tree judge; the bar value is owed #568's samples. Neither this PR nor main carries a command or log for them. The host tests holdtoyos_perfstateto those values, and the metal row re-measures them on ToyOS.What I am unsure of
MSR_PLATFORM_INFOwas not read on the T14. Its ratio 4 is inferred from Linux'scpuinfo_min_freqof 400000 kHz. The per-CPU line prints the register, so the first T14 boot settles it. If it is not 4, the metal row reds onhwp_request.hwp_interrupt=Some(0)orhwp_interrupt=None.perfdivergeboot is the first metal boot to end in a panic raised from a job. Whether the panic record on the page after the reset carries the assert's text is unmeasured./system/bin/perfstate. Deleting its row'sdevicesstays green. A launch test needs a launcher-driven boot. This is recorded in the track's stage 1 as present-state debt.The metal row, owed
perf_requeston the T14. Neither assertion has run yet:testcasesboot. Every CPUcpuNthat the SMP records bring up logscontrol_regs: cpuN pm_enable=1 hwp_request=0x80002a04 hwp_request_pkg=0x8000ff01 epb=6, andtest_rs_perf_stateexits 0. That binary reads every CPU back holding its declaration, twice.perfdivergeboot (perf-request-diverges, jobtest_rs_perf_state). The loader pass after the reset carries, afterPrevious boot's panic:,control_regs: cpu1 holds hwp_request=0x80002a05, the declaration is 0x80002a04. This is also the only measurement ofhwp_check's assert loop.🤖 Generated with Claude Code