Conversation
Size Report
Startup median (7 runs, lower is better):
|
There was a problem hiding this comment.
All reported issues were addressed across 21 files
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
|
The code at 5411c0b looks correct to me, but I could not run two of the checks that matter here, so I am holding the call until they are covered. I did not run the startup-race readiness control, so I only infer that removing the gate fails it, from LIVE_DAEMON_PROBE_RETRIES and the fixture's 12 not-ready probes. I also did not run the CLI import-closure gate for the lazy stopAndRetireDaemon import. Please run both and share the output: the readiness control should fail with the gate removed and pass with it in place, and the import-closure gate should pass. CI is green: 14 checks, 0 not passing, and the daemon-client startup, timeout and host-kit owner liveness routes this diff touches are covered by unit-core and coverage, which pass. No conflicts. The eight resolved threads about daemonPreservedAfterTimeout misreporting a force-killed daemon as preserved are fixed only in dependent #3138 (0780e5e), which I did not review, so they still apply at this head. Please land #3138 in the same stack before this reaches main. Not blocking, and fine to take or leave: isProcessAlive in packages/host-kit/src/internal/host-process.ts still returns false for pids above 0x7fffffff, so callers like owned-process-reaper.ts keep the out-of-range hole this PR closes elsewhere, and a follow-up could make an invalid pid never read as death at that owner. Also, readIdentity in test/integration/support/daemon-test-cleanup.ts reads daemon.json twice, and one status read would do. Could the PID guard be smaller if the invalid state could not be represented? For example, host-kit could parse every pid once into a branded ProcessPid (with isProcessPid as the only constructor), and OwnerIdentity.pid and DaemonProcessIdentity.pid could use that type. The guards in stopDaemonProcess, waitForDaemonExit and classifyOwnerLiveness would then hold by type, and readDaemonInfo could reject the record instead of mapping an invalid pid to 0, which feeds takeover and retirement today. That is not required for this stacked fix. One open question for later: does the real daemon publish daemon.json before /health is ready on every transport? If not, a contender with a non-auto transport preference could still reach takeover after the gate passes on the other transport. That is a policy choice from #3130, not part of this change. |
Summary
Fixes seven review findings in the daemon ownership stack. A startup contender waits for a live winner to become ready before applying takeover policy. Stop, exit confirmation, registration parsing and lock ownership share one native-PID validator; invalid identity cannot authorize deletion. Test cleanup joins both its observed child and a registered successor before deleting their directory.
Timeout diagnostics report retained state from the retirement result. Fixture waits observe their own child's exit and clear their timers. Fixture directories remain owned by runner cleanup until journal writes settle. Startup deadline tests advance directly to the deadline boundary, preserving the 15-second budget while avoiding coverage-dependent polling timeouts. Manual stop loads retirement lazily, preserving CLI startup closure.
Depends on #3132; addresses comments on #3124, #3130 and #3131. Ref #3116. Twenty-one files, 379 gross changed lines.
Validation
Tested
5411c0bb96a070eac1cdcb6266d88d60912fdfbf: