diff --git a/pstack/.cursor-plugin/plugin.json b/pstack/.cursor-plugin/plugin.json index 6ebc34b96..a9a1d13a1 100644 --- a/pstack/.cursor-plugin/plugin.json +++ b/pstack/.cursor-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "pstack", "displayName": "pstack", - "version": "0.15.5", + "version": "0.15.6", "description": "if you want to go fast, go deep first. pstack helps you write less, but higher quality code. rigorous agent workflows you can parallelize with confidence.", "author": { "name": "Lauren Tan" diff --git a/pstack/skills/architect/SKILL.md b/pstack/skills/architect/SKILL.md index afa6eead8..7de4233b3 100644 --- a/pstack/skills/architect/SKILL.md +++ b/pstack/skills/architect/SKILL.md @@ -22,7 +22,7 @@ Open a todolist with one entry per phase before starting. Build a real mental model of every system the new code touches. Run the **how** skill over the relevant subsystems. -Naming a file isn't grounding. Produce the traced model `how` prescribes. If the design redefines ownership or layering, also run the **why** skill on the existing shape so the rationale becomes a constraint, not a guess. +Naming a file isn't grounding. Produce the traced model `how` prescribes. If the design redefines ownership or layering, also run the **why** skill on the existing shape so the rationale becomes a constraint, not a guess. Tag each constraint as operator-stated, code-derived, or agent-proposed. When a fix for an agent-proposed constraint grows or fails live, offer dropping the constraint before adding code. Skip Phase A only when the work is genuinely greenfield with no surrounding system to integrate. @@ -32,6 +32,8 @@ Run the **arena** skill with the design-sketch task and the Phase A grounding ar Take the runners from the `architect runners` line in the `pstack-models.mdc` rule, in place of the `arena runners` line. If the rule or that line is missing, use `claude-opus-5-5-max`, `gpt-5.6-sol-max`, `grok-4.7-xhigh-fast`. Alias and rejected entries follow the runner rules in the **arena** skill's Phase A. +When an OS permission system or a live run decides the open question, run the Prototype playbook of the **poteto-mode** skill first and design around the observed result. + Design it twice. Require at least two structurally distinct candidates before synthesis, even when the first looks sufficient. This is the **exhaust-the-design-space** principle skill made concrete. Whole-shape alternatives, not point fixes inside one shape. Screen every candidate against [`references/design-red-flags.md`](references/design-red-flags.md) before synthesis. Reject or revise shallow modules, information leakage, temporal decomposition, and pass-through methods. diff --git a/pstack/skills/poteto-mode/SKILL.md b/pstack/skills/poteto-mode/SKILL.md index f8ddcae37..47df4be27 100644 --- a/pstack/skills/poteto-mode/SKILL.md +++ b/pstack/skills/poteto-mode/SKILL.md @@ -16,13 +16,13 @@ The Principles section below grounds every trigger. In your reply, name each pri Remaining triggers: -- Nontrivial change, architecture decision, or "are we sure?" → the **how** skill. -- About to `AskQuestion` on a "which approach", "how should I", or "what should this do" fork → classify it before you ask. If the answer is a fact you could observe by running something (behavior, timing, layout, output, perf, even whether an eval separates), it is not the human's to answer. Sketch it via the Prototype playbook (`playbooks/prototype.md`) and let the result decide. If the task is a read-only Investigation whose deliverable is a cited answer, stay in it and answer from the evidence rather than building a sketch. Reserve the question for a genuine product or preference call no experiment can settle. Under a full-autonomy grant, decide a call that the grant covers, act on it, and report it, with no reply word and no offer. Under the grant, apply a default for a call that only the operator can make. Report the default with a full explanation and the one word that reverses it. Gates that the operator named and the Always-pause list in Autonomy still need the operator. +- Nontrivial change, architecture decision, or "are we sure?" → the **how** skill. When the operator is in a live test loop, ship the smallest build first. Run `how` once, after the loop closes or when nobody is waiting on a build. +- About to `AskQuestion` on a "which approach", "how should I", or "what should this do" fork → classify it before you ask. If the answer is a fact you could observe by running something (behavior, timing, layout, output, perf, even whether an eval separates), it is not the human's to answer. Sketch it via the Prototype playbook (`playbooks/prototype.md`) and let the result decide. If the task is a read-only Investigation whose deliverable is a cited answer, stay in it and answer from the evidence rather than building a sketch. Reserve the question for a genuine product or preference call no experiment can settle. Under a full-autonomy grant, decide a call that the grant covers, act on it, and report it, with no reply word and no offer. Under the grant, apply a default for a call that only the operator can make. Report the default with a full explanation and the one word that reverses it. Gates that the operator named and the Always-pause list in Autonomy still need the operator. A new operator prohibition that changes the set of checks earns one immediate question. Name its narrowest scope and what it protects (operator time, the machine, or the app under test). Tag each constraint as operator-stated, code-derived, or agent-proposed. When a fix for an agent-proposed constraint grows or fails live, offer dropping the constraint before adding code. - Any code → name the data shape first, and choose its organizing structure per **principle-model-the-domain**. -- Code crossing a function boundary → the **architect** skill, parallel design exploration before implementing. +- Code crossing a function boundary → the **architect** skill, parallel design exploration before implementing. In a live test loop, the same deferral as `how` applies. Smallest build first, design rounds once after the loop closes. - Parallel fan-out → the **swarm** skill for coverage matrices, races, gauntlets, and exploration partitions. Use **arena** for design or code bakeoffs with base selection and grafting. - Contested design → the **interrogate** skill (multi-model adversarial) before shipping. -- Nontrivial multi-step → write the throughput checkpoint (Feature step 3). +- Nontrivial multi-step → write the throughput checkpoint (Feature step 4). - Any prose surface → the **unslop** skill. Your reply is a prose surface. Write it per **Writing the reply**. Agent-facing prose also follows the **create-skill** skill (Cursor's built-in for authoring SKILL.md files). - Docs, RFCs, readmes, PR descriptions, or commit messages → the **technical-writing** skill (`/technical-writing`). - Before commit → the `deslop` skill from the `cursor-team-kit` plugin (`/deslop`). @@ -90,9 +90,9 @@ Read the leaf skill in full for any principle you apply. Each entry names when i **Use `subagent_type: "poteto-agent"` for any subagent you spawn inside a playbook step** (code-writing delegates, ad-hoc helpers). `/poteto-mode` and `poteto-agent` route through the same wrapper. Routed workflow skills (`how`, `why`, `interrogate`, `reflect`, `swarm`) set their own `subagent_type` for diverse-model review. Respect what the skill prescribes, don't override to `poteto-agent`. -**Defaults for every `Task` call.** `run_in_background: true`, agent mode (readonly strips MCP), file pointers not inlined context, explicit model per role (configurable via `/setup-pstack`. Defaults `grok-4.7-xhigh-fast` for code, `claude-opus-5-5-max` for prose and judgment). Code delegates tier by difficulty. The hardest changes (cross-cutting design, gnarly concurrency, subtle algorithms) go to your strongest judgment model (`claude-opus-5-5-max`), whether the task needs judgment on vague intent or is a precisely specified sequence of steps to execute to the letter. Trivial mechanical edits go to your fast code model. Per-role lines in the `/setup-pstack` rule override these defaults and the model choices in the routed skills (`how`, `why`, `arena`, `swarm`, `architect`, `interrogate`, `reflect`). A role with no line keeps its default, and a role line of `inherit-parent` or `auto` runs that role on the parent chat model (omit Task `model`). Each code playbook's configured model comes from its line (`feature, refactoring`, `bug-fix`, `perf-issue`, or `hillclimb`), and the hardest changes read `hardest tasks`. Prose and judgment read `judgment and prose`. +**Defaults for every `Task` call.** `run_in_background: true`, agent mode (readonly strips MCP), file pointers not inlined context, explicit model per role (configurable via `/setup-pstack`. Defaults `grok-4.7-xhigh-fast` for code, `claude-opus-5-5-max` for prose and judgment). Code delegates tier by difficulty. The hardest changes (cross-cutting design, gnarly concurrency, subtle algorithms) go to your strongest judgment model (`claude-opus-5-5-max`), whether the task needs judgment on vague intent or is a precisely specified sequence of steps to execute to the letter. Trivial mechanical edits go to your fast code model. Per-role lines in the `/setup-pstack` rule override these defaults and the model choices in the routed skills (`how`, `why`, `arena`, `swarm`, `architect`, `interrogate`, `reflect`). A role with no line keeps its default, and a role line of `inherit-parent` or `auto` runs that role on the parent chat model (omit Task `model`). Each code playbook's configured model comes from its line (`feature, refactoring`, `bug-fix`, `perf-issue`, or `hillclimb`), and the hardest changes read `hardest tasks`. Prose and judgment read `judgment and prose`. Read the `/setup-pstack` rule file from disk at each fan-out, not from a copy taken at session start. -You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal. +You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. An interrupt also drops the check in flight. Interrupt text says to finish the running check, then start nothing new. A second opinion is the same prompt against a different model. Agreement is high-signal. ## Writing the reply @@ -119,7 +119,7 @@ Open a todolist whose first items are the matched playbook's steps, copied in ve A large or cross-cutting effort (a migration across many call sites, an ambitious multi-part change), or work the user steps away from to trust later, routes to the **figure-it-out** skill even when a narrower playbook like Feature fits. Use **figure-it-out** whenever no bundled playbook fits. It designs a bespoke, rigorous playbook for the task. A standing project-scale program (multi-day, many stacked PRs, a fleet of subagents under one coordinator) routes to **Orchestrate** instead. figure-it-out designs one bespoke run, orchestrate runs the program. - **Investigation.** Read-only question: how does X work, why was Y built this way, are we sure about Z, should we do X or Y. `playbooks/investigation.md`. -- **Bug fix.** A reported defect to reproduce, root-cause, and fix with runtime evidence. `playbooks/bug-fix.md`. +- **Bug fix.** A reported defect to reproduce, root-cause, and fix with runtime evidence. A defect the operator reports mid-run (a screenshot, "this is broken") routes here even inside Feature. `playbooks/bug-fix.md`. - **Perf issue.** A measured slowness to trace and improve against a baseline. `playbooks/perf-issue.md`. - **Hillclimb.** Sustained, scientific improvement of one metric against a target: loop hypotheses with before/after measurement, a decision log, and one commit per accepted win. Distinct from Perf issue, which is a one-off fix. `playbooks/hillclimb.md`. - **Runtime forensics.** Diagnose a runtime symptom (leak, idle-CPU spin, glitch) from live instrumentation. The deliverable is a diagnosis, not a fix. `playbooks/runtime-forensics.md`. diff --git a/pstack/skills/poteto-mode/playbooks/bug-fix.md b/pstack/skills/poteto-mode/playbooks/bug-fix.md index eeebc42b7..65495c272 100644 --- a/pstack/skills/poteto-mode/playbooks/bug-fix.md +++ b/pstack/skills/poteto-mode/playbooks/bug-fix.md @@ -1,11 +1,11 @@ ### Bug fix -**You own this task. Plan, review, verify.** Delegate investigation and the fix to subagents, stay in the lead. +**You own this task. Plan, review, verify.** Delegate investigation and the fix to subagents, stay in the lead. A delegate brief carries the symptom and the evidence, never a candidate cause. Be scientific. Every shipped line traces to runtime evidence. Belt-and-suspenders that "might help" is a hypothesis, not a fix. It does not ship. When evidence refutes a hypothesis, revert what it motivated. The smallest change the evidence justifies ships, nothing more. 1. Reproduce it yourself on the matching surface via the control skill (Non-negotiables), even when a debug or instrumentation protocol says to ask the user to reproduce. Ask the user only with a stated, specific reason the control surface cannot reach the target, and only after driving it as far as it goes. If it won't reproduce directly, synthesize the trigger, tighten conditions, or instrument until it fires. -2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with Cursor's `/loop` command. Confirm the surviving *mechanism* with runtime evidence before the step-3 architect/interrogate fan-out. +2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with Cursor's `/loop` command. A missing log line is evidence only after you found the line the success path writes. A status code means what its consuming handler does with it, not what its name suggests. A hypothesis you did not eliminate stays listed as untested when the symptom stops. A fix that worked with one still open is unproven. Confirm the surviving *mechanism* with runtime evidence before the step-3 architect/interrogate fan-out. 3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation to a subagent using your configured bug-fix model (default `grok-4.7-xhigh-fast`) with a specific scope. 4. Verify on the same surface. The original repro now passes. "Inconclusive" or wrong-surface is not a pass. Flag it. Unit tests show branch behavior, not bug absence. 5. Stage the commits so the failing repro lands before the fix in git history. See the **tdd** skill for the failing-test-first cadence when the bug has a cheap local test path. Skip it when the test would be expensive, integration-heavy, or unclear. diff --git a/pstack/skills/poteto-mode/playbooks/feature.md b/pstack/skills/poteto-mode/playbooks/feature.md index 79a6bbbe0..1128f79b2 100644 --- a/pstack/skills/poteto-mode/playbooks/feature.md +++ b/pstack/skills/poteto-mode/playbooks/feature.md @@ -2,19 +2,20 @@ **You own the design. Plan, review, verify.** Delegate implementation. Stay in the lead. -1. `how` over the affected subsystem. -2. `architect` for parallel design exploration. -3. Write the throughput checkpoint as four todo items. A dimension that genuinely does not apply (single file, no fan-out) keeps its item with `n/a: ` rather than being dropped: - - **Blocking first steps.** Gates run before fan-out. +1. Check whether the operator is in a live test loop, testing each build by hand as it lands. If so, ship the smallest build that keeps the loop moving before anything else. Run the design rounds, the review passes, and full verification once, after the loop closes or when nobody is waiting on a build. +2. `how` over the affected subsystem. +3. `architect` for parallel design exploration. +4. Write the throughput checkpoint as four todo items. A dimension that genuinely does not apply (single file, no fan-out) keeps its item with `n/a: ` rather than being dropped: + - **Blocking first steps.** Gates run before fan-out. A reviewer's pre-brief of the delegate brief is a gate. It runs before the delegate, not beside it. - **Independent workstreams.** Disjoint files, services, or layers parallelize. Shared writes serialize. - **Shared mutable state.** Default to splitting the target (the **separate-before-serializing-shared-state** principle skill). Serialize only for real invariants. - **Smallest safe decomposition.** If one worker is best, name why. -4. Delegate code-writing to a subagent using your configured feature model (default `grok-4.7-xhigh-fast`) with a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). A subagent forbidden to spawn satisfies this by owning the diff directly with the same review separation. No "standing by" reply that waits on a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally. -5. Verify on the matching surface. "Inconclusive" or wrong-surface is not a pass. Flag it. -6. Rebase into small, ordered commits. Stack follow-ups. +5. Delegate code-writing to a subagent using your configured feature model (default `grok-4.7-xhigh-fast`) with a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). A subagent forbidden to spawn satisfies this by owning the diff directly with the same review separation. No "standing by" reply that waits on a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally. +6. Verify on the matching surface. "Inconclusive" or wrong-surface is not a pass. Flag it. Run the primary path yourself once and cite the log line that proves it before any hand-off to the operator. +7. Rebase into small, ordered commits. Stack follow-ups. Use the **sequence-verifiable-units** principle skill, building, verifying, and committing each small unit before the next. -7. If the design is contested, `interrogate` before shipping. -8. Run **Opening a PR**. +8. If the design is contested, `interrogate` before shipping. +9. Run **Opening a PR**. Code-coupled work (one feature, one migration) goes to a single owner with the checkpoint inline. That owner fans out internally after the blocking phase. Parent-level fan-out is for slices that produce independent artifacts (audits, cross-subsystem investigations, competing experiments). Rewrite the checkpoint at phase boundaries. Spawn a fresh owner rather than chaining interrupts. diff --git a/pstack/skills/poteto-mode/playbooks/opening-a-pr.md b/pstack/skills/poteto-mode/playbooks/opening-a-pr.md index 626f376bd..662095446 100644 --- a/pstack/skills/poteto-mode/playbooks/opening-a-pr.md +++ b/pstack/skills/poteto-mode/playbooks/opening-a-pr.md @@ -2,15 +2,15 @@ Invoked at the end of every other playbook. -**Worktree.** Work from a git worktree off main. Subagents inherit it. Multiple `Task` calls on the same branch each get their own worktree, or `git fetch && git reset --hard origin/` between them. Dirty branch with unrelated work: patch out, fresh worktree, apply. Snarled worktree: reset from main, redo minimally. +**Worktree.** Work from a git worktree off main. Subagents inherit it. Multiple `Task` calls on the same branch each get their own worktree, or `git fetch && git reset --hard origin/` between them. Dirty branch with unrelated work: patch out, fresh worktree, apply. Snarled worktree: reset from main, redo minimally. Never rebase, cherry-pick, or edit in the worktree the operator runs the app from. Do history work in a sibling worktree and push it with `--force-with-lease`. -**Commits.** Commit liberally. Rebase into small, ordered commits before opening PRs. Each commit is a future PR: landable, ordered to tell the story. Amend when the fix belongs in a just-made commit. New commit when separable. +**Commits.** Commit liberally. Rebase into small, ordered commits before opening PRs. Each commit is a future PR: landable, ordered to tell the story. Amend when the fix belongs in a just-made commit. New commit when separable. While history is still being rewritten, task files reference commits by subject. Add hashes once, after the last rewrite. **PRs.** Run `/deslop` from `cursor-team-kit` over the diff before commit. Run `/no-comments` before review. Write every PR title, PR description, and commit body with `/technical-writing`, then apply `/unslop`. Apply every technical-writing layer except Diátaxis. Use one word for each action, keep articles, and avoid `-ing` when a plain verb works. **Titles.** Use Conventional Commits in the form `type(scope): subject`. Use `feat`, `fix`, `docs`, `refactor`, `test`, `chore`, or `perf` as the type. Use the changed area, such as `pstack` or `poteto-mode`, as the scope. Keep the subject short and imperative. Name a real symbol when one carries the change. For example, `fix(pstack): retarget opening-a-pr babysit trigger`. Do not add a trailing period. -**Descriptions.** The PR body is a briefing, not the lab notebook. A reviewer who has the diff should learn why the change exists, what is out of scope, and how you proved the change works. The squash commit body is the PR body. If the body would make the squash commit longer than about 40 lines, cut the body. +**Descriptions.** The PR body is a briefing, not the lab notebook. A reviewer who has the diff should learn why the change exists, what is out of scope, and how you proved the change works. The squash commit body is the PR body. If the body would make the squash commit longer than about 40 lines, cut the body. After each scope addition, re-read Why, Blast Radius, and Verification against the new diff. Use these sections in order. Drop a section when it has nothing to say. @@ -24,9 +24,9 @@ After these sections, attach videos or screenshots when they prove a claim. Do n **Forge.** Resolve the forge before the first PR operation and keep that choice for create, edit, view, watch, and merge. GitHub CLI (`gh`) is the default. If `command -v origin` succeeds and Origin can resolve the repository, prefer `origin pr ...`. If Origin is absent or cannot resolve the repository, stay on `gh` and record the fallback. Do not require Graphite (`gt`). -**Size and stacks.** Prefer five narrow PRs to one large PR. A stack is a base-branch chain. The root PR targets trunk. Each child branch rebases onto its parent's exact tip and its PR targets the parent branch. Create a child with `origin pr create --status open --base ` or `gh pr create --base ` according to the resolved forge. Retarget an existing child with `origin pr edit --base ` or `gh pr edit --base `. Branch from trunk only for independent work. Rebase on trunk before substantial stack work. +**Size and stacks.** Prefer five narrow PRs to one large PR. A stack is a base-branch chain. The root PR targets trunk. Each child branch rebases onto its parent's exact tip and its PR targets the parent branch. Create a child with `origin pr create --status open --base ` or `gh pr create --base ` according to the resolved forge. Retarget an existing child with `origin pr edit --base ` or `gh pr edit --base `. Branch from trunk only for independent work. Rebase on trunk before substantial stack work. A bug found during a live test that predates the branch goes to its own PR unless it blocks the test. -**Readiness.** Open every PR ready, never as a draft. With Origin, pass `--status open`. With `gh`, omit `--draft`. Cloud-agent PR tools default to draft, so set `draft: false` on every PR creation call. If a PR still opens as a draft, run `origin pr ready ` or `gh pr ready ` according to the resolved forge. Run `origin pr view ` or `gh pr view ` before you refer to PR status. +**Readiness.** Open every PR ready, never as a draft. With Origin, pass `--status open`. With `gh`, omit `--draft`. Cloud-agent PR tools default to draft, so set `draft: false` on every PR creation call. If a PR still opens as a draft, run `origin pr ready ` or `gh pr ready ` according to the resolved forge. Run `origin pr view ` or `gh pr view ` before you refer to PR status. Before any CI claim, also run `origin pr view --checks` or `gh pr checks `. **Babysit.** Opening a PR does not start a babysit. Post the URL and keep building. Finish the phase or stack first. Run a separate babysit pass only when the user asks for one after the whole stack exists. A babysit for each new PR stalls the build and spends checks on commits that later waves restart. Push back when feedback drifts from intent. diff --git a/pstack/skills/principle-prove-it-works/SKILL.md b/pstack/skills/principle-prove-it-works/SKILL.md index 5dc14a3b0..e1e1fd3f7 100644 --- a/pstack/skills/principle-prove-it-works/SKILL.md +++ b/pstack/skills/principle-prove-it-works/SKILL.md @@ -13,6 +13,7 @@ Verify every task output by checking the real thing directly. Do not infer from Check the real thing, not a proxy: - Check process liveness directly, not indirectly through derived state - Read the actual value, not a cached or derived representation +- Treat the production launch path and the process identity (parent process, signature, granted permissions) as part of the real artifact, not only the build output - When verification fails, suspect the observation method before suspecting the system ## Script the check when you can diff --git a/pstack/skills/principle-sequence-verifiable-units/SKILL.md b/pstack/skills/principle-sequence-verifiable-units/SKILL.md index 986371260..5f272821d 100644 --- a/pstack/skills/principle-sequence-verifiable-units/SKILL.md +++ b/pstack/skills/principle-sequence-verifiable-units/SKILL.md @@ -10,7 +10,7 @@ Order work as a sequence of small units, each ending in a state you can check, a **Why:** A break caught at the unit that caused it is cheap to localize. A break caught after a batch is buried, and you have already built further on a broken base. Sequencing those same units into a delivery a reviewer can replay turns "trust me" into "watch it go red, then green." -**Execution.** In a sweep, migration, or any run of similar edits, verify each change before starting the next. Each unit is a before/after bracket: known-good state, one change, run the check, then proceed. Rebase onto clean trunk first so every check measures against the real baseline. When a lever does the edits, the per-unit check is nearly free. Run it anyway. +**Execution.** In a sweep, migration, or any run of similar edits, verify each change before starting the next. Each unit is a before/after bracket: known-good state, one change, run the check, then proceed. Rebase onto clean trunk first so every check measures against the real baseline. When a lever does the edits, the per-unit check is nearly free. Run it anyway. The per-unit check is the check the pipeline fails on, not a cheaper stand-in such as a compile when the pipeline also lints. While a human waits on the result, run it once at HEAD. Replay it per commit only when nobody is blocked. **Delivery.** Stack commits and PRs in the order that proves the work. The canonical shape is the failing test first, then the fix on top. Other story orders are a subtraction before the reshape, a baseline capture before the treatment, the scaffold before the feature. Each commit lands on its own and the sequence reads as an argument.