Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 46 additions & 0 deletions docs/upstream-review-priority.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Priority workflow ports

This change selectively ports workflow fixes from pstack at
[`12d587df`](https://github.com/cursor/plugins/commit/12d587dfb20741cafc376c42c696c5f6e2a64487).
The source pins remain unchanged because model defaults, prompt pruning, and
verification reuse belong to a separate review. Target digests record these
reviewed adaptations against the existing pinned sources.

## Owner lifecycle

An autopilot owner's full-lifecycle brief authorizes babysitting. Ordinary
PR-opening workers still return without starting a watch. Independent
babysitters report conflicts to the topology owner. Autopilot-full owners
control their own branches; the Autopilot-stack root controls stack topology.

A code-ready report starts independent verification while self-proof, CI, and
babysitting continue. Merge-ready or STACK-READY reports include receipts and
the exact SHA. Each patch-changing push starts a new round. Merge preparation
requires current-head CI and the verdict checks in the shipping playbook.

Owners record delegated children. Replacement work uses isolated write targets
unless the old writer's access has been revoked. Audit ticks continue until all
delegated work is finished, and chat updates report only previously unreported
changes.

## Evidence and logs

Verification results identify the commit SHAs named in the brief. Measurement
results also identify sample count, sample definition, and execution order.
Missing attribution causes one retry, then an explicit gap rather than a pass.
Workers report every proven defect.

Decision trails remain append-only across handoffs. A `start` row identifies a
new run and the preceding rows it did not write. Each run audits its own
stretches and corrects inaccurate records with a superseding row. The portable
Node logger already creates new files exclusively; the upstream Bash header
append fix does not replace it.

Reflect reviewers report evidence-backed learnings without a required count.

## Verification boundary

The repository tests validate packaging, installation, skill integrity, and
source baselines. They do not prove that a future model will follow these
instructions. Review the owner and babysitter boundary together when changing
this workflow.
20 changes: 10 additions & 10 deletions profiles/skill-manifest.json
Original file line number Diff line number Diff line change
Expand Up @@ -206,19 +206,19 @@
"upstream": "pstack",
"source": "skills/poteto-mode/playbooks/autopilot-full.md",
"sourceDigest": "2042a89a0399b5fe700110bc03746899536c10f0df186e9fb15bbd4a2f6dfce5",
"targetDigest": "be9eeb3df8df26671d5e134b14b337d9b2fa311e096d1ee6c66c06c6c55c047b"
"targetDigest": "dcc69e752a90147f9097a8462d77dec327e42957b12122311b304c95b3e62ede"
},
"skills/meta-mode/playbooks/autopilot-stack.md": {
"upstream": "pstack",
"source": "skills/poteto-mode/playbooks/autopilot-stack.md",
"sourceDigest": "60e0d997f7b369004a02089c71b9091d3f8e9bc0b4a3c1142598380faa7722fd",
"targetDigest": "7a3dd0ab570560e1051e38fd16711914dacd476202be358e25c0f34c917a2b45"
"targetDigest": "33c6e9c2e2db8b62bbe3f394eb7161a2895cbe4bcd915cbbc28db69250d4fafa"
},
"skills/meta-mode/playbooks/babysit.md": {
"upstream": "pstack",
"source": "skills/poteto-mode/playbooks/babysit.md",
"sourceDigest": "51be21872152f97612d37258273269b88212b6696432d25d678cac30f02f4157",
"targetDigest": "a62385a3c3c3f9f7a54a04b313f76cc5176c188d7b25c32e510208451b3ed026"
"targetDigest": "d52fed9c782e79df75e2e0f4f6344bc764e7aded1e5b8ac880ffdcc69dc9768f"
},
"skills/meta-mode/playbooks/bug-fix.md": {
"upstream": "pstack",
Expand Down Expand Up @@ -254,13 +254,13 @@
"upstream": "pstack",
"source": "skills/poteto-mode/playbooks/multi-phase-plan.md",
"sourceDigest": "b823e05b64bbef84ae5169717c5fab1d9ab5079625996271e3aa7abcdce7b8ad",
"targetDigest": "51f824fa3caae6bd0fd69852f66387ef2e203e6c15f8f8982cf20d0d22965bf1"
"targetDigest": "b4771a313aa64ea72edec9133aa4f88795ee755f78240952fdc1f19d450f62a4"
},
"skills/meta-mode/playbooks/opening-a-pr.md": {
"upstream": "pstack",
"source": "skills/poteto-mode/playbooks/opening-a-pr.md",
"sourceDigest": "b173ff9638ff62f91f6813c1a0ed3417cdf9f6545a659be0c1fb1f42130e0e66",
"targetDigest": "143945e5ded37c8d5b2bc0b3889d2af95ecf884dca1edfe597c59abf387a4200"
"targetDigest": "81cc0748aa97accbee24b743fd03852f20f31722c857569e2b7030d8dcea071f"
},
"skills/meta-mode/playbooks/orchestrate.md": {
"upstream": "pstack",
Expand Down Expand Up @@ -509,13 +509,13 @@
"upstream": "pstack",
"source": "skills/reflect/references/divergent-reviewer.md",
"sourceDigest": "7da0334ce282e6d8d7349c0cff14ffbd10a91a756c0b5c3ccc97291f57b8bed6",
"targetDigest": "1e3f57c622107887b6507e1c7fa599367a629ad05273322814cca87614ef5f02"
"targetDigest": "6b21c64302826cea77c539f2be4bab2188be2821c1caaec536c16ef94f4ad99d"
},
"skills/reflect/references/judgment-reviewer.md": {
"upstream": "pstack",
"source": "skills/reflect/references/judgment-reviewer.md",
"sourceDigest": "917bb2ad21240a42b440aeb2a17766092e09f70d0982417b2f4caa6d74f27164",
"targetDigest": "7fec7389146c2d13fbe1e08382041420cfb43df9ac09ec54867eb6f4a7942044"
"targetDigest": "eafa6b19333572c21daeebe1506bda066d7dd44f0b0e0f9185ba8b54b2d10a93"
},
"skills/reflect/references/synthesizer.md": {
"upstream": "pstack",
Expand All @@ -527,7 +527,7 @@
"upstream": "pstack",
"source": "skills/reflect/references/tooling-reviewer.md",
"sourceDigest": "8c74ebf7811801d634db7ab609b097d77ea00ab77bb4b258d69bd96e671cfa39",
"targetDigest": "9766d27b377ca63a2c3dc44ebd8b8ccbde47e5f7c063c19a400e3398ffa6b516"
"targetDigest": "c154b4f948cf470d099acd1770925aa1506edf17ba54de833af16c8d8528bc34"
},
"skills/reflect/SKILL.md": {
"upstream": "pstack",
Expand Down Expand Up @@ -558,13 +558,13 @@
"upstream": "pstack",
"source": "skills/show-me-your-work/SKILL.md",
"sourceDigest": "831e85ba3f84f38bf338cd03e6af050fea357752e25a825e334c99f59d7f2077",
"targetDigest": "d189454d45306b8b2ace4456684fc9dbaf9d3c5305e2fe3c35afed8977f1552f"
"targetDigest": "1464cc0dbe44a03df3250456260de8fdfa8535b80e10abaaf9e5b121637919cd"
},
"skills/swarm/SKILL.md": {
"upstream": "pstack",
"source": "skills/swarm/SKILL.md",
"sourceDigest": "d578d877c3b8b9de7163f52579c4683337a940acc0fbe4496a498f5d3b27df30",
"targetDigest": "015c31b71a020a585b628ed8308245530c008b42facedd3040b385628b3ea599"
"targetDigest": "c24493812f03e0165cdbb3dd19e16d376742afe3b07f92cc48575f27883ae729"
},
"skills/tdd/SKILL.md": {
"upstream": "pstack",
Expand Down
4 changes: 2 additions & 2 deletions profiles/upstream-manifest.json
Original file line number Diff line number Diff line change
Expand Up @@ -221,11 +221,11 @@
},
"show-me-your-work": {
"source": "831e85ba3f84f38bf338cd03e6af050fea357752e25a825e334c99f59d7f2077",
"target": "d189454d45306b8b2ace4456684fc9dbaf9d3c5305e2fe3c35afed8977f1552f"
"target": "1464cc0dbe44a03df3250456260de8fdfa8535b80e10abaaf9e5b121637919cd"
},
"swarm": {
"source": "32440570076df5a124d679e0419e5b03c1d405365bb443552789338427a8ce7b",
"target": "015c31b71a020a585b628ed8308245530c008b42facedd3040b385628b3ea599"
"target": "c24493812f03e0165cdbb3dd19e16d376742afe3b07f92cc48575f27883ae729"
},
"tdd": {
"source": "011cab0ecc04a3632121efb493ae4d60aa282a9b72c66de74dd8ad7e4313e05a",
Expand Down
8 changes: 4 additions & 4 deletions skills/meta-mode/playbooks/autopilot-full.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,11 +3,11 @@
**You own the verdicts, never the PRs. One owner runs each PR from build to merge, and nothing merges without your clean swarm verdict.** For "autopilot this queue", "full autopilot", and one-owner-per-PR programs. Orchestrate runs a standing program whose coordinator lands verified work itself and whose workers never merge. Here each PR's owner carries the whole lifecycle through the merge, and the root keeps only verification, countersigns, and audits.

1. **Mark the operator's items and honor state-then-wait.** Items the operator names stay with the operator. The operator reviews and clicks, and no owner merges one. When the operator asks for the protocol or the plan to be stated, deliver the statement and stop. Execution starts only on the operator's explicit go. On that go, arm a `/goal` with the full program objective. The goal continues across turns until the queue is done.
2. **Spawn one owner per PR with the full lifecycle and an early trail.** Resolve the forge once for the program. GitHub CLI (`gh`) is the default. If `command -v origin` succeeds and Origin can resolve the repository, use `origin pr ...` for PR create, edit, view, watch, and merge operations. Otherwise stay on `gh` and record the fallback. Never require Graphite (`gt`). One remote worker per PR owns the build, first push, ready PR, proof on the real artifact, review-comment triage, prose cleanup with `unslop`, the **no-comments** skill, a rebase onto current trunk, the babysit loop to green (`playbooks/babysit.md`), and the merge itself. Every owner starts a `decisions.tsv` trail per the **show-me-your-work** skill, pushes its first branch snapshot, and opens the PR ready, never draft. Open the PR before self-proof so the URL, decisions, and checks form a durable trail. Keep `decisions.tsv` uncommitted and return it with the reports. The rebase always precedes babysit and never waits for drift or conflicts. The merge is the one step an owner may not take alone. Step 4 gates it.
2. **Spawn one owner per PR with the full lifecycle and an early trail.** Resolve the forge once for the program. GitHub CLI (`gh`) is the default. If `command -v origin` succeeds and Origin can resolve the repository, use `origin pr ...` for PR create, edit, view, watch, and merge operations. Otherwise stay on `gh` and record the fallback. Never require Graphite (`gt`). One remote worker per PR owns the build, first push, ready PR, proof on the real artifact, review-comment triage, prose cleanup with `unslop`, the **no-comments** skill, a rebase onto current trunk, the babysit loop to green (`playbooks/babysit.md`), and the merge itself. Every owner starts a `decisions.tsv` trail per the **show-me-your-work** skill, pushes its first branch snapshot, and opens the PR ready, never draft. Open the PR before self-proof so the URL, decisions, and checks form a durable trail. Keep `decisions.tsv` uncommitted and return it with the reports. As soon as a subagent starts, the owner records its ID, expected runtime of at least the longest past run of that kind, and state in a local `children.tsv`. The owner rebases before the code-ready report and babysit, whether or not trunk has drifted. Fix rounds keep that merge base. Rebase again only at merge prep, on a `git merge-tree` conflict with trunk, or on a CI failure caused by a change on trunk. After cleanup, report the code-ready head SHA and the SHA of every later push that changes the patch. Self-proof, CI, and babysit run in parallel with the swarm. Report merge-ready with the head SHA when they finish. Before a push starts a round, run the repository's required pre-review checks on the committed head. A hook pass is not proof. Publish a rebase only to the owner's branch with `git push --force-with-lease` after an `ls-remote` check. Never force-push a shared branch. The merge is the one step an owner may not take alone. Step 4 gates it.
3. **Run owners in true parallel and never stack.** Many owners at once when PRs are self-contained: one writer per branch, disjoint files, cross-PR drift absorbed by rebase. Only genuinely overlapping work serializes. Self-contained PRs branch straight off main, and sequenced work is merge-then-branch. One exception: an owner that must split a genuinely dependent change may hold a short private base-branch stack.
4. **Swarm-verify every merge-ready head before its merge.** At the owner's merge-ready head SHA, fan out parallel independent verifiers per the **swarm** skill and aggregate to one verdict. The lanes re-run the gates at that SHA and prove the load-bearing behavior on the real surface. Audit the receipts and the diff, distrusting the PR body. **Regression lane against trunk.** Run the same load-bearing scenario on current trunk. If trunk does not have the feature, record that fact and gate the behavior the diff adds plus the end state the user waits for instead of pretending trunk can produce it. The live lane is the floor, and a verdict without it is not clean. No merge without the root's clean verdict. Findings go back to the owner for fix-forward, and the new head gets a fresh swarm and a fresh verdict.
5. **On a clean verdict the owner merges and takes the next item.** The owner merges only from a head freshly rebased onto trunk. The merge-ready report is made at a trunk-current head, and the swarm verdict pins that SHA. If trunk moves again before the merge, the patch-id rule in `playbooks/shipping.md` governs re-verification. A new head voids the verdict unless the patch-id is unchanged. The owner squash-merges its own PR through the resolved forge and picks up its next self-contained item from the queue. The operator's full-autonomy grant plus the root's clean verdict is the merge authorization that babysitting alone never has. Operator-named items stop at merge-ready and wait for the operator's click.
6. **Run the root layer.** A genuinely new raise of a pinned gate or budget value (a limit CI only lets tighten) needs your fresh countersign, granted only after verifier proof. Absorbing values that already landed on main is drift, not a raise. Run an audit tick over all owners roughly every 30 minutes. A local root arms each tick as a real terminal `/loop`. The loop uses a monitored-shell 30-minute sleep and emits an output-notification sentinel. A cloud root uses the existing cloud-sleeper wake chain instead. Never leave the cadence to memory or lossy completion notifications. At each tick, re-read `<mstack-skills>/meta-mode/playbooks/autopilot-full.md` from the active skill installation, then re-read the armed `/goal`. Audit the operation against both. Fix drift during that tick. Probe each owner with a generic liveness or status check, and collect the decision trails. Count only side effects as progress: commits, pushes, PR or check deltas, and store reports. Treat a lane that passes its expected runtime without a side effect as stuck. Stand it down and dispatch a replacement at once. Do not wait for a polite return. When merges batch, run a retro pass and a post-merge bot-comment sweep.
4. **Swarm-verify each round before its merge.** A round starts at the code-ready head SHA and at every later push that changes the patch. At that SHA, fan out parallel independent verifiers per the **swarm** skill and aggregate to one verdict. Merge requires a clean verdict whose patch matches the merge-ready head. Audit the merge-ready receipts before the verdict. The lanes re-run the gates at that SHA and prove the load-bearing behavior on the real surface through its runtime driver. Audit the diff in two or more independent review lanes, each with the full brief and a main focus such as consumer parity, lifetimes and races, or data and config safety. Distrust the PR body. **Regression lane against trunk.** Run the same load-bearing scenario on current trunk. If trunk does not have the feature, record that fact and gate the behavior the diff adds plus the end state the user waits for instead of pretending trunk can produce it. The live lane is the floor, and a verdict without it is not clean. No merge without the root's clean verdict. Send every proven finding to the owner in one fix-forward, including defects filed as notes. For each behavior finding, request a red test covering every site with the same defect, or a repro receipt when no test can show it. Carry that defect into the next round's brief. The new head gets a fresh swarm and verdict, except for results that remain valid under the patch-id rule in `playbooks/shipping.md`.
5. **On a clean verdict the owner merges and takes the next item.** The owner merges only from a head freshly rebased onto trunk. Merge prep starts only after the round's lanes have started and ends with a rebase onto current trunk immediately before merge. The owner reports the new head SHA. CI must pass on that head, and the patch-id rule in `playbooks/shipping.md` decides whether the round's verdict still holds. If trunk moves again before the merge, the patch-id rule in `playbooks/shipping.md` governs re-verification. A new head voids the verdict unless the patch-id is unchanged. The owner squash-merges its own PR through the resolved forge and picks up its next self-contained item from the queue. The operator's full-autonomy grant plus the root's clean verdict is the merge authorization that babysitting alone never has. Operator-named items stop at merge-ready and wait for the operator's click.
6. **Run the root layer.** A genuinely new raise of a pinned gate or budget value (a limit CI only lets tighten) needs your fresh countersign, granted only after verifier proof. When the operator's grant covers approvals, the countersign is the approval. The owner records it in the form the tool's approval contract permits, with a pointer to the countersign. A lane verifies that record. The root never bypasses a forge-enforced approval. Absorbing values that already landed on main is drift, not a raise. Run an audit tick over all owners roughly every 30 minutes. A local root arms each tick as a real terminal `/loop`. The loop uses a monitored-shell 30-minute sleep and emits an output-notification sentinel. A cloud root uses the existing cloud-sleeper wake chain instead. Never leave the cadence to memory or lossy completion notifications. At each tick, re-read `<mstack-skills>/meta-mode/playbooks/autopilot-full.md` from the active skill installation, then re-read the armed `/goal`. Audit the operation against both. Fix drift during that tick. Probe each owner with a generic liveness or status check, and collect the decision trails. Count only side effects as progress: commits, pushes, PR or check deltas, and store reports. Treat a lane that errors or passes its expected runtime without a side effect as stuck. Stand it down and dispatch a replacement at once. Do not wait for a polite return. Each tick checks the platform's agent list when available and every owner's `children.tsv`. Whether or not stopping succeeds, record the child as stuck and replace it if its work is still needed. Apply the same rule to stalled replacements. The root takes both steps when the owner cannot. Before a replacement writes, revoke the old writer's access or give the replacement an isolated write target. A stall neither proves nor drops the work. When merges batch, run a retro pass and a post-merge bot-comment sweep. Stop the recurring audit only when no delegated work remains, even after the last merge.
7. **Stand down instantly on the operator's stop.** The operator's hold or stand-down reaches every owner as a zero-writes order immediately. Owners hold their briefs until the operator releases them.

**Reply:** the queue with each PR's owner, state, and head SHA. Each verdict and the swarm that produced it. What merged and what each owner took next. Countersigns granted and why. Open operator gates. Where the collected decision trails live.
Loading
Loading