From 37eae2c75f95aaa47ebd997196e3d8d44b0034f1 Mon Sep 17 00:00:00 2001 From: Rasmus Hjulskov Date: Wed, 23 Sep 2026 08:44:20 +0200 Subject: [PATCH] feat(models): move Sol row to gpt-6-sol and Grok row to grok-4.7 GPT-6 Sol (Sep 22) supersedes GPT-5.6 Sol in Codex, which now labels 5.6 Sol as an older model. Grok 4.7 (Sep 21) is the Grok CLI default and matches Grok 4.6 on price and speed. Update the provider-dispatch matrix, first-run sheet, skill and playbook defaults, and docs. Fable and Opus already use rolling aliases and are unchanged. --- CHANGES.md | 2 +- README.md | 2 +- UPSTREAM.md | 2 +- docs/reference.md | 6 +++--- plugins/pstack/skills/architect/SKILL.md | 2 +- plugins/pstack/skills/arena/SKILL.md | 2 +- plugins/pstack/skills/how/SKILL.md | 2 +- plugins/pstack/skills/interrogate/SKILL.md | 4 ++-- plugins/pstack/skills/poteto-mode/SKILL.md | 2 +- .../skills/poteto-mode/playbooks/bug-fix.md | 2 +- .../skills/poteto-mode/playbooks/feature.md | 2 +- .../skills/poteto-mode/playbooks/hillclimb.md | 2 +- .../poteto-mode/playbooks/perf-issue.md | 2 +- .../poteto-mode/playbooks/refactoring.md | 2 +- .../poteto-mode/references/codex-tools.md | 2 +- .../references/provider-dispatch.md | 6 +++--- plugins/pstack/skills/setup-pstack/SKILL.md | 20 +++++++++---------- plugins/pstack/skills/swarm/SKILL.md | 2 +- 18 files changed, 32 insertions(+), 32 deletions(-) diff --git a/CHANGES.md b/CHANGES.md index 92906368..3a3b6cb7 100644 --- a/CHANGES.md +++ b/CHANGES.md @@ -40,7 +40,7 @@ Origin merge commands omit GitHub's unsupported `--squash` flag and stop when Or **No-op upstream revisions.** Upstream moved several Fable references to a newer revision slug. Open Pstack already stores rolling `claude:fable` and `claude:opus` aliases, so those edits require no port change. -**Upstream-only exclusions.** `make-bot-ui` from `799151d` and `6fecddb` is not ported because it is built entirely from Cursor routine, webhook, and UI primitives. The `disable-model-invocation` additions from `73f8be4` are not applied to `how`, `why`, `unslop`, or `typescript-best-practices`; the flag would break poteto-mode's named invocation path on Claude Code. The `23a56e2` defaults that move `bug-fix`, `perf-issue`, and `hillclimb` from Sol to Fable are not applied because Fable costs much more for these frequent delegated code roles; all three stay on `codex:gpt-5.6-sol@max`. The Claude manifest logo field from `efa2a53` is also omitted because Claude Code has no schema for it. +**Upstream-only exclusions.** `make-bot-ui` from `799151d` and `6fecddb` is not ported because it is built entirely from Cursor routine, webhook, and UI primitives. The `disable-model-invocation` additions from `73f8be4` are not applied to `how`, `why`, `unslop`, or `typescript-best-practices`; the flag would break poteto-mode's named invocation path on Claude Code. The `23a56e2` defaults that move `bug-fix`, `perf-issue`, and `hillclimb` from Sol to Fable are not applied because Fable costs much more for these frequent delegated code roles; all three stay on `codex:gpt-6-sol@max`. The Claude manifest logo field from `efa2a53` is also omitted because Claude Code has no schema for it. ## 1.2.1 keeps Fable and Opus on their latest Claude revisions diff --git a/README.md b/README.md index 64901115..32b25040 100644 --- a/README.md +++ b/README.md @@ -80,7 +80,7 @@ In Codex, ask: Use pstack:setup-pstack to configure pstack. ``` -Setup checks the models you can actually run, shows how each one will start, and asks before saving the choices. The current default group uses Fable, GPT-5.6 Sol, Grok 4.6, and Opus. +Setup checks the models you can actually run, shows how each one will start, and asks before saving the choices. The current default group uses Fable, GPT-6 Sol, Grok 4.7, and Opus. An older model sheet starts using the rolling aliases in memory as soon as this release is installed. Run setup once after updating to persist that migration. It replaces versioned Fable and Opus entries while preserving every role assignment and effort selection. diff --git a/UPSTREAM.md b/UPSTREAM.md index 6d725c48..c6b0c08e 100644 --- a/UPSTREAM.md +++ b/UPSTREAM.md @@ -18,7 +18,7 @@ The table above is the current Cursor sync point. Open Pstack 1.4.1 imports this - Commits `799151d` and `6fecddb` add and relocate `make-bot-ui`. It depends on Cursor routines, webhook events, and UI primitives that Claude Code and Codex do not share. - Four `disable-model-invocation: true` lines from `73f8be4` are not applied to `how`, `why`, `unslop`, or `typescript-best-practices`. Poteto-mode invokes those skills by name, and the flag blocks that route on Claude Code. -- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on `codex:gpt-5.6-sol@max` for cost. +- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on `codex:gpt-6-sol@max` for cost. - The Claude manifest does not take the logo field from `efa2a53` because Claude Code has no schema for it. The shared asset is exposed through the Codex manifest instead. ## Check for changes diff --git a/docs/reference.md b/docs/reference.md index 23015146..475cc581 100644 --- a/docs/reference.md +++ b/docs/reference.md @@ -86,7 +86,7 @@ The Codex build shares one `skills/` tree with the Claude Code build. Nothing is - **Tool and built-in mapping.** Claude tool names and built-in skills resolve through [`codex-tools.md`](../plugins/pstack/skills/poteto-mode/references/codex-tools.md). Model execution resolves separately through [`provider-dispatch.md`](../plugins/pstack/skills/poteto-mode/references/provider-dispatch.md), so Codex can keep Sol native while invoking Claude and Grok externally. - **Subagents.** The `Agent` tool maps to Codex `spawn_agent` / `wait_agent`, enabled by `multi_agent = true`. Parallel fan-out is multiple `spawn_agent` calls in one turn. If the native Codex lane is unavailable, record that lane as a dropout; external Claude and Grok lanes still run, and no provider is silently substituted. There is no `poteto-agent` subagent type on Codex; route ad-hoc subagents by dispatching a `spawn_agent` told to read `poteto-mode` first. - **Auto-fire.** The `hooks/` SessionStart injection is Claude Code-only; Codex has no plugin hook runtime. Enter `pstack:poteto-mode` by name, or add a standing instruction to `~/.codex/AGENTS.md` if you want the same always-on routing. -- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per frontier family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-5.6 Sol max, Grok 4.6 xhigh, and Opus xhigh. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. In Codex, Sol uses native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; Sol and Grok use the runner. Children never detect the parent or reroute themselves. The `bug-fix`, `perf-issue`, and `hillclimb` roles stay on GPT-5.6 Sol max instead of upstream's Fable default because Sol costs less for these frequent delegated code roles. +- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per frontier family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-6 Sol max, Grok 4.7 xhigh, and Opus xhigh. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. In Codex, Sol uses native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; Sol and Grok use the runner. Children never detect the parent or reroute themselves. The `bug-fix`, `perf-issue`, and `hillclimb` roles stay on GPT-6 Sol max instead of upstream's Fable default because Sol costs less for these frequent delegated code roles. Verified in fresh installed Claude Code and Codex sessions: the user-facing skills are discovered and namespaced under `pstack`; both parents fan out the frontier quad through the documented native/external route table, retain long-running handles without a default timeout, and cross-judge only after every candidate is terminal. The `principle-*` leaves remain available for `poteto-mode` to read by path. Claude honors their `user-invocable: false` metadata; Codex 0.149.0 does not ([#8](https://github.com/ericlitman/open-pstack/issues/8)). @@ -193,7 +193,7 @@ The port is editorial, not mechanical. Anywhere upstream pstack assumed Cursor-s | Cursor's `/goal` (standing objective across turns) | The program objective written into the run's standing orders and restated in the todolist | | The Cursor agent store (path in the system prompt) | `~/.claude/orchestrate//`, which survives the session restarts a multi-day program expects | | Model rule `~/.cursor/rules/pstack-models.mdc` | Override sheet `~/.claude/pstack-models.md`, included from `CLAUDE.md` | -| Multi-model panels (arena, architect, interrogate) | Provider dispatch restores the upstream frontier quad: `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Same-provider lanes stay native; external lanes use the bundled runner. | +| Multi-model panels (arena, architect, interrogate) | Provider dispatch restores the upstream frontier quad: `claude:fable@max`, `codex:gpt-6-sol@max`, `grok:grok-4.7@xhigh`, `claude:opus@xhigh`. Same-provider lanes stay native; external lanes use the bundled runner. | ### Cross-vendor dispatch @@ -211,7 +211,7 @@ The earlier port collapsed panels to Claude-only models. The bundled runner rest - **`automations/benny/`** (upstream `0452e08`, the only pstack change between `e46364b` and v0.10.0) — a dormant Slack issue-triage and reproduce-and-fix automation pack built on Cursor's event-triggered automations. It registers no slash skills even upstream, so excluding it changes nothing about the ported plugin's behavior. Porting it would require Cursor's event-trigger runtime, Slack, and tracker plumbing that Open Pstack does not provide. - **`docs/guide/`** (upstream `02c03a9`, `0b7ef5b`, `424829e`) — the ten-chapter usage tutorial and its six screenshots (2.3 MB). It teaches pstack through Cursor's UI, sticky mode, and cloud agents, so a faithful port would be a rewrite rather than a sync, and none of it ships as skill content. Read it upstream at [cursor/plugins/pstack/docs/guide](https://github.com/cursor/plugins/tree/main/pstack/docs/guide); the concepts map through the substitution table above. - **`make-bot-ui`** (upstream `799151d`, relocated by `6fecddb`) uses Cursor routines, webhook events, hosted bot state, and Cursor UI primitives that have no shared Claude Code and Codex mapping. A provider-specific rewrite would be a separate feature, not an upstream sync. -- **Fable solo code defaults** (upstream `23a56e2`) move `bug-fix`, `perf-issue`, and `hillclimb` from GPT-5.6 Sol to Fable. Open Pstack keeps these frequent delegated code roles on `codex:gpt-5.6-sol@max` because Fable costs much more per task. +- **Fable solo code defaults** (upstream `23a56e2`) move `bug-fix`, `perf-issue`, and `hillclimb` from GPT-5.6 Sol to Fable. Open Pstack keeps these frequent delegated code roles on `codex:gpt-6-sol@max` because Fable costs much more per task. - **Sticky mode** (upstream `#144`) — Cursor-only `mode`/`icon`/`color`/`reminder` frontmatter with no Claude Code equivalent. The port's 0.9.5 SessionStart hook is the analog and already carries the non-trivial / trivial / opt-out logic. - **`is_background: true` on `poteto-agent`** (upstream `99559f2`) — Cursor names this key differently. Claude-native frontier definitions use `background: true`; ad-hoc `poteto-agent` calls remain background dispatches at the call site. - **`cursor-team-kit` beyond the seven imported skills** — the rest either duplicate Claude Code built-ins (`verify-this` → the `verify` skill and built-in verification discipline; `check-compiler-errors` → LSP diagnostics; `control-cli`/`control-ui` → `run`/`verify`, already the substitution targets) or overlap skills this port ships (`loop-on-ci`, `review-and-ship`, `weekly-review` vs `babysit`, `fix-ci`, `make-pr-easy-to-review`, `what-did-i-get-done`). `pr-review-canvas` is Cursor-UI-specific. diff --git a/plugins/pstack/skills/architect/SKILL.md b/plugins/pstack/skills/architect/SKILL.md index f6893207..b36c0976 100644 --- a/plugins/pstack/skills/architect/SKILL.md +++ b/plugins/pstack/skills/architect/SKILL.md @@ -31,7 +31,7 @@ Skip Phase A only when the work is genuinely greenfield with no surrounding syst Run the **arena** skill with the design-sketch task and the Phase A grounding artifacts. Pass `references/runner-prompt.md` as each runner's prompt. Each candidate produces a design package shaped per `references/rationale-template.md`. -Use your configured architect runners (defaults `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`). +Use your configured architect runners (defaults `claude:fable@max`, `codex:gpt-6-sol@max`, `grok:grok-4.7@xhigh`, `claude:opus@xhigh`). Design it twice. Require at least two structurally distinct candidates before synthesis, even when the first looks sufficient. This is the **exhaust-the-design-space** principle skill made concrete. Whole-shape alternatives, not point fixes inside one shape. diff --git a/plugins/pstack/skills/arena/SKILL.md b/plugins/pstack/skills/arena/SKILL.md index 0e3b5f38..4a0efb0c 100644 --- a/plugins/pstack/skills/arena/SKILL.md +++ b/plugins/pstack/skills/arena/SKILL.md @@ -26,7 +26,7 @@ The N candidates will receive the same prompt, so the prompt is the contract. 1. State the artifact each candidate is producing. 2. Derive the rubric. State what success looks like for *this* task, then turn it into 3-6 concrete gradeable criteria. The rubric is the picker's tool in Phase D. Candidates only see the task. -3. Pick the runners. Use `arena runners` from the current harness's pstack model sheet when present. Otherwise default to one each on `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Spawn more when the arena covers multiple design directions. Same descriptor N times when the work is generation-bound rather than judgment-sensitive. +3. Pick the runners. Use `arena runners` from the current harness's pstack model sheet when present. Otherwise default to one each on `claude:fable@max`, `codex:gpt-6-sol@max`, `grok:grok-4.7@xhigh`, `claude:opus@xhigh`. Spawn more when the arena covers multiple design directions. Same descriptor N times when the work is generation-bound rather than judgment-sensitive. 4. Assign output paths. Each candidate writes to its own location (a git worktree where possible, otherwise `/tmp/arena-/candidate-/`), per the **separate-before-serializing-shared-state** principle skill. ## Phase B: Fan out diff --git a/plugins/pstack/skills/how/SKILL.md b/plugins/pstack/skills/how/SKILL.md index 84803c9d..5467a7b2 100644 --- a/plugins/pstack/skills/how/SKILL.md +++ b/plugins/pstack/skills/how/SKILL.md @@ -20,7 +20,7 @@ When in doubt, take the simple path. ## Step 2a. Explore (complex questions only) -Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Start all explorers in one fan-out phase through provider dispatch. Use your configured how-explorer descriptor (default `grok:grok-4.6@xhigh`) in `read-only` mode. A native lane uses the parent subagent primitive; an external lane uses the launcher directly. +Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Start all explorers in one fan-out phase through provider dispatch. Use your configured how-explorer descriptor (default `grok:grok-4.7@xhigh`) in `read-only` mode. A native lane uses the parent subagent primitive; an external lane uses the launcher directly. Each explorer gets the prompt in `references/explorer-prompt.md` with its angle filled in. Then go to Step 3. diff --git a/plugins/pstack/skills/interrogate/SKILL.md b/plugins/pstack/skills/interrogate/SKILL.md index ee7514ff..01297736 100644 --- a/plugins/pstack/skills/interrogate/SKILL.md +++ b/plugins/pstack/skills/interrogate/SKILL.md @@ -39,8 +39,8 @@ Start all reviewers in one fan-out phase. Use `interrogate reviewers` from the c | Subagent | Default model | |----------|---------------| | Reviewer A | `claude:fable@max` | -| Reviewer B | `codex:gpt-5.6-sol@max` | -| Reviewer C | `grok:grok-4.6@xhigh` | +| Reviewer B | `codex:gpt-6-sol@max` | +| Reviewer C | `grok:grok-4.7@xhigh` | | Reviewer D | `claude:opus@xhigh` | For each reviewer, route the configured descriptor with `read-only` access and a unique output/receipt path. If the descriptor is `inherit-parent` or `auto`, use the parent subagent primitive without a model override. If a provider, login, or model is unavailable, record a dropout and continue with the completed reviewers. Never pick the closest model or silently fall back; that destroys the meaning of cross-provider agreement. diff --git a/plugins/pstack/skills/poteto-mode/SKILL.md b/plugins/pstack/skills/poteto-mode/SKILL.md index 5e51bfc6..7d1e6d04 100644 --- a/plugins/pstack/skills/poteto-mode/SKILL.md +++ b/plugins/pstack/skills/poteto-mode/SKILL.md @@ -89,7 +89,7 @@ Read the leaf skill in full for any principle you apply. Each entry names when i For `inherit-parent`, `auto`, or an unconfigured native ad-hoc helper, prefer `poteto-agent`. `/poteto-mode` and `poteto-agent` route through the same wrapper. A provider-qualified role instead follows provider dispatch: Claude's shipped frontier agent definitions select the model alias and requested effort, Codex passes both to `spawn_agent`, and external providers run through the deterministic launcher. Routed workflow skills set the task and access mode. Do not override their choices. -**Defaults for every delegation.** Start independent lanes together, use file pointers rather than inlined dumps, preserve only the tools or MCPs the task needs, and assign every writer a worktree or unique output directory. `/setup-pstack` configures the descriptor per role. Upstream defaults use Grok 4.6 xhigh for feature/refactoring, exploration, and swarm work; GPT-5.6 Sol max for bug fixes, performance work, hillclimbing, and tooling review; Fable max for judgment, prose, explanation, synthesis, and hardest tasks; and the four-provider frontier panel for model-diverse judgment. The panel defaults are enumerated in `arena`, `architect`, and `interrogate`. `inherit-parent` and `auto` use the parent model natively and reduce provider diversity when used in a panel. +**Defaults for every delegation.** Start independent lanes together, use file pointers rather than inlined dumps, preserve only the tools or MCPs the task needs, and assign every writer a worktree or unique output directory. `/setup-pstack` configures the descriptor per role. Upstream defaults use Grok 4.7 xhigh for feature/refactoring, exploration, and swarm work; GPT-6 Sol max for bug fixes, performance work, hillclimbing, and tooling review; Fable max for judgment, prose, explanation, synthesis, and hardest tasks; and the four-provider frontier panel for model-diverse judgment. The panel defaults are enumerated in `arena`, `architect`, and `interrogate`. `inherit-parent` and `auto` use the parent model natively and reduce provider diversity when used in a panel. You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md b/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md index 18412149..55b9a7b8 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md @@ -6,7 +6,7 @@ Be scientific. Every shipped line traces to runtime evidence. Belt-and-suspender 1. Reproduce it yourself on the matching surface via the driver skill (`run` for CLIs/TUIs, `verify` for UIs) (Non-negotiables). Don't hand the repro to the user. A debug or instrumentation protocol that says to ask the user does not override this. You drive the instrumented runtime. Ask the user only with a stated, specific reason the control surface cannot reach the target, and only after driving it as far as it goes. Won't reproduce directly, force it: synthesize the trigger, tighten conditions, or instrument until it fires. 2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with Claude Code's `loop` skill. Confirm the surviving *mechanism* with runtime evidence before the step-3 architect/interrogate fan-out. -3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured bug-fix descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write`, a dedicated worktree, and a specific scope. Review the diff. +3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured bug-fix descriptor (default `codex:gpt-6-sol@max`) with `isolated-write`, a dedicated worktree, and a specific scope. Review the diff. 4. Verify on the same surface. The original repro now passes. "Inconclusive" or wrong-surface is not a pass. Flag it. Unit tests show branch behavior, not bug absence. 5. Stage the commits so the failing repro lands before the fix in git history. See the **tdd** skill for the failing-test-first cadence when the bug has a cheap local test path. Skip it when the test would be expensive, integration-heavy, or unclear. This is the canonical **sequence-verifiable-units** principle skill, the failing test first and the fix on top. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/feature.md b/plugins/pstack/skills/poteto-mode/playbooks/feature.md index 1cc22990..e5fe7df8 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/feature.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/feature.md @@ -9,7 +9,7 @@ - **Independent workstreams.** Disjoint files, services, or layers parallelize. Shared writes serialize. - **Shared mutable state.** Default to splitting the target (the **separate-before-serializing-shared-state** principle skill). Serialize only for real invariants. - **Smallest safe decomposition.** If one worker is best, name why. -4. Delegate code-writing through provider dispatch using your configured feature descriptor (default `grok:grok-4.6@xhigh`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). Review its diff yourself. When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). The delegate owns the diff directly and never waits on or launches a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally. +4. Delegate code-writing through provider dispatch using your configured feature descriptor (default `grok:grok-4.7@xhigh`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). Review its diff yourself. When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). The delegate owns the diff directly and never waits on or launches a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally. 5. Verify on the matching surface. "Inconclusive" or wrong-surface is not a pass. Flag it. 6. Rebase into small, ordered commits. Stack follow-ups. Use the **sequence-verifiable-units** principle skill, building, verifying, and committing each small unit before the next. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md b/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md index c38089b7..01dd8381 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md @@ -9,7 +9,7 @@ Core discipline: one change, one measurement, keep or revert. Never stack untest 3. Open the decision log via the **show-me-your-work** skill. A `decision.tsv`, one row per attempt: id, hypothesis, change, before, after, delta, tests, verdict (kept or reverted), note. Read it before each attempt. Keep it out of the tree (gitignored). 4. Ground each hypothesis in the architecture model from step 1, so it names a specific mechanism ("defer X off the boot path because it blocks first paint"), not "try memoizing something". 5. Loop, one hypothesis per iteration: - - Hand the change through provider dispatch using your configured hillclimb descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write` and a tight worktree scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel lanes, each in its own worktree (the **separate-before-serializing-shared-state** principle skill). + - Hand the change through provider dispatch using your configured hillclimb descriptor (default `codex:gpt-6-sol@max`) with `isolated-write` and a tight worktree scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel lanes, each in its own worktree (the **separate-before-serializing-shared-state** principle skill). - Measure before and after with the frozen harness, and run the regression gate. - Accept only when the metric moves past noise and the gate stays green. Otherwise revert the change in full. A tweak that "might help" is not kept. - One commit per accepted fix, staging only the files you changed (`git add `, never `-A`). Log the row either way, kept or reverted. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md b/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md index e08ac677..647a94a1 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md @@ -13,7 +13,7 @@ - **Redundancy.** The wait hangs on one slow instance or attempt. Duplicate the work (replicas, hedged requests, speculative execution) and take the fastest result. The trace has to show the wait dominates and the system has headroom. - **Lazy evaluation.** Cost lands on results that are never used or not needed yet (eager init on the boot path, rendering offscreen items). Defer the work until first use. - **Scheduling.** The work must happen, but not during the interactive moment. Move it to where nobody is waiting: idle callbacks, a background warmup after boot, precompute before the user arrives, cleanup after the frame commits. The win is perceived latency, so measure the interactive path, not total work done. -3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured perf-issue descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write` in a dedicated worktree. Review the diff. Capture a post-fix trace. +3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured perf-issue descriptor (default `codex:gpt-6-sol@max`) with `isolated-write` in a dedicated worktree. Review the diff. Capture a post-fix trace. Apply the **sequence-verifiable-units** principle skill, verifying each attempt before trying the next. 4. Parse and compare the artifacts (JSON to sqlite, diff). "Inconclusive" or wrong-surface is not a pass. Flag it. 5. Cite the measurement in the PR. diff --git a/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md b/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md index 5106c39a..7b7cd654 100644 --- a/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md +++ b/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md @@ -8,7 +8,7 @@ If the cleanup reveals a missing feature or a real bug, split it out and ship th 2. Name the structure the code is missing per **principle-model-the-domain**. Boring code stays when the shape is already clear and local. The reshape must delete branches or invalid states, not add indirection. 3. Name the target shape. State what the module layout, types, and call graph should be if built today (**principle-foundational-thinking**, **principle-redesign-from-first-principles**). If the target crosses a function boundary, run the **architect** skill for parallel design exploration of the shape before the move. 4. Subtract before you add. Delete dead code, collapse one-caller wrappers, drop redundant validators, and remove orphan references before introducing the new shape (**principle-subtract-before-you-add**). The smallest change that reaches the target shape ships (**principle-laziness-protocol**). A speculative cleanup that "might help" gets reverted. -5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits through provider dispatch using your configured refactoring descriptor (default `grok:grok-4.6@xhigh`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, the names being moved, the behavior to hold). Review the diff yourself. +5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits through provider dispatch using your configured refactoring descriptor (default `grok:grok-4.7@xhigh`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, the names being moved, the behavior to hold). Review the diff yourself. 6. Prove behavior is unchanged on the real artifact, not "it compiles" (**principle-prove-it-works**). For larger reshapes, run an equivalence check: a script that diffs old-vs-new outputs, a recorded baseline replayed against the new code, or a smoke run on the matching surface via the driver skill (`run` for CLIs/TUIs, `verify` for UIs). Own the verification yourself. Do not trust a delegate's "looks good" summary. 7. Confirm the change is worth keeping. The success measure is reduced reader load (**principle-minimize-reader-load**). If the diff does not lower reader load somewhere, revert it. 8. Rebase into small ordered commits. A subtraction commit, then the reshape, then any follow-on cleanup. Shape them with the **sequence-verifiable-units** principle skill, so each behavior-preserving slice stays green before the next. Run **Opening a PR**. diff --git a/plugins/pstack/skills/poteto-mode/references/codex-tools.md b/plugins/pstack/skills/poteto-mode/references/codex-tools.md index b967458d..4574638f 100644 --- a/plugins/pstack/skills/poteto-mode/references/codex-tools.md +++ b/plugins/pstack/skills/poteto-mode/references/codex-tools.md @@ -42,7 +42,7 @@ poteto-mode's Subagents section sets Claude-specific defaults (`subagent_type: " ## Models and providers -Do not replace every configured entry with a Codex model. `/setup-pstack` writes portable descriptors such as `claude:fable@max`, `codex:gpt-5.6-sol@max`, and `grok:grok-4.6@xhigh`. In a Codex parent, only `codex:*` is native. Route Claude and Grok descriptors through the external launcher exactly as `provider-dispatch.md` specifies. The current default panel intentionally keeps four-provider frontier diversity and contains no older GPT or Claude substitute. +Do not replace every configured entry with a Codex model. `/setup-pstack` writes portable descriptors such as `claude:fable@max`, `codex:gpt-6-sol@max`, and `grok:grok-4.7@xhigh`. In a Codex parent, only `codex:*` is native. Route Claude and Grok descriptors through the external launcher exactly as `provider-dispatch.md` specifies. The current default panel intentionally keeps four-provider frontier diversity and contains no older GPT or Claude substitute. ## Claude built-in skills pstack references diff --git a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md index 74ab90da..638d01cb 100644 --- a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md +++ b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md @@ -11,8 +11,8 @@ pstack model choices are provider-qualified descriptors: | Family | Upstream pstack choice | Provider | Model | Default effort | Selectable efforts | Claude-native agent stem | |---|---|---|---|---|---|---| | fable | fable | claude | fable | max | low medium high xhigh max | fable | -| sol | gpt-5.6-sol-max | codex | gpt-5.6-sol | max | low medium high xhigh max | - | -| grok | grok-4.6-fast-xhigh | grok | grok-4.6 | xhigh | low medium high xhigh max | - | +| sol | gpt-5.6-sol-max | codex | gpt-6-sol | max | low medium high xhigh max | - | +| grok | grok-4.7-xhigh-fast | grok | grok-4.7 | xhigh | low medium high xhigh max | - | | opus | opus | claude | opus | xhigh | low medium high xhigh max | opus | The allowed effort universe is exactly `low`, `medium`, `high`, `xhigh`, `max`. First-run requested efforts are the Default effort cell of each row. A Claude-native agent stem of `-` means the family has no Claude-native agent. Otherwise the shipped agent name is `pstack--`. @@ -25,7 +25,7 @@ Normalize configured descriptors before matching them to the matrix or choosing This read-time rule makes an older installed sheet use the latest family revision immediately without writing user files. Once per parent run, report that the persisted sheet is stale and that `/setup-pstack` will rewrite it after its normal probes and confirmation. Unknown versioned Claude models remain invalid. The external runner rejects a missed Fable or Opus version pin instead of silently executing it. -`fast` is part of Cursor's Grok selector, not a Grok Build CLI model or effort flag. The portable Grok route pins the current CLI model `grok-4.6`. The first-run Grok effort is `xhigh`. +`fast` is part of Cursor's Grok selector, not a Grok Build CLI model or effort flag. The portable Grok route pins the current CLI model `grok-4.7`. The first-run Grok effort is `xhigh`. ## The parent owns the route diff --git a/plugins/pstack/skills/setup-pstack/SKILL.md b/plugins/pstack/skills/setup-pstack/SKILL.md index 9a0e7441..7e8ceef6 100644 --- a/plugins/pstack/skills/setup-pstack/SKILL.md +++ b/plugins/pstack/skills/setup-pstack/SKILL.md @@ -88,21 +88,21 @@ After the operator confirms, write the in-memory render from step 6. Never paste Provider-qualified per-role choices. Read the installed pstack provider-dispatch reference before dispatching a configured role. Every documented role remains present. `inherit-parent` and `auto` use the parent model natively and still count as one panel lane. -feature, refactoring: grok:grok-4.6@xhigh -bug-fix: codex:gpt-5.6-sol@max -perf-issue: codex:gpt-5.6-sol@max -hillclimb: codex:gpt-5.6-sol@max +feature, refactoring: grok:grok-4.7@xhigh +bug-fix: codex:gpt-6-sol@max +perf-issue: codex:gpt-6-sol@max +hillclimb: codex:gpt-6-sol@max judgment and prose: claude:fable@max hardest tasks: claude:fable@max -how explorer: grok:grok-4.6@xhigh +how explorer: grok:grok-4.7@xhigh how explainer: claude:fable@max why investigators, synthesizer: inherit-parent reflect tooling, judgment, divergent, synthesizer: inherit-parent -arena runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh -arena cross-judge pool: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh -swarm workers: grok:grok-4.6@xhigh -architect runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh -interrogate reviewers: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh +arena runners: claude:fable@max, codex:gpt-6-sol@max, grok:grok-4.7@xhigh, claude:opus@xhigh +arena cross-judge pool: claude:fable@max, codex:gpt-6-sol@max, grok:grok-4.7@xhigh, claude:opus@xhigh +swarm workers: grok:grok-4.7@xhigh +architect runners: claude:fable@max, codex:gpt-6-sol@max, grok:grok-4.7@xhigh, claude:opus@xhigh +interrogate reviewers: claude:fable@max, codex:gpt-6-sol@max, grok:grok-4.7@xhigh, claude:opus@xhigh ``` ### 8. Wire it in diff --git a/plugins/pstack/skills/swarm/SKILL.md b/plugins/pstack/skills/swarm/SKILL.md index 92c8f48a..814bfe81 100644 --- a/plugins/pstack/skills/swarm/SKILL.md +++ b/plugins/pstack/skills/swarm/SKILL.md @@ -23,7 +23,7 @@ Open a todolist with one entry per phase before launching anything. 1. State the done predicate and the artifact or report the swarm must return. 2. Choose the shape. Partition into slices, race N workers on identical briefs, or mix both. For a race or mixed shape, declare `first pass`, `rank all`, or `best-of` before spawning. 3. Set N from the user or derive it from the shape. N is total workers, not the number that run at once. -4. Pick the worker descriptor from `swarm workers` in the current harness's pstack model sheet when present. Otherwise use `grok:grok-4.6@xhigh`. For a model race, name each arm's descriptor up front. +4. Pick the worker descriptor from `swarm workers` in the current harness's pstack model sheet when present. Otherwise use `grok:grok-4.7@xhigh`. For a model race, name each arm's descriptor up front. 5. Give each worker its own writable output when it writes. ## Phase B: Fan out