diff --git a/CHANGES.md b/CHANGES.md
index ff9a2957..11330ddf 100644
--- a/CHANGES.md
+++ b/CHANGES.md
@@ -1,5 +1,11 @@
# CHANGES — applied substitutions
+## Unreleased: GPT-6 Codex families become stock
+
+- Move `astra`, `sol-6`, and `luna` from the additional matrix into the stock model matrix. The first-run sheet now assigns GPT-6 Sol high to `feature, refactoring`, `bug-fix`, `perf-issue`, and `hillclimb`; Luna high to `how explorer` and `swarm workers`; and Astra high in place of GPT-5.6 Sol on every panel. `codex:gpt-5.6-sol` stays selectable; existing sheets are untouched until a role is reassigned in setup.
+- `provider-dispatch.md` gains a "Default panel" section that is the single source for the four panel lanes. `setup-pstack`, `arena`, `architect`, and `interrogate` copy it; the static invariant reads that line instead of deriving the panel from every matrix row, and the solo-code invariant checks the `sol-6` row.
+- The matrix test asserts seven stock rows, the GPT-6 descriptors in the first-run sheet, and the panel contract. `UPSTREAM-FLEX.md` records the new permanent conflict surface on upstream syncs. Tracked in [pstack-flex #17](https://github.com/thisguymartin/pstack-flex/issues/17).
+
## Unreleased: multiple gateway model choices
- Add DeepSeek V4 Pro and MiniMax M3.1 Flash Preview as independent setup families alongside existing Flash and M3 choices. Keep provider-owned routing and existing descriptors.
diff --git a/README.md b/README.md
index 745ea4eb..8f61368c 100644
--- a/README.md
+++ b/README.md
@@ -1,24 +1,16 @@
-# open-pstack
+# pstack-flex
-[](https://github.com/ericlitman/open-pstack/actions/workflows/ci.yml)
-[](https://github.com/ericlitman/open-pstack/releases/latest)
-[](LICENSE)
+[](https://github.com/thisguymartin/pstack-flex/actions/workflows/ci.yml)
+[](https://github.com/ericlitman/open-pstack/releases/tag/v1.4.1)
+[](LICENSE)
-**Open Pstack brings [Lauren Tan (@poteto)](https://x.com/poteto)'s [pstack](https://github.com/cursor/plugins/tree/main/pstack) to Claude Code and Codex.** Its job is to stay as close to her original work as possible while translating the parts that depend on Cursor.
+**pstack-flex runs [Lauren Tan (@poteto)](https://x.com/poteto)'s [pstack](https://github.com/cursor/plugins/tree/main/pstack) in Claude Code and Codex on the models you actually have.** It is a fork of [ericlitman/open-pstack](https://github.com/ericlitman/open-pstack), which translates pstack's Cursor-specific parts for Claude Code and Codex. open-pstack assumes four frontier subscriptions. This fork keeps its skills and workflows and changes one thing: which models setup accepts and how they are reached.
Lauren built pstack from the skills she uses to ship code at Cursor. In a [55-minute interview with Denis Labelle](https://x.com/DenisLabelle/status/2091337807939706928), she says that she shipped 1,000 pull requests in one month after steadily improving how her agents work and verify their results.
> If you want to go fast, go deep first.
-Open Pstack is an unofficial community project that makes pstack work in Claude Code and Codex. If Cursor is your main coding environment, use [Lauren's original pstack](https://github.com/cursor/plugins/tree/main/pstack). If Claude Code or Codex is your main coding environment, use this repository.
-
-## This fork: pstack-flex
-
-**pstack-flex** is a fork of [ericlitman/open-pstack](https://github.com/ericlitman/open-pstack) at v1.4.1. Setup can assign only the model families you have. The `deepseek:*` and `minimax:*` routes run the stock `claude` binary against each lab's Anthropic-compatible endpoint with that lab's API key. The runner isolates Claude configuration, strips inherited provider routing and Anthropic headers, and rejects OAuth credentials found in the gateway config directory. Gateway receipts keep token usage but set `costUsd` to null because Claude Code's cost estimate uses Anthropic prices.
-
-The [lane guide](docs/LANES.md) covers setup, costs, and safety notes. The fork's provenance and sync process are recorded in [UPSTREAM-FLEX.md](UPSTREAM-FLEX.md). Anthropic does not support pointing Claude Code at non-Anthropic endpoints; use synthetic data for gateway testing and keep API keys in your local environment.
-
-New here? **[docs/USAGE.md](docs/USAGE.md)** is the walkthrough: diagrams of how work flows through the lanes, three setup configurations (full frontier, hybrid saver, zero-subscription), copy-paste examples for the daily skills, and troubleshooting.
+If Cursor is your main coding environment, use [Lauren's original pstack](https://github.com/cursor/plugins/tree/main/pstack). If you hold all four subscriptions and want the closest translation, use [open-pstack](https://github.com/ericlitman/open-pstack). If you want to pick your own models, pay per token where it makes sense, or run with no subscription at all, use this repository.
## What pstack does
@@ -31,16 +23,56 @@ The normal entry point is `poteto-mode`. You give it a task in plain language. I
- compares designs when the choice matters;
- favors small, simple changes over extra machinery;
- asks several models to challenge important decisions when useful;
-- runs the code and checks real behavior instead of stopping at “the tests pass”; and
+- runs the code and checks real behavior instead of stopping at "the tests pass"; and
- carries the work through review, continuous integration (CI), and a ready-to-merge pull request when asked.

pstack does not ask you to trust an agent on day one. It helps the agent leave evidence you can inspect. Start with supervised work. Let it run more work in parallel only after its checks have earned that trust in your own repositories.
+## The models
+
+Every pstack role (who writes code, who explores, who sits on a review panel) maps to one family. A family is one `(provider, model)` pair with its own requested effort and its own live probe in setup. These are the families pstack-flex ships:
+
+| Family | Descriptor at default effort | Needs | First-run role |
+| --- | --- | --- | --- |
+| `fable` | `claude:fable@max` | Claude Code login | judgment, prose, explanation, hardest tasks, panels |
+| `opus` | `claude:opus@xhigh` | Claude Code login | panels |
+| `astra` | `codex:gpt-6-astra@high` | Codex (ChatGPT) login | panels |
+| `sol-6` | `codex:gpt-6-sol@high` | Codex (ChatGPT) login | feature, refactoring, bug-fix, perf-issue, hillclimb |
+| `luna` | `codex:gpt-6-luna@high` | Codex (ChatGPT) login | how explorer, swarm workers |
+| `sol` | `codex:gpt-5.6-sol@max` | Codex (ChatGPT) login | none; selectable |
+| `grok` | `grok:grok-4.6@xhigh` | Grok CLI login | panels |
+| `deepseek` | `deepseek:deepseek-flash@high` | `DEEPSEEK_API_KEY` | none; selectable |
+| `deepseek-pro` | `deepseek:deepseek-v4-pro@high` | `DEEPSEEK_API_KEY` | none; selectable |
+| `minimax` | `minimax:MiniMax-M3@high` | `MINIMAX_API_KEY` | none; selectable |
+| `minimax-preview` | `minimax:MiniMax-M3.1-Flash-Preview@high` | `MINIMAX_API_KEY` (Token Plan) | none; selectable |
+
+The default review panel is `claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh`: four lanes across three providers. Any family can take any role. Panels must span at least two providers, and two models from one provider count as one, because the adversarial signal comes from model diversity.
+
+The DeepSeek and MiniMax lanes run the stock `claude` binary against the lab's Anthropic-compatible endpoint with that lab's key, in an isolated config directory, with inherited Anthropic routing stripped. A lane refuses to start if it finds a claude.ai login in that directory, so a subscription credential can never reach a third-party endpoint. Their receipts keep real token usage but set `costUsd` to null (Claude Code prices at Anthropic rates); the price table is in [docs/LANES.md](docs/LANES.md). Anthropic does not support pointing Claude Code at non-Anthropic endpoints; use synthetic data for gateway testing and keep keys in your local environment.
+
+### How a role becomes a lane
+
+```mermaid
+flowchart LR
+ S["pstack-models.md
role -> provider:model@effort"] --> P["Parent harness
(Claude Code or Codex)"]
+ P -->|"parent's own provider"| N["Native subagent
Agent / spawn_agent"]
+ P -->|"any other provider"| R["pstack-runner
one process per lane"]
+ R --> C1["codex CLI"]
+ R --> C2["grok CLI"]
+ R --> C3["claude CLI + env
DeepSeek or MiniMax endpoint"]
+ N --> O["Output + receipt
model, effort, tokens, status"]
+ C1 --> O
+ C2 --> O
+ C3 --> O
+```
+
+The parent resolves every route once, before fan-out. Children never detect the harness or pick a model. A lane that cannot start drops out with a named receipt; nothing substitutes a weaker model or invents a timeout.
+
## Install
-You need a current Claude Code or Codex installation. For the full four-model review, install and sign in to the Claude Code, Codex, and Grok command-line tools. [Bun](https://bun.sh) runs the small local tool that starts models outside the app you are using. You can still use the core workflows with fewer models.
+You need a current Claude Code or Codex installation and [Bun](https://bun.sh) for the lane runner. Sign in only to the CLIs whose plans you have (Claude Code, Codex, Grok), and export `DEEPSEEK_API_KEY` or `MINIMAX_API_KEY` in the shell that starts your session for the gateway lanes. Any subset works, down to a zero-subscription setup on two keys.
### Claude Code
@@ -72,8 +104,6 @@ Start a new Codex task after installation so it can discover the new skills and
## Get started
-Lauren's original setup has two steps. Open Pstack keeps the same flow.
-
### 1. Set up the models
In Claude Code, run:
@@ -88,9 +118,9 @@ In Codex, ask:
Use pstack:setup-pstack to configure pstack.
```
-Setup checks the models you can actually run, shows how each one will start, and asks before saving the choices. The current default group uses Fable, GPT-5.6 Sol, Grok 4.6, and Opus. You can also assign GPT-6 Astra, Sol, or Luna to any role through Codex.
+Setup is assignment-first. It shows the role map, asks which roles to change, asks one effort per assigned family, probes only those families with a real one-turn run, and writes nothing until every probe passes and you confirm. A fresh run proposes the defaults in the table above. An existing sheet keeps its assignments until you change a named role.
-An older model sheet starts using the rolling aliases in memory as soon as this release is installed. Run setup once after updating to persist that migration. It replaces versioned Fable and Opus entries while preserving every role assignment and effort selection.
+The sheet lives at `~/.claude/pstack-models.md` (Claude Code) or `~/.codex/pstack-models.md` (Codex). It is global, not per project. Change it by rerunning setup rather than editing it by hand, so every choice is probed before it is saved.
### 2. Use poteto-mode
@@ -110,7 +140,7 @@ Use pstack:poteto-mode. Add saved filters to search. Keep the design simple, ver
For that feature, poteto-mode should first understand how search works today. It should decide how the data should be represented before writing code, implement the smallest complete version, run the feature the way a user would, review the result, and prepare the pull request.
-That is the main workflow. The other skills are there when poteto-mode needs them or when you want to call one directly.
+That is the main workflow. The other skills are there when poteto-mode needs them or when you want to call one directly. **[docs/USAGE.md](docs/USAGE.md)** is the longer walkthrough: three setup configurations (full frontier, hybrid saver, zero-subscription), copy-paste examples for the daily skills, how to read receipts, and troubleshooting.
## Useful skills
@@ -128,11 +158,9 @@ That is the main workflow. The other skills are there when poteto-mode needs the
Plugin skills include `pstack:` in their name. In Claude Code, invoke a native skill such as `/pstack:architect`. In Codex, ask for the skill, such as `Use pstack:architect for this design.` See the [technical reference](docs/reference.md) for the full list.
-## Models and token use
-
-Some pstack workflows use one model. Skills such as `architect`, `arena`, and `interrogate` can run several models in parallel. Each model run uses the subscription and token allowance of its own command-line tool.
+## Cost
-`setup-pstack` lets you choose the models, one requested effort per model family, and how many run in parallel. A model from the app you are using runs inside that app. Other models run through their own command-line tools. Open Pstack does not quietly replace a failed model with a weaker one.
+Some workflows use one model. `architect`, `arena`, and `interrogate` run several in parallel. Subscription lanes spend that CLI's plan; gateway lanes bill per token on the lab's account. The cost playbook in [docs/USAGE.md](docs/USAGE.md#cost-playbook) shows where the gateway lanes pay off: high-volume code-writing roles on DeepSeek Flash, long-context reading on MiniMax M3, and one frontier lane plus two gateway lanes for a three-provider panel at a fraction of three subscriptions. Keep `judgment and prose` and `hardest tasks` on your strongest lane; they are the last roles to economize. pstack-flex never replaces a failed model with a cheaper one; a lane that fails is reported, not swapped.
## Claude Code and Codex
@@ -141,38 +169,27 @@ Both apps read the same pstack skills. Only the way they start those skills and
| | Claude Code | Codex |
| --- | --- | --- |
| Start poteto-mode | Claude loads a small startup instruction that can route non-trivial work into it. You can also run `/pstack:poteto-mode` yourself. | Ask for `pstack:poteto-mode` by name. Codex does not load the Claude startup instruction. |
-| Runs inside the app | Claude models stay inside Claude Code. | The Sol model stays inside Codex. |
-| Other models | Codex and Grok run through their signed-in command-line tools. | Claude and Grok run through their signed-in command-line tools. |
+| Runs inside the app | Claude models stay inside Claude Code. | The Codex families stay inside Codex. |
+| Other models | The Codex families and Grok run through their signed-in command-line tools. | Claude and Grok run through their signed-in command-line tools. |
+| Gateway models | DeepSeek and MiniMax always run through the external runner with an isolated config directory, never as a native agent. | Same. |
| Skills and workflows | Shared with Codex. | Shared with Claude Code. |
-Grok can take part in a multi-model review. You cannot use Grok as the main app running pstack.
+Grok, DeepSeek, and MiniMax can take part in a multi-model review. You cannot use any of them as the main app running pstack.
-## Learn from the original
+## Upstream
Lauren's [pstack guide](https://github.com/cursor/plugins/tree/main/pstack/docs/guide) walks through a real task, verification, and longer unattended runs. It uses Cursor's interface, but the ideas are the same. Use the translated skill invocations above in Claude Code or Codex.
-This repository also keeps:
-
-- [the original README](README-UPSTREAM.md), unchanged;
-- [the technical reference](docs/reference.md) for every skill, dependency, and Claude Code or Codex detail;
-- [the upstream sync record](UPSTREAM.md) and update process;
-- [the change record](CHANGES.md) for every adaptation; and
-- [the attribution record](NOTICE.md) for pstack and the imported Cursor Team Kit skills.
-
-## Staying close to Lauren's pstack
-
-Open Pstack 1.4.1 tracks pstack 0.15.1 at Cursor commit [`f8abeddd1862dc73704e3d719dd73df0d51b8c71`](https://github.com/cursor/plugins/commit/f8abeddd1862dc73704e3d719dd73df0d51b8c71).
-
-The two projects have separate version numbers. The pstack version identifies Lauren's upstream content. The Open Pstack version identifies the Claude Code and Codex package built from it.
+This repository tracks two upstreams. [UPSTREAM.md](UPSTREAM.md) records the Cursor pstack commit open-pstack imported (0.15.1 at [`f8abedd`](https://github.com/cursor/plugins/commit/f8abeddd1862dc73704e3d719dd73df0d51b8c71)) and how new pstack releases are brought over. [UPSTREAM-FLEX.md](UPSTREAM-FLEX.md) records the open-pstack fork point (v1.4.1), which files this fork owns, and the merge procedure. The fork keeps every upstream skill body as-is except the default model descriptors; its own changes are the model matrix, the first-run sheet, the gateway providers in the runner, setup's assignment-first flow, and the docs.
-In this repository, “upstream” means Lauren's original pstack. Open Pstack does not promise instant updates. It records the exact version it follows, reviews new changes in order, and changes only what Claude Code and Codex require. New pstack behavior belongs in Lauren's project first whenever possible.
+Also kept here: [the original README](README-UPSTREAM.md), unchanged; [the technical reference](docs/reference.md) for every skill and harness detail; [the change record](CHANGES.md); and [the attribution record](NOTICE.md).
## Contributing
-Fixes for Claude Code or Codex and help bringing over new pstack releases are welcome. Search [GitHub Issues](https://github.com/ericlitman/open-pstack/issues) before opening a new issue. For larger behavior changes, explain why the change belongs in Open Pstack instead of Lauren's original project.
+Fixes for Claude Code or Codex, new lanes, and help bringing over new pstack releases are welcome. Search this repository's [GitHub Issues](https://github.com/thisguymartin/pstack-flex/issues) before opening a new one. For changes to upstream-derived content, explain why the change belongs here instead of in open-pstack or Lauren's original project.
-Read [UPSTREAM.md](UPSTREAM.md) before changing content brought over from Lauren's pstack. Pull requests must keep one shared skill tree for Claude Code and Codex and pass the repository's tests, type checks, plugin validation, and static checks.
+Read [UPSTREAM.md](UPSTREAM.md) and [UPSTREAM-FLEX.md](UPSTREAM-FLEX.md) before changing content brought over from either upstream. Pull requests must keep one shared skill tree for Claude Code and Codex and pass the repository's tests, type checks, plugin validation, and static checks. Nothing merges until the exact candidate is installed and the changed behavior passes a live test from the real user surface in every affected harness; the [pull request template](.github/pull_request_template.md) records that evidence, and a PR without it stays a draft. Adding a gateway provider has its own checklist in [docs/LANES.md](docs/LANES.md#adding-a-gateway-provider).
## License
-MIT. pstack was created by Lauren Tan. Open Pstack builds on Michael Denyer's [pstack-claude](https://github.com/michael-denyer/pstack-claude) port and includes attributed MIT-licensed work from Cursor Team Kit and Superpowers. See [NOTICE.md](NOTICE.md) and the preserved license files for details.
+MIT. pstack was created by Lauren Tan. open-pstack builds on Michael Denyer's [pstack-claude](https://github.com/michael-denyer/pstack-claude) port and includes attributed MIT-licensed work from Cursor Team Kit and Superpowers. pstack-flex is a fork of open-pstack. See [NOTICE.md](NOTICE.md) and the preserved license files for details.
diff --git a/UPSTREAM-FLEX.md b/UPSTREAM-FLEX.md
index 7ecbf211..9d4b234d 100644
--- a/UPSTREAM-FLEX.md
+++ b/UPSTREAM-FLEX.md
@@ -22,11 +22,12 @@ All flex changes are additive and live in port-owned files so upstream merges st
- `plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts` and `flex-providers.test.ts` (new)
- Gateway-provider hooks in `runner/{types,commands,run,parse-output,cli}.ts` and their tests
-- The "Additional model matrix" and "Flex model matrix" sections and route-table columns in `references/provider-dispatch.md`
+- The three GPT-6 rows in the stock model matrix, the "Default panel" section, the "Flex model matrix" section, and the route-table columns in `references/provider-dispatch.md`
+- The first-run sheet in `skills/setup-pstack/SKILL.md` and the default descriptors named in `arena`, `architect`, `interrogate`, `how`, `swarm`, and the `feature`, `refactoring`, `bug-fix`, `perf-issue`, and `hillclimb` playbooks
- The assignment-first restructure of `skills/setup-pstack/SKILL.md`
- `docs/LANES.md`, this file, the README fork section, and the NOTICE/LICENSE/CHANGES additions
-The stock model matrix, the first-run sheet, every upstream skill body, and the static quad invariants are byte-unchanged. Flex rows and setup instructions are additive.
+Every upstream skill body is byte-unchanged except for the default-descriptor mentions listed above. Since [#17](https://github.com/thisguymartin/pstack-flex/issues/17), the stock matrix carries three fork-owned GPT-6 rows and the first-run sheet uses them, so those two surfaces conflict on every upstream sync and are resolved by hand: keep the fork's rows and defaults, take upstream's wording for everything else.
## Merge procedure
@@ -41,6 +42,6 @@ Expected conflict surface on future upstream releases:
- `plugins/pstack/skills/setup-pstack/SKILL.md` — upstream issue #88 (1.5.0, syncing Cursor pstack 0.15.5) folds upstream PR #73, which moves setup to the same assignment-first, probe-only-assigned shape this fork already uses. Resolve toward upstream's wording wherever it covers the same rule; keep the flex families and the diversity rule.
- `plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts` — upstream 1.5.0 changes the stock panel to three lanes. Take upstream's stock assertions verbatim; the flex-matrix describe block is fork-owned and should survive as-is.
-- `plugins/pstack/skills/poteto-mode/references/provider-dispatch.md` — stock matrix and default-panel prose are upstream's; the flex section is fork-owned.
+- `plugins/pstack/skills/poteto-mode/references/provider-dispatch.md` — the four upstream rows and their prose are upstream's; the GPT-6 rows, the "Default panel" section, and the flex section are fork-owned. Upstream's own panel changes land in the "Default panel" line only if the fork wants them.
After every merge: run the full local gate (`bun install --frozen-lockfile`, `bun run test`, `bun run typecheck`, manifest JSON parse, `PSTACK_STATIC_ONLY=1 bash tests/skill-collision-repro.sh`), then record the installed version, action, and observed result for each affected harness in the pull request before tagging.
diff --git a/UPSTREAM.md b/UPSTREAM.md
index 6d725c48..71b5b08d 100644
--- a/UPSTREAM.md
+++ b/UPSTREAM.md
@@ -18,7 +18,7 @@ The table above is the current Cursor sync point. Open Pstack 1.4.1 imports this
- Commits `799151d` and `6fecddb` add and relocate `make-bot-ui`. It depends on Cursor routines, webhook events, and UI primitives that Claude Code and Codex do not share.
- Four `disable-model-invocation: true` lines from `73f8be4` are not applied to `how`, `why`, `unslop`, or `typescript-best-practices`. Poteto-mode invokes those skills by name, and the flag blocks that route on Claude Code.
-- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on `codex:gpt-5.6-sol@max` for cost.
+- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on a Codex model (`codex:gpt-6-sol@high` in pstack-flex) for cost.
- The Claude manifest does not take the logo field from `efa2a53` because Claude Code has no schema for it. The shared asset is exposed through the Codex manifest instead.
## Check for changes
diff --git a/docs/LANES.md b/docs/LANES.md
index cf725498..9cd07efb 100644
--- a/docs/LANES.md
+++ b/docs/LANES.md
@@ -8,22 +8,22 @@ Prices and endpoints below were verified 2026-09-25 and drift. Re-verify against
| Kind | Lanes | Auth | Billing | Route |
| --- | --- | --- | --- | --- |
-| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-5.6-sol`, `grok:grok-4.6`; optional `codex:gpt-6-astra`, `codex:gpt-6-sol`, `codex:gpt-6-luna` | each CLI's own login | that CLI's plan | native or external per the route table |
+| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-6-astra`, `codex:gpt-6-sol`, `codex:gpt-6-luna`, `codex:gpt-5.6-sol`, `grok:grok-4.6` | each CLI's own login | that CLI's plan | native or external per the route table |
| Gateway (flex) | DeepSeek Flash / V4 Pro; MiniMax M3 / M3.1 Flash Preview | API key in the environment | provider billing; preview requires Token Plan | always the external runner |
A gateway lane is the stock `claude` binary env-pointed at the lab's Anthropic-compatible endpoint. There is no custom agent loop and no separate harness: the same runner that spawns Codex and Grok lanes spawns gateway lanes with injected environment. Both labs document this Claude Code setup themselves (DeepSeek: `deepseek-ai/awesome-deepseek-agent`, `docs/claude_code.md`; MiniMax: platform.minimax.io, Claude Code guide).
-## Optional GPT-6 Codex families
+## GPT-6 Codex families
-The additional model matrix adds three Codex families. They use the same ChatGPT login as `codex:gpt-5.6-sol`:
+Three stock Codex families carry the first-run defaults. They use the same ChatGPT login as `codex:gpt-5.6-sol`:
-| Family | Descriptor at default requested effort | Codex's own description |
-| --- | --- | --- |
-| astra | `codex:gpt-6-astra@high` | Frontier tier for the most demanding work |
-| sol-6 | `codex:gpt-6-sol@high` | Coding and everyday workhorse |
-| luna | `codex:gpt-6-luna@high` | Fast, low-cost tier for easier tasks |
+| Family | Descriptor at default requested effort | First-run roles | Codex's own description |
+| --- | --- | --- | --- |
+| astra | `codex:gpt-6-astra@high` | every panel (`arena runners`, `arena cross-judge pool`, `architect runners`, `interrogate reviewers`) | Frontier tier for the most demanding work |
+| sol-6 | `codex:gpt-6-sol@high` | `feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb` | Coding and everyday workhorse |
+| luna | `codex:gpt-6-luna@high` | `how explorer`, `swarm workers` | Fast, low-cost tier for easier tasks |
-These are choices, not replacements. The first-run sheet still uses `codex:gpt-5.6-sol@max`. An existing sheet changes only when you assign a role to a GPT-6 family in `/setup-pstack`. The `sol-6` family is separate from the stock `sol` family, so each keeps its own effort. All Codex families count as one provider for panel diversity, so Astra plus Sol does not satisfy the two-provider rule. The route matches Sol: native `spawn_agent` in a Codex parent, and the external runner (`codex exec`) in a Claude Code parent. Codex also lists an `ultra` effort for Astra and GPT-6 Sol. It is outside the pstack effort universe and is not selectable. The descriptions and effort lists come from the Codex CLI 0.157.1 model list, checked 2026-09-27.
+A fresh `/setup-pstack` run proposes these. An existing sheet keeps its assignments until you change a named role in setup; `codex:gpt-5.6-sol` remains a selectable family for that. The `sol-6` family is separate from the `sol` family, so each keeps its own effort. All Codex families count as one provider for panel diversity, so Astra plus GPT-6 Sol does not satisfy the two-provider rule; the default panel spans Claude, Codex, and Grok. The route matches Sol: native `spawn_agent` in a Codex parent, and the external runner (`codex exec`) in a Claude Code parent. Codex also lists an `ultra` effort for Astra and GPT-6 Sol. It is outside the pstack effort universe and is not selectable. The descriptions and effort lists come from the Codex CLI 0.157.1 model list, checked 2026-09-27.
## Multiple models per provider
diff --git a/docs/USAGE.md b/docs/USAGE.md
index 9819c964..5a49334a 100644
--- a/docs/USAGE.md
+++ b/docs/USAGE.md
@@ -19,7 +19,7 @@ flowchart TD
PB --> S[Skills fire per step
how, tdd, interrogate, arena, ...]
S --> F{Lane fan-out}
F --> N1["claude:fable / claude:opus
native Agent (Claude sub)"]
- F --> N2["codex:gpt-5.6-sol
native or codex CLI (ChatGPT sub)"]
+ F --> N2["codex:gpt-6-astra / gpt-6-sol / gpt-6-luna
native or codex CLI (ChatGPT sub)"]
F --> N3["grok:grok-4.6
grok CLI (Grok sub)"]
F --> G1["deepseek:deepseek-flash
runner + env -> DeepSeek API (key)"]
F --> G2["minimax:MiniMax-M3
runner + env -> MiniMax API (key)"]
@@ -83,10 +83,14 @@ Daily flow: `pstack-keys -> claude -> /pstack:poteto-mode`. Alternatives, the th
Setup is assignment-first: pick which roles run on which families, answer one effort question per **assigned** family, and only assigned families get probed. Unassigned families are skipped, not errors. Every probe is a real one-turn run — a failed probe writes nothing. Three configurations that make sense:
-**A. Full frontier** (Claude + ChatGPT + Grok subs) — accept the defaults; behaves exactly like stock upstream:
+**A. Full frontier** (Claude + ChatGPT + Grok subs) — accept the defaults. GPT-6 Sol writes code, Luna explores and verifies, and the panel spans three providers:
```text
-arena runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh
+feature, refactoring: codex:gpt-6-sol@high
+bug-fix: codex:gpt-6-sol@high
+how explorer: codex:gpt-6-luna@high
+swarm workers: codex:gpt-6-luna@high
+arena runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh
```
**B. Hybrid saver** (Claude sub + two API keys) — frontier judgment, cheap volume:
@@ -232,14 +236,14 @@ Every external lane writes a JSON receipt next to its output. The fields that ma
| A panel ran with fewer lanes than configured | a lane dropped out with a named receipt | read that receipt; pstack proceeds N-1 and never silently substitutes a model |
| Everything gateway broke after a claude CLI update | Anthropic doesn't support third-party endpoints; compatibility can shift | pin the CLI version on machines that depend on gateway lanes; see [LANES.md](LANES.md#safety-and-policy) |
-## Selecting the GPT-6 Codex models
+## The GPT-6 Codex models
-Run `/setup-pstack` and assign `astra` (`codex:gpt-6-astra@high`), `sol-6` (`codex:gpt-6-sol@high`), or `luna` (`codex:gpt-6-luna@high`) to named roles, such as `architect runners`. Every role you do not change keeps its current descriptor, including the stock `codex:gpt-5.6-sol@max` defaults. Each GPT-6 family gets its own effort question and live probe. They need only your Codex login. See [optional GPT-6 Codex families](LANES.md#optional-gpt-6-codex-families).
+`astra` (`codex:gpt-6-astra@high`), `sol-6` (`codex:gpt-6-sol@high`), and `luna` (`codex:gpt-6-luna@high`) are stock families and the first-run defaults: Astra on every panel, GPT-6 Sol on the solo code-writing roles, Luna on exploration and swarm work. They need only your Codex login. Each gets its own effort question and live probe. See [GPT-6 Codex families](LANES.md#gpt-6-codex-families).
-For example, this row puts Astra on the architect panel and keeps the other providers:
+A sheet written before this release keeps its assignments. To move a role, run `/setup-pstack` and name it; every role you do not change keeps its descriptor, and `codex:gpt-5.6-sol` stays selectable. For example, this row keeps GPT-5.6 Sol on bug fixes while the rest of the sheet takes the new defaults:
```text
-architect runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh
+bug-fix: codex:gpt-5.6-sol@max
```
## Selecting the additional gateway models
diff --git a/docs/reference.md b/docs/reference.md
index f070eb3e..5b7ef3de 100644
--- a/docs/reference.md
+++ b/docs/reference.md
@@ -86,7 +86,7 @@ The Codex build shares one `skills/` tree with the Claude Code build. Nothing is
- **Tool and built-in mapping.** Claude tool names and built-in skills resolve through [`codex-tools.md`](../plugins/pstack/skills/poteto-mode/references/codex-tools.md). Model execution resolves separately through [`provider-dispatch.md`](../plugins/pstack/skills/poteto-mode/references/provider-dispatch.md), so Codex can keep Sol native while invoking Claude and Grok externally.
- **Subagents.** The `Agent` tool maps to Codex `spawn_agent` / `wait_agent`, enabled by `multi_agent = true`. Parallel fan-out is multiple `spawn_agent` calls in one turn. If the native Codex lane is unavailable, record that lane as a dropout; external Claude and Grok lanes still run, and no provider is silently substituted. There is no `poteto-agent` subagent type on Codex; route ad-hoc subagents by dispatching a `spawn_agent` told to read `poteto-mode` first.
- **Auto-fire.** The `hooks/` SessionStart injection is Claude Code-only; Codex has no plugin hook runtime. Enter `pstack:poteto-mode` by name, or add a standing instruction to `~/.codex/AGENTS.md` if you want the same always-on routing.
-- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per frontier family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-5.6 Sol max, Grok 4.6 xhigh, and Opus xhigh. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. Optional GPT-6 Astra, Sol, and Luna Codex families can be assigned to any role and never change the first-run sheet. In Codex, Sol and the GPT-6 families use native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; Sol, the GPT-6 families, and Grok use the runner. Children never detect the parent or reroute themselves. The `bug-fix`, `perf-issue`, and `hillclimb` roles stay on GPT-5.6 Sol max instead of upstream's Fable default because Sol costs less for these frequent delegated code roles.
+- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per frontier family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-6 Astra high, Grok 4.6 xhigh, and Opus xhigh. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. The GPT-6 Astra, Sol, and Luna Codex families are stock: GPT-6 Sol high carries `feature, refactoring`, `bug-fix`, `perf-issue`, and `hillclimb`; Luna high carries `how explorer` and `swarm workers`; Astra high sits on every panel. GPT-5.6 Sol remains a selectable family. In Codex, every Codex family uses native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; the Codex families and Grok use the runner. Children never detect the parent or reroute themselves. The solo code roles stay on a Codex model instead of upstream's Fable default because it costs less for these frequent delegated code roles.
Verified in fresh installed Claude Code and Codex sessions: the user-facing skills are discovered and namespaced under `pstack`; both parents fan out the frontier quad through the documented native/external route table, retain long-running handles without a default timeout, and cross-judge only after every candidate is terminal. The `principle-*` leaves remain available for `poteto-mode` to read by path. Claude honors their `user-invocable: false` metadata; Codex 0.149.0 does not ([#8](https://github.com/ericlitman/open-pstack/issues/8)).
@@ -193,7 +193,7 @@ The port is editorial, not mechanical. Anywhere upstream pstack assumed Cursor-s
| Cursor's `/goal` (standing objective across turns) | The program objective written into the run's standing orders and restated in the todolist |
| The Cursor agent store (path in the system prompt) | `~/.claude/orchestrate//`, which survives the session restarts a multi-day program expects |
| Model rule `~/.cursor/rules/pstack-models.mdc` | Override sheet `~/.claude/pstack-models.md`, included from `CLAUDE.md` |
-| Multi-model panels (arena, architect, interrogate) | Provider dispatch restores the upstream frontier quad: `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Same-provider lanes stay native; external lanes use the bundled runner. |
+| Multi-model panels (arena, architect, interrogate) | Provider dispatch owns the default panel: `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Same-provider lanes stay native; external lanes use the bundled runner. |
### Cross-vendor dispatch
@@ -211,7 +211,7 @@ The earlier port collapsed panels to Claude-only models. The bundled runner rest
- **`automations/benny/`** (upstream `0452e08`, the only pstack change between `e46364b` and v0.10.0) — a dormant Slack issue-triage and reproduce-and-fix automation pack built on Cursor's event-triggered automations. It registers no slash skills even upstream, so excluding it changes nothing about the ported plugin's behavior. Porting it would require Cursor's event-trigger runtime, Slack, and tracker plumbing that Open Pstack does not provide.
- **`docs/guide/`** (upstream `02c03a9`, `0b7ef5b`, `424829e`) — the ten-chapter usage tutorial and its six screenshots (2.3 MB). It teaches pstack through Cursor's UI, sticky mode, and cloud agents, so a faithful port would be a rewrite rather than a sync, and none of it ships as skill content. Read it upstream at [cursor/plugins/pstack/docs/guide](https://github.com/cursor/plugins/tree/main/pstack/docs/guide); the concepts map through the substitution table above.
- **`make-bot-ui`** (upstream `799151d`, relocated by `6fecddb`) uses Cursor routines, webhook events, hosted bot state, and Cursor UI primitives that have no shared Claude Code and Codex mapping. A provider-specific rewrite would be a separate feature, not an upstream sync.
-- **Fable solo code defaults** (upstream `23a56e2`) move `bug-fix`, `perf-issue`, and `hillclimb` from GPT-5.6 Sol to Fable. Open Pstack keeps these frequent delegated code roles on `codex:gpt-5.6-sol@max` because Fable costs much more per task.
+- **Fable solo code defaults** (upstream `23a56e2`) move `bug-fix`, `perf-issue`, and `hillclimb` from GPT-5.6 Sol to Fable. pstack-flex keeps these frequent delegated code roles on a Codex model (`codex:gpt-6-sol@high`) because Fable costs much more per task.
- **Sticky mode** (upstream `#144`) — Cursor-only `mode`/`icon`/`color`/`reminder` frontmatter with no Claude Code equivalent. The port's 0.9.5 SessionStart hook is the analog and already carries the non-trivial / trivial / opt-out logic.
- **`is_background: true` on `poteto-agent`** (upstream `99559f2`) — Cursor names this key differently. Claude-native frontier definitions use `background: true`; ad-hoc `poteto-agent` calls remain background dispatches at the call site.
- **`cursor-team-kit` beyond the seven imported skills** — the rest either duplicate Claude Code built-ins (`verify-this` → the `verify` skill and built-in verification discipline; `check-compiler-errors` → LSP diagnostics; `control-cli`/`control-ui` → `run`/`verify`, already the substitution targets) or overlap skills this port ships (`loop-on-ci`, `review-and-ship`, `weekly-review` vs `babysit`, `fix-ci`, `make-pr-easy-to-review`, `what-did-i-get-done`). `pr-review-canvas` is Cursor-UI-specific.
diff --git a/plugins/pstack/skills/architect/SKILL.md b/plugins/pstack/skills/architect/SKILL.md
index f6893207..00031bc7 100644
--- a/plugins/pstack/skills/architect/SKILL.md
+++ b/plugins/pstack/skills/architect/SKILL.md
@@ -31,7 +31,7 @@ Skip Phase A only when the work is genuinely greenfield with no surrounding syst
Run the **arena** skill with the design-sketch task and the Phase A grounding artifacts. Pass `references/runner-prompt.md` as each runner's prompt. Each candidate produces a design package shaped per `references/rationale-template.md`.
-Use your configured architect runners (defaults `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`).
+Use your configured architect runners (defaults `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`).
Design it twice. Require at least two structurally distinct candidates before synthesis, even when the first looks sufficient. This is the **exhaust-the-design-space** principle skill made concrete. Whole-shape alternatives, not point fixes inside one shape.
diff --git a/plugins/pstack/skills/arena/SKILL.md b/plugins/pstack/skills/arena/SKILL.md
index 0e3b5f38..e03c7ae6 100644
--- a/plugins/pstack/skills/arena/SKILL.md
+++ b/plugins/pstack/skills/arena/SKILL.md
@@ -26,7 +26,7 @@ The N candidates will receive the same prompt, so the prompt is the contract.
1. State the artifact each candidate is producing.
2. Derive the rubric. State what success looks like for *this* task, then turn it into 3-6 concrete gradeable criteria. The rubric is the picker's tool in Phase D. Candidates only see the task.
-3. Pick the runners. Use `arena runners` from the current harness's pstack model sheet when present. Otherwise default to one each on `claude:fable@max`, `codex:gpt-5.6-sol@max`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Spawn more when the arena covers multiple design directions. Same descriptor N times when the work is generation-bound rather than judgment-sensitive.
+3. Pick the runners. Use `arena runners` from the current harness's pstack model sheet when present. Otherwise default to one each on `claude:fable@max`, `codex:gpt-6-astra@high`, `grok:grok-4.6@xhigh`, `claude:opus@xhigh`. Spawn more when the arena covers multiple design directions. Same descriptor N times when the work is generation-bound rather than judgment-sensitive.
4. Assign output paths. Each candidate writes to its own location (a git worktree where possible, otherwise `/tmp/arena-/candidate-/`), per the **separate-before-serializing-shared-state** principle skill.
## Phase B: Fan out
diff --git a/plugins/pstack/skills/how/SKILL.md b/plugins/pstack/skills/how/SKILL.md
index 84803c9d..7022727d 100644
--- a/plugins/pstack/skills/how/SKILL.md
+++ b/plugins/pstack/skills/how/SKILL.md
@@ -20,7 +20,7 @@ When in doubt, take the simple path.
## Step 2a. Explore (complex questions only)
-Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Start all explorers in one fan-out phase through provider dispatch. Use your configured how-explorer descriptor (default `grok:grok-4.6@xhigh`) in `read-only` mode. A native lane uses the parent subagent primitive; an external lane uses the launcher directly.
+Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Start all explorers in one fan-out phase through provider dispatch. Use your configured how-explorer descriptor (default `codex:gpt-6-luna@high`) in `read-only` mode. A native lane uses the parent subagent primitive; an external lane uses the launcher directly.
Each explorer gets the prompt in `references/explorer-prompt.md` with its angle filled in. Then go to Step 3.
diff --git a/plugins/pstack/skills/interrogate/SKILL.md b/plugins/pstack/skills/interrogate/SKILL.md
index ee7514ff..5f388930 100644
--- a/plugins/pstack/skills/interrogate/SKILL.md
+++ b/plugins/pstack/skills/interrogate/SKILL.md
@@ -39,7 +39,7 @@ Start all reviewers in one fan-out phase. Use `interrogate reviewers` from the c
| Subagent | Default model |
|----------|---------------|
| Reviewer A | `claude:fable@max` |
-| Reviewer B | `codex:gpt-5.6-sol@max` |
+| Reviewer B | `codex:gpt-6-astra@high` |
| Reviewer C | `grok:grok-4.6@xhigh` |
| Reviewer D | `claude:opus@xhigh` |
diff --git a/plugins/pstack/skills/poteto-mode/SKILL.md b/plugins/pstack/skills/poteto-mode/SKILL.md
index 5e51bfc6..67aa271a 100644
--- a/plugins/pstack/skills/poteto-mode/SKILL.md
+++ b/plugins/pstack/skills/poteto-mode/SKILL.md
@@ -89,7 +89,7 @@ Read the leaf skill in full for any principle you apply. Each entry names when i
For `inherit-parent`, `auto`, or an unconfigured native ad-hoc helper, prefer `poteto-agent`. `/poteto-mode` and `poteto-agent` route through the same wrapper. A provider-qualified role instead follows provider dispatch: Claude's shipped frontier agent definitions select the model alias and requested effort, Codex passes both to `spawn_agent`, and external providers run through the deterministic launcher. Routed workflow skills set the task and access mode. Do not override their choices.
-**Defaults for every delegation.** Start independent lanes together, use file pointers rather than inlined dumps, preserve only the tools or MCPs the task needs, and assign every writer a worktree or unique output directory. `/setup-pstack` configures the descriptor per role. Upstream defaults use Grok 4.6 xhigh for feature/refactoring, exploration, and swarm work; GPT-5.6 Sol max for bug fixes, performance work, hillclimbing, and tooling review; Fable max for judgment, prose, explanation, synthesis, and hardest tasks; and the four-provider frontier panel for model-diverse judgment. The panel defaults are enumerated in `arena`, `architect`, and `interrogate`. `inherit-parent` and `auto` use the parent model natively and reduce provider diversity when used in a panel.
+**Defaults for every delegation.** Start independent lanes together, use file pointers rather than inlined dumps, preserve only the tools or MCPs the task needs, and assign every writer a worktree or unique output directory. `/setup-pstack` configures the descriptor per role. First-run defaults use GPT-6 Sol high for feature/refactoring, bug fixes, performance work, and hillclimbing; GPT-6 Luna high for exploration and swarm work; Fable max for judgment, prose, explanation, synthesis, and hardest tasks; and the Fable, GPT-6 Astra, Grok 4.6, Opus panel for model-diverse judgment. The panel defaults are enumerated in `arena`, `architect`, and `interrogate`. `inherit-parent` and `auto` use the parent model natively and reduce provider diversity when used in a panel.
You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal.
diff --git a/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md b/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md
index 18412149..ea0bc5d9 100644
--- a/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md
+++ b/plugins/pstack/skills/poteto-mode/playbooks/bug-fix.md
@@ -6,7 +6,7 @@ Be scientific. Every shipped line traces to runtime evidence. Belt-and-suspender
1. Reproduce it yourself on the matching surface via the driver skill (`run` for CLIs/TUIs, `verify` for UIs) (Non-negotiables). Don't hand the repro to the user. A debug or instrumentation protocol that says to ask the user does not override this. You drive the instrumented runtime. Ask the user only with a stated, specific reason the control surface cannot reach the target, and only after driving it as far as it goes. Won't reproduce directly, force it: synthesize the trigger, tighten conditions, or instrument until it fires.
2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with Claude Code's `loop` skill. Confirm the surviving *mechanism* with runtime evidence before the step-3 architect/interrogate fan-out.
-3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured bug-fix descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write`, a dedicated worktree, and a specific scope. Review the diff.
+3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured bug-fix descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope. Review the diff.
4. Verify on the same surface. The original repro now passes. "Inconclusive" or wrong-surface is not a pass. Flag it. Unit tests show branch behavior, not bug absence.
5. Stage the commits so the failing repro lands before the fix in git history. See the **tdd** skill for the failing-test-first cadence when the bug has a cheap local test path. Skip it when the test would be expensive, integration-heavy, or unclear.
This is the canonical **sequence-verifiable-units** principle skill, the failing test first and the fix on top.
diff --git a/plugins/pstack/skills/poteto-mode/playbooks/feature.md b/plugins/pstack/skills/poteto-mode/playbooks/feature.md
index 1cc22990..ab11a919 100644
--- a/plugins/pstack/skills/poteto-mode/playbooks/feature.md
+++ b/plugins/pstack/skills/poteto-mode/playbooks/feature.md
@@ -9,7 +9,7 @@
- **Independent workstreams.** Disjoint files, services, or layers parallelize. Shared writes serialize.
- **Shared mutable state.** Default to splitting the target (the **separate-before-serializing-shared-state** principle skill). Serialize only for real invariants.
- **Smallest safe decomposition.** If one worker is best, name why.
-4. Delegate code-writing through provider dispatch using your configured feature descriptor (default `grok:grok-4.6@xhigh`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). Review its diff yourself. When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). The delegate owns the diff directly and never waits on or launches a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally.
+4. Delegate code-writing through provider dispatch using your configured feature descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). Review its diff yourself. When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). The delegate owns the diff directly and never waits on or launches a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally.
5. Verify on the matching surface. "Inconclusive" or wrong-surface is not a pass. Flag it.
6. Rebase into small, ordered commits. Stack follow-ups.
Use the **sequence-verifiable-units** principle skill, building, verifying, and committing each small unit before the next.
diff --git a/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md b/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md
index c38089b7..bdde9697 100644
--- a/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md
+++ b/plugins/pstack/skills/poteto-mode/playbooks/hillclimb.md
@@ -9,7 +9,7 @@ Core discipline: one change, one measurement, keep or revert. Never stack untest
3. Open the decision log via the **show-me-your-work** skill. A `decision.tsv`, one row per attempt: id, hypothesis, change, before, after, delta, tests, verdict (kept or reverted), note. Read it before each attempt. Keep it out of the tree (gitignored).
4. Ground each hypothesis in the architecture model from step 1, so it names a specific mechanism ("defer X off the boot path because it blocks first paint"), not "try memoizing something".
5. Loop, one hypothesis per iteration:
- - Hand the change through provider dispatch using your configured hillclimb descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write` and a tight worktree scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel lanes, each in its own worktree (the **separate-before-serializing-shared-state** principle skill).
+ - Hand the change through provider dispatch using your configured hillclimb descriptor (default `codex:gpt-6-sol@high`) with `isolated-write` and a tight worktree scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel lanes, each in its own worktree (the **separate-before-serializing-shared-state** principle skill).
- Measure before and after with the frozen harness, and run the regression gate.
- Accept only when the metric moves past noise and the gate stays green. Otherwise revert the change in full. A tweak that "might help" is not kept.
- One commit per accepted fix, staging only the files you changed (`git add `, never `-A`). Log the row either way, kept or reverted.
diff --git a/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md b/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md
index e08ac677..0f755d3e 100644
--- a/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md
+++ b/plugins/pstack/skills/poteto-mode/playbooks/perf-issue.md
@@ -13,7 +13,7 @@
- **Redundancy.** The wait hangs on one slow instance or attempt. Duplicate the work (replicas, hedged requests, speculative execution) and take the fastest result. The trace has to show the wait dominates and the system has headroom.
- **Lazy evaluation.** Cost lands on results that are never used or not needed yet (eager init on the boot path, rendering offscreen items). Defer the work until first use.
- **Scheduling.** The work must happen, but not during the interactive moment. Move it to where nobody is waiting: idle callbacks, a background warmup after boot, precompute before the user arrives, cleanup after the frame commits. The win is perceived latency, so measure the interactive path, not total work done.
-3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured perf-issue descriptor (default `codex:gpt-5.6-sol@max`) with `isolated-write` in a dedicated worktree. Review the diff. Capture a post-fix trace.
+3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation through provider dispatch using your configured perf-issue descriptor (default `codex:gpt-6-sol@high`) with `isolated-write` in a dedicated worktree. Review the diff. Capture a post-fix trace.
Apply the **sequence-verifiable-units** principle skill, verifying each attempt before trying the next.
4. Parse and compare the artifacts (JSON to sqlite, diff). "Inconclusive" or wrong-surface is not a pass. Flag it.
5. Cite the measurement in the PR.
diff --git a/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md b/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md
index 5106c39a..ef73bce0 100644
--- a/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md
+++ b/plugins/pstack/skills/poteto-mode/playbooks/refactoring.md
@@ -8,7 +8,7 @@ If the cleanup reveals a missing feature or a real bug, split it out and ship th
2. Name the structure the code is missing per **principle-model-the-domain**. Boring code stays when the shape is already clear and local. The reshape must delete branches or invalid states, not add indirection.
3. Name the target shape. State what the module layout, types, and call graph should be if built today (**principle-foundational-thinking**, **principle-redesign-from-first-principles**). If the target crosses a function boundary, run the **architect** skill for parallel design exploration of the shape before the move.
4. Subtract before you add. Delete dead code, collapse one-caller wrappers, drop redundant validators, and remove orphan references before introducing the new shape (**principle-subtract-before-you-add**). The smallest change that reaches the target shape ships (**principle-laziness-protocol**). A speculative cleanup that "might help" gets reverted.
-5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits through provider dispatch using your configured refactoring descriptor (default `grok:grok-4.6@xhigh`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, the names being moved, the behavior to hold). Review the diff yourself.
+5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits through provider dispatch using your configured refactoring descriptor (default `codex:gpt-6-sol@high`) with `isolated-write`, a dedicated worktree, and a specific scope (file paths, the names being moved, the behavior to hold). Review the diff yourself.
6. Prove behavior is unchanged on the real artifact, not "it compiles" (**principle-prove-it-works**). For larger reshapes, run an equivalence check: a script that diffs old-vs-new outputs, a recorded baseline replayed against the new code, or a smoke run on the matching surface via the driver skill (`run` for CLIs/TUIs, `verify` for UIs). Own the verification yourself. Do not trust a delegate's "looks good" summary.
7. Confirm the change is worth keeping. The success measure is reduced reader load (**principle-minimize-reader-load**). If the diff does not lower reader load somewhere, revert it.
8. Rebase into small ordered commits. A subtraction commit, then the reshape, then any follow-on cleanup. Shape them with the **sequence-verifiable-units** principle skill, so each behavior-preserving slice stays green before the next. Run **Opening a PR**.
diff --git a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md
index 729c096c..bec620b8 100644
--- a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md
+++ b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md
@@ -14,26 +14,27 @@ pstack model choices are provider-qualified descriptors:
| sol | gpt-5.6-sol-max | codex | gpt-5.6-sol | max | low medium high xhigh max | - |
| grok | grok-4.6-fast-xhigh | grok | grok-4.6 | xhigh | low medium high xhigh max | - |
| opus | opus | claude | opus | xhigh | low medium high xhigh max | opus |
+| astra | - | codex | gpt-6-astra | high | low medium high xhigh max | - |
+| sol-6 | - | codex | gpt-6-sol | high | low medium high xhigh max | - |
+| luna | - | codex | gpt-6-luna | high | low medium high xhigh max | - |
-The allowed effort universe is exactly `low`, `medium`, `high`, `xhigh`, `max`. First-run requested efforts are the Default effort cell of each row. A Claude-native agent stem of `-` means the family has no Claude-native agent. Otherwise the shipped agent name is `pstack--`.
+The allowed effort universe is exactly `low`, `medium`, `high`, `xhigh`, `max`. First-run requested efforts are the Default effort cell of each row. A Claude-native agent stem of `-` means the family has no Claude-native agent. Otherwise the shipped agent name is `pstack--`. `-` in Upstream pstack choice means Cursor's pstack has no default for that family; the row is fork-owned.
`fable` and `opus` are Claude Code's rolling aliases. Claude resolves each alias to the latest available family revision. A runner receipt keeps the requested alias in `model` and the concrete provider-reported revision in `reportedModel`; verification accepts only a numeric `claude-fable-*` or `claude-opus-*` revision from the matching family.
-## Additional model matrix
+The `astra`, `sol-6`, and `luna` rows are the GPT-6 Codex families (pstack-flex addition). These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent. Each is its own family with its own requested effort and probe; `sol-6` is independent of `sol`, so an existing GPT-5.6 Sol assignment stays unchanged until setup reassigns the role. All Codex families count as one provider for panel diversity.
-pstack-flex additions using the existing provider routes. These optional families do not change the stock matrix or first-run role assignments. Default effort is pstack's proposed requested effort when assigning a family, not the provider's default. `-` in Upstream pstack choice means there is no upstream default to replace.
+## Default panel
-| Family | Upstream pstack choice | Provider | Model | Default effort | Selectable efforts | Claude-native agent stem |
-|---|---|---|---|---|---|---|
-| astra | - | codex | gpt-6-astra | high | low medium high xhigh max | - |
-| sol-6 | - | codex | gpt-6-sol | high | low medium high xhigh max | - |
-| luna | - | codex | gpt-6-luna | high | low medium high xhigh max | - |
+The first-run panel roles (`arena runners`, `arena cross-judge pool`, `architect runners`, `interrogate reviewers`) use these four lanes, one per entry, at each family's default effort:
+
+`claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh`
-These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent. Setup may assign them to any configurable role, including `architect runners`, after each requested model and effort passes the parent-specific probe. The `sol-6` family is independent of the stock `sol` family; existing GPT-5.6 Sol assignments stay unchanged.
+This line is the single source for the panel default. `setup-pstack`'s first-run sheet and the `arena`, `architect`, and `interrogate` skills copy it verbatim; the static invariant check fails when they drift. Solo code-writing roles (`feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb`) default to the `sol-6` row; exploration and swarm roles default to the `luna` row.
## Flex model matrix
-pstack-flex addition. The stock matrix above is upstream-owned and unchanged; these lanes are additive. A flex lane runs the stock `claude` binary env-pointed at the provider's Anthropic-compatible endpoint, with the provider's own API key and an isolated `CLAUDE_CONFIG_DIR`, so it uses no Anthropic account, no claude.ai login, and no subscription.
+pstack-flex addition. These lanes are additive. A flex lane runs the stock `claude` binary env-pointed at the provider's Anthropic-compatible endpoint, with the provider's own API key and an isolated `CLAUDE_CONFIG_DIR`, so it uses no Anthropic account, no claude.ai login, and no subscription.
| Family | Provider | Model | Default effort | Selectable efforts | API key variable | Base URL default |
|---|---|---|---|---|---|---|
diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts
index b82be28e..3cedcb2f 100644
--- a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts
+++ b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts
@@ -30,7 +30,8 @@ const MATRIX_HEADER = [
"Claude-native agent stem",
] as const;
-const FAMILY_ORDER = ["fable", "sol", "grok", "opus"] as const;
+const FAMILY_ORDER = ["fable", "sol", "grok", "opus", "astra", "sol-6", "luna"] as const;
+const GPT6_FAMILIES = ["astra", "sol-6", "luna"] as const;
const PROVIDERS = ["claude", "codex", "grok"] as const;
const DESCRIPTOR_RE =
/(claude|codex|grok):[a-z0-9.-]+@(low|medium|high|xhigh|max)/g;
@@ -182,10 +183,25 @@ function parseModelMatrix(
});
}
-function defaultDescriptors(rows: MatrixRow[]): string[] {
- return rows.map(
- (row) => `${row.provider}:${row.model}@${row.defaultEffort}`
- );
+function defaultDescriptor(row: MatrixRow): string {
+ return `${row.provider}:${row.model}@${row.defaultEffort}`;
+}
+
+function parseDefaultPanel(markdown: string): string[] {
+ const lines = markdown.split(/\r?\n/);
+ const start = lines.findIndex((line) => line.trim() === "## Default panel");
+ if (start < 0) {
+ throw new Error("missing ## Default panel");
+ }
+ for (let i = start + 1; i < lines.length; i++) {
+ if (lines[i].startsWith("## ")) {
+ break;
+ }
+ if (lines[i].startsWith("`")) {
+ return lines[i].match(DESCRIPTOR_RE) ?? [];
+ }
+ }
+ throw new Error("## Default panel has no descriptor line");
}
function parseFrontmatter(text: string): {
@@ -223,9 +239,8 @@ function firstRunSheet(setup: string): string {
describe("model matrix", () => {
const dispatch = readFileSync(DISPATCH_PATH, "utf8");
const rows = parseModelMatrix(dispatch);
- const additionalRows = parseModelMatrix(dispatch, "## Additional model matrix", 3);
const setup = readFileSync(SETUP_PATH, "utf8");
- const quad = defaultDescriptors(rows);
+ const panel = parseDefaultPanel(dispatch);
it("owns the effort universe and first-run defaults", () => {
expect([...EFFORTS]).toEqual(["low", "medium", "high", "xhigh", "max"]);
@@ -245,6 +260,9 @@ describe("model matrix", () => {
["sol", "max"],
["grok", "xhigh"],
["opus", "xhigh"],
+ ["astra", "high"],
+ ["sol-6", "high"],
+ ["luna", "high"],
]);
expect(
rows
@@ -259,7 +277,7 @@ describe("model matrix", () => {
it("ships exactly the declared Claude-native frontier agents", () => {
const expected = new Set();
const familyBodies = new Map();
- for (const row of [...rows, ...additionalRows]) {
+ for (const row of rows) {
const stem = row.claudeNativeAgentStem;
if (stem === null) {
continue;
@@ -300,47 +318,76 @@ describe("model matrix", () => {
expect(shipped).toEqual([...expected].sort());
});
- it("adds GPT-6 families without changing the stock matrix or first-run assignments", () => {
- expect(additionalRows.map((row) => [row.family, row.model])).toEqual([
+ it("ships the GPT-6 Codex families as stock rows and puts them in the first-run sheet", () => {
+ const gpt6Rows = rows.filter((row) =>
+ (GPT6_FAMILIES as readonly string[]).includes(row.family)
+ );
+ expect(gpt6Rows.map((row) => [row.family, row.model])).toEqual([
["astra", "gpt-6-astra"],
["sol-6", "gpt-6-sol"],
["luna", "gpt-6-luna"],
]);
- for (const row of additionalRows) {
+ const sheet = firstRunSheet(setup);
+ for (const row of gpt6Rows) {
expect(row.upstreamChoice).toBe("-");
expect(row.provider).toBe("codex");
expect(row.defaultEffort).toBe("high");
expect(row.selectableEfforts).toEqual([...EFFORTS]);
expect(row.claudeNativeAgentStem).toBeNull();
- expect(firstRunSheet(setup)).not.toContain(row.model);
+ expect(sheet).toContain(defaultDescriptor(row));
+ }
+ expect(new Set(rows.map((row) => row.family)).size).toBe(rows.length);
+ expect(new Set(rows.map((row) => `${row.provider}:${row.model}`)).size)
+ .toBe(rows.length);
+ // Solo code roles ride the sol-6 row; exploration and swarm ride luna.
+ const sol6 = defaultDescriptor(rows.find((row) => row.family === "sol-6")!);
+ const luna = defaultDescriptor(rows.find((row) => row.family === "luna")!);
+ for (const role of ["feature, refactoring", "bug-fix", "perf-issue", "hillclimb"]) {
+ expect(sheet).toContain(`${role}: ${sol6}\n`);
}
- const allRows = [...rows, ...additionalRows];
- expect(new Set(allRows.map((row) => row.family)).size).toBe(allRows.length);
- expect(new Set(allRows.map((row) => `${row.provider}:${row.model}`)).size)
- .toBe(allRows.length);
- expect(setup).toContain("Its model matrices (stock, additional, and flex)");
- expect(setup).toContain("Read the model matrices, stock, additional, and flex.");
- expect(setup).toContain("any stock, additional, or flex matrix family");
- expect(setup).toContain("Offer Astra, GPT-6 Sol, and Luna from the additional matrix when changing `architect runners`");
+ for (const role of ["how explorer", "swarm workers"]) {
+ expect(sheet).toContain(`${role}: ${luna}\n`);
+ }
+ expect(setup).toContain("Its model matrices (stock and flex)");
+ expect(setup).toContain("Read the model matrices, stock and flex.");
+ expect(setup).toContain("any stock or flex matrix family");
+ expect(setup).toContain("Offer every stock family, including Astra, GPT-6 Sol, and Luna, when changing `architect runners`");
expect(setup).toContain("Read each model, proposed effort, and selectable efforts from its row.");
- expect(setup).toContain("outside the stock, additional, and flex matrix families");
+ expect(setup).toContain("outside the stock and flex matrix families");
expect(setup).toContain(
- "| Astra | Astra additional row + selected effort | external runner | native `spawn_agent` |"
+ "| Astra | Astra matrix row + selected effort | external runner | native `spawn_agent` |"
);
expect(setup).toContain(
- "| GPT-6 Sol | sol-6 additional row + selected effort | external runner | native `spawn_agent` |"
+ "| GPT-6 Sol | sol-6 matrix row + selected effort | external runner | native `spawn_agent` |"
);
expect(setup).toContain(
- "| Luna | Luna additional row + selected effort | external runner | native `spawn_agent` |"
+ "| Luna | Luna matrix row + selected effort | external runner | native `spawn_agent` |"
);
expect(setup).toContain("each assigned Codex family gets a native `spawn_agent` probe");
+ expect(setup).not.toContain("additional matrix");
+ expect(dispatch).not.toContain("## Additional model matrix");
expect(dispatch).toContain(
"These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent."
);
});
- it("passes each additional family's selected model and effort to the existing runner", () => {
- for (const row of additionalRows) {
+ it("owns the default panel: four lanes, three providers, matrix default efforts", () => {
+ expect(panel).toEqual([
+ "claude:fable@max",
+ "codex:gpt-6-astra@high",
+ "grok:grok-4.6@xhigh",
+ "claude:opus@xhigh",
+ ]);
+ const byDescriptor = new Set(rows.map(defaultDescriptor));
+ for (const descriptor of panel) {
+ expect(byDescriptor.has(descriptor)).toBe(true);
+ }
+ const providers = new Set(panel.map((descriptor) => descriptor.split(":")[0]));
+ expect(providers.size).toBeGreaterThanOrEqual(2);
+ });
+
+ it("passes each GPT-6 family's selected model and effort to the existing runner", () => {
+ for (const row of rows.filter((row) => (GPT6_FAMILIES as readonly string[]).includes(row.family))) {
for (const effort of row.selectableEfforts) {
const options = parseArgs([
"--parent", "claude",
@@ -390,7 +437,7 @@ describe("model matrix", () => {
}
expect(effort).toBe(row.defaultEffort);
}
- const expectedPanel = quad.join(", ");
+ const expectedPanel = panel.join(", ");
for (const role of PANEL_ROLES) {
const line = sheet
.split("\n")
diff --git a/plugins/pstack/skills/setup-pstack/SKILL.md b/plugins/pstack/skills/setup-pstack/SKILL.md
index 737452c7..5290a7a7 100644
--- a/plugins/pstack/skills/setup-pstack/SKILL.md
+++ b/plugins/pstack/skills/setup-pstack/SKILL.md
@@ -5,7 +5,7 @@ description: Configure pstack's provider-qualified models, per-family requested
# Setup pstack
-Configure one portable model sheet for the current parent harness. Read [`provider-dispatch.md`](../poteto-mode/references/provider-dispatch.md) before probing or writing anything. Its model matrices (stock, additional, and flex), descriptor grammar, and route table are the contract. Role assignments are selected first; then choose one requested effort per assigned matrix family. Do not add a second configuration file, a runtime resolver, or a weaker-model fallback.
+Configure one portable model sheet for the current parent harness. Read [`provider-dispatch.md`](../poteto-mode/references/provider-dispatch.md) before probing or writing anything. Its model matrices (stock and flex), descriptor grammar, and route table are the contract. Role assignments are selected first; then choose one requested effort per assigned matrix family. Do not add a second configuration file, a runtime resolver, or a weaker-model fallback.
Claude Code writes `~/.claude/pstack-models.md` and loads it from `~/.claude/CLAUDE.md` with:
@@ -35,7 +35,7 @@ Treat the normalized values as current role-to-family assignments. Overlay those
### 3. Parse per-family efforts
-Read the model matrices, stock, additional, and flex. Every non-alias value must match `:@`. Map it to exactly one matrix family by `(provider, model)`, require its effort to appear in that row's Selectable efforts cell, and collect the effort. `inherit-parent` and `auto` rows carry no family effort.
+Read the model matrices, stock and flex. Every non-alias value must match `:@`. Map it to exactly one matrix family by `(provider, model)`, require its effort to appear in that row's Selectable efforts cell, and collect the effort. `inherit-parent` and `auto` rows carry no family effort.
An unmatched provider/model, out-of-domain effort, duplicate role, or unknown role is inconsistent state. Stop, show the conflicting rows verbatim, and ask for an explicit matrix family or alias replacement. If one or more families have mixed efforts, show every conflicting family and role row, then ask for one normalized effort per family from its Selectable efforts cell. Do not invent a precedence rule. Do not probe or write while any inconsistency is unresolved.
@@ -45,9 +45,9 @@ One distinct effort per family is the current value. A family with no non-alias
### 4. Choose role assignments, then collect efforts
-Role assignments come first. Show the current role-to-family map (loaded and normalized from step 2, or the first-run map from step 7) and ask whether to keep it or change named roles. Keeping it is the default. A changed role may use any stock, additional, or flex matrix family, `inherit-parent`, or `auto`. Apply only role changes the operator names; never offer a reset of a customized sheet to the first-run assignments.
+Role assignments come first. Show the current role-to-family map (loaded and normalized from step 2, or the first-run map from step 7) and ask whether to keep it or change named roles. Keeping it is the default. A changed role may use any stock or flex matrix family, `inherit-parent`, or `auto`. Apply only role changes the operator names; never offer a reset of a customized sheet to the first-run assignments.
-Offer Astra, GPT-6 Sol, and Luna from the additional matrix when changing `architect runners` or another configurable role. Read each model, proposed effort, and selectable efforts from its row. The additional families and stock Sol are separate families even though they share the Codex provider; changing one family's effort does not change another's. GPT-6 Sol uses the `sol-6` family; the stock `sol` family keeps GPT-5.6 Sol.
+Offer every stock family, including Astra, GPT-6 Sol, and Luna, when changing `architect runners` or another configurable role. Read each model, proposed effort, and selectable efforts from its row. The Codex families are separate families even though they share the Codex provider; changing one family's effort does not change another's. GPT-6 Sol uses the `sol-6` family; the `sol` family keeps GPT-5.6 Sol for sheets that still assign it.
The assigned families are exactly the matrix families that appear in the resulting role map. An unassigned family gets no effort question and no probe. There is no requirement to assign every matrix family.
@@ -61,9 +61,9 @@ Probe only the assigned families' selected `provider:model@effort` pairs. Run on
|---|---|---|---|---|
| Fable | Fable matrix row + selected effort | native Agent `pstack-fable-` | Claude CLI | native one-turn probe or `claude auth status --json` plus one-turn probe |
| Sol | Sol matrix row + selected effort | `codex exec` | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe |
-| Astra | Astra additional row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe |
-| GPT-6 Sol | sol-6 additional row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe |
-| Luna | Luna additional row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe |
+| Astra | Astra matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe |
+| GPT-6 Sol | sol-6 matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe |
+| Luna | Luna matrix row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe |
| Grok | Grok matrix row + selected effort | Grok CLI | Grok CLI | `grok models` must list the requested model; one-turn probe |
| Opus | Opus matrix row + selected effort | native Agent `pstack-opus-` | Claude CLI | native one-turn probe or `claude auth status --json` plus one-turn probe |
| DeepSeek Flash / Pro | Each assigned DeepSeek flex row + selected effort | external runner | external runner | `DEEPSEEK_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint |
@@ -88,7 +88,7 @@ Different models sharing a provider count as one provider, even when their effor
Validate panel diversity: `arena runners` and `interrogate reviewers` must span at least two distinct providers. A single-provider panel is written only after the operator explicitly confirms the reduced diversity; record that confirmation in the setup report.
-Rewrite every matrix-family descriptor to `provider:model@`. Leave `inherit-parent` and `auto` unchanged. An effort-only rerun cannot change a role's family. Changing Grok's effort updates every Grok occurrence and does not move a Sol role onto Grok. Refuse an unqualified slug, an unavailable route, a model outside the stock, additional, and flex matrix families, or a provider/model mismatch.
+Rewrite every matrix-family descriptor to `provider:model@`. Leave `inherit-parent` and `auto` unchanged. An effort-only rerun cannot change a role's family. Changing Grok's effort updates every Grok occurrence and does not move a Sol role onto Grok. Refuse an unqualified slug, an unavailable route, a model outside the stock and flex matrix families, or a provider/model mismatch.
### 7. Confirm and commit
@@ -105,21 +105,21 @@ After the operator confirms, write the in-memory render from step 6. Never paste
Provider-qualified per-role choices. Read the installed pstack provider-dispatch reference before dispatching a configured role. Every documented role remains present. `inherit-parent` and `auto` use the parent model natively and still count as one panel lane.
-feature, refactoring: grok:grok-4.6@xhigh
-bug-fix: codex:gpt-5.6-sol@max
-perf-issue: codex:gpt-5.6-sol@max
-hillclimb: codex:gpt-5.6-sol@max
+feature, refactoring: codex:gpt-6-sol@high
+bug-fix: codex:gpt-6-sol@high
+perf-issue: codex:gpt-6-sol@high
+hillclimb: codex:gpt-6-sol@high
judgment and prose: claude:fable@max
hardest tasks: claude:fable@max
-how explorer: grok:grok-4.6@xhigh
+how explorer: codex:gpt-6-luna@high
how explainer: claude:fable@max
why investigators, synthesizer: inherit-parent
reflect tooling, judgment, divergent, synthesizer: inherit-parent
-arena runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh
-arena cross-judge pool: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh
-swarm workers: grok:grok-4.6@xhigh
-architect runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh
-interrogate reviewers: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh
+arena runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh
+arena cross-judge pool: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh
+swarm workers: codex:gpt-6-luna@high
+architect runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh
+interrogate reviewers: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh
```
### 8. Wire it in
diff --git a/plugins/pstack/skills/swarm/SKILL.md b/plugins/pstack/skills/swarm/SKILL.md
index 92c8f48a..d010532d 100644
--- a/plugins/pstack/skills/swarm/SKILL.md
+++ b/plugins/pstack/skills/swarm/SKILL.md
@@ -23,7 +23,7 @@ Open a todolist with one entry per phase before launching anything.
1. State the done predicate and the artifact or report the swarm must return.
2. Choose the shape. Partition into slices, race N workers on identical briefs, or mix both. For a race or mixed shape, declare `first pass`, `rank all`, or `best-of` before spawning.
3. Set N from the user or derive it from the shape. N is total workers, not the number that run at once.
-4. Pick the worker descriptor from `swarm workers` in the current harness's pstack model sheet when present. Otherwise use `grok:grok-4.6@xhigh`. For a model race, name each arm's descriptor up front.
+4. Pick the worker descriptor from `swarm workers` in the current harness's pstack model sheet when present. Otherwise use `codex:gpt-6-luna@high`. For a model race, name each arm's descriptor up front.
5. Give each worker its own writable output when it writes.
## Phase B: Fan out
diff --git a/tests/skill-collision-repro.sh b/tests/skill-collision-repro.sh
index 22db0d8e..99bb7320 100755
--- a/tests/skill-collision-repro.sh
+++ b/tests/skill-collision-repro.sh
@@ -73,32 +73,15 @@ else
fi
# Static invariant (CHANGES maintenance note): provider-dispatch owns the default
-# provider/model quad and the three panel skills plus setup-pstack copy it verbatim.
+# panel quad (its "## Default panel" line) and the three panel skills plus setup-pstack copy it verbatim.
setup="$repo/plugins/pstack/skills/setup-pstack/SKILL.md"
dispatch="$repo/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md"
quad_of() { { grep -oE '(claude|codex|grok):[a-z0-9.-]+@(low|medium|high|xhigh|max)' || true; } | tr '\n' ' ' | sed 's/ $//'; }
canon_quad="$(awk '
- $0 == "## Model matrix" { in_matrix = 1; next }
- in_matrix && /^## / { exit }
- in_matrix && /^\|/ {
- line = $0
- sub(/^\|/, "", line)
- sub(/\|$/, "", line)
- n = split(line, cells, "|")
- for (i = 1; i <= n; i++) {
- gsub(/^ +| +$/, "", cells[i])
- gsub(/`/, "", cells[i])
- }
- family = cells[1]
- if (family == "Family" || family ~ /^:?-+:?$/) next
- provider = cells[3]
- model = cells[4]
- effort = cells[5]
- if (out != "") out = out " "
- out = out provider ":" model "@" effort
- }
- END { print out }
-' "$dispatch")"
+ $0 == "## Default panel" { in_panel = 1; next }
+ in_panel && /^## / { exit }
+ in_panel && /^`/ { print; exit }
+' "$dispatch" | quad_of)"
quad_bad=""
[ -n "$canon_quad" ] || quad_bad="could not read the canonical quad from $dispatch"$'\n'
# Anchor on the quad's last slug rather than a hard-coded one, so a model swap in
@@ -350,14 +333,14 @@ else
fi
sol_descriptor="$(awk -F '|' '
- $2 ~ /^[[:space:]]*sol[[:space:]]*$/ {
+ $2 ~ /^[[:space:]]*sol-6[[:space:]]*$/ {
for (i = 4; i <= 6; i++) gsub(/^[[:space:]]+|[[:space:]]+$/, "", $i)
print $4 ":" $5 "@" $6
}
' "$dispatch")"
solo_code_bad=""
if [ -z "$sol_descriptor" ]; then
- solo_code_bad="could not read the sol row from $dispatch"$'\n'
+ solo_code_bad="could not read the sol-6 row from $dispatch"$'\n'
fi
for role in bug-fix perf-issue hillclimb; do
setup_descriptor="$(sed -n "s/^${role}: //p" "$setup")"
@@ -371,11 +354,11 @@ for role in bug-fix perf-issue hillclimb; do
fi
done
if [ -n "$solo_code_bad" ]; then
- note "FAIL: solo code roles must use the sol row:"
+ note "FAIL: solo code roles must use the sol-6 row:"
note "$solo_code_bad"
fail=1
else
- note "ok: solo code roles stay on the sol row ($sol_descriptor)"
+ note "ok: solo code roles stay on the sol-6 row ($sol_descriptor)"
fi
codex_manifest="$plugin/.codex-plugin/plugin.json"