Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGES.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# CHANGES — applied substitutions

## Unreleased: GPT-6 Codex families become stock

- Move `astra`, `sol-6`, and `luna` from the additional matrix into the stock model matrix. The first-run sheet now assigns GPT-6 Sol high to `feature, refactoring`, `bug-fix`, `perf-issue`, and `hillclimb`; Luna high to `how explorer` and `swarm workers`; and Astra high in place of GPT-5.6 Sol on every panel. `codex:gpt-5.6-sol` stays selectable; existing sheets are untouched until a role is reassigned in setup.
- `provider-dispatch.md` gains a "Default panel" section that is the single source for the four panel lanes. `setup-pstack`, `arena`, `architect`, and `interrogate` copy it; the static invariant reads that line instead of deriving the panel from every matrix row, and the solo-code invariant checks the `sol-6` row.
- The matrix test asserts seven stock rows, the GPT-6 descriptors in the first-run sheet, and the panel contract. `UPSTREAM-FLEX.md` records the new permanent conflict surface on upstream syncs. Tracked in [pstack-flex #17](https://github.com/thisguymartin/pstack-flex/issues/17).

## Unreleased: multiple gateway model choices

- Add DeepSeek V4 Pro and MiniMax M3.1 Flash Preview as independent setup families alongside existing Flash and M3 choices. Keep provider-owned routing and existing descriptors.
Expand Down
109 changes: 63 additions & 46 deletions README.md

Large diffs are not rendered by default.

7 changes: 4 additions & 3 deletions UPSTREAM-FLEX.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,11 +22,12 @@ All flex changes are additive and live in port-owned files so upstream merges st

- `plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts` and `flex-providers.test.ts` (new)
- Gateway-provider hooks in `runner/{types,commands,run,parse-output,cli}.ts` and their tests
- The "Additional model matrix" and "Flex model matrix" sections and route-table columns in `references/provider-dispatch.md`
- The three GPT-6 rows in the stock model matrix, the "Default panel" section, the "Flex model matrix" section, and the route-table columns in `references/provider-dispatch.md`
- The first-run sheet in `skills/setup-pstack/SKILL.md` and the default descriptors named in `arena`, `architect`, `interrogate`, `how`, `swarm`, and the `feature`, `refactoring`, `bug-fix`, `perf-issue`, and `hillclimb` playbooks
- The assignment-first restructure of `skills/setup-pstack/SKILL.md`
- `docs/LANES.md`, this file, the README fork section, and the NOTICE/LICENSE/CHANGES additions

The stock model matrix, the first-run sheet, every upstream skill body, and the static quad invariants are byte-unchanged. Flex rows and setup instructions are additive.
Every upstream skill body is byte-unchanged except for the default-descriptor mentions listed above. Since [#17](https://github.com/thisguymartin/pstack-flex/issues/17), the stock matrix carries three fork-owned GPT-6 rows and the first-run sheet uses them, so those two surfaces conflict on every upstream sync and are resolved by hand: keep the fork's rows and defaults, take upstream's wording for everything else.

## Merge procedure

Expand All @@ -41,6 +42,6 @@ Expected conflict surface on future upstream releases:

- `plugins/pstack/skills/setup-pstack/SKILL.md` — upstream issue #88 (1.5.0, syncing Cursor pstack 0.15.5) folds upstream PR #73, which moves setup to the same assignment-first, probe-only-assigned shape this fork already uses. Resolve toward upstream's wording wherever it covers the same rule; keep the flex families and the diversity rule.
- `plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts` — upstream 1.5.0 changes the stock panel to three lanes. Take upstream's stock assertions verbatim; the flex-matrix describe block is fork-owned and should survive as-is.
- `plugins/pstack/skills/poteto-mode/references/provider-dispatch.md` — stock matrix and default-panel prose are upstream's; the flex section is fork-owned.
- `plugins/pstack/skills/poteto-mode/references/provider-dispatch.md` — the four upstream rows and their prose are upstream's; the GPT-6 rows, the "Default panel" section, and the flex section are fork-owned. Upstream's own panel changes land in the "Default panel" line only if the fork wants them.

After every merge: run the full local gate (`bun install --frozen-lockfile`, `bun run test`, `bun run typecheck`, manifest JSON parse, `PSTACK_STATIC_ONLY=1 bash tests/skill-collision-repro.sh`), then record the installed version, action, and observed result for each affected harness in the pull request before tagging.
2 changes: 1 addition & 1 deletion UPSTREAM.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ The table above is the current Cursor sync point. Open Pstack 1.4.1 imports this

- Commits `799151d` and `6fecddb` add and relocate `make-bot-ui`. It depends on Cursor routines, webhook events, and UI primitives that Claude Code and Codex do not share.
- Four `disable-model-invocation: true` lines from `73f8be4` are not applied to `how`, `why`, `unslop`, or `typescript-best-practices`. Poteto-mode invokes those skills by name, and the flag blocks that route on Claude Code.
- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on `codex:gpt-5.6-sol@max` for cost.
- The `23a56e2` default-model hunks for `bug-fix`, `perf-issue`, and `hillclimb` are not applied. Those frequent code-writing roles stay on a Codex model (`codex:gpt-6-sol@high` in pstack-flex) for cost.
- The Claude manifest does not take the logo field from `efa2a53` because Claude Code has no schema for it. The shared asset is exposed through the Codex manifest instead.

## Check for changes
Expand Down
18 changes: 9 additions & 9 deletions docs/LANES.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,22 +8,22 @@ Prices and endpoints below were verified 2026-09-25 and drift. Re-verify against

| Kind | Lanes | Auth | Billing | Route |
| --- | --- | --- | --- | --- |
| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-5.6-sol`, `grok:grok-4.6`; optional `codex:gpt-6-astra`, `codex:gpt-6-sol`, `codex:gpt-6-luna` | each CLI's own login | that CLI's plan | native or external per the route table |
| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-6-astra`, `codex:gpt-6-sol`, `codex:gpt-6-luna`, `codex:gpt-5.6-sol`, `grok:grok-4.6` | each CLI's own login | that CLI's plan | native or external per the route table |
| Gateway (flex) | DeepSeek Flash / V4 Pro; MiniMax M3 / M3.1 Flash Preview | API key in the environment | provider billing; preview requires Token Plan | always the external runner |

A gateway lane is the stock `claude` binary env-pointed at the lab's Anthropic-compatible endpoint. There is no custom agent loop and no separate harness: the same runner that spawns Codex and Grok lanes spawns gateway lanes with injected environment. Both labs document this Claude Code setup themselves (DeepSeek: `deepseek-ai/awesome-deepseek-agent`, `docs/claude_code.md`; MiniMax: platform.minimax.io, Claude Code guide).

## Optional GPT-6 Codex families
## GPT-6 Codex families

The additional model matrix adds three Codex families. They use the same ChatGPT login as `codex:gpt-5.6-sol`:
Three stock Codex families carry the first-run defaults. They use the same ChatGPT login as `codex:gpt-5.6-sol`:

| Family | Descriptor at default requested effort | Codex's own description |
| --- | --- | --- |
| astra | `codex:gpt-6-astra@high` | Frontier tier for the most demanding work |
| sol-6 | `codex:gpt-6-sol@high` | Coding and everyday workhorse |
| luna | `codex:gpt-6-luna@high` | Fast, low-cost tier for easier tasks |
| Family | Descriptor at default requested effort | First-run roles | Codex's own description |
| --- | --- | --- | --- |
| astra | `codex:gpt-6-astra@high` | every panel (`arena runners`, `arena cross-judge pool`, `architect runners`, `interrogate reviewers`) | Frontier tier for the most demanding work |
| sol-6 | `codex:gpt-6-sol@high` | `feature, refactoring`, `bug-fix`, `perf-issue`, `hillclimb` | Coding and everyday workhorse |
| luna | `codex:gpt-6-luna@high` | `how explorer`, `swarm workers` | Fast, low-cost tier for easier tasks |

These are choices, not replacements. The first-run sheet still uses `codex:gpt-5.6-sol@max`. An existing sheet changes only when you assign a role to a GPT-6 family in `/setup-pstack`. The `sol-6` family is separate from the stock `sol` family, so each keeps its own effort. All Codex families count as one provider for panel diversity, so Astra plus Sol does not satisfy the two-provider rule. The route matches Sol: native `spawn_agent` in a Codex parent, and the external runner (`codex exec`) in a Claude Code parent. Codex also lists an `ultra` effort for Astra and GPT-6 Sol. It is outside the pstack effort universe and is not selectable. The descriptions and effort lists come from the Codex CLI 0.157.1 model list, checked 2026-09-27.
A fresh `/setup-pstack` run proposes these. An existing sheet keeps its assignments until you change a named role in setup; `codex:gpt-5.6-sol` remains a selectable family for that. The `sol-6` family is separate from the `sol` family, so each keeps its own effort. All Codex families count as one provider for panel diversity, so Astra plus GPT-6 Sol does not satisfy the two-provider rule; the default panel spans Claude, Codex, and Grok. The route matches Sol: native `spawn_agent` in a Codex parent, and the external runner (`codex exec`) in a Claude Code parent. Codex also lists an `ultra` effort for Astra and GPT-6 Sol. It is outside the pstack effort universe and is not selectable. The descriptions and effort lists come from the Codex CLI 0.157.1 model list, checked 2026-09-27.

## Multiple models per provider

Expand Down
18 changes: 11 additions & 7 deletions docs/USAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ flowchart TD
PB --> S[Skills fire per step<br/>how, tdd, interrogate, arena, ...]
S --> F{Lane fan-out}
F --> N1["claude:fable / claude:opus<br/>native Agent (Claude sub)"]
F --> N2["codex:gpt-5.6-sol<br/>native or codex CLI (ChatGPT sub)"]
F --> N2["codex:gpt-6-astra / gpt-6-sol / gpt-6-luna<br/>native or codex CLI (ChatGPT sub)"]
F --> N3["grok:grok-4.6<br/>grok CLI (Grok sub)"]
F --> G1["deepseek:deepseek-flash<br/>runner + env -> DeepSeek API (key)"]
F --> G2["minimax:MiniMax-M3<br/>runner + env -> MiniMax API (key)"]
Expand Down Expand Up @@ -83,10 +83,14 @@ Daily flow: `pstack-keys -> claude -> /pstack:poteto-mode`. Alternatives, the th

Setup is assignment-first: pick which roles run on which families, answer one effort question per **assigned** family, and only assigned families get probed. Unassigned families are skipped, not errors. Every probe is a real one-turn run — a failed probe writes nothing. Three configurations that make sense:

**A. Full frontier** (Claude + ChatGPT + Grok subs) — accept the defaults; behaves exactly like stock upstream:
**A. Full frontier** (Claude + ChatGPT + Grok subs) — accept the defaults. GPT-6 Sol writes code, Luna explores and verifies, and the panel spans three providers:

```text
arena runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh
feature, refactoring: codex:gpt-6-sol@high
bug-fix: codex:gpt-6-sol@high
how explorer: codex:gpt-6-luna@high
swarm workers: codex:gpt-6-luna@high
arena runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh
```

**B. Hybrid saver** (Claude sub + two API keys) — frontier judgment, cheap volume:
Expand Down Expand Up @@ -232,14 +236,14 @@ Every external lane writes a JSON receipt next to its output. The fields that ma
| A panel ran with fewer lanes than configured | a lane dropped out with a named receipt | read that receipt; pstack proceeds N-1 and never silently substitutes a model |
| Everything gateway broke after a claude CLI update | Anthropic doesn't support third-party endpoints; compatibility can shift | pin the CLI version on machines that depend on gateway lanes; see [LANES.md](LANES.md#safety-and-policy) |

## Selecting the GPT-6 Codex models
## The GPT-6 Codex models

Run `/setup-pstack` and assign `astra` (`codex:gpt-6-astra@high`), `sol-6` (`codex:gpt-6-sol@high`), or `luna` (`codex:gpt-6-luna@high`) to named roles, such as `architect runners`. Every role you do not change keeps its current descriptor, including the stock `codex:gpt-5.6-sol@max` defaults. Each GPT-6 family gets its own effort question and live probe. They need only your Codex login. See [optional GPT-6 Codex families](LANES.md#optional-gpt-6-codex-families).
`astra` (`codex:gpt-6-astra@high`), `sol-6` (`codex:gpt-6-sol@high`), and `luna` (`codex:gpt-6-luna@high`) are stock families and the first-run defaults: Astra on every panel, GPT-6 Sol on the solo code-writing roles, Luna on exploration and swarm work. They need only your Codex login. Each gets its own effort question and live probe. See [GPT-6 Codex families](LANES.md#gpt-6-codex-families).

For example, this row puts Astra on the architect panel and keeps the other providers:
A sheet written before this release keeps its assignments. To move a role, run `/setup-pstack` and name it; every role you do not change keeps its descriptor, and `codex:gpt-5.6-sol` stays selectable. For example, this row keeps GPT-5.6 Sol on bug fixes while the rest of the sheet takes the new defaults:

```text
architect runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh
bug-fix: codex:gpt-5.6-sol@max
```

## Selecting the additional gateway models
Expand Down
Loading
Loading