diff --git a/CHANGES.md b/CHANGES.md index f475c63..2de084b 100644 --- a/CHANGES.md +++ b/CHANGES.md @@ -1,5 +1,11 @@ # CHANGES — applied substitutions +## Unreleased: multiple gateway model choices + +- Add DeepSeek V4 Pro and MiniMax M3.1 Flash Preview as independent setup families alongside existing Flash and M3 choices. Keep provider-owned routing and existing descriptors. +- Validate unique model families and provider/model pairs instead of requiring one row per provider; cover both parent routes and substituted-model rejection. +- Document preview Token Plan access, model-specific thinking semantics, and required installed live validation. Tracked in [pstack-flex #5](https://github.com/thisguymartin/pstack-flex/issues/5). + This port applies the Cursor → Claude Code substitutions in skill bodies. Earlier drafts left them flagged; this revision resolves them. A later pass added a Codex build that shares the same skills; see [Codex port](#codex-port) below. ## pstack-flex (unreleased) — gateway lanes and optional families diff --git a/docs/LANES.md b/docs/LANES.md index 82566f5..2ed9884 100644 --- a/docs/LANES.md +++ b/docs/LANES.md @@ -9,10 +9,29 @@ Prices and endpoints below were verified 2026-09-25 and drift. Re-verify against | Kind | Lanes | Auth | Billing | Route | | --- | --- | --- | --- | --- | | Subscription | `claude:fable`, `claude:opus`, `codex:gpt-5.6-sol`, `grok:grok-4.6` | each CLI's own login | that CLI's plan | native or external per the route table | -| Gateway (flex) | `deepseek:deepseek-flash`, `minimax:MiniMax-M3` | API key in the environment | pay per token on the lab's key | always the external runner | +| Gateway (flex) | DeepSeek Flash / V4 Pro; MiniMax M3 / M3.1 Flash Preview | API key in the environment | provider billing; preview requires Token Plan | always the external runner | A gateway lane is the stock `claude` binary env-pointed at the lab's Anthropic-compatible endpoint. There is no custom agent loop and no separate harness: the same runner that spawns Codex and Grok lanes spawns gateway lanes with injected environment. Both labs document this Claude Code setup themselves (DeepSeek: `deepseek-ai/awesome-deepseek-agent`, `docs/claude_code.md`; MiniMax: platform.minimax.io, Claude Code guide). +## Multiple models per provider + +The flex matrix now includes four independently assignable model families: + +| Family | Descriptor at default requested effort | Selection guidance | +| --- | --- | --- | +| deepseek | `deepseek:deepseek-flash@high` | Existing everyday option | +| deepseek-pro | `deepseek:deepseek-v4-pro@high` | Candidate for difficult debugging, architecture, and review | +| minimax | `minimax:MiniMax-M3@high` | Existing MiniMax option | +| minimax-preview | `minimax:MiniMax-M3.1-Flash-Preview@high` | Preview coding option with tunable thinking | + +These are choices, not automatic replacements or a performance ranking. Existing sheets keep their assignments. In `/setup-pstack`, assign named roles to the desired model family; efforts and probes are independent per model, even for models sharing a key. Two models from one provider count as one provider for panel diversity. No runtime routing change or new configuration file is needed. + +As of 2026-09-27, [MiniMax's model guide](https://platform.minimax.io/docs/guides/models-intro) restricts M3.1 Flash Preview to Token Plan and MiniMax Code. For gateway access, supply the eligible Token Plan key as `MINIMAX_API_KEY`; the live probe must confirm entitlement. It is not a zero-subscription option. A working M3 call does not establish preview access. + +[MiniMax's Anthropic API](https://platform.minimax.io/docs/api-reference/text-anthropic-api) documents always-on thinking for the preview and `output_config.effort` from `low` to `max`. Higher effort increases thinking latency; the matrix proposes `high`, while the API defaults to `max` when omitted. M3 defaults to thinking off at the API and needs adaptive thinking to enable it. Its requested effort flag is not evidence of the preview's depth controls. Verify the installed CLI forwards the intended parameters; receipts prove requested effort, not hidden applied depth. [DeepSeek documents V4 Pro through its Anthropic endpoint](https://api-docs.deepseek.com/guides/anthropic_api). + +Before recommending a fastest or strongest default, compare the same synthetic coding tasks for correctness, completion time, tool-call reliability, token usage, and actual provider billing. Preview pricing and plan limits must be checked against the active plan rather than inferred from M3 rates. + ## Gateway environment reference Set by you: @@ -108,10 +127,11 @@ Any lab that serves an Anthropic-compatible `/v1/messages` endpoint can become a 4. Add its probe row to the table in `plugins/pstack/skills/setup-pstack/SKILL.md`, its variables to the gateway environment reference above, and its prices to the price table. 5. Run the live validation checklist below for the new lane before merging. -## Live validation checklist (post-merge, real keys, never in CI) +## Live validation checklist (before merge or rollout, real keys, never in CI) - V1: one DeepSeek probe through the runner (`--provider deepseek --model deepseek-flash --effort high`, read-only). Expect a `complete` receipt with `costUsd: null`; record the `reportedModel` string and confirm the base-URL default against DeepSeek's current guide; confirm `--effort` is accepted end-to-end. - V2: same for MiniMax (`MiniMax-M3`); record the served-model casing. +- New-model gate: install the exact candidate and run `/setup-pstack` from both real Claude Code and Codex surfaces. Select Flash plus Pro and M3 plus Preview, verify independent efforts and probes, then run a read-only mixed panel. Record installed version/commit, surface, action, requested model/effort, served model, and observed result. Verify a failed preview entitlement probe leaves the sheet unchanged and does not select M3. A fake CLI regression test is not this gate. - V3: run `claude auth status --json` inside a fresh flex config dir with `ANTHROPIC_AUTH_TOKEN` set and record the output here. On macOS, confirm whether `claude login` under an explicit `CLAUDE_CONFIG_DIR` writes `.credentials.json` or the Keychain. - V4: the zero-subscription walkthrough above, end to end, on a machine with no stored provider logins. - V5: OAuth guard live: `claude login` inside a scratch flex config dir, run a lane, confirm the refusal receipt, then delete that login. diff --git a/docs/USAGE.md b/docs/USAGE.md index 7ce4e2d..b4f1ac6 100644 --- a/docs/USAGE.md +++ b/docs/USAGE.md @@ -231,3 +231,7 @@ Every external lane writes a JSON receipt next to its output. The fields that ma | Exit 69 `unavailable-cli` | the `claude` binary isn't on PATH for the runner | install it or fix PATH | | A panel ran with fewer lanes than configured | a lane dropped out with a named receipt | read that receipt; pstack proceeds N-1 and never silently substitutes a model | | Everything gateway broke after a claude CLI update | Anthropic doesn't support third-party endpoints; compatibility can shift | pin the CLI version on machines that depend on gateway lanes; see [LANES.md](LANES.md#safety-and-policy) | + +## Selecting the additional gateway models + +Run `/setup-pstack` and assign `deepseek-pro` (`deepseek:deepseek-v4-pro@high`) or `minimax-preview` (`minimax:MiniMax-M3.1-Flash-Preview@high`) to named roles. Existing `deepseek` and `minimax` choices remain available. Each model has its own effort selection and live probe. MiniMax preview requires an eligible Token Plan key in `MINIMAX_API_KEY`; see [model choices and thinking controls](LANES.md#multiple-models-per-provider). No existing assignment changes until setup succeeds and you confirm the rendered sheet. diff --git a/docs/gateway-model-probes.md b/docs/gateway-model-probes.md new file mode 100644 index 0000000..ab03514 --- /dev/null +++ b/docs/gateway-model-probes.md @@ -0,0 +1,30 @@ +# Gateway model probe evidence + +Tracking: [issue #5](https://github.com/thisguymartin/pstack-flex/issues/5). + +## Candidate and scope + +- Source candidate: branch `flex/multiple-gateway-models`, based on `92dc0bc`, with uncommitted implementation changes. +- Implementation diff SHA-256 before this evidence file: `84b4c3ee7fc66902532e1d457048a487e6a63629b31d2d214e52585f72863c75`. +- Packaged version: 1.4.1. This candidate has not been installed as a plugin. +- Actual parent: Codex session, invoking the candidate's external runner with `--parent codex`. +- CLI: Claude Code 2.1.283. +- Each probe used `--effort high`, read-only mode, a separate empty synthetic workspace and isolated Claude configuration, and a synthetic text file. No repository or customer data was used in the prompt. +- Keys were supplied through hidden terminal input, injected into child environments, and were not included in commands, this repository, or evidence below. + +## Observed results + +Each probe exited 0, returned the exact requested marker, recorded `modelVerified: true` with `modelEvidence: provider-report`, and retained `costUsd: null`. + +| Requested model | Reported model | Elapsed milliseconds | Receipt status | +| --- | --- | --- | --- | +| `deepseek-flash` | `deepseek-flash` | 2684 | `complete` | +| `deepseek-v4-pro` | `deepseek-v4-pro` | 7737 | `complete` | +| `MiniMax-M3` | `MiniMax-M3` | 11416 | `complete` | +| `MiniMax-M3.1-Flash-Preview` | `MiniMax-M3.1-Flash-Preview` | 5090 | `complete` | + +These single short probes establish authentication, model selection, and successful completion through the runner. They do not rank coding quality or speed, prove hidden reasoning depth, or verify CLI request-body effort forwarding. The prompt included the expected marker, so completion does not independently prove a file tool was used. + +## Remaining release gate + +Install the exact candidate and run setup from both real Claude Code and Codex user surfaces. Verify independent model effort choices, per-model probes, mixed-provider panels, saved-sheet readback, and unchanged configuration on failed access. Record installed version, surface, action, and observed result before merge or rollout. Changing only the runner's `--parent` flag would not satisfy this gate. diff --git a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md index f30e88c..ae60d07 100644 --- a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md +++ b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md @@ -26,7 +26,13 @@ pstack-flex addition. The stock matrix above is upstream-owned and unchanged; th | Family | Provider | Model | Default effort | Selectable efforts | API key variable | Base URL default | |---|---|---|---|---|---|---| | deepseek | deepseek | deepseek-flash | high | low medium high xhigh max | DEEPSEEK_API_KEY | https://api.deepseek.com/anthropic | +| deepseek-pro | deepseek | deepseek-v4-pro | high | low medium high xhigh max | DEEPSEEK_API_KEY | https://api.deepseek.com/anthropic | | minimax | minimax | MiniMax-M3 | high | low medium high xhigh max | MINIMAX_API_KEY | https://api.minimax.io/anthropic | +| minimax-preview | minimax | MiniMax-M3.1-Flash-Preview | high | low medium high xhigh max | MINIMAX_API_KEY | https://api.minimax.io/anthropic | + +A family identifies one `(provider, model)` pair, not an entire provider. The existing `deepseek` and `minimax` family names and descriptors remain valid. `deepseek-pro` and `minimax-preview` are additional choices with independent requested efforts. Multiple models from one provider still count as one provider for panel diversity. + +MiniMax preview requires Token Plan access; set `MINIMAX_API_KEY` to the eligible subscription key. A pay-as-you-go key is not proof of preview access. The preview always thinks and supports `low` through `max`; do not disable thinking. M3 thinking is off by default at the API and requires adaptive thinking to enable it; its effort flag does not imply preview-style depth control. Selectable efforts are runner requests, not a claim that every provider applies five distinct reasoning levels. Verify CLI forwarding and model access with live probes. Sources: [MiniMax models](https://platform.minimax.io/docs/guides/models-intro), [MiniMax thinking controls](https://platform.minimax.io/docs/api-reference/text-anthropic-api), [DeepSeek Anthropic compatibility](https://api-docs.deepseek.com/guides/anthropic_api) (checked 2026-09-27). Flex lanes have no Claude-native agent stem and always take the external runner in both parents. The base URL is a documented default; override it with `DEEPSEEK_BASE_URL` or `MINIMAX_BASE_URL`, and confirm it against the provider's current Claude Code guide during setup's live probe. The config dir defaults to `~/.pstack-flex/` (override: `PSTACK_FLEX__CONFIG_DIR`). Secrets stay in the environment: nothing in the sheet, the receipts, or this repository carries a key. diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts index af138b1..af7258f 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts @@ -366,10 +366,12 @@ describe("model matrix", () => { .slice(start + 1, end) .map((line) => line.trim()) .filter((line) => line.startsWith("|")); - expect(table.length).toBe(2 + GATEWAY_PROVIDERS.length); + expect(table.length).toBeGreaterThan(2 + GATEWAY_PROVIDERS.length); expect(splitRow(table[0]).join("|")).toBe(FLEX_MATRIX_HEADER.join("|")); expect(isSeparator(splitRow(table[1]))).toBe(true); - const seen: GatewayProvider[] = []; + const seen = new Set(); + const families = new Set(); + const pairs = new Set(); for (const line of table.slice(2)) { const cells = splitRow(line); expect(cells.length).toBe(FLEX_MATRIX_HEADER.length); @@ -377,8 +379,13 @@ describe("model matrix", () => { cells; expect(GATEWAY_PROVIDERS as readonly string[]).toContain(provider); const gateway = provider as GatewayProvider; - seen.push(gateway); - expect(family).toBe(gateway); + seen.add(gateway); + expect(/^[a-z0-9-]+$/.test(family)).toBe(true); + expect(families.has(family)).toBe(false); + families.add(family); + const pair = `${provider}:${model}`; + expect(pairs.has(pair)).toBe(false); + pairs.add(pair); expect(/^[A-Za-z0-9.-]+$/.test(model)).toBe(true); const selectable = selectableRaw.split(/\s+/).map(asEffort); expect(selectable).toContain(asEffort(defaultEffortRaw)); @@ -386,7 +393,15 @@ describe("model matrix", () => { expect(baseUrl).toBe(GATEWAY_SPECS[gateway].baseUrlDefault); expect(baseUrl.startsWith("https://")).toBe(true); } - expect(seen).toEqual([...GATEWAY_PROVIDERS]); + expect([...seen]).toEqual([...GATEWAY_PROVIDERS]); + for (const pair of [ + "deepseek:deepseek-flash", + "deepseek:deepseek-v4-pro", + "minimax:MiniMax-M3", + "minimax:MiniMax-M3.1-Flash-Preview", + ]) expect(pairs.has(pair)).toBe(true); + expect(setup).toContain("Never group efforts or deduplicate probes by provider alone."); + expect(setup).toContain("Different models sharing a provider count as one provider"); // The stock quad and first-run sheet must not carry flex descriptors: // upstream's own checks parse descriptors with a lowercase-only, // three-provider grammar and must never see a flex lane. diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts index 41e9a5f..8c2c091 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts @@ -1032,6 +1032,46 @@ describe("gateway lanes", () => { expect(readFileSync(input.outputPath, "utf8")).toBe("CLAUDE_OK"); }); + for (const parent of ["claude", "codex"] as const) { + for (const [provider, model] of [ + ["deepseek", "deepseek-v4-pro"], + ["minimax", "MiniMax-M3.1-Flash-Preview"], + ] as const) { + it(`pins ${model} and effort through the ${parent} parent route`, async () => { + const dumpPath = join(scratch, "new-model-env.json"); + process.env.FAKE_DUMP_ENV_PATH = dumpPath; + process.env.FAKE_REPORT_MODEL = model.toLowerCase(); + const input: RunnerOptions = { + ...gatewayOptions(provider, "new-model"), parent, model, effort: "max", + }; + expect((await runLane(input)).exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written).toMatchObject({ + status: "complete", parent, provider, model, effort: "max", + modelVerified: true, modelEvidence: "provider-report", costUsd: null, + }); + expect(written.argv[written.argv.indexOf("--model") + 1]).toBe(model); + expect(written.argv[written.argv.indexOf("--effort") + 1]).toBe("max"); + const child = JSON.parse(readFileSync(dumpPath, "utf8")); + for (const key of [ + "ANTHROPIC_MODEL", "ANTHROPIC_DEFAULT_OPUS_MODEL", + "ANTHROPIC_DEFAULT_SONNET_MODEL", "ANTHROPIC_DEFAULT_HAIKU_MODEL", + "CLAUDE_CODE_SUBAGENT_MODEL", + ]) expect(child[key]).toBe(model); + }); + + it(`rejects a substituted ${model} in the ${parent} parent route`, async () => { + const input: RunnerOptions = { + ...gatewayOptions(provider, "substituted-model"), parent, model, + }; + process.env.FAKE_REPORT_MODEL = provider === "deepseek" ? "deepseek-flash" : "MiniMax-M3"; + expect((await runLane(input)).exitCode).toBe(65); + expect(receipt(input.receiptPath).status).toBe("malformed-output"); + expect(existsSync(input.outputPath)).toBe(false); + }); + } + } + it("verifies a case-shifted served model for MiniMax", async () => { process.env.FAKE_REPORT_MODEL = "minimax-m3"; const input = gatewayOptions("minimax", "case-shift"); diff --git a/plugins/pstack/skills/setup-pstack/SKILL.md b/plugins/pstack/skills/setup-pstack/SKILL.md index 95c3962..06c8ed9 100644 --- a/plugins/pstack/skills/setup-pstack/SKILL.md +++ b/plugins/pstack/skills/setup-pstack/SKILL.md @@ -39,6 +39,8 @@ Read the model matrices, stock and flex. Every non-alias value must match `` | Claude CLI | native one-turn probe or `claude auth status --json` plus one-turn probe | -| DeepSeek | DeepSeek flex row + selected effort | external runner | external runner | `DEEPSEEK_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint | -| MiniMax | MiniMax flex row + selected effort | external runner | external runner | `MINIMAX_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint | +| DeepSeek Flash / Pro | Each assigned DeepSeek flex row + selected effort | external runner | external runner | `DEEPSEEK_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint | +| MiniMax M3 / M3.1 Flash Preview | Each assigned MiniMax flex row + selected effort | external runner | external runner | `MINIMAX_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint | + +For MiniMax M3.1 Flash Preview, disclose the Token Plan requirement before probing. Use the eligible subscription key through `MINIMAX_API_KEY`; do not assume a working M3 key grants preview access. A failed preview probe must not silently select M3. Keep preview thinking enabled and verify requested effort forwarding; distinguish request evidence from hidden applied reasoning depth. Use a tiny read-only probe that returns a unique marker. A login-status command alone proves credentials, not that the requested model and effort flags run. Record native and external results separately. Never call the external launcher for the parent's own provider. On a Claude parent, the Fable and Opus probes are one-turn runs of the mapped `pstack--` agent. On a Codex parent, the Sol probe is native `spawn_agent` with the selected `reasoning_effort`. Every other pair, flex families always included, uses the external runner with the selected effort flag. A flex probe doubles as the base-URL confirmation: it proves the documented default (or the operator's override) actually serves the lane's model. @@ -75,6 +79,8 @@ Build the new sheet in memory. Do not write it yet. The role assignments were already chosen in step 4; do not re-open them here. Require every documented role to remain present and non-empty, `architect runners` to keep at least two entries, and the final role map to contain at least one assigned matrix family. There is no requirement to assign every matrix family. The sheet stores effort only in role descriptors, so an unassigned family's selection cannot persist without adding a second source of truth. +Different models sharing a provider count as one provider, even when their efforts differ. + Validate panel diversity: `arena runners` and `interrogate reviewers` must span at least two distinct providers. A single-provider panel is written only after the operator explicitly confirms the reduced diversity; record that confirmation in the setup report. Rewrite every matrix-family descriptor to `provider:model@`. Leave `inherit-parent` and `auto` unchanged. An effort-only rerun cannot change a role's family. Changing Grok's effort updates every Grok occurrence and does not move a Sol role onto Grok. Refuse an unqualified slug, an unavailable route, a model outside the stock and flex matrix families, or a provider/model mismatch.