Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGES.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# CHANGES — applied substitutions

## Unreleased: multiple gateway model choices

- Add DeepSeek V4 Pro and MiniMax M3.1 Flash Preview as independent setup families alongside existing Flash and M3 choices. Keep provider-owned routing and existing descriptors.
- Validate unique model families and provider/model pairs instead of requiring one row per provider; cover both parent routes and substituted-model rejection.
- Document preview Token Plan access, model-specific thinking semantics, and required installed live validation. Tracked in [pstack-flex #5](https://github.com/thisguymartin/pstack-flex/issues/5).

This port applies the Cursor → Claude Code substitutions in skill bodies. Earlier drafts left them flagged; this revision resolves them. A later pass added a Codex build that shares the same skills; see [Codex port](#codex-port) below.

## pstack-flex (unreleased) — gateway lanes and optional families
Expand Down
24 changes: 22 additions & 2 deletions docs/LANES.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,10 +9,29 @@ Prices and endpoints below were verified 2026-09-25 and drift. Re-verify against
| Kind | Lanes | Auth | Billing | Route |
| --- | --- | --- | --- | --- |
| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-5.6-sol`, `grok:grok-4.6` | each CLI's own login | that CLI's plan | native or external per the route table |
| Gateway (flex) | `deepseek:deepseek-flash`, `minimax:MiniMax-M3` | API key in the environment | pay per token on the lab's key | always the external runner |
| Gateway (flex) | DeepSeek Flash / V4 Pro; MiniMax M3 / M3.1 Flash Preview | API key in the environment | provider billing; preview requires Token Plan | always the external runner |

A gateway lane is the stock `claude` binary env-pointed at the lab's Anthropic-compatible endpoint. There is no custom agent loop and no separate harness: the same runner that spawns Codex and Grok lanes spawns gateway lanes with injected environment. Both labs document this Claude Code setup themselves (DeepSeek: `deepseek-ai/awesome-deepseek-agent`, `docs/claude_code.md`; MiniMax: platform.minimax.io, Claude Code guide).

## Multiple models per provider

The flex matrix now includes four independently assignable model families:

| Family | Descriptor at default requested effort | Selection guidance |
| --- | --- | --- |
| deepseek | `deepseek:deepseek-flash@high` | Existing everyday option |
| deepseek-pro | `deepseek:deepseek-v4-pro@high` | Candidate for difficult debugging, architecture, and review |
| minimax | `minimax:MiniMax-M3@high` | Existing MiniMax option |
| minimax-preview | `minimax:MiniMax-M3.1-Flash-Preview@high` | Preview coding option with tunable thinking |

These are choices, not automatic replacements or a performance ranking. Existing sheets keep their assignments. In `/setup-pstack`, assign named roles to the desired model family; efforts and probes are independent per model, even for models sharing a key. Two models from one provider count as one provider for panel diversity. No runtime routing change or new configuration file is needed.

As of 2026-09-27, [MiniMax's model guide](https://platform.minimax.io/docs/guides/models-intro) restricts M3.1 Flash Preview to Token Plan and MiniMax Code. For gateway access, supply the eligible Token Plan key as `MINIMAX_API_KEY`; the live probe must confirm entitlement. It is not a zero-subscription option. A working M3 call does not establish preview access.

[MiniMax's Anthropic API](https://platform.minimax.io/docs/api-reference/text-anthropic-api) documents always-on thinking for the preview and `output_config.effort` from `low` to `max`. Higher effort increases thinking latency; the matrix proposes `high`, while the API defaults to `max` when omitted. M3 defaults to thinking off at the API and needs adaptive thinking to enable it. Its requested effort flag is not evidence of the preview's depth controls. Verify the installed CLI forwards the intended parameters; receipts prove requested effort, not hidden applied depth. [DeepSeek documents V4 Pro through its Anthropic endpoint](https://api-docs.deepseek.com/guides/anthropic_api).

Before recommending a fastest or strongest default, compare the same synthetic coding tasks for correctness, completion time, tool-call reliability, token usage, and actual provider billing. Preview pricing and plan limits must be checked against the active plan rather than inferred from M3 rates.

## Gateway environment reference

Set by you:
Expand Down Expand Up @@ -108,10 +127,11 @@ Any lab that serves an Anthropic-compatible `/v1/messages` endpoint can become a
4. Add its probe row to the table in `plugins/pstack/skills/setup-pstack/SKILL.md`, its variables to the gateway environment reference above, and its prices to the price table.
5. Run the live validation checklist below for the new lane before merging.

## Live validation checklist (post-merge, real keys, never in CI)
## Live validation checklist (before merge or rollout, real keys, never in CI)

- V1: one DeepSeek probe through the runner (`--provider deepseek --model deepseek-flash --effort high`, read-only). Expect a `complete` receipt with `costUsd: null`; record the `reportedModel` string and confirm the base-URL default against DeepSeek's current guide; confirm `--effort` is accepted end-to-end.
- V2: same for MiniMax (`MiniMax-M3`); record the served-model casing.
- New-model gate: install the exact candidate and run `/setup-pstack` from both real Claude Code and Codex surfaces. Select Flash plus Pro and M3 plus Preview, verify independent efforts and probes, then run a read-only mixed panel. Record installed version/commit, surface, action, requested model/effort, served model, and observed result. Verify a failed preview entitlement probe leaves the sheet unchanged and does not select M3. A fake CLI regression test is not this gate.
- V3: run `claude auth status --json` inside a fresh flex config dir with `ANTHROPIC_AUTH_TOKEN` set and record the output here. On macOS, confirm whether `claude login` under an explicit `CLAUDE_CONFIG_DIR` writes `.credentials.json` or the Keychain.
- V4: the zero-subscription walkthrough above, end to end, on a machine with no stored provider logins.
- V5: OAuth guard live: `claude login` inside a scratch flex config dir, run a lane, confirm the refusal receipt, then delete that login.
4 changes: 4 additions & 0 deletions docs/USAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -231,3 +231,7 @@ Every external lane writes a JSON receipt next to its output. The fields that ma
| Exit 69 `unavailable-cli` | the `claude` binary isn't on PATH for the runner | install it or fix PATH |
| A panel ran with fewer lanes than configured | a lane dropped out with a named receipt | read that receipt; pstack proceeds N-1 and never silently substitutes a model |
| Everything gateway broke after a claude CLI update | Anthropic doesn't support third-party endpoints; compatibility can shift | pin the CLI version on machines that depend on gateway lanes; see [LANES.md](LANES.md#safety-and-policy) |

## Selecting the additional gateway models

Run `/setup-pstack` and assign `deepseek-pro` (`deepseek:deepseek-v4-pro@high`) or `minimax-preview` (`minimax:MiniMax-M3.1-Flash-Preview@high`) to named roles. Existing `deepseek` and `minimax` choices remain available. Each model has its own effort selection and live probe. MiniMax preview requires an eligible Token Plan key in `MINIMAX_API_KEY`; see [model choices and thinking controls](LANES.md#multiple-models-per-provider). No existing assignment changes until setup succeeds and you confirm the rendered sheet.
30 changes: 30 additions & 0 deletions docs/gateway-model-probes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# Gateway model probe evidence

Tracking: [issue #5](https://github.com/thisguymartin/pstack-flex/issues/5).

## Candidate and scope

- Source candidate: branch `flex/multiple-gateway-models`, based on `92dc0bc`, with uncommitted implementation changes.
- Implementation diff SHA-256 before this evidence file: `84b4c3ee7fc66902532e1d457048a487e6a63629b31d2d214e52585f72863c75`.
- Packaged version: 1.4.1. This candidate has not been installed as a plugin.
- Actual parent: Codex session, invoking the candidate's external runner with `--parent codex`.
- CLI: Claude Code 2.1.283.
- Each probe used `--effort high`, read-only mode, a separate empty synthetic workspace and isolated Claude configuration, and a synthetic text file. No repository or customer data was used in the prompt.
- Keys were supplied through hidden terminal input, injected into child environments, and were not included in commands, this repository, or evidence below.

## Observed results

Each probe exited 0, returned the exact requested marker, recorded `modelVerified: true` with `modelEvidence: provider-report`, and retained `costUsd: null`.

| Requested model | Reported model | Elapsed milliseconds | Receipt status |
| --- | --- | --- | --- |
| `deepseek-flash` | `deepseek-flash` | 2684 | `complete` |
| `deepseek-v4-pro` | `deepseek-v4-pro` | 7737 | `complete` |
| `MiniMax-M3` | `MiniMax-M3` | 11416 | `complete` |
| `MiniMax-M3.1-Flash-Preview` | `MiniMax-M3.1-Flash-Preview` | 5090 | `complete` |

These single short probes establish authentication, model selection, and successful completion through the runner. They do not rank coding quality or speed, prove hidden reasoning depth, or verify CLI request-body effort forwarding. The prompt included the expected marker, so completion does not independently prove a file tool was used.

## Remaining release gate

Install the exact candidate and run setup from both real Claude Code and Codex user surfaces. Verify independent model effort choices, per-model probes, mixed-provider panels, saved-sheet readback, and unchanged configuration on failed access. Record installed version, surface, action, and observed result before merge or rollout. Changing only the runner's `--parent` flag would not satisfy this gate.
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,13 @@ pstack-flex addition. The stock matrix above is upstream-owned and unchanged; th
| Family | Provider | Model | Default effort | Selectable efforts | API key variable | Base URL default |
|---|---|---|---|---|---|---|
| deepseek | deepseek | deepseek-flash | high | low medium high xhigh max | DEEPSEEK_API_KEY | https://api.deepseek.com/anthropic |
| deepseek-pro | deepseek | deepseek-v4-pro | high | low medium high xhigh max | DEEPSEEK_API_KEY | https://api.deepseek.com/anthropic |
| minimax | minimax | MiniMax-M3 | high | low medium high xhigh max | MINIMAX_API_KEY | https://api.minimax.io/anthropic |
| minimax-preview | minimax | MiniMax-M3.1-Flash-Preview | high | low medium high xhigh max | MINIMAX_API_KEY | https://api.minimax.io/anthropic |

A family identifies one `(provider, model)` pair, not an entire provider. The existing `deepseek` and `minimax` family names and descriptors remain valid. `deepseek-pro` and `minimax-preview` are additional choices with independent requested efforts. Multiple models from one provider still count as one provider for panel diversity.

MiniMax preview requires Token Plan access; set `MINIMAX_API_KEY` to the eligible subscription key. A pay-as-you-go key is not proof of preview access. The preview always thinks and supports `low` through `max`; do not disable thinking. M3 thinking is off by default at the API and requires adaptive thinking to enable it; its effort flag does not imply preview-style depth control. Selectable efforts are runner requests, not a claim that every provider applies five distinct reasoning levels. Verify CLI forwarding and model access with live probes. Sources: [MiniMax models](https://platform.minimax.io/docs/guides/models-intro), [MiniMax thinking controls](https://platform.minimax.io/docs/api-reference/text-anthropic-api), [DeepSeek Anthropic compatibility](https://api-docs.deepseek.com/guides/anthropic_api) (checked 2026-09-27).

Flex lanes have no Claude-native agent stem and always take the external runner in both parents. The base URL is a documented default; override it with `DEEPSEEK_BASE_URL` or `MINIMAX_BASE_URL`, and confirm it against the provider's current Claude Code guide during setup's live probe. The config dir defaults to `~/.pstack-flex/<provider>` (override: `PSTACK_FLEX_<PROVIDER>_CONFIG_DIR`). Secrets stay in the environment: nothing in the sheet, the receipts, or this repository carries a key.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -366,27 +366,42 @@ describe("model matrix", () => {
.slice(start + 1, end)
.map((line) => line.trim())
.filter((line) => line.startsWith("|"));
expect(table.length).toBe(2 + GATEWAY_PROVIDERS.length);
expect(table.length).toBeGreaterThan(2 + GATEWAY_PROVIDERS.length);
expect(splitRow(table[0]).join("|")).toBe(FLEX_MATRIX_HEADER.join("|"));
expect(isSeparator(splitRow(table[1]))).toBe(true);
const seen: GatewayProvider[] = [];
const seen = new Set<GatewayProvider>();
const families = new Set<string>();
const pairs = new Set<string>();
for (const line of table.slice(2)) {
const cells = splitRow(line);
expect(cells.length).toBe(FLEX_MATRIX_HEADER.length);
const [family, provider, model, defaultEffortRaw, selectableRaw, keyVar, baseUrl] =
cells;
expect(GATEWAY_PROVIDERS as readonly string[]).toContain(provider);
const gateway = provider as GatewayProvider;
seen.push(gateway);
expect(family).toBe(gateway);
seen.add(gateway);
expect(/^[a-z0-9-]+$/.test(family)).toBe(true);
expect(families.has(family)).toBe(false);
families.add(family);
const pair = `${provider}:${model}`;
expect(pairs.has(pair)).toBe(false);
pairs.add(pair);
expect(/^[A-Za-z0-9.-]+$/.test(model)).toBe(true);
const selectable = selectableRaw.split(/\s+/).map(asEffort);
expect(selectable).toContain(asEffort(defaultEffortRaw));
expect(keyVar).toBe(GATEWAY_SPECS[gateway].apiKeyVar);
expect(baseUrl).toBe(GATEWAY_SPECS[gateway].baseUrlDefault);
expect(baseUrl.startsWith("https://")).toBe(true);
}
expect(seen).toEqual([...GATEWAY_PROVIDERS]);
expect([...seen]).toEqual([...GATEWAY_PROVIDERS]);
for (const pair of [
"deepseek:deepseek-flash",
"deepseek:deepseek-v4-pro",
"minimax:MiniMax-M3",
"minimax:MiniMax-M3.1-Flash-Preview",
]) expect(pairs.has(pair)).toBe(true);
expect(setup).toContain("Never group efforts or deduplicate probes by provider alone.");
expect(setup).toContain("Different models sharing a provider count as one provider");
// The stock quad and first-run sheet must not carry flex descriptors:
// upstream's own checks parse descriptors with a lowercase-only,
// three-provider grammar and must never see a flex lane.
Expand Down
40 changes: 40 additions & 0 deletions plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -1032,6 +1032,46 @@ describe("gateway lanes", () => {
expect(readFileSync(input.outputPath, "utf8")).toBe("CLAUDE_OK");
});

for (const parent of ["claude", "codex"] as const) {
for (const [provider, model] of [
["deepseek", "deepseek-v4-pro"],
["minimax", "MiniMax-M3.1-Flash-Preview"],
] as const) {
it(`pins ${model} and effort through the ${parent} parent route`, async () => {
const dumpPath = join(scratch, "new-model-env.json");
process.env.FAKE_DUMP_ENV_PATH = dumpPath;
process.env.FAKE_REPORT_MODEL = model.toLowerCase();
const input: RunnerOptions = {
...gatewayOptions(provider, "new-model"), parent, model, effort: "max",
};
expect((await runLane(input)).exitCode).toBe(0);
const written = receipt(input.receiptPath);
expect(written).toMatchObject({
status: "complete", parent, provider, model, effort: "max",
modelVerified: true, modelEvidence: "provider-report", costUsd: null,
});
expect(written.argv[written.argv.indexOf("--model") + 1]).toBe(model);
expect(written.argv[written.argv.indexOf("--effort") + 1]).toBe("max");
const child = JSON.parse(readFileSync(dumpPath, "utf8"));
for (const key of [
"ANTHROPIC_MODEL", "ANTHROPIC_DEFAULT_OPUS_MODEL",
"ANTHROPIC_DEFAULT_SONNET_MODEL", "ANTHROPIC_DEFAULT_HAIKU_MODEL",
"CLAUDE_CODE_SUBAGENT_MODEL",
]) expect(child[key]).toBe(model);
});

it(`rejects a substituted ${model} in the ${parent} parent route`, async () => {
const input: RunnerOptions = {
...gatewayOptions(provider, "substituted-model"), parent, model,
};
process.env.FAKE_REPORT_MODEL = provider === "deepseek" ? "deepseek-flash" : "MiniMax-M3";
expect((await runLane(input)).exitCode).toBe(65);
expect(receipt(input.receiptPath).status).toBe("malformed-output");
expect(existsSync(input.outputPath)).toBe(false);
});
}
}

it("verifies a case-shifted served model for MiniMax", async () => {
process.env.FAKE_REPORT_MODEL = "minimax-m3";
const input = gatewayOptions("minimax", "case-shift");
Expand Down
Loading
Loading