diff --git a/CHANGES.md b/CHANGES.md index 92906368..ff9a2957 100644 --- a/CHANGES.md +++ b/CHANGES.md @@ -1,7 +1,25 @@ # CHANGES — applied substitutions +## Unreleased: multiple gateway model choices + +- Add DeepSeek V4 Pro and MiniMax M3.1 Flash Preview as independent setup families alongside existing Flash and M3 choices. Keep provider-owned routing and existing descriptors. +- Validate unique model families and provider/model pairs instead of requiring one row per provider; cover both parent routes and substituted-model rejection. +- Document preview Token Plan access, model-specific thinking semantics, and required installed live validation. Tracked in [pstack-flex #5](https://github.com/thisguymartin/pstack-flex/issues/5). + This port applies the Cursor → Claude Code substitutions in skill bodies. Earlier drafts left them flagged; this revision resolves them. A later pass added a Codex build that shares the same skills; see [Codex port](#codex-port) below. +## pstack-flex (unreleased) — gateway lanes and optional families + +Fork of open-pstack v1.4.1. Additive changes, all in port-owned files: + +- Runner: new gateway providers `deepseek` and `minimax` (`runner/flex-providers.ts`). Each spawns the stock `claude` binary with the exact claude argv, plus injected environment: the lab's Anthropic-compatible endpoint, `ANTHROPIC_AUTH_TOKEN` from `DEEPSEEK_API_KEY`/`MINIMAX_API_KEY`, model pins, and an isolated `CLAUDE_CONFIG_DIR` (`~/.pstack-flex/`). Inherited `ANTHROPIC_*` values are deleted before injection so a parent's credentials or endpoint never bleed into a gateway child. +- OAuth-leak guard: a gateway lane refuses to start (in-process, `unauthenticated` receipt, exit 77) when its API key variable is missing or when its config dir carries a claude.ai OAuth credentials file, so a claude.ai login can never be pointed at a third-party endpoint. +- Gateway preflight is `claude --version`; the one-shot invocation is the real auth test. Gateway receipts force `costUsd` to null (the CLI prices at Anthropic rates) and match served models case-insensitively, falling back to `modelEvidence: "pinned-argv"` like Codex. +- `provider-dispatch.md`: new additive "Flex model matrix" section, extended route table, gateway preflight semantics, and the panel-diversity rule (arena runners and interrogate reviewers span at least two providers unless the operator explicitly confirms otherwise). The stock model matrix is byte-unchanged. +- Optional GPT-6 families: the additional model matrix declares `codex:gpt-6-astra`, `codex:gpt-6-sol`, and `codex:gpt-6-luna` with default effort `high` and selectable `low`, `medium`, `high`, `xhigh`, and `max`. Setup can assign and probe each for `architect runners` or another configurable role. Codex parents use native `spawn_agent`; Claude Code parents use the external Codex runner. GPT-6 Sol has its own `sol-6` family. Stock models and first-run role assignments stay unchanged. +- `setup-pstack`: role assignments are selected first, and only assigned families get effort questions and probes; there is no requirement to assign every matrix family (mirrors upstream PR #73 / issue #72). The first-run sheet, its stock quad, and the fail-closed write rules are unchanged. +- Tests: the model-matrix contract gains a flex-matrix section check cross-validated against the runner's gateway specs; runner, commands, parse-output, and CLI tests cover env injection, the guard, cost nulling, and case-insensitive verification. Nothing in the suite performs network I/O. + ## 1.4.1 syncs to Cursor pstack 0.15.1 Open Pstack 1.4.1 tracks Cursor pstack 0.15.1 at `f8abeddd1862dc73704e3d719dd73df0d51b8c71`. Poteto-mode now requires each claim to include its evidence or a measured, inferred, or guess label in the same sentence. Agents also run any check they can run themselves instead of handing that check to the user. No playbook, model, runtime, or dependency changed. diff --git a/LICENSE b/LICENSE index 6b540023..643dd5d0 100644 --- a/LICENSE +++ b/LICENSE @@ -1,6 +1,7 @@ MIT License Copyright (c) 2026 Lauren Tan +Copyright (c) 2026 Martin Patino (pstack-flex modifications) Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal diff --git a/NOTICE.md b/NOTICE.md index 0ec8a23e..a1be00dc 100644 --- a/NOTICE.md +++ b/NOTICE.md @@ -2,6 +2,10 @@ This plugin is a port of upstream MIT-licensed work. All upstream copyright notices and license terms are preserved. The open-pstack history begins from `michael-denyer/pstack-claude` through proven import commit `053ed78732e3b71826933170eafe7f7782dda844`. +## pstack-flex provenance + +This repository, **pstack-flex** (Martin Patino), is a fork of [ericlitman/open-pstack](https://github.com/ericlitman/open-pstack) at v1.4.1 (`de67e6b40511814171e5e4c8ad7af3b79f07c9ee`), which ports [Lauren Tan's pstack](https://github.com/cursor/plugins/tree/main/pstack) (Cursor) to Claude Code and Codex. Provenance chain: pstack-flex <- ericlitman/open-pstack <- cursor/plugins/pstack. All licenses remain MIT; every upstream license and notice file is preserved. The flex gateway providers, optional families, docs, and tests are (c) 2026 Martin Patino, MIT, and are inventoried in [UPSTREAM-FLEX.md](UPSTREAM-FLEX.md). + ## Upstream sources | Component | Upstream | Copyright | License | License file | diff --git a/README.md b/README.md index 64901115..745ea4eb 100644 --- a/README.md +++ b/README.md @@ -12,6 +12,14 @@ Lauren built pstack from the skills she uses to ship code at Cursor. In a [55-mi Open Pstack is an unofficial community project that makes pstack work in Claude Code and Codex. If Cursor is your main coding environment, use [Lauren's original pstack](https://github.com/cursor/plugins/tree/main/pstack). If Claude Code or Codex is your main coding environment, use this repository. +## This fork: pstack-flex + +**pstack-flex** is a fork of [ericlitman/open-pstack](https://github.com/ericlitman/open-pstack) at v1.4.1. Setup can assign only the model families you have. The `deepseek:*` and `minimax:*` routes run the stock `claude` binary against each lab's Anthropic-compatible endpoint with that lab's API key. The runner isolates Claude configuration, strips inherited provider routing and Anthropic headers, and rejects OAuth credentials found in the gateway config directory. Gateway receipts keep token usage but set `costUsd` to null because Claude Code's cost estimate uses Anthropic prices. + +The [lane guide](docs/LANES.md) covers setup, costs, and safety notes. The fork's provenance and sync process are recorded in [UPSTREAM-FLEX.md](UPSTREAM-FLEX.md). Anthropic does not support pointing Claude Code at non-Anthropic endpoints; use synthetic data for gateway testing and keep API keys in your local environment. + +New here? **[docs/USAGE.md](docs/USAGE.md)** is the walkthrough: diagrams of how work flows through the lanes, three setup configurations (full frontier, hybrid saver, zero-subscription), copy-paste examples for the daily skills, and troubleshooting. + ## What pstack does pstack is a plugin for coding agents. It is not a new model or a hosted service. It gives your agent engineering rules, step-by-step workflows for different kinds of work, focused skills, and small local tools. @@ -39,7 +47,7 @@ You need a current Claude Code or Codex installation. For the full four-model re Run these commands inside Claude Code: ```text -/plugin marketplace add ericlitman/open-pstack +/plugin marketplace add thisguymartin/pstack-flex /plugin install pstack@open-pstack /reload-plugins ``` @@ -49,7 +57,7 @@ Run these commands inside Claude Code: Run these commands in your shell: ```shell -codex plugin marketplace add ericlitman/open-pstack --ref main +codex plugin marketplace add thisguymartin/pstack-flex --ref main codex plugin add pstack@open-pstack ``` @@ -80,7 +88,7 @@ In Codex, ask: Use pstack:setup-pstack to configure pstack. ``` -Setup checks the models you can actually run, shows how each one will start, and asks before saving the choices. The current default group uses Fable, GPT-5.6 Sol, Grok 4.6, and Opus. +Setup checks the models you can actually run, shows how each one will start, and asks before saving the choices. The current default group uses Fable, GPT-5.6 Sol, Grok 4.6, and Opus. You can also assign GPT-6 Astra, Sol, or Luna to any role through Codex. An older model sheet starts using the rolling aliases in memory as soon as this release is installed. Run setup once after updating to persist that migration. It replaces versioned Fable and Opus entries while preserving every role assignment and effort selection. diff --git a/UPSTREAM-FLEX.md b/UPSTREAM-FLEX.md new file mode 100644 index 00000000..7ecbf211 --- /dev/null +++ b/UPSTREAM-FLEX.md @@ -0,0 +1,46 @@ +# Flex fork synchronization + +pstack-flex layers on top of open-pstack's own upstream tracking. Two sync relationships exist: + +1. `cursor/plugins/pstack` -> `ericlitman/open-pstack` — documented in [UPSTREAM.md](UPSTREAM.md), unchanged by this fork. +2. `ericlitman/open-pstack` -> `thisguymartin/pstack-flex` — this document. + +## Fork point + +| Source | Value | +| --- | --- | +| Repository | `https://github.com/ericlitman/open-pstack.git` | +| Tag | `v1.4.1` | +| Commit | `de67e6b40511814171e5e4c8ad7af3b79f07c9ee` | +| Tracks Cursor pstack | `0.15.1` (`f8abedd`) | + +The fork keeps full upstream history. The `upstream` remote points at ericlitman/open-pstack. + +## What the fork owns + +All flex changes are additive and live in port-owned files so upstream merges stay cheap: + +- `plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts` and `flex-providers.test.ts` (new) +- Gateway-provider hooks in `runner/{types,commands,run,parse-output,cli}.ts` and their tests +- The "Additional model matrix" and "Flex model matrix" sections and route-table columns in `references/provider-dispatch.md` +- The assignment-first restructure of `skills/setup-pstack/SKILL.md` +- `docs/LANES.md`, this file, the README fork section, and the NOTICE/LICENSE/CHANGES additions + +The stock model matrix, the first-run sheet, every upstream skill body, and the static quad invariants are byte-unchanged. Flex rows and setup instructions are additive. + +## Merge procedure + +```shell +git fetch upstream +git switch -c merge-rehearsal +git merge --no-ff --no-commit upstream/main +# inspect, resolve, run the full local gate, then merge for real or abort +``` + +Expected conflict surface on future upstream releases: + +- `plugins/pstack/skills/setup-pstack/SKILL.md` — upstream issue #88 (1.5.0, syncing Cursor pstack 0.15.5) folds upstream PR #73, which moves setup to the same assignment-first, probe-only-assigned shape this fork already uses. Resolve toward upstream's wording wherever it covers the same rule; keep the flex families and the diversity rule. +- `plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts` — upstream 1.5.0 changes the stock panel to three lanes. Take upstream's stock assertions verbatim; the flex-matrix describe block is fork-owned and should survive as-is. +- `plugins/pstack/skills/poteto-mode/references/provider-dispatch.md` — stock matrix and default-panel prose are upstream's; the flex section is fork-owned. + +After every merge: run the full local gate (`bun install --frozen-lockfile`, `bun run test`, `bun run typecheck`, manifest JSON parse, `PSTACK_STATIC_ONLY=1 bash tests/skill-collision-repro.sh`), then record the installed version, action, and observed result for each affected harness in the pull request before tagging. diff --git a/docs/LANES.md b/docs/LANES.md new file mode 100644 index 00000000..cf725498 --- /dev/null +++ b/docs/LANES.md @@ -0,0 +1,149 @@ +# Lanes: models, providers, and cost control + +pstack-flex's reason to exist: you choose which models run and what they cost. This document covers the lane concepts, the gateway environment reference, prices, the zero-subscription walkthrough, and the safety rules. + +Prices and endpoints below were verified 2026-09-25 and drift. Re-verify against each provider's own docs before relying on a number. + +## Lane kinds + +| Kind | Lanes | Auth | Billing | Route | +| --- | --- | --- | --- | --- | +| Subscription | `claude:fable`, `claude:opus`, `codex:gpt-5.6-sol`, `grok:grok-4.6`; optional `codex:gpt-6-astra`, `codex:gpt-6-sol`, `codex:gpt-6-luna` | each CLI's own login | that CLI's plan | native or external per the route table | +| Gateway (flex) | DeepSeek Flash / V4 Pro; MiniMax M3 / M3.1 Flash Preview | API key in the environment | provider billing; preview requires Token Plan | always the external runner | + +A gateway lane is the stock `claude` binary env-pointed at the lab's Anthropic-compatible endpoint. There is no custom agent loop and no separate harness: the same runner that spawns Codex and Grok lanes spawns gateway lanes with injected environment. Both labs document this Claude Code setup themselves (DeepSeek: `deepseek-ai/awesome-deepseek-agent`, `docs/claude_code.md`; MiniMax: platform.minimax.io, Claude Code guide). + +## Optional GPT-6 Codex families + +The additional model matrix adds three Codex families. They use the same ChatGPT login as `codex:gpt-5.6-sol`: + +| Family | Descriptor at default requested effort | Codex's own description | +| --- | --- | --- | +| astra | `codex:gpt-6-astra@high` | Frontier tier for the most demanding work | +| sol-6 | `codex:gpt-6-sol@high` | Coding and everyday workhorse | +| luna | `codex:gpt-6-luna@high` | Fast, low-cost tier for easier tasks | + +These are choices, not replacements. The first-run sheet still uses `codex:gpt-5.6-sol@max`. An existing sheet changes only when you assign a role to a GPT-6 family in `/setup-pstack`. The `sol-6` family is separate from the stock `sol` family, so each keeps its own effort. All Codex families count as one provider for panel diversity, so Astra plus Sol does not satisfy the two-provider rule. The route matches Sol: native `spawn_agent` in a Codex parent, and the external runner (`codex exec`) in a Claude Code parent. Codex also lists an `ultra` effort for Astra and GPT-6 Sol. It is outside the pstack effort universe and is not selectable. The descriptions and effort lists come from the Codex CLI 0.157.1 model list, checked 2026-09-27. + +## Multiple models per provider + +The flex matrix now includes four independently assignable model families: + +| Family | Descriptor at default requested effort | Selection guidance | +| --- | --- | --- | +| deepseek | `deepseek:deepseek-flash@high` | Existing everyday option | +| deepseek-pro | `deepseek:deepseek-v4-pro@high` | Candidate for difficult debugging, architecture, and review | +| minimax | `minimax:MiniMax-M3@high` | Existing MiniMax option | +| minimax-preview | `minimax:MiniMax-M3.1-Flash-Preview@high` | Preview coding option with tunable thinking | + +These are choices, not automatic replacements or a performance ranking. Existing sheets keep their assignments. In `/setup-pstack`, assign named roles to the desired model family; efforts and probes are independent per model, even for models sharing a key. Two models from one provider count as one provider for panel diversity. No runtime routing change or new configuration file is needed. + +As of 2026-09-27, [MiniMax's model guide](https://platform.minimax.io/docs/guides/models-intro) restricts M3.1 Flash Preview to Token Plan and MiniMax Code. For gateway access, supply the eligible Token Plan key as `MINIMAX_API_KEY`; the live probe must confirm entitlement. It is not a zero-subscription option. A working M3 call does not establish preview access. + +[MiniMax's Anthropic API](https://platform.minimax.io/docs/api-reference/text-anthropic-api) documents always-on thinking for the preview and `output_config.effort` from `low` to `max`. Higher effort increases thinking latency; the matrix proposes `high`, while the API defaults to `max` when omitted. M3 defaults to thinking off at the API and needs adaptive thinking to enable it. Its requested effort flag is not evidence of the preview's depth controls. Verify the installed CLI forwards the intended parameters; receipts prove requested effort, not hidden applied depth. [DeepSeek documents V4 Pro through its Anthropic endpoint](https://api-docs.deepseek.com/guides/anthropic_api). + +Before recommending a fastest or strongest default, compare the same synthetic coding tasks for correctness, completion time, tool-call reliability, token usage, and actual provider billing. Preview pricing and plan limits must be checked against the active plan rather than inferred from M3 rates. + +## Gateway environment reference + +Set by you: + +| Variable | Required | Meaning | +| --- | --- | --- | +| `DEEPSEEK_API_KEY` / `MINIMAX_API_KEY` | yes, per lane | the lab's API key; the lane refuses to start without it | +| `DEEPSEEK_BASE_URL` / `MINIMAX_BASE_URL` | no | endpoint override; defaults are in the flex model matrix | +| `PSTACK_FLEX_DEEPSEEK_CONFIG_DIR` / `PSTACK_FLEX_MINIMAX_CONFIG_DIR` | no | config-dir override; default `~/.pstack-flex/` | +| `DEEPSEEK_MAX_CONTEXT_TOKENS` / `MINIMAX_MAX_CONTEXT_TOKENS` | no | context-cap override for the claude CLI | + +Injected by the runner at spawn time (never written to disk, never in receipts): `ANTHROPIC_BASE_URL`, `ANTHROPIC_AUTH_TOKEN`, the model pins (`ANTHROPIC_MODEL`, the opus/sonnet/haiku alias defaults, `CLAUDE_CODE_SUBAGENT_MODEL`), `CLAUDE_CODE_ATTRIBUTION_HEADER=0`, `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1`, and `CLAUDE_CONFIG_DIR`. The runner first removes inherited `ANTHROPIC_*` values and Claude Code cloud-provider flags from the parent session. + +## Storing keys + +Keys reach a lane through the environment only; the runner never writes them to disk, receipts, or sheets. So key hygiene is entirely about how your shell gets them. Do not put raw keys in dotfiles or committed `.env` files. + +Recommended: your OS keychain, loaded on demand. + +- **macOS** (built in, encrypted at rest, unlocks with login): + + ```zsh + # once per key — prompts for the value, nothing lands in shell history + security add-generic-password -a "$USER" -s pstack-deepseek -w + security add-generic-password -a "$USER" -s pstack-minimax -w + + # in .zshrc: a function, not an export — keys enter env only when called + pstack-keys() { + export DEEPSEEK_API_KEY=$(security find-generic-password -a "$USER" -s pstack-deepseek -w) + export MINIMAX_API_KEY=$(security find-generic-password -a "$USER" -s pstack-minimax -w) + } + ``` + +- **Linux**: `pass` (GPG-encrypted, git-syncable) or `secret-tool` (libsecret) with the same load-on-demand function shape. +- **1Password CLI**: `op run --env-file=.env.tpl -- claude` injects the keys at process start with biometric unlock and exports nothing into the shell permanently. +- **direnv**: fine for per-project scoping (gateway lanes are per-project opt-in anyway), but a raw `.envrc` is plaintext — have it call the keychain instead of holding the key. + +Honest threat model: encryption at rest protects against dotfile repos, backups, and file theft. Once a key is in process env, any process running as your user can read it — the same exposure your CLI OAuth credential files already have. Keychain storage plus two ops controls is the right amount: **set spend caps on the DeepSeek and MiniMax dashboards** (the real blast-radius limiter) and rotate keys if a machine is ever compromised. + +## Prices (verified 2026-09-25 — re-check before budgeting) + +| Lane | Price per million tokens | Notes | +| --- | --- | --- | +| DeepSeek V4.1-Flash (`deepseek-flash`) | $0.30 in / $1.20 out peak; $0.15 / $0.60 off-peak; cache hits near-free | Off-peak windows: 01:00-04:00 and 06:00-10:00 UTC on weekdays. The discount is automatic on DeepSeek's side; pstack-flex surfaces the window but never delays your work to hit it. MIT open weights. | +| DeepSeek V4-Pro | $1.32 / $3.96 peak; half off-peak | Stronger model for hard lanes; assign it per role if wanted. | +| MiniMax M3 (`MiniMax-M3`) | $0.30 / $1.20 at up to 512K input; higher above | 1M context. Custom community model license (irrelevant for API use). | +| Claude / Codex / Grok subscription lanes | plan-dependent | Billed by each provider's plan, not per token here. | + +Gateway receipts always report `costUsd: null`: the claude CLI computes `total_cost_usd` at Anthropic list prices, which would be fiction for third-party traffic. Token usage in receipts is real — multiply it by the table above. + +## Zero-subscription walkthrough + +Goal: run poteto-mode and its panels with no Claude, ChatGPT, or Grok plan — only two API keys. The Claude Code binary is a free download; a subscription is only needed to reach Anthropic's servers. + +1. Install the claude CLI, Bun, and this plugin as usual. Do not run `claude login` anywhere in this setup. +2. Export `DEEPSEEK_API_KEY` and `MINIMAX_API_KEY`. +3. Make the parent session itself a DeepSeek session — same mechanism as a lane, applied to your interactive shell: + + ```shell + export ANTHROPIC_BASE_URL="" + export ANTHROPIC_AUTH_TOKEN="$DEEPSEEK_API_KEY" + export ANTHROPIC_MODEL="deepseek-flash" + export CLAUDE_CODE_SUBAGENT_MODEL="deepseek-flash" + export CLAUDE_CONFIG_DIR="$HOME/.pstack-flex/parent-deepseek" + claude + ``` + + Native `claude:*` lanes spawned by this parent inherit the endpoint, so the fable/opus role slots ride DeepSeek too. +4. Run `/setup-pstack`. Assign roles across the `deepseek` and `minimax` families (a `budget-duo` style panel), skip the unassigned stock families, and let the probes confirm both endpoints. +5. Panels keep real diversity: DeepSeek and MiniMax are two distinct providers, which satisfies the two-provider panel rule without any override. + +Quality note: this trades peak capability for cost control. The hardest-task role on a frontier subscription lane is a config choice you can add later without touching anything else. + +## Safety and policy + +- **Unsupported, not prohibited.** Anthropic's docs state that routing Claude Code to non-Claude models through gateways is not supported. No terms clause or enforcement against pointing the unmodified binary at a third-party endpoint was found (2026-09-25), but a CLI update can break compatibility without notice. Pin the claude CLI version on machines that depend on gateway lanes and bump it deliberately. +- **Never a claude.ai login on a gateway path.** Do not run `claude login` or `claude setup-token` inside any `~/.pstack-flex/` config dir. The runner enforces this: a gateway lane refuses to start when its config dir carries an OAuth credentials file. Caveat: on macOS the CLI may store credentials in the Keychain where the file check cannot see them — the rule above is the real defense; the check is a backstop. +- **Privacy: gateway lanes are opt-in per project.** Do not send client or customer code to third-party providers by default. Keep sensitive repositories on subscription lanes, and enable gateway lanes deliberately, per project. +- **No runner fallback.** A failed gateway lane is a named dropout receipt. A reported model mismatch fails the lane. If the endpoint reports no model, the receipt says `modelVerified: false` and `modelEvidence: "pinned-argv"`; this cannot prove which model the gateway served. Confirm supported model slugs during the live probe. + +## Optional lanes + +- **OpenRouter (off by default).** OpenRouter has no Anthropic-format endpoint, so a lane needs a local translator that serves `/v1/messages` — musistudio/claude-code-router or a version-pinned LiteLLM — with `ANTHROPIC_BASE_URL` pointed at it. That is one extra long-running local process, which is why it is documented rather than shipped. Expect roughly a 5.5% credit fee on top of provider list prices. If you build it, model it as another gateway provider in `flex-providers.ts`. +- **Local via Ollama (planned).** Ollama serves an Anthropic-compatible API since v0.14, so a `local` gateway provider pointed at it is the natural next lane: full compute control, zero per-token cost, your hardware. Not wired in yet. + +## Adding a gateway provider + +Any lab that serves an Anthropic-compatible `/v1/messages` endpoint can become a gateway lane. The runner, parser, and preflight branch on `isGatewayProvider`, so no `switch` needs a new case. + +1. Add the provider name to `GATEWAY_PROVIDERS` in `plugins/pstack/skills/poteto-mode/scripts/runner/types.ts`. +2. Add its row to `GATEWAY_SPECS` in `runner/flex-providers.ts`: API key variable, base URL default, override variables, and context-window default. Typecheck fails until this row exists. +3. Add its row to the "Flex model matrix" in `plugins/pstack/skills/poteto-mode/references/provider-dispatch.md`. `model-matrix.test.ts` fails until the key variable and base URL match the spec. +4. Add its probe row to the table in `plugins/pstack/skills/setup-pstack/SKILL.md`, its variables to the gateway environment reference above, and its prices to the price table. +5. Run the live validation checklist below for the new lane before merging. + +## Live validation checklist (before merge or rollout, real keys, never in CI) + +- V1: one DeepSeek probe through the runner (`--provider deepseek --model deepseek-flash --effort high`, read-only). Expect a `complete` receipt with `costUsd: null`; record the `reportedModel` string and confirm the base-URL default against DeepSeek's current guide; confirm `--effort` is accepted end-to-end. +- V2: same for MiniMax (`MiniMax-M3`); record the served-model casing. +- New-model gate: install the exact candidate and run `/setup-pstack` from both real Claude Code and Codex surfaces. Select Flash plus Pro and M3 plus Preview, verify independent efforts and probes, then run a read-only mixed panel. Record installed version/commit, surface, action, requested model/effort, served model, and observed result. Verify a failed preview entitlement probe leaves the sheet unchanged and does not select M3. A fake CLI regression test is not this gate. +- V3: run `claude auth status --json` inside a fresh flex config dir with `ANTHROPIC_AUTH_TOKEN` set and record the output here. On macOS, confirm whether `claude login` under an explicit `CLAUDE_CONFIG_DIR` writes `.credentials.json` or the Keychain. +- V4: the zero-subscription walkthrough above, end to end, on a machine with no stored provider logins. +- V5: OAuth guard live: `claude login` inside a scratch flex config dir, run a lane, confirm the refusal receipt, then delete that login. diff --git a/docs/USAGE.md b/docs/USAGE.md new file mode 100644 index 00000000..9819c964 --- /dev/null +++ b/docs/USAGE.md @@ -0,0 +1,247 @@ +# Using pstack-flex + +The walkthrough: what this plugin is, how work flows through it, how to set it up on the models you actually have, and copy-paste examples for the skills you will use daily. Lane mechanics and pricing live in [LANES.md](LANES.md); the fork's delta over upstream is in [UPSTREAM-FLEX.md](../UPSTREAM-FLEX.md). + +## What this is + +pstack is a plugin of engineering skills, playbooks, and small local tools for coding agents — not a model, not a service. You hand `poteto-mode` a task; it matches the task to a playbook, works the steps, and leaves evidence (diffs, runs, receipts) you can inspect instead of asking for trust. Its sharpest edge is multi-model adversarial review: several different model families challenge important work, because the adversarial signal comes from model diversity, not assigned personas. + +pstack-flex adds one thing on top: **you choose the models and the compute**. Any subset of families works, and two open labs — DeepSeek and MiniMax — are first-class lanes on plain API keys, down to a zero-subscription setup. + +If you also use my [thisguyskills](https://github.com/thisguymartin/skills) collection: that repo decides **what** to build (shaping, spec, Linear, handoff) and its handoff ends with "Use `pstack:poteto-mode`" — which is exactly where this repo picks up. + +## The big picture + +```mermaid +flowchart TD + T([Your task]) --> P["/pstack:poteto-mode"] + P --> PB[Playbook match
feature, bug-fix, refactoring, perf, ...] + PB --> S[Skills fire per step
how, tdd, interrogate, arena, ...] + S --> F{Lane fan-out} + F --> N1["claude:fable / claude:opus
native Agent (Claude sub)"] + F --> N2["codex:gpt-5.6-sol
native or codex CLI (ChatGPT sub)"] + F --> N3["grok:grok-4.6
grok CLI (Grok sub)"] + F --> G1["deepseek:deepseek-flash
runner + env -> DeepSeek API (key)"] + F --> G2["minimax:MiniMax-M3
runner + env -> MiniMax API (key)"] + N1 --> R[Outputs + receipts] + N2 --> R + N3 --> R + G1 --> R + G2 --> R + R --> V[Verification: run it, judge it,
cross-model consensus] + V --> PR([Review-ready PR]) +``` + +Every lane is a real agent process with tools and file access. The parent harness (your Claude Code or Codex session) resolves the route once; children never pick their own models. + +## Install this fork + +The marketplace keeps upstream's name (`open-pstack`), so only the source changes. + +Claude Code: + +```text +/plugin marketplace add thisguymartin/pstack-flex +/plugin install pstack@open-pstack +/reload-plugins +``` + +Codex: + +```shell +codex plugin marketplace add thisguymartin/pstack-flex --ref main +codex plugin add pstack@open-pstack +``` + +Plus [Bun](https://bun.sh) for the lane runner, and `multi_agent = true` under `[features]` in `~/.codex/config.toml` if Codex is your parent. Sign in only to the CLIs whose subscriptions you actually have — missing families are fine now. + +## Keys for the gateway lanes + +DeepSeek and MiniMax have no login flow here; their lanes read an API key from your environment at spawn time. The runner never writes keys to disk or receipts, so the only question is how the env gets populated. Don't paste keys into `.zshrc` — store them encrypted and load on demand. macOS Keychain, built in and free: + +```zsh +# once: store each key (prompts for the value, nothing in shell history) +security add-generic-password -a "$USER" -s pstack-deepseek -w +security add-generic-password -a "$USER" -s pstack-minimax -w + +# in .zshrc: a function, not an export — keys enter env only when you call it +pstack-keys() { + export DEEPSEEK_API_KEY=$(security find-generic-password -a "$USER" -s pstack-deepseek -w) + export MINIMAX_API_KEY=$(security find-generic-password -a "$USER" -s pstack-minimax -w) +} +``` + +Daily flow: `pstack-keys -> claude -> /pstack:poteto-mode`. Alternatives, the threat model, and the spend-cap advice are in [LANES.md](LANES.md#storing-keys). Set spend caps on both provider dashboards; that is the real blast-radius control. + +## First-time setup: /setup-pstack + +```text +/pstack:setup-pstack +``` + +(Codex: `Use pstack:setup-pstack to configure pstack.`) + +Setup is assignment-first: pick which roles run on which families, answer one effort question per **assigned** family, and only assigned families get probed. Unassigned families are skipped, not errors. Every probe is a real one-turn run — a failed probe writes nothing. Three configurations that make sense: + +**A. Full frontier** (Claude + ChatGPT + Grok subs) — accept the defaults; behaves exactly like stock upstream: + +```text +arena runners: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xhigh, claude:opus@xhigh +``` + +**B. Hybrid saver** (Claude sub + two API keys) — frontier judgment, cheap volume: + +```text +feature, refactoring: deepseek:deepseek-flash@high +bug-fix: deepseek:deepseek-flash@high +judgment and prose: claude:fable@max +hardest tasks: claude:fable@max +swarm workers: deepseek:deepseek-flash@high +arena runners: claude:fable@max, deepseek:deepseek-flash@high, minimax:MiniMax-M3@high +interrogate reviewers: claude:fable@max, deepseek:deepseek-flash@high, minimax:MiniMax-M3@high +``` + +**C. Zero-subscription budget duo** (nothing but two keys) — start your parent session env-pointed at DeepSeek (walkthrough in [LANES.md](LANES.md#zero-subscription-walkthrough)), then assign everything across the two flex families: + +```text +arena runners: deepseek:deepseek-flash@high, minimax:MiniMax-M3@high +interrogate reviewers: deepseek:deepseek-flash@high, minimax:MiniMax-M3@high +``` + +Two labs are two distinct families, so panels keep real diversity without any override. A single-provider panel needs your explicit confirmation — by design. + +## Daily driving: the skills, with examples + +**poteto-mode** — the default entry point for any real task. It stays sticky across turns and pairs well with long autonomous sessions. + +```text +/pstack:poteto-mode + +Take ENG-142: saved reports lose their date-range filter after rename. +Repro is in the issue. Fix it, prove it in the running app, and prep the PR. +``` + +**interrogate** — multi-model review of a decision, design, or diff. Reviewers come from different families; the parent sorts their findings. + +```text +/pstack:interrogate + +Review this migration plan in docs/plans/report-store.md. Attack the +premise, the rollout order, and anything that loses data on rollback. +``` + +```mermaid +flowchart LR + Q[Decision or diff] --> A[Reviewer A
family 1] + Q --> B[Reviewer B
family 2] + Q --> C[Reviewer C
family 3] + A --> S[Parent synthesizes] + B --> S + C --> S + S --> O["consensus (2+ models) -> act on
lone findings -> consider
disagreements -> resolve explicitly"] +``` + +**arena** — N parallel attempts at the same task, an independent cross-judge, then graft the best parts onto a base. + +```text +/pstack:arena + +Implement the rate limiter from the spec in docs/spec.md. Run the +configured arena panel and keep the winner's tests regardless of base. +``` + +```mermaid +flowchart LR + T[Task] --> C1[Candidate 1] + T --> C2[Candidate 2] + T --> C3[Candidate 3] + C1 --> J[Cross-judge
different provider] + C2 --> J + C3 --> J + J --> G[Pick base + graft
best pieces] +``` + +**swarm** — same-shaped work fanned across N workers, one combined report. Good for sweeps: "apply this codemod across packages," "audit every endpoint for X." + +```text +/pstack:swarm + +Audit every handler under src/api/ for missing input validation. +One worker per file group, combined findings ranked by severity. +``` + +**architect** — competing designs from different families, scored by a judge on yet another family, before any code. + +```text +/pstack:architect + +Design the offline sync layer: local-first edits, conflict policy, +and migration from the current always-online store. +``` + +Worth knowing by name: `how` (explain how something works before touching it), `why` (root-cause an incident with your MCP context), `tdd`, `unslop` (de-slop prose and code), `fix-ci`, `babysit` (drive a PR to green). The 23 `principle-*` leaves are loaded by poteto-mode as needed — you rarely invoke them directly. + +## What actually happens on a gateway lane + +No new harness. The same runner that launches Codex and Grok lanes spawns the stock `claude` binary with swapped environment: + +```mermaid +sequenceDiagram + participant P as Parent session + participant R as pstack-runner + participant C as claude -p (subprocess) + participant D as DeepSeek / MiniMax API + P->>R: lane: deepseek:deepseek-flash@high + R->>R: guard: DEEPSEEK_API_KEY set?
config dir free of OAuth creds? + Note over R: refusal = unauthenticated receipt,
no subprocess ever spawned + R->>C: spawn with ANTHROPIC_BASE_URL,
ANTHROPIC_AUTH_TOKEN, isolated CLAUDE_CONFIG_DIR + C->>D: every model request in the agent loop + D-->>C: completions + C-->>R: JSON result + R-->>P: output file + receipt +``` + +The guard order matters: key check and OAuth check happen in-process **before** anything runs, so a claude.ai login can never be pointed at a third-party endpoint. Inherited `ANTHROPIC_*` values from your parent session are stripped before injection. + +## Reading receipts + +Every external lane writes a JSON receipt next to its output. The fields that matter: + +| Field | Meaning | +| --- | --- | +| `status` | `complete`, or a named dropout (`unauthenticated`, `unavailable-cli`, `timed-out`, ...) | +| `modelVerified` + `modelEvidence` | `provider-report` = the endpoint echoed the requested model (case-insensitive for gateways). `pinned-argv` = it didn't, but the argv pinned it — normal for Codex and sometimes gateways | +| `usage` | real token counts — trust these | +| `costUsd` | real for claude/grok subscription lanes; **always `null` on gateway lanes** (the CLI would price at Anthropic rates). Multiply `usage` by the [LANES.md](LANES.md) table instead | + +## Cost playbook + +- High-volume code-writing roles (`feature`, `bug-fix`, `swarm workers`) -> `deepseek:deepseek-flash` — cheapest tokens, near-free cache hits, and half price in the off-peak window. +- Long-context research and big-repo reading -> `minimax:MiniMax-M3` — 1M context. +- `judgment and prose` and `hardest tasks` -> your best frontier lane if you have one; this is the last role to economize. +- Panels: one frontier + two flex lanes gets you three-family diversity at a fraction of three subscriptions. + +## Troubleshooting + +| Symptom | Meaning | Fix | +| --- | --- | --- | +| Receipt `unauthenticated`, exit 77, "KEY is not set" | lane env missing | run your `pstack-keys` function (or export the key) in the shell that starts the parent | +| Receipt `unauthenticated`, "OAuth credentials found" | a claude.ai login sits in the lane's config dir | that's the leak guard working; remove the login from `~/.pstack-flex/` — never `claude login` there | +| Receipt `unauthenticated` after the model ran | the endpoint rejected the key (401) | check the key and the base URL against the provider's current guide | +| Exit 69 `unavailable-cli` | the `claude` binary isn't on PATH for the runner | install it or fix PATH | +| A panel ran with fewer lanes than configured | a lane dropped out with a named receipt | read that receipt; pstack proceeds N-1 and never silently substitutes a model | +| Everything gateway broke after a claude CLI update | Anthropic doesn't support third-party endpoints; compatibility can shift | pin the CLI version on machines that depend on gateway lanes; see [LANES.md](LANES.md#safety-and-policy) | + +## Selecting the GPT-6 Codex models + +Run `/setup-pstack` and assign `astra` (`codex:gpt-6-astra@high`), `sol-6` (`codex:gpt-6-sol@high`), or `luna` (`codex:gpt-6-luna@high`) to named roles, such as `architect runners`. Every role you do not change keeps its current descriptor, including the stock `codex:gpt-5.6-sol@max` defaults. Each GPT-6 family gets its own effort question and live probe. They need only your Codex login. See [optional GPT-6 Codex families](LANES.md#optional-gpt-6-codex-families). + +For example, this row puts Astra on the architect panel and keeps the other providers: + +```text +architect runners: claude:fable@max, codex:gpt-6-astra@high, grok:grok-4.6@xhigh, claude:opus@xhigh +``` + +## Selecting the additional gateway models + +Run `/setup-pstack` and assign `deepseek-pro` (`deepseek:deepseek-v4-pro@high`) or `minimax-preview` (`minimax:MiniMax-M3.1-Flash-Preview@high`) to named roles. Existing `deepseek` and `minimax` choices remain available. Each model has its own effort selection and live probe. MiniMax preview requires an eligible Token Plan key in `MINIMAX_API_KEY`; see [model choices and thinking controls](LANES.md#multiple-models-per-provider). No existing assignment changes until setup succeeds and you confirm the rendered sheet. diff --git a/docs/gateway-model-probes.md b/docs/gateway-model-probes.md new file mode 100644 index 00000000..ab03514a --- /dev/null +++ b/docs/gateway-model-probes.md @@ -0,0 +1,30 @@ +# Gateway model probe evidence + +Tracking: [issue #5](https://github.com/thisguymartin/pstack-flex/issues/5). + +## Candidate and scope + +- Source candidate: branch `flex/multiple-gateway-models`, based on `92dc0bc`, with uncommitted implementation changes. +- Implementation diff SHA-256 before this evidence file: `84b4c3ee7fc66902532e1d457048a487e6a63629b31d2d214e52585f72863c75`. +- Packaged version: 1.4.1. This candidate has not been installed as a plugin. +- Actual parent: Codex session, invoking the candidate's external runner with `--parent codex`. +- CLI: Claude Code 2.1.283. +- Each probe used `--effort high`, read-only mode, a separate empty synthetic workspace and isolated Claude configuration, and a synthetic text file. No repository or customer data was used in the prompt. +- Keys were supplied through hidden terminal input, injected into child environments, and were not included in commands, this repository, or evidence below. + +## Observed results + +Each probe exited 0, returned the exact requested marker, recorded `modelVerified: true` with `modelEvidence: provider-report`, and retained `costUsd: null`. + +| Requested model | Reported model | Elapsed milliseconds | Receipt status | +| --- | --- | --- | --- | +| `deepseek-flash` | `deepseek-flash` | 2684 | `complete` | +| `deepseek-v4-pro` | `deepseek-v4-pro` | 7737 | `complete` | +| `MiniMax-M3` | `MiniMax-M3` | 11416 | `complete` | +| `MiniMax-M3.1-Flash-Preview` | `MiniMax-M3.1-Flash-Preview` | 5090 | `complete` | + +These single short probes establish authentication, model selection, and successful completion through the runner. They do not rank coding quality or speed, prove hidden reasoning depth, or verify CLI request-body effort forwarding. The prompt included the expected marker, so completion does not independently prove a file tool was used. + +## Remaining release gate + +Install the exact candidate and run setup from both real Claude Code and Codex user surfaces. Verify independent model effort choices, per-model probes, mixed-provider panels, saved-sheet readback, and unchanged configuration on failed access. Record installed version, surface, action, and observed result before merge or rollout. Changing only the runner's `--parent` flag would not satisfy this gate. diff --git a/docs/reference.md b/docs/reference.md index 23015146..f070eb3e 100644 --- a/docs/reference.md +++ b/docs/reference.md @@ -86,7 +86,7 @@ The Codex build shares one `skills/` tree with the Claude Code build. Nothing is - **Tool and built-in mapping.** Claude tool names and built-in skills resolve through [`codex-tools.md`](../plugins/pstack/skills/poteto-mode/references/codex-tools.md). Model execution resolves separately through [`provider-dispatch.md`](../plugins/pstack/skills/poteto-mode/references/provider-dispatch.md), so Codex can keep Sol native while invoking Claude and Grok externally. - **Subagents.** The `Agent` tool maps to Codex `spawn_agent` / `wait_agent`, enabled by `multi_agent = true`. Parallel fan-out is multiple `spawn_agent` calls in one turn. If the native Codex lane is unavailable, record that lane as a dropout; external Claude and Grok lanes still run, and no provider is silently substituted. There is no `poteto-agent` subagent type on Codex; route ad-hoc subagents by dispatching a `spawn_agent` told to read `poteto-mode` first. - **Auto-fire.** The `hooks/` SessionStart injection is Claude Code-only; Codex has no plugin hook runtime. Enter `pstack:poteto-mode` by name, or add a standing instruction to `~/.codex/AGENTS.md` if you want the same always-on routing. -- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per frontier family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-5.6 Sol max, Grok 4.6 xhigh, and Opus xhigh. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. In Codex, Sol uses native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; Sol and Grok use the runner. Children never detect the parent or reroute themselves. The `bug-fix`, `perf-issue`, and `hillclimb` roles stay on GPT-5.6 Sol max instead of upstream's Fable default because Sol costs less for these frequent delegated code roles. +- **Models.** `/setup-pstack` writes provider-qualified descriptors and asks one requested effort per frontier family (`low`, `medium`, `high`, `xhigh`, `max`). The first-run panel is Fable max, GPT-5.6 Sol max, Grok 4.6 xhigh, and Opus xhigh. Fable and Opus use Claude's rolling aliases. Runtime dispatch normalizes older versioned descriptors in memory, so an installed sheet stops pinning immediately. A setup rerun persists that migration while keeping each role's family and effort. Optional GPT-6 Astra, Sol, and Luna Codex families can be assigned to any role and never change the first-run sheet. In Codex, Sol and the GPT-6 families use native `spawn_agent`; Claude and Grok use the deterministic external runner. In Claude Code, Fable and Opus use native agents; Sol, the GPT-6 families, and Grok use the runner. Children never detect the parent or reroute themselves. The `bug-fix`, `perf-issue`, and `hillclimb` roles stay on GPT-5.6 Sol max instead of upstream's Fable default because Sol costs less for these frequent delegated code roles. Verified in fresh installed Claude Code and Codex sessions: the user-facing skills are discovered and namespaced under `pstack`; both parents fan out the frontier quad through the documented native/external route table, retain long-running handles without a default timeout, and cross-judge only after every candidate is terminal. The `principle-*` leaves remain available for `poteto-mode` to read by path. Claude honors their `user-invocable: false` metadata; Codex 0.149.0 does not ([#8](https://github.com/ericlitman/open-pstack/issues/8)). diff --git a/plugins/pstack/skills/poteto-mode/references/codex-tools.md b/plugins/pstack/skills/poteto-mode/references/codex-tools.md index b967458d..1d8621c3 100644 --- a/plugins/pstack/skills/poteto-mode/references/codex-tools.md +++ b/plugins/pstack/skills/poteto-mode/references/codex-tools.md @@ -42,7 +42,7 @@ poteto-mode's Subagents section sets Claude-specific defaults (`subagent_type: " ## Models and providers -Do not replace every configured entry with a Codex model. `/setup-pstack` writes portable descriptors such as `claude:fable@max`, `codex:gpt-5.6-sol@max`, and `grok:grok-4.6@xhigh`. In a Codex parent, only `codex:*` is native. Route Claude and Grok descriptors through the external launcher exactly as `provider-dispatch.md` specifies. The current default panel intentionally keeps four-provider frontier diversity and contains no older GPT or Claude substitute. +Do not replace every configured entry with a Codex model. `/setup-pstack` writes portable descriptors such as `claude:fable@max`, `codex:gpt-5.6-sol@max`, and `grok:grok-4.6@xhigh`. In a Codex parent, only `codex:*` is native. Route Claude and Grok descriptors through the external launcher exactly as `provider-dispatch.md` specifies. The current default panel intentionally keeps four-provider frontier diversity and contains no older GPT or Claude substitute. pstack-flex gateway descriptors (`deepseek:*`, `minimax:*`) also always route through the external launcher in a Codex parent; they are never `spawn_agent` lanes. ## Claude built-in skills pstack references diff --git a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md index 74ab90da..729c096c 100644 --- a/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md +++ b/plugins/pstack/skills/poteto-mode/references/provider-dispatch.md @@ -19,6 +19,39 @@ The allowed effort universe is exactly `low`, `medium`, `high`, `xhigh`, `max`. `fable` and `opus` are Claude Code's rolling aliases. Claude resolves each alias to the latest available family revision. A runner receipt keeps the requested alias in `model` and the concrete provider-reported revision in `reportedModel`; verification accepts only a numeric `claude-fable-*` or `claude-opus-*` revision from the matching family. +## Additional model matrix + +pstack-flex additions using the existing provider routes. These optional families do not change the stock matrix or first-run role assignments. Default effort is pstack's proposed requested effort when assigning a family, not the provider's default. `-` in Upstream pstack choice means there is no upstream default to replace. + +| Family | Upstream pstack choice | Provider | Model | Default effort | Selectable efforts | Claude-native agent stem | +|---|---|---|---|---|---|---| +| astra | - | codex | gpt-6-astra | high | low medium high xhigh max | - | +| sol-6 | - | codex | gpt-6-sol | high | low medium high xhigh max | - | +| luna | - | codex | gpt-6-luna | high | low medium high xhigh max | - | + +These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent. Setup may assign them to any configurable role, including `architect runners`, after each requested model and effort passes the parent-specific probe. The `sol-6` family is independent of the stock `sol` family; existing GPT-5.6 Sol assignments stay unchanged. + +## Flex model matrix + +pstack-flex addition. The stock matrix above is upstream-owned and unchanged; these lanes are additive. A flex lane runs the stock `claude` binary env-pointed at the provider's Anthropic-compatible endpoint, with the provider's own API key and an isolated `CLAUDE_CONFIG_DIR`, so it uses no Anthropic account, no claude.ai login, and no subscription. + +| Family | Provider | Model | Default effort | Selectable efforts | API key variable | Base URL default | +|---|---|---|---|---|---|---| +| deepseek | deepseek | deepseek-flash | high | low medium high xhigh max | DEEPSEEK_API_KEY | https://api.deepseek.com/anthropic | +| deepseek-pro | deepseek | deepseek-v4-pro | high | low medium high xhigh max | DEEPSEEK_API_KEY | https://api.deepseek.com/anthropic | +| minimax | minimax | MiniMax-M3 | high | low medium high xhigh max | MINIMAX_API_KEY | https://api.minimax.io/anthropic | +| minimax-preview | minimax | MiniMax-M3.1-Flash-Preview | high | low medium high xhigh max | MINIMAX_API_KEY | https://api.minimax.io/anthropic | + +A family identifies one `(provider, model)` pair, not an entire provider. The existing `deepseek` and `minimax` family names and descriptors remain valid. `deepseek-pro` and `minimax-preview` are additional choices with independent requested efforts. Multiple models from one provider still count as one provider for panel diversity. + +MiniMax preview requires Token Plan access; set `MINIMAX_API_KEY` to the eligible subscription key. A pay-as-you-go key is not proof of preview access. The preview always thinks and supports `low` through `max`; do not disable thinking. M3 thinking is off by default at the API and requires adaptive thinking to enable it; its effort flag does not imply preview-style depth control. Selectable efforts are runner requests, not a claim that every provider applies five distinct reasoning levels. Verify CLI forwarding and model access with live probes. Sources: [MiniMax models](https://platform.minimax.io/docs/guides/models-intro), [MiniMax thinking controls](https://platform.minimax.io/docs/api-reference/text-anthropic-api), [DeepSeek Anthropic compatibility](https://api-docs.deepseek.com/guides/anthropic_api) (checked 2026-09-27). + +Flex lanes have no Claude-native agent stem and always take the external runner in both parents. The base URL is a documented default; override it with `DEEPSEEK_BASE_URL` or `MINIMAX_BASE_URL`, and confirm it against the provider's current Claude Code guide during setup's live probe. The config dir defaults to `~/.pstack-flex/` (override: `PSTACK_FLEX__CONFIG_DIR`). Secrets stay in the environment: nothing in the sheet, the receipts, or this repository carries a key. + +Gateway receipt semantics differ from stock claude lanes in two documented ways. `costUsd` is always `null`: the claude CLI prices `total_cost_usd` at Anthropic rates, which would be fiction for third-party traffic; real prices live in [LANES.md](../../../../../docs/LANES.md), and token usage in the receipt stays accurate. Model verification accepts a case-insensitive matching provider report. A mismatched report fails the lane. When the endpoint reports no model, the receipt uses `modelEvidence: "pinned-argv"` and `modelVerified: false`. + +Panel diversity rule (pstack-flex): `arena runners` and `interrogate reviewers` must span at least two distinct providers. DeepSeek plus MiniMax satisfies it. A single-provider panel is written only after the operator explicitly confirms the reduced diversity during setup, and the setup report records that confirmation. The adversarial signal comes from model diversity, so treat the override as an exception, not a configuration style. + ## Read-time normalization Normalize configured descriptors before matching them to the matrix or choosing a route. If a provider-qualified Claude model starts with `claude-fable-` or `claude-opus-` and its remaining revision contains only digits and hyphens, replace that model component in memory with `fable` or `opus`. Preserve provider, effort, role, and lane order. Use only the normalized descriptor for native dispatch or runner argv. Never pass the versioned predecessor to Claude. @@ -31,10 +64,12 @@ This read-time rule makes an older installed sheet use the latest family revisio The top-level harness resolves the route once. A child receives an assigned provider, model, effort, access mode, prompt, working directory, and output path. A child never detects the harness, chooses a provider, or launches another model. Environment markers may corroborate the top-level harness before fan-out, but nested processes inherit parent markers and must not use them for routing. -| Parent | `claude:*` | `codex:*` | `grok:*` | -|---|---|---|---| -| Claude Code | native `Agent` | external runner | external runner | -| Codex | external runner | native `spawn_agent` | external runner | +| Parent | `claude:*` | `codex:*` | `grok:*` | `deepseek:*` | `minimax:*` | +|---|---|---|---|---|---| +| Claude Code | native `Agent` | external runner | external runner | external runner | external runner | +| Codex | external runner | native `spawn_agent` | external runner | external runner | external runner | + +Flex gateway descriptors are never native, even under a Claude Code parent: the gateway lane must run in its own process with injected endpoint, token, and isolated config dir, which the parent's native `Agent` primitive cannot provide. `inherit-parent` and `auto` remain aliases. They use the parent's current model and effort through its native subagent primitive. In a panel they still consume one lane, but they reduce provider diversity; say so in the synthesis record. @@ -54,7 +89,7 @@ The launcher lives at `skills/poteto-mode/scripts/runner/pstack-runner` under th ```text pstack-runner \ --parent \ - --provider \ + --provider \ --model \ --effort \ --mode \ @@ -67,6 +102,8 @@ pstack-runner \ Pass arguments as an argv array or quote every path. Never interpolate prompt text into a shell command. The launcher preflights the assigned CLI and authentication, invokes the model exactly once, disables recursive agents and ambient skill dispatch where the CLI supports it, restricts the built-in tool surface, and records the exact provider/model/effort flags. External lanes do not receive the parent's MCP surface. Keep MCP-dependent Why and Reflect roles on `inherit-parent` or `auto`. The launcher never falls back. +Gateway lanes (`deepseek`, `minimax`) run three checks before the model executes, all fail-closed. First, in-process: the lane's API key variable must be set, and the lane's isolated `CLAUDE_CONFIG_DIR` must be free of OAuth credentials — a `.credentials.json` carrying a claude.ai login, or one that cannot be parsed, refuses the lane with an `unauthenticated` receipt before any subprocess runs, so a claude.ai credential can never be sent to a third-party endpoint. Second, the spawned preflight is `claude --version`, which proves the binary executes; `claude auth status` is deliberately not used because its behavior under token auth is undocumented. Third, the one-shot invocation is the real authentication and model test; an endpoint authentication error classifies as `unauthenticated` like any other lane. + Grok authentication preflight has one bounded retry. If the first `grok models` result would be classified as unauthenticated, the runner waits five seconds and tries the same preflight once more. A second failure is terminal. The delay and second attempt share the runner's absolute deadline and cancellation latch, and the receipt keeps evidence from both attempts. Model execution is never retried. The parent tool sandbox still governs whether a subscribed child CLI can reach its credentials and network. Run setup's live probe from the actual parent profile. A blocked external CLI is a loud dropout, not a reason to elevate permissions or substitute a model silently. @@ -90,7 +127,7 @@ Success requires all of these: 1. Exit status `0`. 2. Receipt status `complete`. -3. Either `modelVerified: true` with `modelEvidence: "provider-report"`, or a Codex receipt with `reportedModel: null`, `modelVerified: false`, and `modelEvidence: "pinned-argv"`. For Claude's `fable` and `opus` aliases, the concrete provider report must belong to the requested family. Codex 0.149.0 accepts the exact `--model` argument but does not report the served model in its JSONL stream. +3. Either `modelVerified: true` with `modelEvidence: "provider-report"`, or a Codex receipt with `reportedModel: null`, `modelVerified: false`, and `modelEvidence: "pinned-argv"`, or a gateway (`deepseek`/`minimax`) receipt with `modelVerified: false` and `modelEvidence: "pinned-argv"` when the endpoint does not echo the requested slug. For Claude's `fable` and `opus` aliases, the concrete provider report must belong to the requested family. Codex 0.149.0 accepts the exact `--model` argument but does not report the served model in its JSONL stream. Gateway reports match case-insensitively because third-party endpoints are inconsistent about slug casing. 4. A non-empty output file. The receipt also carries elapsed time, token usage when the CLI exposes it, and cost when available. Keep it with the arena or review artifacts so parent-harness comparisons are evidence-based. diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/cli.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/cli.test.ts index 05f79b5a..e76b52dd 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/cli.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/cli.test.ts @@ -40,4 +40,28 @@ describe("runner CLI parsing", () => { "greater than zero" ); }); + + it("accepts gateway providers", () => { + const parsed = parseArgs([ + ...argv().map((value, index, all) => + all[index - 1] === "--provider" + ? "minimax" + : all[index - 1] === "--model" + ? "MiniMax-M3" + : value + ), + ]); + expect(parsed?.provider).toBe("minimax"); + expect(parsed?.model).toBe("MiniMax-M3"); + }); + + it("names the gateway providers in the provider rejection", () => { + expect(() => + parseArgs( + argv().map((value, index, all) => + all[index - 1] === "--provider" ? "gemini" : value + ) + ) + ).toThrow("deepseek, minimax"); + }); }); diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/cli.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/cli.ts index 4fcce242..eb326f85 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/cli.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/cli.ts @@ -13,7 +13,7 @@ import { UsageError, } from "./types.ts"; -const HELP = `Usage: pstack-runner --parent --provider \\ +const HELP = `Usage: pstack-runner --parent --provider <${PROVIDERS.join("|")}> \\ --model --effort --mode \\ --prompt --cwd --output --receipt [--timeout ] diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/commands.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/commands.test.ts index ea697ff2..827a1b66 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/commands.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/commands.test.ts @@ -1,5 +1,5 @@ import { describe, expect, it } from "bun:test"; -import { invocationCommand } from "./commands.ts"; +import { invocationCommand, preflightCommand } from "./commands.ts"; import type { RunnerOptions } from "./types.ts"; function options(overrides: Partial = {}): RunnerOptions { @@ -146,6 +146,30 @@ describe("invocationCommand", () => { ); }); + it("runs gateway lanes with the exact claude argv for the lane's model", () => { + for (const [provider, model] of [ + ["deepseek", "deepseek-flash"], + ["minimax", "MiniMax-M3"], + ] as const) { + const gateway = invocationCommand(options({ provider, model })); + const claude = invocationCommand( + options({ provider: "claude", model }) + ); + expect(gateway.command).toBe("claude"); + expect(gateway.stdin).toBe("prompt"); + expect(gateway.args).toEqual(claude.args); + } + }); + + it("preflights gateway lanes with a version probe, not an auth check", () => { + for (const provider of ["deepseek", "minimax"] as const) { + const spec = preflightCommand(provider); + expect(spec.command).toBe("claude"); + expect(spec.args).toEqual(["--version"]); + expect(spec.stdin).toBe("none"); + } + }); + it("covers low, medium, and high for every external provider", () => { const cases = [ { diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/commands.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/commands.ts index 5f2b10c6..e61951d6 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/commands.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/commands.ts @@ -4,6 +4,7 @@ import type { Provider, RunnerOptions, } from "./types.ts"; +import { isGatewayProvider } from "./types.ts"; export interface CommandSpec { readonly command: string; @@ -12,6 +13,14 @@ export interface CommandSpec { } export function preflightCommand(provider: Provider): CommandSpec { + if (isGatewayProvider(provider)) { + // Gateway lanes run the claude binary with token auth against a + // third-party endpoint. `claude auth status` semantics under token + // auth are undocumented, so the preflight only proves the binary + // executes; credentials are checked in-process by the gateway guard + // and the one-shot invocation is the real auth test. + return { command: "claude", args: ["--version"], stdin: "none" }; + } switch (provider) { case "claude": return { @@ -63,33 +72,40 @@ function effortOverride(effort: Effort): string { return `model_reasoning_effort=${JSON.stringify(effort)}`; } +function claudeInvocation(options: RunnerOptions): CommandSpec { + return { + command: "claude", + args: [ + "-p", + "--model", + options.model, + "--effort", + options.effort, + "--permission-mode", + permissionMode(options.mode), + "--setting-sources", + "project", + "--strict-mcp-config", + "--tools", + claudeTools(options.mode), + "--no-session-persistence", + "--disable-slash-commands", + "--disallowed-tools", + claudeDeniedTools(options.mode), + "--output-format", + "json", + ], + stdin: "prompt", + }; +} + export function invocationCommand(options: RunnerOptions): CommandSpec { + // Gateway lanes use the same binary and argv as claude; the difference is + // injected environment (endpoint, token, isolated CLAUDE_CONFIG_DIR). + if (isGatewayProvider(options.provider)) return claudeInvocation(options); switch (options.provider) { case "claude": - return { - command: "claude", - args: [ - "-p", - "--model", - options.model, - "--effort", - options.effort, - "--permission-mode", - permissionMode(options.mode), - "--setting-sources", - "project", - "--strict-mcp-config", - "--tools", - claudeTools(options.mode), - "--no-session-persistence", - "--disable-slash-commands", - "--disallowed-tools", - claudeDeniedTools(options.mode), - "--output-format", - "json", - ], - stdin: "prompt", - }; + return claudeInvocation(options); case "codex": return { command: "codex", diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.test.ts new file mode 100644 index 00000000..755b70d3 --- /dev/null +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.test.ts @@ -0,0 +1,183 @@ +import { afterEach, beforeEach, describe, expect, it } from "bun:test"; +import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { homedir, tmpdir } from "node:os"; +import { join } from "node:path"; +import { + GATEWAY_INHERITED_CONFLICTS, + GATEWAY_SPECS, + gatewayConfigDir, + gatewayEnvironment, + gatewayGuard, +} from "./flex-providers.ts"; +import { GATEWAY_PROVIDERS } from "./types.ts"; + +let scratch = ""; + +beforeEach(() => { + scratch = mkdtempSync(join(tmpdir(), "flex-providers-")); +}); + +afterEach(() => { + rmSync(scratch, { recursive: true, force: true }); +}); + +describe("GATEWAY_SPECS", () => { + it("covers every gateway provider with an https default endpoint", () => { + for (const provider of GATEWAY_PROVIDERS) { + const spec = GATEWAY_SPECS[provider]; + expect(spec.apiKeyVar.length).toBeGreaterThan(0); + expect(spec.baseUrlDefault.startsWith("https://")).toBe(true); + } + }); +}); + +describe("gatewayConfigDir", () => { + it("defaults under the home directory per provider", () => { + expect(gatewayConfigDir("deepseek", {})).toBe( + join(homedir(), ".pstack-flex", "deepseek") + ); + expect(gatewayConfigDir("minimax", {})).toBe( + join(homedir(), ".pstack-flex", "minimax") + ); + }); + + it("honors the override variable and ignores blank overrides", () => { + expect( + gatewayConfigDir("deepseek", { PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: "/opt/lane" }) + ).toBe("/opt/lane"); + expect( + gatewayConfigDir("deepseek", { PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: " " }) + ).toBe(join(homedir(), ".pstack-flex", "deepseek")); + }); +}); + +describe("gatewayEnvironment", () => { + it("injects the full endpoint, token, model, and isolation map", () => { + const source = { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: scratch, + }; + expect(gatewayEnvironment("deepseek", "deepseek-flash", source)).toEqual({ + ANTHROPIC_BASE_URL: "https://api.deepseek.com/anthropic", + ANTHROPIC_AUTH_TOKEN: "sk-test", + ANTHROPIC_MODEL: "deepseek-flash", + ANTHROPIC_DEFAULT_OPUS_MODEL: "deepseek-flash", + ANTHROPIC_DEFAULT_SONNET_MODEL: "deepseek-flash", + ANTHROPIC_DEFAULT_HAIKU_MODEL: "deepseek-flash", + CLAUDE_CODE_SUBAGENT_MODEL: "deepseek-flash", + CLAUDE_CODE_ATTRIBUTION_HEADER: "0", + CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC: "1", + CLAUDE_CODE_MAX_CONTEXT_TOKENS: "128000", + CLAUDE_CONFIG_DIR: scratch, + }); + }); + + it("omits the context cap when the provider has no default", () => { + const env = gatewayEnvironment("minimax", "MiniMax-M3", { + MINIMAX_API_KEY: "mm-test", + PSTACK_FLEX_MINIMAX_CONFIG_DIR: scratch, + }); + expect(env.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBeUndefined(); + expect(env.ANTHROPIC_BASE_URL).toBe("https://api.minimax.io/anthropic"); + expect(env.ANTHROPIC_MODEL).toBe("MiniMax-M3"); + }); + + it("honors base URL and context overrides", () => { + const env = gatewayEnvironment("deepseek", "deepseek-flash", { + DEEPSEEK_API_KEY: "sk-test", + DEEPSEEK_BASE_URL: "https://proxy.internal/anthropic", + DEEPSEEK_MAX_CONTEXT_TOKENS: "64000", + }); + expect(env.ANTHROPIC_BASE_URL).toBe("https://proxy.internal/anthropic"); + expect(env.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe("64000"); + }); + + it("never leaks a value from a non-token source variable", () => { + const env = gatewayEnvironment("deepseek", "deepseek-flash", { + DEEPSEEK_API_KEY: "sk-secret", + UNRELATED_SECRET: "do-not-copy", + }); + const values = Object.entries(env) + .filter(([key]) => key !== "ANTHROPIC_AUTH_TOKEN") + .map(([, value]) => value); + expect(values).not.toContain("sk-secret"); + expect(values).not.toContain("do-not-copy"); + }); + + it("lists every alternative Claude provider selector as an inherited conflict", () => { + const conflicts = new Set(GATEWAY_INHERITED_CONFLICTS); + for (const key of [ + "CLAUDE_CODE_USE_ANTHROPIC_AWS", + "CLAUDE_CODE_USE_BEDROCK", + "CLAUDE_CODE_USE_FOUNDRY", + "CLAUDE_CODE_USE_MANTLE", + "CLAUDE_CODE_USE_VERTEX", + "CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST", + "CLAUDE_CONFIG_DIR", + ]) { + expect(conflicts.has(key)).toBe(true); + } + }); +}); + +describe("gatewayGuard", () => { + it("refuses when the API key variable is missing or blank", () => { + expect(gatewayGuard("deepseek", {})?.message).toBe("DEEPSEEK_API_KEY is not set"); + expect(gatewayGuard("minimax", { MINIMAX_API_KEY: " " })?.message).toBe( + "MINIMAX_API_KEY is not set" + ); + }); + + it("passes when the config dir does not exist yet", () => { + expect( + gatewayGuard("deepseek", { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: join(scratch, "never-created"), + }) + ).toBeNull(); + }); + + it("refuses an OAuth credentials file and cites the path, not the contents", () => { + const dir = join(scratch, "oauth"); + mkdirSync(dir); + const credentials = join(dir, ".credentials.json"); + writeFileSync( + credentials, + JSON.stringify({ claudeAiOauth: { accessToken: "oauth-secret" } }), + { mode: 0o600 } + ); + const refusal = gatewayGuard("deepseek", { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: dir, + }); + expect(refusal?.message).toContain("OAuth credentials found"); + expect(refusal?.evidence).toBe(credentials); + expect(refusal?.evidence).not.toContain("oauth-secret"); + expect(refusal?.message).not.toContain("oauth-secret"); + }); + + it("refuses an unparseable credentials file", () => { + const dir = join(scratch, "garbage"); + mkdirSync(dir); + writeFileSync(join(dir, ".credentials.json"), "not json", { mode: 0o600 }); + const refusal = gatewayGuard("deepseek", { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: dir, + }); + expect(refusal?.message).toContain("unknown credential state"); + }); + + it("passes a credentials file that carries no OAuth markers", () => { + const dir = join(scratch, "clean"); + mkdirSync(dir); + writeFileSync(join(dir, ".credentials.json"), JSON.stringify({ note: "empty" }), { + mode: 0o600, + }); + expect( + gatewayGuard("deepseek", { + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: dir, + }) + ).toBeNull(); + }); +}); diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts new file mode 100644 index 00000000..1a247860 --- /dev/null +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/flex-providers.ts @@ -0,0 +1,135 @@ +import { existsSync, readFileSync } from "node:fs"; +import { homedir } from "node:os"; +import { join } from "node:path"; +import type { GatewayProvider } from "./types.ts"; + +// pstack-flex addition. Gateway providers run the stock `claude` binary +// against a third-party Anthropic-compatible endpoint. Everything a lane +// needs is injected as environment at spawn time; secrets come from the +// operator's environment and are never written to disk or receipts. + +export interface GatewaySpec { + readonly apiKeyVar: string; + readonly baseUrlDefault: string; + readonly baseUrlOverrideVar: string; + readonly configDirOverrideVar: string; + readonly maxContextTokensDefault: string | null; + readonly maxContextTokensOverrideVar: string; +} + +export const GATEWAY_SPECS: Record = { + deepseek: { + apiKeyVar: "DEEPSEEK_API_KEY", + baseUrlDefault: "https://api.deepseek.com/anthropic", + baseUrlOverrideVar: "DEEPSEEK_BASE_URL", + configDirOverrideVar: "PSTACK_FLEX_DEEPSEEK_CONFIG_DIR", + maxContextTokensDefault: "128000", + maxContextTokensOverrideVar: "DEEPSEEK_MAX_CONTEXT_TOKENS", + }, + minimax: { + apiKeyVar: "MINIMAX_API_KEY", + baseUrlDefault: "https://api.minimax.io/anthropic", + baseUrlOverrideVar: "MINIMAX_BASE_URL", + configDirOverrideVar: "PSTACK_FLEX_MINIMAX_CONFIG_DIR", + maxContextTokensDefault: null, + maxContextTokensOverrideVar: "MINIMAX_MAX_CONTEXT_TOKENS", + }, +}; + +// Provider selection and Claude configuration from the parent must not +// override the gateway's endpoint, token, or isolated config directory. +export const GATEWAY_INHERITED_CONFLICTS = [ + "CLAUDE_CODE_USE_ANTHROPIC_AWS", + "CLAUDE_CODE_USE_BEDROCK", + "CLAUDE_CODE_USE_FOUNDRY", + "CLAUDE_CODE_USE_MANTLE", + "CLAUDE_CODE_USE_VERTEX", + "CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST", + "CLAUDE_CODE_SUBAGENT_MODEL", + "CLAUDE_CODE_MAX_CONTEXT_TOKENS", + "CLAUDE_CONFIG_DIR", +] as const; + +function overridden(source: NodeJS.ProcessEnv, name: string): string | null { + const value = source[name]; + return value !== undefined && value.trim().length > 0 ? value : null; +} + +export function gatewayConfigDir( + provider: GatewayProvider, + source: NodeJS.ProcessEnv = process.env +): string { + return ( + overridden(source, GATEWAY_SPECS[provider].configDirOverrideVar) ?? + join(homedir(), ".pstack-flex", provider) + ); +} + +export function gatewayEnvironment( + provider: GatewayProvider, + model: string, + source: NodeJS.ProcessEnv = process.env +): NodeJS.ProcessEnv { + const spec = GATEWAY_SPECS[provider]; + const injected: NodeJS.ProcessEnv = { + ANTHROPIC_BASE_URL: overridden(source, spec.baseUrlOverrideVar) ?? spec.baseUrlDefault, + ANTHROPIC_MODEL: model, + ANTHROPIC_DEFAULT_OPUS_MODEL: model, + ANTHROPIC_DEFAULT_SONNET_MODEL: model, + ANTHROPIC_DEFAULT_HAIKU_MODEL: model, + CLAUDE_CODE_SUBAGENT_MODEL: model, + CLAUDE_CODE_ATTRIBUTION_HEADER: "0", + CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC: "1", + CLAUDE_CONFIG_DIR: gatewayConfigDir(provider, source), + }; + const token = overridden(source, spec.apiKeyVar); + if (token !== null) injected.ANTHROPIC_AUTH_TOKEN = token; + const maxContext = + overridden(source, spec.maxContextTokensOverrideVar) ?? spec.maxContextTokensDefault; + if (maxContext !== null) injected.CLAUDE_CODE_MAX_CONTEXT_TOKENS = maxContext; + return injected; +} + +export interface GatewayRefusal { + readonly message: string; + readonly evidence: string; +} + +// Runs in-process before any subprocess is spawned, so no request can leave +// the machine first. Refusals surface as `unauthenticated` receipts. +export function gatewayGuard( + provider: GatewayProvider, + source: NodeJS.ProcessEnv = process.env +): GatewayRefusal | null { + const spec = GATEWAY_SPECS[provider]; + if (overridden(source, spec.apiKeyVar) === null) { + return { + message: `${spec.apiKeyVar} is not set`, + evidence: `gateway lane ${provider} requires ${spec.apiKeyVar} in the environment`, + }; + } + const credentialsPath = join(gatewayConfigDir(provider, source), ".credentials.json"); + if (!existsSync(credentialsPath)) return null; + let raw: unknown; + try { + raw = JSON.parse(readFileSync(credentialsPath, "utf8")); + } catch { + return { + message: + "unreadable credentials file in gateway config dir; refusing to run with unknown credential state", + evidence: credentialsPath, + }; + } + const record = + raw !== null && typeof raw === "object" && !Array.isArray(raw) + ? (raw as Record) + : null; + if (record === null || "claudeAiOauth" in record || "accessToken" in record) { + return { + message: + "OAuth credentials found in gateway config dir; refusing to point a claude.ai login at a third-party endpoint", + evidence: credentialsPath, + }; + } + return null; +} diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts index e0d6df1e..b82be28e 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/model-matrix.test.ts @@ -1,7 +1,16 @@ import { describe, expect, it } from "bun:test"; import { readdirSync, readFileSync } from "node:fs"; import { join } from "node:path"; -import { EFFORTS, type Effort } from "./types.ts"; +import { parseArgs } from "./cli.ts"; +import { invocationCommand } from "./commands.ts"; +import { GATEWAY_SPECS } from "./flex-providers.ts"; +import { validateOptions } from "./run.ts"; +import { + EFFORTS, + GATEWAY_PROVIDERS, + type Effort, + type GatewayProvider, +} from "./types.ts"; const PLUGIN_ROOT = join(import.meta.dir, "../../../.."); const DISPATCH_PATH = join( @@ -51,12 +60,22 @@ const SHEET_ROLES = [ const SETUP_SECTION_ORDER = [ "### 2. Load current state", "### 3. Parse per-family efforts", - "### 4. Collect one requested effort per family", - "### 5. Probe the four requested pairs", + "### 4. Choose role assignments, then collect efforts", + "### 5. Probe the assigned pairs", "### 6. Render, preserving role families", "### 7. Confirm and commit", ] as const; +const FLEX_MATRIX_HEADER = [ + "Family", + "Provider", + "Model", + "Default effort", + "Selectable efforts", + "API key variable", + "Base URL default", +] as const; + interface MatrixRow { family: string; upstreamChoice: string; @@ -89,11 +108,15 @@ function asEffort(value: string): Effort { throw new Error(`not an effort: ${value}`); } -function parseModelMatrix(markdown: string): MatrixRow[] { +function parseModelMatrix( + markdown: string, + heading = "## Model matrix", + rowCount: number = FAMILY_ORDER.length +): MatrixRow[] { const lines = markdown.split(/\r?\n/); - const start = lines.findIndex((line) => line.trim() === "## Model matrix"); + const start = lines.findIndex((line) => line.trim() === heading); if (start < 0) { - throw new Error("missing ## Model matrix"); + throw new Error(`missing ${heading}`); } let end = lines.length; for (let i = start + 1; i < lines.length; i++) { @@ -106,9 +129,9 @@ function parseModelMatrix(markdown: string): MatrixRow[] { .slice(start + 1, end) .map((line) => line.trim()) .filter((line) => line.startsWith("|")); - if (table.length !== 6) { + if (table.length !== rowCount + 2) { throw new Error( - `model matrix must be header, separator, and 4 data rows, got ${table.length}` + `${heading} must be header, separator, and ${rowCount} data rows, got ${table.length}` ); } const header = splitRow(table[0]); @@ -198,7 +221,9 @@ function firstRunSheet(setup: string): string { } describe("model matrix", () => { - const rows = parseModelMatrix(readFileSync(DISPATCH_PATH, "utf8")); + const dispatch = readFileSync(DISPATCH_PATH, "utf8"); + const rows = parseModelMatrix(dispatch); + const additionalRows = parseModelMatrix(dispatch, "## Additional model matrix", 3); const setup = readFileSync(SETUP_PATH, "utf8"); const quad = defaultDescriptors(rows); @@ -234,7 +259,7 @@ describe("model matrix", () => { it("ships exactly the declared Claude-native frontier agents", () => { const expected = new Set(); const familyBodies = new Map(); - for (const row of rows) { + for (const row of [...rows, ...additionalRows]) { const stem = row.claudeNativeAgentStem; if (stem === null) { continue; @@ -275,6 +300,76 @@ describe("model matrix", () => { expect(shipped).toEqual([...expected].sort()); }); + it("adds GPT-6 families without changing the stock matrix or first-run assignments", () => { + expect(additionalRows.map((row) => [row.family, row.model])).toEqual([ + ["astra", "gpt-6-astra"], + ["sol-6", "gpt-6-sol"], + ["luna", "gpt-6-luna"], + ]); + for (const row of additionalRows) { + expect(row.upstreamChoice).toBe("-"); + expect(row.provider).toBe("codex"); + expect(row.defaultEffort).toBe("high"); + expect(row.selectableEfforts).toEqual([...EFFORTS]); + expect(row.claudeNativeAgentStem).toBeNull(); + expect(firstRunSheet(setup)).not.toContain(row.model); + } + const allRows = [...rows, ...additionalRows]; + expect(new Set(allRows.map((row) => row.family)).size).toBe(allRows.length); + expect(new Set(allRows.map((row) => `${row.provider}:${row.model}`)).size) + .toBe(allRows.length); + expect(setup).toContain("Its model matrices (stock, additional, and flex)"); + expect(setup).toContain("Read the model matrices, stock, additional, and flex."); + expect(setup).toContain("any stock, additional, or flex matrix family"); + expect(setup).toContain("Offer Astra, GPT-6 Sol, and Luna from the additional matrix when changing `architect runners`"); + expect(setup).toContain("Read each model, proposed effort, and selectable efforts from its row."); + expect(setup).toContain("outside the stock, additional, and flex matrix families"); + expect(setup).toContain( + "| Astra | Astra additional row + selected effort | external runner | native `spawn_agent` |" + ); + expect(setup).toContain( + "| GPT-6 Sol | sol-6 additional row + selected effort | external runner | native `spawn_agent` |" + ); + expect(setup).toContain( + "| Luna | Luna additional row + selected effort | external runner | native `spawn_agent` |" + ); + expect(setup).toContain("each assigned Codex family gets a native `spawn_agent` probe"); + expect(dispatch).toContain( + "These Codex families use native `spawn_agent` under a Codex parent and the external Codex runner under a Claude Code parent." + ); + }); + + it("passes each additional family's selected model and effort to the existing runner", () => { + for (const row of additionalRows) { + for (const effort of row.selectableEfforts) { + const options = parseArgs([ + "--parent", "claude", + "--provider", row.provider, + "--model", row.model, + "--effort", effort, + "--mode", "read-only", + "--prompt", DISPATCH_PATH, + "--cwd", PLUGIN_ROOT, + "--output", join(PLUGIN_ROOT, `${row.family}-probe.md`), + "--receipt", join(PLUGIN_ROOT, `${row.family}-probe.json`), + ]); + if (options === null) { + throw new Error("model probe arguments must produce runner options"); + } + validateOptions(options); + expect(options.timeoutMs).toBeNull(); + const command = invocationCommand(options); + expect(command.command).toBe("codex"); + expect(command.args.slice(0, 5)).toEqual([ + "exec", "--model", row.model, + "--config", `model_reasoning_effort="${effort}"`, + ]); + expect(() => validateOptions({ ...options, parent: "codex" })) + .toThrow("provider codex is native to parent codex"); + } + } + }); + it("keeps setup's first-run default panel copy aligned with the matrix", () => { const sheet = firstRunSheet(setup); const roles = sheet @@ -317,7 +412,13 @@ describe("model matrix", () => { expect(setup).toContain("Do not invent a precedence rule."); expect(setup).toContain("Do not probe or write while any inconsistency is unresolved."); expect(setup).toContain("A failed probe writes nothing:"); - expect(setup).toContain("Run one probe per family"); + expect(setup).toContain("Run one probe per assigned family"); + expect(setup).toContain("There is no requirement to assign every matrix family."); + expect(setup).toContain("`architect runners` to keep at least two entries"); + expect(setup).toContain("span at least two distinct providers"); + expect(setup).toContain( + "A failed model demands explicit repair or role reassignment before saving." + ); expect(setup).toContain("normalized complete role map from step 2"); expect(setup).toContain("starts with `claude-fable-` or `claude-opus-`"); expect(setup).toContain("preserving the provider, effort, role, and lane order"); @@ -328,6 +429,64 @@ describe("model matrix", () => { expect(setup).toContain(""); }); + it("keeps the flex matrix additive, parseable, and aligned with the runner", () => { + const dispatch = readFileSync(DISPATCH_PATH, "utf8"); + const lines = dispatch.split(/\r?\n/); + const start = lines.findIndex((line) => line.trim() === "## Flex model matrix"); + expect(start).toBeGreaterThan(-1); + let end = lines.length; + for (let i = start + 1; i < lines.length; i++) { + if (lines[i].startsWith("## ")) { + end = i; + break; + } + } + const table = lines + .slice(start + 1, end) + .map((line) => line.trim()) + .filter((line) => line.startsWith("|")); + expect(table.length).toBeGreaterThan(2 + GATEWAY_PROVIDERS.length); + expect(splitRow(table[0]).join("|")).toBe(FLEX_MATRIX_HEADER.join("|")); + expect(isSeparator(splitRow(table[1]))).toBe(true); + const seen = new Set(); + const families = new Set(); + const pairs = new Set(); + for (const line of table.slice(2)) { + const cells = splitRow(line); + expect(cells.length).toBe(FLEX_MATRIX_HEADER.length); + const [family, provider, model, defaultEffortRaw, selectableRaw, keyVar, baseUrl] = + cells; + expect(GATEWAY_PROVIDERS as readonly string[]).toContain(provider); + const gateway = provider as GatewayProvider; + seen.add(gateway); + expect(/^[a-z0-9-]+$/.test(family)).toBe(true); + expect(families.has(family)).toBe(false); + families.add(family); + const pair = `${provider}:${model}`; + expect(pairs.has(pair)).toBe(false); + pairs.add(pair); + expect(/^[A-Za-z0-9.-]+$/.test(model)).toBe(true); + const selectable = selectableRaw.split(/\s+/).map(asEffort); + expect(selectable).toContain(asEffort(defaultEffortRaw)); + expect(keyVar).toBe(GATEWAY_SPECS[gateway].apiKeyVar); + expect(baseUrl).toBe(GATEWAY_SPECS[gateway].baseUrlDefault); + expect(baseUrl.startsWith("https://")).toBe(true); + } + expect([...seen]).toEqual([...GATEWAY_PROVIDERS]); + for (const pair of [ + "deepseek:deepseek-flash", + "deepseek:deepseek-v4-pro", + "minimax:MiniMax-M3", + "minimax:MiniMax-M3.1-Flash-Preview", + ]) expect(pairs.has(pair)).toBe(true); + expect(setup).toContain("Never group efforts or deduplicate probes by provider alone."); + expect(setup).toContain("Different models sharing a provider count as one provider"); + // The stock quad and first-run sheet must not carry flex descriptors: + // upstream's own checks parse descriptors with a lowercase-only, + // three-provider grammar and must never see a flex lane. + expect(firstRunSheet(setup)).not.toMatch(/deepseek:|minimax:/i); + }); + it("binds Claude-native dispatch to the matrix mapping", () => { const dispatch = readFileSync(DISPATCH_PATH, "utf8"); const nativeStart = dispatch.indexOf("## Native lanes"); diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.test.ts index b4ebc041..270a5863 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.test.ts @@ -110,6 +110,38 @@ describe("parseProviderOutput", () => { expect(parsed.reportedModel).toBe("claude-fable-9-9"); }); + it("parses gateway output as claude-shaped JSON with cost forced null", () => { + const parsed = parseProviderOutput( + "minimax", + JSON.stringify({ + result: "GATEWAY_OK", + session_id: "mm-session", + usage: { input_tokens: 12, output_tokens: 5 }, + total_cost_usd: 0.42, + modelUsage: { "minimax-m3": {} }, + }), + "", + "MiniMax-M3" + ); + expect(parsed).toMatchObject({ + text: "GATEWAY_OK", + reportedModel: "minimax-m3", + sessionId: "mm-session", + usage: { inputTokens: 12, outputTokens: 5 }, + costUsd: null, + }); + }); + + it("matches gateway model slugs case-insensitively", () => { + expect(reportedModelMatches("minimax", "MiniMax-M3", "minimax-m3")).toBe(true); + expect(reportedModelMatches("deepseek", "deepseek-flash", "DeepSeek-Flash")).toBe(true); + expect( + reportedModelMatches("deepseek", "deepseek-flash", "deepseek-flash-0731") + ).toBe(true); + expect(reportedModelMatches("minimax", "MiniMax-M3", "some-other-model")).toBe(false); + expect(reportedModelMatches("claude", "MiniMax-M3", "minimax-m3")).toBe(false); + }); + it("matches only concrete Claude revisions from the requested rolling family", () => { expect(reportedModelMatches("claude", "fable", "claude-fable-9-9")).toBe(true); expect(reportedModelMatches("claude", "opus", "claude-opus-9")).toBe(true); diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.ts index 81ed53d4..64d51b75 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/parse-output.ts @@ -3,6 +3,7 @@ import type { ParsedOutput, Provider, } from "./types.ts"; +import { isGatewayProvider } from "./types.ts"; import { concreteModelMatchesRollingAlias, isRollingClaudeAlias, @@ -61,7 +62,11 @@ function modelFromUsage( ?? null; } -function parseClaude(stdout: string, requestedModel: string): ParsedOutput { +function parseClaude( + stdout: string, + requestedModel: string, + provider: Provider = "claude" +): ParsedOutput { let raw: unknown; try { raw = JSON.parse(stdout); @@ -77,7 +82,7 @@ function parseClaude(stdout: string, requestedModel: string): ParsedOutput { return { text, - reportedModel: modelFromUsage(value.modelUsage, "claude", requestedModel), + reportedModel: modelFromUsage(value.modelUsage, provider, requestedModel), sessionId: nullableString(value.session_id ?? value.sessionId), usage: normalizedUsage(value.usage), costUsd: finiteNumber(value.total_cost_usd) ?? null, @@ -163,6 +168,13 @@ export function parseProviderOutput( stderr: string, requestedModel: string ): ParsedOutput { + if (isGatewayProvider(provider)) { + // Gateway lanes emit claude-shaped JSON, but the CLI's + // total_cost_usd is computed at Anthropic rates and would be + // fiction for third-party traffic. Token usage stays; cost is null. + const parsed = parseClaude(stdout, requestedModel, provider); + return { ...parsed, costUsd: null }; + } switch (provider) { case "claude": return parseClaude(stdout, requestedModel); @@ -182,6 +194,13 @@ export function reportedModelMatches( if (provider === "claude" && isRollingClaudeAlias(requested)) { return concreteModelMatchesRollingAlias(requested, reported); } + if (isGatewayProvider(provider)) { + // Third-party endpoints are inconsistent about slug casing + // (e.g. MiniMax-M3 vs minimax-m3); compare case-insensitively. + const wanted = requested.toLowerCase(); + const got = reported.toLowerCase(); + return got === wanted || got.startsWith(`${wanted}-`); + } if (reported === requested || reported.startsWith(`${requested}-`)) { return true; } diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts index 20743b52..8c2c0914 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/run.test.ts @@ -25,7 +25,7 @@ import { appendFileSync, existsSync, unlinkSync, writeFileSync } from "node:fs"; const args = process.argv.slice(2); const name = process.argv[1].split("/").at(-1); const isPreflight = - (name === "claude" && args[0] === "auth") || + (name === "claude" && (args[0] === "auth" || args[0] === "--version")) || (name === "codex" && args[0] === "login") || (name === "grok" && args[0] === "models"); const stage = isPreflight ? "preflight" : "model"; @@ -61,6 +61,10 @@ if (name === "claude" && args[0] === "auth") { console.log(JSON.stringify({loggedIn:true})); process.exit(0); } +if (name === "claude" && args[0] === "--version") { + console.log("9.9.9 (fake)"); + process.exit(0); +} if (name === "codex" && args[0] === "login") { console.log("Logged in using ChatGPT"); process.exit(0); @@ -89,11 +93,18 @@ if (name === "grok" && args[0] === "models") { } const modelIndex = args.findIndex((value) => value === "--model"); const model = modelIndex >= 0 ? args[modelIndex + 1] : "unknown"; -const reportedModel = model === "fable" +const reportedModel = process.env.FAKE_REPORT_MODEL ?? (model === "fable" ? "claude-fable-9-9" : model === "opus" ? "claude-opus-9" - : model; + : model); +if (stage === "model" && process.env.FAKE_DUMP_ENV_PATH) { + writeFileSync(process.env.FAKE_DUMP_ENV_PATH, JSON.stringify(process.env)); +} +if (stage === "model" && process.env.FAKE_AUTH_ERROR === "1") { + console.error("API error: authentication_error - invalid api key"); + process.exit(1); +} if (process.env.FAKE_INVALID_MODEL === "1") { console.error("The requested model is not supported with this account."); process.exit(1); @@ -115,7 +126,7 @@ if (stage === "model" && process.env.FAKE_SELF_SIGNAL) { await Bun.sleep(5_000); } if (name === "claude") { - console.log(JSON.stringify({result:"CLAUDE_OK",session_id:"c1",usage:{input_tokens:10,output_tokens:2},total_cost_usd:0.01,modelUsage:{[reportedModel]:{}}})); + console.log(JSON.stringify({result:"CLAUDE_OK",session_id:"c1",usage:{input_tokens:10,output_tokens:2},total_cost_usd:0.01,...(process.env.FAKE_OMIT_MODEL_USAGE === "1" ? {} : {modelUsage:{[reportedModel]:{}}})})); } else if (name === "codex") { console.log(JSON.stringify({type:"thread.started",thread_id:"o1"})); console.log(JSON.stringify({type:"item.completed",item:{type:"agent_message",text:"CODEX_OK"}})); @@ -906,7 +917,237 @@ describe("runLane", () => { }); }); +describe("gateway lanes", () => { + const GATEWAY_TEST_KEYS = [ + "DEEPSEEK_API_KEY", + "MINIMAX_API_KEY", + "PSTACK_FLEX_DEEPSEEK_CONFIG_DIR", + "PSTACK_FLEX_MINIMAX_CONFIG_DIR", + "ANTHROPIC_API_KEY", + "ANTHROPIC_CUSTOM_HEADERS", + "CLAUDE_CODE_USE_BEDROCK", + "FAKE_DUMP_ENV_PATH", + "FAKE_AUTH_ERROR", + "FAKE_REPORT_MODEL", + "FAKE_OMIT_MODEL_USAGE", + ] as const; + + function gatewayOptions( + provider: "deepseek" | "minimax", + suffix: string + ): RunnerOptions { + return { + ...options(provider === "deepseek" ? "claude" : "codex", suffix), + provider, + parent: "claude", + model: provider === "deepseek" ? "deepseek-flash" : "MiniMax-M3", + effort: "high", + }; + } + + beforeEach(() => { + for (const key of GATEWAY_TEST_KEYS) delete process.env[key]; + process.env.DEEPSEEK_API_KEY = "sk-deepseek-test"; + process.env.MINIMAX_API_KEY = "sk-minimax-test"; + process.env.PSTACK_FLEX_DEEPSEEK_CONFIG_DIR = join(scratch, "flex-deepseek"); + process.env.PSTACK_FLEX_MINIMAX_CONFIG_DIR = join(scratch, "flex-minimax"); + }); + + afterEach(() => { + for (const key of GATEWAY_TEST_KEYS) delete process.env[key]; + }); + + it("refuses without spawning anything when the API key is missing", async () => { + delete process.env.DEEPSEEK_API_KEY; + process.env.FAKE_PREFLIGHT_STARTED_PATH = join(scratch, "preflight-started"); + process.env.FAKE_MODEL_STARTED_PATH = join(scratch, "model-started"); + const input = gatewayOptions("deepseek", "missing-key"); + const result = await runLane(input); + expect(result.exitCode).toBe(77); + const written = receipt(input.receiptPath); + expect(written.status).toBe("unauthenticated"); + expect(written.preflight.status).toBe("not-run"); + expect(written.error?.message).toBe("DEEPSEEK_API_KEY is not set"); + expect(existsSync(join(scratch, "preflight-started"))).toBe(false); + expect(existsSync(join(scratch, "model-started"))).toBe(false); + expect(existsSync(input.outputPath)).toBe(false); + }); + + it("refuses to run over an OAuth login without leaking its contents", async () => { + const dir = join(scratch, "flex-deepseek"); + mkdirSync(dir, { recursive: true }); + writeFileSync( + join(dir, ".credentials.json"), + JSON.stringify({ claudeAiOauth: { accessToken: "oauth-secret" } }), + { mode: 0o600 } + ); + const input = gatewayOptions("deepseek", "oauth-refused"); + const result = await runLane(input); + expect(result.exitCode).toBe(77); + const written = receipt(input.receiptPath); + expect(written.status).toBe("unauthenticated"); + expect(written.error?.message).toContain("OAuth credentials found"); + expect(written.error?.evidence).toBe(join(dir, ".credentials.json")); + expect(JSON.stringify(written)).not.toContain("oauth-secret"); + }); + + it("injects the gateway environment and never the parent's Anthropic identity", async () => { + process.env.ANTHROPIC_API_KEY = "parent-anthropic-secret"; + process.env.ANTHROPIC_CUSTOM_HEADERS = "Authorization: Bearer parent-header-secret"; + process.env.CLAUDE_CODE_USE_BEDROCK = "1"; + const dumpPath = join(scratch, "env-dump.json"); + process.env.FAKE_DUMP_ENV_PATH = dumpPath; + const input = gatewayOptions("deepseek", "env-dump"); + const result = await runLane(input); + expect(result.exitCode).toBe(0); + const child = JSON.parse(readFileSync(dumpPath, "utf8")) as Record; + expect(child.ANTHROPIC_BASE_URL).toBe("https://api.deepseek.com/anthropic"); + expect(child.ANTHROPIC_AUTH_TOKEN).toBe("sk-deepseek-test"); + expect(child.ANTHROPIC_MODEL).toBe("deepseek-flash"); + expect(child.CLAUDE_CODE_SUBAGENT_MODEL).toBe("deepseek-flash"); + expect(child.CLAUDE_CODE_ATTRIBUTION_HEADER).toBe("0"); + expect(child.CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC).toBe("1"); + expect(child.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe("128000"); + expect(child.CLAUDE_CONFIG_DIR).toBe(join(scratch, "flex-deepseek")); + expect(child.ANTHROPIC_API_KEY).toBeUndefined(); + expect(child.ANTHROPIC_CUSTOM_HEADERS).toBeUndefined(); + expect(child.CLAUDE_CODE_USE_BEDROCK).toBeUndefined(); + expect(child.CLAUDECODE).toBeUndefined(); + }); + + it("completes with cost null and a verified provider report on the happy path", async () => { + const input = gatewayOptions("deepseek", "happy"); + const result = await runLane(input); + expect(result.exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written.status).toBe("complete"); + expect(written.costUsd).toBeNull(); + expect(written.usage).toMatchObject({ inputTokens: 10, outputTokens: 2 }); + expect(written.modelVerified).toBe(true); + expect(written.modelEvidence).toBe("provider-report"); + expect(written.preflight.status).toBe("passed"); + expect(written.preflight.evidence).toBe( + "claude binary responded; gateway credentials verified in-process" + ); + expect(readFileSync(input.outputPath, "utf8")).toBe("CLAUDE_OK"); + }); + + for (const parent of ["claude", "codex"] as const) { + for (const [provider, model] of [ + ["deepseek", "deepseek-v4-pro"], + ["minimax", "MiniMax-M3.1-Flash-Preview"], + ] as const) { + it(`pins ${model} and effort through the ${parent} parent route`, async () => { + const dumpPath = join(scratch, "new-model-env.json"); + process.env.FAKE_DUMP_ENV_PATH = dumpPath; + process.env.FAKE_REPORT_MODEL = model.toLowerCase(); + const input: RunnerOptions = { + ...gatewayOptions(provider, "new-model"), parent, model, effort: "max", + }; + expect((await runLane(input)).exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written).toMatchObject({ + status: "complete", parent, provider, model, effort: "max", + modelVerified: true, modelEvidence: "provider-report", costUsd: null, + }); + expect(written.argv[written.argv.indexOf("--model") + 1]).toBe(model); + expect(written.argv[written.argv.indexOf("--effort") + 1]).toBe("max"); + const child = JSON.parse(readFileSync(dumpPath, "utf8")); + for (const key of [ + "ANTHROPIC_MODEL", "ANTHROPIC_DEFAULT_OPUS_MODEL", + "ANTHROPIC_DEFAULT_SONNET_MODEL", "ANTHROPIC_DEFAULT_HAIKU_MODEL", + "CLAUDE_CODE_SUBAGENT_MODEL", + ]) expect(child[key]).toBe(model); + }); + + it(`rejects a substituted ${model} in the ${parent} parent route`, async () => { + const input: RunnerOptions = { + ...gatewayOptions(provider, "substituted-model"), parent, model, + }; + process.env.FAKE_REPORT_MODEL = provider === "deepseek" ? "deepseek-flash" : "MiniMax-M3"; + expect((await runLane(input)).exitCode).toBe(65); + expect(receipt(input.receiptPath).status).toBe("malformed-output"); + expect(existsSync(input.outputPath)).toBe(false); + }); + } + } + + it("verifies a case-shifted served model for MiniMax", async () => { + process.env.FAKE_REPORT_MODEL = "minimax-m3"; + const input = gatewayOptions("minimax", "case-shift"); + const result = await runLane(input); + expect(result.exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written.modelVerified).toBe(true); + expect(written.modelEvidence).toBe("provider-report"); + expect(written.reportedModel).toBe("minimax-m3"); + }); + + it("fails when the endpoint reports a different model", async () => { + process.env.FAKE_REPORT_MODEL = "unrelated-model"; + const input = gatewayOptions("minimax", "pinned"); + const result = await runLane(input); + expect(result.exitCode).toBe(65); + const written = receipt(input.receiptPath); + expect(written.status).toBe("malformed-output"); + expect(written.modelVerified).toBe(false); + expect(written.modelEvidence).toBeNull(); + expect(written.error?.message).toContain("requested model MiniMax-M3 was not reported"); + expect(existsSync(input.outputPath)).toBe(false); + }); + + it("uses the pinned argv only when the endpoint reports no model", async () => { + process.env.FAKE_OMIT_MODEL_USAGE = "1"; + const input = gatewayOptions("minimax", "unreported-model"); + const result = await runLane(input); + expect(result.exitCode).toBe(0); + const written = receipt(input.receiptPath); + expect(written.status).toBe("complete"); + expect(written.reportedModel).toBeNull(); + expect(written.modelVerified).toBe(false); + expect(written.modelEvidence).toBe("pinned-argv"); + }); + + it("classifies an endpoint authentication error as unauthenticated", async () => { + process.env.FAKE_AUTH_ERROR = "1"; + const input = gatewayOptions("deepseek", "endpoint-401"); + const result = await runLane(input); + expect(result.exitCode).toBe(77); + expect(receipt(input.receiptPath).status).toBe("unauthenticated"); + }); +}); + describe("childEnvironment", () => { + it("strips identity and Anthropic inheritance before gateway injection", () => { + const source = { + PATH: "/bin", + CLAUDECODE: "1", + CODEX_CI: "1", + ANTHROPIC_API_KEY: "parent-secret", + ANTHROPIC_BASE_URL: "https://api.anthropic.com", + ANTHROPIC_CUSTOM_HEADERS: "Authorization: Bearer parent-secret", + ANTHROPIC_BEDROCK_BASE_URL: "https://parent-bedrock.example", + CLAUDE_CODE_USE_BEDROCK: "1", + CLAUDE_CODE_USE_VERTEX: "1", + DEEPSEEK_API_KEY: "sk-test", + PSTACK_FLEX_DEEPSEEK_CONFIG_DIR: "/tmp/flex-deepseek", + KEEP_ME: "yes", + }; + const env = childEnvironment("deepseek", source, "deepseek-flash"); + expect(env.CLAUDECODE).toBeUndefined(); + expect(env.CODEX_CI).toBeUndefined(); + expect(env.ANTHROPIC_API_KEY).toBeUndefined(); + expect(env.ANTHROPIC_CUSTOM_HEADERS).toBeUndefined(); + expect(env.ANTHROPIC_BEDROCK_BASE_URL).toBeUndefined(); + expect(env.CLAUDE_CODE_USE_BEDROCK).toBeUndefined(); + expect(env.CLAUDE_CODE_USE_VERTEX).toBeUndefined(); + expect(env.ANTHROPIC_BASE_URL).toBe("https://api.deepseek.com/anthropic"); + expect(env.ANTHROPIC_AUTH_TOKEN).toBe("sk-test"); + expect(env.CLAUDE_CONFIG_DIR).toBe("/tmp/flex-deepseek"); + expect(env.KEEP_ME).toBe("yes"); + expect(env.PATH).toBe("/bin"); + }); + it("removes only inherited runtime identity needed to avoid nested detection", () => { const source = { PATH: "/bin", diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/run.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/run.ts index 054564a4..b3e2ef03 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/run.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/run.ts @@ -10,6 +10,11 @@ import { } from "node:fs"; import { dirname, resolve } from "node:path"; import { invocationCommand, preflightCommand, type CommandSpec } from "./commands.ts"; +import { + GATEWAY_INHERITED_CONFLICTS, + gatewayEnvironment, + gatewayGuard, +} from "./flex-providers.ts"; import { versionedClaudeAlias } from "./model-aliases.ts"; import { parseProviderOutput, reportedModelMatches } from "./parse-output.ts"; import type { @@ -18,7 +23,7 @@ import type { RunnerOptions, RunnerReceipt, } from "./types.ts"; -import { UsageError } from "./types.ts"; +import { isGatewayProvider, UsageError } from "./types.ts"; const ERROR_EVIDENCE_LIMIT = 4_000; const GROK_PREFLIGHT_RETRY_DELAY_MS = 5_000; @@ -131,7 +136,8 @@ const CLAUDE_IDENTITY = [ export function childEnvironment( provider: Provider, - source: NodeJS.ProcessEnv = process.env + source: NodeJS.ProcessEnv = process.env, + model: string = "" ): NodeJS.ProcessEnv { const result = { ...source }; const remove = provider === "claude" @@ -140,6 +146,13 @@ export function childEnvironment( ? CLAUDE_IDENTITY : [...CODEX_IDENTITY, ...CLAUDE_IDENTITY]; for (const key of remove) delete result[key]; + if (isGatewayProvider(provider)) { + for (const key of Object.keys(result)) { + if (key.startsWith("ANTHROPIC_")) delete result[key]; + } + for (const key of GATEWAY_INHERITED_CONFLICTS) delete result[key]; + Object.assign(result, gatewayEnvironment(provider, model, source)); + } return result; } @@ -352,6 +365,9 @@ async function waitForGrokPreflightRetry( function preflightPassed(provider: Provider, model: string, result: ProcessResult): boolean { if (result.exitCode !== 0 || result.timedOut) return false; + // `claude --version` succeeded; gateway credentials were already verified + // in-process by the gateway guard before any subprocess ran. + if (isGatewayProvider(provider)) return true; const combined = `${result.stdout}\n${result.stderr}`; switch (provider) { case "claude": { @@ -374,6 +390,9 @@ function preflightPassed(provider: Provider, model: string, result: ProcessResul } function successfulPreflightEvidence(provider: Provider, model: string): string { + if (isGatewayProvider(provider)) { + return "claude binary responded; gateway credentials verified in-process"; + } return provider === "grok" ? `authenticated; model ${model} available` : "authenticated"; @@ -396,6 +415,11 @@ function preflightFailureStatus( ): ReceiptStatus { const status = unavailableStatus(value); if (status !== "child-failed") return status; + if (isGatewayProvider(provider)) { + // The gateway preflight is a version probe, not an auth check; a + // failure here means the binary misbehaved, not that auth failed. + return "child-failed"; + } return provider === "grok" && !value.includes(model) ? "unavailable-model" : "unauthenticated"; @@ -457,6 +481,16 @@ function modelProof( modelEvidence: "pinned-argv", }; } + if (isGatewayProvider(provider) && reported === null) { + // Third-party Anthropic-compatible endpoints do not reliably echo the + // requested model slug. A reported mismatch is a failure, since some + // gateways silently substitute a default model for unknown slugs. + return { + reportedModel: null, + modelVerified: false, + modelEvidence: "pinned-argv", + }; + } return { reportedModel: reported, modelVerified: false, @@ -534,7 +568,7 @@ async function executeLane( ): Promise { const startedAt = new Date(started).toISOString(); const prompt = readFileSync(options.promptPath, "utf8"); - const env = childEnvironment(options.provider); + const env = childEnvironment(options.provider, process.env, options.model); const executable = Bun.which(invocation.command, { PATH: env.PATH, cwd: options.cwd, @@ -589,6 +623,37 @@ async function executeLane( return finishWithoutChild("timed-out", "before authentication preflight"); } + if (isGatewayProvider(options.provider)) { + const refusal = gatewayGuard(options.provider); + if (refusal !== null) { + const completed = Date.now(); + receipt = completeReceipt(options, { + status: "unauthenticated", + startedAt, + completedAt: new Date(completed).toISOString(), + elapsedMs: completed - started, + executable, + preflight: preflightState, + argv: [executable ?? invocation.command, ...invocation.args], + exitCode: null, + signal: null, + reportedModel: null, + modelVerified: false, + modelEvidence: null, + sessionId: null, + usage: null, + costUsd: null, + error: { + message: refusal.message, + evidence: refusal.evidence, + }, + }); + removeIfExists(options.outputPath); + writeReceipt(options.receiptPath, receipt); + return { exitCode: statusExitCode("unauthenticated"), receipt }; + } + } + if (executable === null) { const completed = Date.now(); receipt = completeReceipt(options, { diff --git a/plugins/pstack/skills/poteto-mode/scripts/runner/types.ts b/plugins/pstack/skills/poteto-mode/scripts/runner/types.ts index 11c6dfb9..418b2bb6 100644 --- a/plugins/pstack/skills/poteto-mode/scripts/runner/types.ts +++ b/plugins/pstack/skills/poteto-mode/scripts/runner/types.ts @@ -1,10 +1,19 @@ export const PARENTS = ["claude", "codex"] as const; -export const PROVIDERS = ["claude", "codex", "grok"] as const; +// pstack-flex: gateway providers run the stock `claude` binary against a +// third-party Anthropic-compatible endpoint with injected environment. Adding +// one here requires a matching row in flex-providers.ts GATEWAY_SPECS. +export const GATEWAY_PROVIDERS = ["deepseek", "minimax"] as const; +export const PROVIDERS = ["claude", "codex", "grok", ...GATEWAY_PROVIDERS] as const; export const EFFORTS = ["low", "medium", "high", "xhigh", "max"] as const; export const ACCESS_MODES = ["read-only", "isolated-write"] as const; export type Parent = (typeof PARENTS)[number]; export type Provider = (typeof PROVIDERS)[number]; +export type GatewayProvider = (typeof GATEWAY_PROVIDERS)[number]; + +export function isGatewayProvider(provider: Provider): provider is GatewayProvider { + return (GATEWAY_PROVIDERS as readonly string[]).includes(provider); +} export type Effort = (typeof EFFORTS)[number]; export type AccessMode = (typeof ACCESS_MODES)[number]; diff --git a/plugins/pstack/skills/setup-pstack/SKILL.md b/plugins/pstack/skills/setup-pstack/SKILL.md index 9a0e7441..737452c7 100644 --- a/plugins/pstack/skills/setup-pstack/SKILL.md +++ b/plugins/pstack/skills/setup-pstack/SKILL.md @@ -1,11 +1,11 @@ --- name: setup-pstack -description: Configure pstack's provider-qualified models, per-family requested effort, and parent-owned routes per role. Verifies native and external Claude, Codex, and Grok lanes before writing the override sheet. Use for /setup-pstack, "configure pstack models", or changing pstack's model choices. +description: Configure pstack's provider-qualified models, per-family requested effort, and parent-owned routes per role. Verifies native and external Claude, Codex, Grok, DeepSeek, and MiniMax lanes before writing the override sheet. Use for /setup-pstack, "configure pstack models", or changing pstack's model choices. --- # Setup pstack -Configure one portable model sheet for the current parent harness. Read [`provider-dispatch.md`](../poteto-mode/references/provider-dispatch.md) before probing or writing anything. Its model matrix, descriptor grammar, and route table are the contract. Choose one requested effort per matrix family. Do not add a second configuration file, a runtime resolver, or a weaker-model fallback. +Configure one portable model sheet for the current parent harness. Read [`provider-dispatch.md`](../poteto-mode/references/provider-dispatch.md) before probing or writing anything. Its model matrices (stock, additional, and flex), descriptor grammar, and route table are the contract. Role assignments are selected first; then choose one requested effort per assigned matrix family. Do not add a second configuration file, a runtime resolver, or a weaker-model fallback. Claude Code writes `~/.claude/pstack-models.md` and loads it from `~/.claude/CLAUDE.md` with: @@ -35,28 +35,43 @@ Treat the normalized values as current role-to-family assignments. Overlay those ### 3. Parse per-family efforts -Read the model matrix. Every non-alias value must match `:@`. Map it to exactly one matrix family by `(provider, model)`, require its effort to appear in that row's Selectable efforts cell, and collect the effort. `inherit-parent` and `auto` rows carry no family effort. +Read the model matrices, stock, additional, and flex. Every non-alias value must match `:@`. Map it to exactly one matrix family by `(provider, model)`, require its effort to appear in that row's Selectable efforts cell, and collect the effort. `inherit-parent` and `auto` rows carry no family effort. An unmatched provider/model, out-of-domain effort, duplicate role, or unknown role is inconsistent state. Stop, show the conflicting rows verbatim, and ask for an explicit matrix family or alias replacement. If one or more families have mixed efforts, show every conflicting family and role row, then ask for one normalized effort per family from its Selectable efforts cell. Do not invent a precedence rule. Do not probe or write while any inconsistency is unresolved. +A family is a single `(provider, model)` matrix row. DeepSeek Flash and Pro have independent efforts, as do MiniMax M3 and M3.1 Flash Preview. Never group efforts or deduplicate probes by provider alone. + One distinct effort per family is the current value. A family with no non-alias occurrence is unassigned; use its matrix Default effort as the proposed value and label it unassigned rather than calling it current. -### 4. Collect one requested effort per family +### 4. Choose role assignments, then collect efforts + +Role assignments come first. Show the current role-to-family map (loaded and normalized from step 2, or the first-run map from step 7) and ask whether to keep it or change named roles. Keeping it is the default. A changed role may use any stock, additional, or flex matrix family, `inherit-parent`, or `auto`. Apply only role changes the operator names; never offer a reset of a customized sheet to the first-run assignments. + +Offer Astra, GPT-6 Sol, and Luna from the additional matrix when changing `architect runners` or another configurable role. Read each model, proposed effort, and selectable efforts from its row. The additional families and stock Sol are separate families even though they share the Codex provider; changing one family's effort does not change another's. GPT-6 Sol uses the `sol-6` family; the stock `sol` family keeps GPT-5.6 Sol. -Ask exactly four effort questions, one each for Fable, Sol, Grok, and Opus. Name each model, its current or proposed value, and the Selectable efforts from its matrix row. Empty input keeps a current value or accepts the matrix proposal for an unassigned family. On a first run, state the four matrix defaults before asking. On a rerun, state the four parsed values without offering to reset customized role lanes. +The assigned families are exactly the matrix families that appear in the resulting role map. An unassigned family gets no effort question and no probe. There is no requirement to assign every matrix family. -### 5. Probe the four requested pairs +Then ask one effort question per assigned family. Name each model, its current or proposed value, and the Selectable efforts from its matrix row. Empty input keeps a current value or accepts the matrix proposal for a newly assigned family. On a first run, state the assigned families' matrix defaults before asking. On a rerun, state the parsed values without re-opening the role choices already made above. -Probe only the four selected `provider:model@effort` pairs. Run one probe per family, even when two families share a provider. Do not enumerate or offer older models as substitutes. A failed probe writes nothing: report the failing pair and provider, stop, and keep the active sheet plus parent integration bytes unchanged. A failed first run creates neither artifact. +### 5. Probe the assigned pairs + +Probe only the assigned families' selected `provider:model@effort` pairs. Run one probe per assigned family, even when two families share a provider. Do not enumerate or offer older models as substitutes. A failed probe writes nothing: report the failing pair and provider, stop, and keep the active sheet plus parent integration bytes unchanged. A failed model demands explicit repair or role reassignment before saving. A failed first run creates neither artifact. | Family | Pair source | Claude parent route | Codex parent route | Availability proof | |---|---|---|---|---| | Fable | Fable matrix row + selected effort | native Agent `pstack-fable-` | Claude CLI | native one-turn probe or `claude auth status --json` plus one-turn probe | | Sol | Sol matrix row + selected effort | `codex exec` | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | +| Astra | Astra additional row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | +| GPT-6 Sol | sol-6 additional row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | +| Luna | Luna additional row + selected effort | external runner | native `spawn_agent` | `codex login status` plus one-turn probe or native one-turn probe | | Grok | Grok matrix row + selected effort | Grok CLI | Grok CLI | `grok models` must list the requested model; one-turn probe | | Opus | Opus matrix row + selected effort | native Agent `pstack-opus-` | Claude CLI | native one-turn probe or `claude auth status --json` plus one-turn probe | +| DeepSeek Flash / Pro | Each assigned DeepSeek flex row + selected effort | external runner | external runner | `DEEPSEEK_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint | +| MiniMax M3 / M3.1 Flash Preview | Each assigned MiniMax flex row + selected effort | external runner | external runner | `MINIMAX_API_KEY` present; isolated config dir free of OAuth credentials; one-turn probe confirms the endpoint | + +For MiniMax M3.1 Flash Preview, disclose the Token Plan requirement before probing. Use the eligible subscription key through `MINIMAX_API_KEY`; do not assume a working M3 key grants preview access. A failed preview probe must not silently select M3. Keep preview thinking enabled and verify requested effort forwarding; distinguish request evidence from hidden applied reasoning depth. -Use a tiny read-only probe that returns a unique marker. A login-status command alone proves credentials, not that the requested model and effort flags run. Record native and external results separately. Never call the external launcher for the parent's own provider. On a Claude parent, the Fable and Opus probes are one-turn runs of the mapped `pstack--` agent. On a Codex parent, the Sol probe is native `spawn_agent` with the selected `reasoning_effort`. Every other pair uses the external runner with the selected effort flag. +Use a tiny read-only probe that returns a unique marker. A login-status command alone proves credentials, not that the requested model and effort flags run. Record native and external results separately. Never call the external launcher for the parent's own provider. On a Claude parent, the Fable and Opus probes are one-turn runs of the mapped `pstack--` agent. On a Codex parent, each assigned Codex family gets a native `spawn_agent` probe with its matrix model and selected `reasoning_effort`. Every other pair, flex families always included, uses the external runner with the selected effort flag. A flex probe doubles as the base-URL confirmation: it proves the documented default (or the operator's override) actually serves the lane's model. Receipts and native transcripts prove the requested effort and the route. They do not prove a provider's hidden applied reasoning depth. There is no implicit timeout, weaker-model fallback, same-provider external fallback, or second mutable configuration source. @@ -67,11 +82,13 @@ Build the new sheet in memory. Do not write it yet. - First run: start from the complete role assignments in step 7. - Rerun: start from the normalized complete role map from step 2, preserving each loaded row's lane order and family (or alias) per lane. -After effort selection, ask whether to keep those role-to-family assignments or change named roles. Keeping them is the default. Apply only role changes the operator names; never offer a reset of a customized sheet to the first-run assignments. A changed role may use one of the four probed matrix families, `inherit-parent`, or `auto`. +The role assignments were already chosen in step 4; do not re-open them here. Require every documented role to remain present and non-empty, `architect runners` to keep at least two entries, and the final role map to contain at least one assigned matrix family. There is no requirement to assign every matrix family. The sheet stores effort only in role descriptors, so an unassigned family's selection cannot persist without adding a second source of truth. + +Different models sharing a provider count as one provider, even when their efforts differ. -Require the final role map to contain at least one descriptor from each matrix family. The sheet stores effort only in role descriptors, so an unassigned family's selection cannot persist without adding a second source of truth. +Validate panel diversity: `arena runners` and `interrogate reviewers` must span at least two distinct providers. A single-provider panel is written only after the operator explicitly confirms the reduced diversity; record that confirmation in the setup report. -Rewrite every matrix-family descriptor to `provider:model@`. Leave `inherit-parent` and `auto` unchanged. An effort-only rerun cannot change a role's family. Changing Grok's effort updates every Grok occurrence and does not move a Sol role onto Grok. Refuse an unqualified slug, an unavailable route, a model other than the four matrix families, or a provider/model mismatch. +Rewrite every matrix-family descriptor to `provider:model@`. Leave `inherit-parent` and `auto` unchanged. An effort-only rerun cannot change a role's family. Changing Grok's effort updates every Grok occurrence and does not move a Sol role onto Grok. Refuse an unqualified slug, an unavailable route, a model outside the stock, additional, and flex matrix families, or a provider/model mismatch. ### 7. Confirm and commit @@ -109,12 +126,12 @@ interrogate reviewers: claude:fable@max, codex:gpt-5.6-sol@max, grok:grok-4.6@xh Render the parent integration in memory before either write. On Claude, the integration is the single `@~/.claude/pstack-models.md` include in `~/.claude/CLAUDE.md`. On Codex, it is the exact sheet bytes between one `` and `` pair in `~/.codex/AGENTS.md`. Replace that whole bounded block on a rerun. Insert one block at the end on first run. If either marker is missing, duplicated, or reversed, stop and report inconsistent state instead of guessing a boundary. -Snapshot every target's current bytes. Write the sheet and parent integration only after all four probes pass and the operator confirms. Read both targets back and compare them with the in-memory render. If either write or readback fails, restore every snapshot and report the failure. An unchanged rerun must produce byte-identical sheet and integration content after normalization. +Snapshot every target's current bytes. Write the sheet and parent integration only after every assigned family's probe passes and the operator confirms. Read both targets back and compare them with the in-memory render. If either write or readback fails, restore every snapshot and report the failure. An unchanged rerun must produce byte-identical sheet and integration content after normalization. Do not copy the model sheet between harnesses without rerunning the parent-specific probes; route availability can differ even on the same host. ### 9. Behavioral smoke -Before declaring setup complete, run one small read-only mixed panel from this parent: all four chosen descriptors, distinct output/receipt paths, and an independent cross-judge. Launch Claude-native agents and every external process in the background with retained handles, then drain them. Verify the native transcript entries and every external receipt. A structural config check or unit test is not a substitute. +Before declaring setup complete, run one small read-only mixed panel from this parent: every assigned family's chosen descriptor, distinct output/receipt paths, and an independent cross-judge when at least two providers are assigned. Launch Claude-native agents and every external process in the background with retained handles, then drain them. Verify the native transcript entries and every external receipt. A structural config check or unit test is not a substitute. Report the sheet path, parent route table, requested-effort probe results, smoke results, and external elapsed/token/cost receipts. Re-running this skill re-probes and updates the same sheet. Do not claim the provider exposed hidden applied-effort observability.