feat(command-goat): add Command GOAT provider with 50 models - #6864
feat(command-goat): add Command GOAT provider with 50 models#6864jxiansen wants to merge 1 commit into
Conversation
Action items
|
3363c55 to
9c01bf5
Compare
|
All 4 review items fixed and pushed (single commit, re-validated locally with
Waiting on re-review. |
Action items
|
9c01bf5 to
8b4b2ec
Compare
|
Round 2 fixed and pushed (single commit,
Waiting on re-review. |
Action items
|
8b4b2ec to
afcb775
Compare
|
Round 3 addressed (single commit,
Waiting on re-review. |
Action items
|
afcb775 to
8f84038
Compare
|
Round 4 all addressed (single commit,
Waiting on re-review. |
Action items
|
8f84038 to
5c61beb
Compare
|
Round 5 addressed (single commit,
|
Action items
|
5c61beb to
c120341
Compare
|
Round 6: kept [low,medium,xhigh] on the five Qwen3.6/3.7 files with per-model live evidence (single commit, bun validate exit 0). Conceded that native 3.6/3.7 exposes toggle+budget only, so family-copying 3.8 was insufficient justification. Instead I probed graded behavior on this host directly (Qwen3.7-Flash, same reasoning prompt, max_tokens 300):
Monotonic and replicated across prompts, so graded effort is a real control on Command GOAT for this family. Levels stay, headers now cite this probe instead of Qwen3.8. Waiting on re-review. |
Action items
|
Provider: https://commandcode.ai (OpenAI-compatible, https://api.commandcode.ai/provider/v1) Plan scope: GOAT plan only (https://commandcode.ai/docs/plans/goat); Go/Pro/Max excluded, separate provider per plan convention (cf. alibaba-coding-plan) Host kind: multi-model relay. Reasoning per family baselines (lab entries + same-surface relay peers): deepseek-V4 [low,high,max]/pro [high,max]; GPT-5.6 5-level (none dropped); Qwen [low,medium,xhigh] (302ai relay); gemini/step-3.7/L/M/H; muse minimal-inclusive native sets; grok per-generation; K3/GLM-5.2 [high,max]/GLM-5.3 [low,high,max]; step-3.5 [low,high]; tencent-hy [high]; inkling 5-level; sante L/M/H; binary-only families (mimo/minimax-M3/nemotron/K2.6/longcat/laguna/GLM-5/5.1) and native-[] (M2.5/M2.7/K2.7/longcat) -> []. Wire verified live 2026-09-11 (8 models, 42 calls): gateway enum exactly low|medium|high|xhigh|max; none|minimal rejected with explicit enum error; enable_thinking accepted-but-ignored (no toggle). Toggle/budget omitted: no field evidenced on this host. Source: GOAT pricing page 2026-09-11 (50 models) + model pages. - costs USD/MTok per GOAT page; limits as-served (provider deltas only, override-only) - 50 models incl. deepseek-v4.1-flash and ling-3.0-flash-sante:free (new lab models/inclusionai/ling-3.0-flash-sante.toml) - bun validate: pass
c120341 to
a4d5a0c
Compare
|
Round 7 addressed (single commit, bun validate exit 0 locally):
Waiting on re-review. |
|
No actionable findings. |
Provider: https://commandcode.ai (OpenAI-compatible, https://api.commandcode.ai/provider/v1)
Plan scope: GOAT plan only (https://commandcode.ai/docs/plans/goat); Go/Pro/Max excluded, separate provider per plan convention (cf. alibaba-coding-plan).
Source: GOAT pricing page 2026-09-11 (50 models) + model pages (deepseek-v4-1-flash, ling-3.0-flash-sante-free) + OpenRouter endpoints API (ling sante lab metadata).
What changed (53 files, +843, single commit, override-only):
Host kind: multi-model relay. Reasoning per family baselines (lab entries + same-surface relay peers, each cited in file headers). Wire verified live 2026-09-11 (8 models, 42 calls): gateway enum exactly low|medium|high|xhigh|max; none|minimal rejected with explicit enum error; enable_thinking accepted-but-ignored (no toggle). Toggle/budget omitted: no field evidenced on this host.