Skip to content

feat(orchestration): providers, waiting, models, profiles and final answers for subagents - #107

Open
BIackFIame wants to merge 9 commits into
howdeploy:mainfrom
BIackFIame:pr/1-orchestration
Open

BIackFIame wants to merge 9 commits into
howdeploy:mainfrom
BIackFIame:pr/1-orchestration

Conversation

@BIackFIame

@BIackFIame BIackFIame commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Branch: pr/1-orchestration → base main (a0f23b4).

Depends on: nothing. This is the first PR of the stack.

Why

  • An orchestrator did not know which agents it could start, so it searched the disk (~/.local/bin, CLI config folders) instead of delegating.
  • Subagents started in a project folder whose name has Cyrillic letters and a space asked "Access external directory?" about their own project (NFC vs NFD spelling).
  • spawn_agent could not pass a model or an effort, and subagents always started in the normal profile, so they asked the person about every step.
  • get_agent_result returned the raw tail of a full-screen TUI, not the agent's final reply, and there was no way to wait for a subagent.

What changes for users after merging

An orchestrator agent can now ask CanvasTTY which agents exist, start them with the model the person asked for, wait until they finish and read their final answer. Subagents start in the right folder even when the folder name uses non-Latin letters, and they inherit the orchestrator's launch profile (never YOLO), so they stop asking about every step.

Scenario (hidden end-to-end run, stub agents) Before (1.7.0) After
Orchestrator finds the available agents searches the file system list_providers: id, installed, sign-in state, models, profiles
Subagent in a Cyrillic folder stored as NFD, path typed as NFC asks "Access external directory" starts in the on-disk spelling, PWD matches (2/2 subagents)
Model requested for a subagent ignored (CLI default) passed as --model (2/2 subagents)
Unknown model subagent fails to start refused up front, with the closest names
Waiting for a subagent polling the screen wait_for_agent returns on idle / approval / exit / timeout
Result of a finished subagent screen tail final reply text (answer, max 4,096 chars, masked)

Numbers come from the hidden smoke of the packaged app (see "How to verify"); the stub agents do not measure real CLIs.

What's inside

  • Discovery. list_providers (MCP) and providers (control CLI): exact ids, availability, sign-in state from the last usage read (never starts a read, never reads credentials), model format and effort levels, OpenCode's cached opencode models, plugin launch options. spawn_agent refuses an unknown provider with a reason that names list_providers.
  • Waiting. wait_for_agent({ sessionId, timeoutSeconds ≤ 600 }) for the orchestrator's own subagents; answers with status, exit code and a masked tail; stops on cancel or disconnect.
  • Folders. A path typed in NFC to a folder stored in NFD launches in the on-disk spelling (no symlink resolution); PWD is set per launch; OpenCode allows exactly that folder in its other spelling.
  • Model and effort. spawn_agent and create --model --effort, mapped to each CLI's own flags for that run only; validated per provider; kept on restart and restore; nothing written to CLI config.
  • Profiles. OpenCode Auto as a per-run config (OPENCODE_CONFIG_CONTENT, .env still asks, bash auto only while base protection guards it). Subagents take profile or inherit the orchestrator's; YOLO is never handed down.
  • Final answers. Codex (Stop hook) and OpenCode (session.idle, last assistant message) subagents report their final reply; memory only, cleared on the next turn.
  • Orchestrator skill, docs, MCP instructions and refusal messages describe list_providers → spawn_agent → wait_for_agent → get_agent_result.

How to verify

npm ci
npm run typecheck
node --test tests/orchestration-providers-wait.test.mjs tests/unicode-cwd.test.mjs tests/launch-model.test.mjs \
  tests/opencode-auto.test.mjs tests/opencode-result.test.mjs tests/subagent-profile.test.mjs tests/orchestration-launch.test.mjs
npm run build

Manual: start an OpenCode orchestrator in a folder named with Cyrillic letters, ask it to "split the task between two OpenCode agents on model X and wait for them". It should call list_providers, spawn two cards on model X without folder prompts, wait, and quote both final answers.

Risks / follow-ups

  • Real provider CLIs were not run in the end-to-end check; stubs speak their protocols. OpenCode's rule order was checked against opencode 1.18.33.
  • The final answer is captured for Codex and OpenCode only; other CLIs still return the screen tail.
  • Sign-in state is only as fresh as the last usage-limits read (unknown otherwise).

An orchestrator with canvastty_agents did not know which providers it
could spawn, so it searched ~/.local/bin and ~/.config instead of
delegating, and had to poll for results.

- list_providers: every agent provider CanvasTTY can launch as a
  subagent, with the exact spawn_agent.provider id, name, installed and
  available from the provider CLI registry, sign-in state from the last
  usage-limits read (ok, signed_out, expired, unknown; LimitsService.peek
  never starts a read and nothing reads credentials), subagent and
  orchestrator support, and plugin launch options with the plugin tool
  that picks them (such as list_routes). The control CLI answers the
  same list as `providers`.
- spawn_agent lists the known ids (enum, mirrored from providerCatalog
  and kept equal by a test) and refuses an unknown provider with
  INVALID_REQUEST naming list_providers. The stdio helper answers
  malformed core calls itself, since the bridge drops the connection
  over them.
- wait_for_agent({ sessionId, timeoutSeconds <= 600 }): returns when the
  subagent is idle (after its screen settled), needs_approval, exited
  (done/failed), quiet, closed, or on timeout, with status, exit code,
  waited time and the tail masked before it is cut. Own subagents only,
  never the orchestrator itself; it stops at once on cancel or
  disconnect. It reads only metadata and the output offset while it
  waits.
- The MCP instructions, tool descriptions, orchestrator skill, docs and
  both refusal messages give the workflow list_providers -> spawn_agent
  -> wait_for_agent -> get_agent_result and say not to search the
  filesystem for agent CLIs or their configuration.
OpenCode subagents asked "Access external directory ~/Downloads/<name>"
for their own project when its name had a letter with two Unicode
spellings (Cyrillic "й"). Finder stores names decomposed (NFD); the path
an orchestrator types is composed (NFC). macOS opens either, so the CLI
started in the NFC path, read its working folder back from getcwd in
NFD, and OpenCode's string comparison between its project root and the
paths in its prompt made the project external. Codex trust entries had
the same mismatch.

- TerminalManager.create spells the requested cwd as stored on disk
  (onDiskPath: each non-ASCII part replaced with the parent's entry of
  that name, exact match first, else the one entry equal after
  normalization; symlinks are kept, ambiguous or unreadable parts are
  left as given). This covers spawn_agent, the control CLI, plugin
  sessions.create and the launcher.
- Every launch gets PWD set to its folder (the wrapped one for plugin
  environments) instead of inheriting the app's.
- OpenCode gets permission.external_directory "allow" for exactly its
  project folder in its other Unicode spelling (the folder and its
  contents) in the per-run inline config; a person's own external_directory
  word is kept as "*", YOLO and ASCII paths are unchanged.
…effort

The person asked an orchestrator for subagents on "GLM-5.3 Flash"; they
ran on OpenCode's default GLM-5.3 because spawn_agent had no model.

- CreateSessionRequest/SessionMetadata carry model and effort; the
  launch passes them to the CLI's own flags for that run only (--model:
  OpenCode provider/model, Codex, Claude, Qwen, Kimi alias, Grok, OMP,
  Pi, Cursor; effort: Codex -c model_reasoning_effort, Claude --effort,
  Grok --reasoning-effort, as their --help lists them). Hermes, MiniMax,
  Devin and Antigravity take none. Nothing is written to CLI config.
- Checked per provider (shared/launchModel.ts) with the reason in the
  refusal and a pointer to list_providers / the providers command; kept
  on restart and restore (persisted, and a stored value the CLI would
  not take is dropped instead of failing the restore).
- spawn_agent gets model and effort (effort enum mirrored and tested);
  the control CLI gets --model and --effort.
- list_providers shows each provider's model format and effort levels,
  and for OpenCode the models `opencode models` lists: run in the
  background with a 5 s timeout, cached for 10 minutes, never awaited.
- The skill, docs and tool texts say: if the person names a model,
  pass it as model.
- OpenCode stops with only "Unexpected server error" on a model it does
  not know. A model missing from the cached `opencode models` list is
  refused before launch (spawn_agent, control create, TerminalManager
  for the launcher and plugins) with up to five closest ids; when the
  list cannot be read, the model is passed as given.
- A subagent that exits anyway reports the last lines of its screen as
  plain masked text (exitLines) in wait_for_agent, observe_agent and
  get_agent_result.
OpenCode subagents asked the person for every action. OpenCode 1.18 has
no auto flag, so its auto profile is an inline OPENCODE_CONFIG_CONTENT,
merged with the core's own inline config and never written to
~/.config/opencode.

Checked against the installed opencode 1.18.33: Permission.evaluate is
rules.findLast(permission and pattern match), fromConfig turns
{ tool: action } into a "*" rule and { tool: { pattern: action } } into
one rule per pattern in order, and an agent's permission is appended
after the top-level one. So the rules go under agent.build (OpenCode's
default agent): after the person's own rules, and every tool not named
keeps what the person's configuration says.

- read, glob, grep, list: allow (".env" files still ask, as OpenCode's
  default does); edit (edit, write, apply_patch): allow.
- bash: allow only when base protection is on and CanvasTTY's guard is
  installed in that OpenCode (tool.execute.before, hard denies still
  deny before OpenCode's own check); otherwise it asks, so auto never
  grants more than the person chose. A launch contributor's other model
  keeps bash asking, like accept-edits elsewhere.
- external_directory is untouched: outside the project asks as before.
- hasAutoMode includes OpenCode (launcher, card label, control CLI);
  TerminalManager learns base protection through configureBaseProtection.
…ofile

Subagents were always launched in the normal profile, so a person who
ran the orchestrator in auto was still asked about every step of every
subagent.

- spawn_agent takes an optional profile, "normal" or "auto" (auto only
  where the subagent's CLI has one). Without it the subagent gets its
  orchestrator's profile; an auto its CLI lacks becomes normal.
- YOLO is never handed to a subagent: core allows it only in an
  isolated environment the person chose, which spawn_agent cannot pick.
  An explicit "yolo" is refused with the reason, and a YOLO
  orchestrator's subagents run in auto (or normal).
- The answer reports profile and profileInherited, list_agents shows
  each subagent's profile (and model), and the card shows the profile
  it runs in. The control CLI's create --profile stays its equivalent.
- The skill and docs describe it.
…screen tail

How core decided an OpenCode subagent finished: CanvasTTY's OpenCode
plugin reports working on session.status busy and idle on session.idle
through the runtime gateway, so wait_for_agent returned at the right
time. But get_agent_result only returned the raw PTY tail (the last
8,192 characters of a full-screen TUI's redraw sequences), and state
stayed "running" while the CLI was open, so the orchestrator never got
the answer text.

- The OpenCode plugin, when result capture is on for its session, reads
  the session's last assistant message (text parts, not tool, reasoning
  or synthetic parts; v2 and v1 SDK shapes; 2 s timeout) at
  session.idle and reports it with the turn's end, at most 4,096
  characters with the end kept. The runtime gateway accepts a result on
  OpenCode's session.idle as it does on a Stop hook, still only for a
  lease that captures results.
- spawn_agent turns result capture on for Codex and OpenCode subagents
  only (TerminalManager accepts it for both); ordinary cards never keep
  an answer.
- TerminalManager keeps the last answer in memory, masks it when read,
  and clears it when the next turn starts. get_agent_result returns it
  as answer, plus status; wait_for_agent includes it when the turn ended.
- Tests drive the plugin in a separate process with a fake SDK client
  against a real runtime gateway, and cover masking, truncation, the
  v1/v2 client shapes and the no-capture path.
@BIackFIame

Copy link
Copy Markdown
Contributor Author

Привет! Собрал всё, что накопилось после 1.7.0, в шесть PR-ов стеком. Коммиты не переставлялись: каждая ветка — префикс одной линейной истории поверх текущего main (15983c4, включая 5 Windows-фиксов), поэтому мержить нужно строго по порядку, после каждого мержа следующий PR показывает только свои коммиты.

Порядок ревью и мержа:

  1. feat(orchestration): providers, waiting, models, profiles and final answers for subagents #107 orchestration (6 коммитов) — оркестратор видит список агентов, передаёт модель, ждёт сабагентов и получает их финальный ответ; папки с кириллицей; наследование профиля.
  2. fix: close the gaps found by the code audit (safety, sessions, plugins, IPC, redaction) #108 audit fixes (10) — дыры base protection, согласованность сессий/владельцев, гонки установки плагинов, проверки отправителя IPC, маскирование секретов.
  3. feat(agents): delegation rules, an OS isolation layer and per-CLI modes #109 agent isolation (5) — правила делегирования, изоляция на уровне ОС (sandbox на macOS, bubblewrap на Linux), режимы для каждого CLI, ask → deny там, где CLI не умеет спрашивать. Самый важный для безопасности, стоит смотреть внимательнее всего.
  4. perf(helpers): a native canvastty-helper for the MCP servers and hooks #110 native helpers (8) — нативный хелпер на Go вместо Electron-as-Node для MCP и хуков, тесты паритета, бенч.
  5. fix: two bug sweeps across orchestration, isolation, sessions and helpers, plus a git audit after isolated sessions #111 bug sweeps (16) — два прохода по багам + аудит git-настроек после изолированной сессии с кнопкой «Neutralize».
  6. perf: smaller app, cheaper canvas and terminals, faster startup, hidden surfaces paused #112 performance (44) — размер, канвас и терминалы, старт, пауза для скрытого.

Цифры (A/B против пересобранного 1.7.0 в одной сессии, 3 прогона, стабы вместо CLI): память приложения с 10 агентами и оркестратором −37 %, хелпер на агента 60 → 6 МБ, permission gate Claude 165 → 23 мс на вызов, компонентов на событие пан/зума 73 → 2, первый интерактивный кадр 511 → 434 мс, приложение 376 → 259 МБ (zip 195 → 121 МБ).

Проверка: на финальной ветке полный набор тестов (1507: 1506 pass, 1 skip), typecheck, build, audit:secrets; на промежуточных — typecheck, build и тесты своего диапазона. Скрытый смоук собранного пакета прошёл.

Честно о границах: на Windows нет слоя изоляции и нативный хелпер там по умолчанию выключен (транспорт через named pipe проверен только в CI); на Linux нужен bwrap; Lightpanda не входит. Стек уже перебазирован на 15983c4; конфликты были в index.ts, OrchestrationGateway.ts и двух тестах, обе стороны сохранены.

По твоим Windows-фиксам в 15983c4: посмотрел — правильные, на macOS/Linux поведение не меняют. Одно замечание: если pipe host оркестрации на Windows падает (fatal), шлюз останавливается и сам больше не поднимается — оркестрация не работает до перезапуска приложения. В нашем стеке при слиянии с твоим кодом stop() на Windows теперь дожидается незавершённого старта pipe host, а не рвёт его сразу; на живом Windows это не проверялось, только в CI.

Плагины: в пяти официальных плагинах есть ветки fix/audit-luna под эти изменения, заметки в plugins.md.

@BIackFIame

Copy link
Copy Markdown
Contributor Author

Rebased the whole stack (#107–#112) onto current main (a0f23b4: keyboard presets, chat history, Web companion, Linux Codex TUI). Conflicts were resolved keeping both sides (e.g. history resume and per-launch model/effort in TerminalManager.create, keyboard shortcuts next to base protection in startup). This PR's own commits are unchanged apart from that.

CI on the fork for this tip (28de035), all required jobs green: verify, macos-cli-resolution, windows-pipe-host — https://github.com/BIackFIame/CanvasTTY/actions/runs/36773044875

Tips and runs for the rest of the stack are in the comments on #108–#112.

@BIackFIame

Copy link
Copy Markdown
Contributor Author

Follow-up to the Windows fatal-host issue recorded in the stack summary: the orchestration gateway now replaces a failed pipe host with at most three retries (500 ms, 1 s, 2 s). An initial startup failure still rejects; explicit stop cancels recovery; disable/re-enable pauses and resumes it. Stale events cannot stop the replacement, and an explicit start safely takes ownership of a pending recovery.

All old connections and capability leases are revoked on failure. Recovery allows newly launched orchestrators to register; an existing orchestrator with an expired lease still needs its terminal restarted. Tests cover those boundaries, including the independently reproduced disable/re-enable/start race.

The shared environment test fixtures now stop all persistence owners before removing storage. The history-HUD fixture uses one fixed clock origin so Windows scheduling cannot reverse its timestamps. No assertions were removed.

Updated head: bb4102ed. All three upstream checks (verify, windows-pipe-host, macos-cli-resolution) passed: https://github.com/howdeploy/CanvasTTY/actions/runs/36830222330.

Independent GPT-6 Luna reviews found no remaining concrete defect in the reported cases. Final cumulative local validation: 1605 passed, 3 skipped, 0 failed; production build/typecheck, secret audit and real-Electron hidden-terminal proof passed. Lower branches were additionally checked with their own gateway/environment/history tests and typecheck.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant