Skip to content

perf: smaller app, cheaper canvas and terminals, faster startup, hidden surfaces paused - #112

Open
BIackFIame wants to merge 116 commits into
howdeploy:mainfrom
BIackFIame:pr/6-performance
Open

BIackFIame wants to merge 116 commits into
howdeploy:mainfrom
BIackFIame:pr/6-performance

Conversation

@BIackFIame

@BIackFIame BIackFIame commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Branch: pr/6-performance → base main (stacked on the previous PR).

Depends on #111 (bug sweeps).
Only this PR's commits: BIackFIame/CanvasTTY@pr/5-bug-sweeps...pr/6-performance

Why

  • The macOS app was 376 MB (195 MB zipped): the asar carried renderer dependencies Vite already bundles, and the skins shipped as large PNGs.
  • Every pan or zoom step re-rendered ~73 React components, including every terminal card, because the camera lived in React state.
  • Off-screen and hidden surfaces kept working: DOM terminals rebuilt on scroll, hidden terminals kept painting, hidden browser tabs ticked at full rate, Settings polled Even G2 and plugin services while closed, plugin frames animated where nobody could see them.
  • The window waited for every service before loading the renderer, so the first usable frame came late.

What changes for users after merging

A smaller download, a canvas that stays smooth with many cards, less CPU while terminals print, and a window that is usable sooner. Hidden terminals stop painting and keep their complete parser state/history. A flooding hidden terminal still incurs the full-stream parsing cost, delivered in bounded pieces; quiet hidden cards receive no output. Hidden browser tabs, closed Settings and off-screen plugin cards reduce their background work as described below.

Measured on the whole stack vs 1.7.0 before the rebase onto 15983c4 (Windows-only upstream changes) (interleaved A/B in one session, 3 runs each, stub CLIs). These rows come from paths this PR changes:

Metric Before (1.7.0) After Change
App size (electron-builder --dir, arm64) 375.6 MB 258.5 MB −31 %
Zip of the app 194.9 MB 120.8 MB −38 %
app.asar 92 MB, 915 files 19.2 MB, 535 files −79 %
React components rendered per pan/zoom event (10 shells) 72.7 2.0 −97 %
Renderer main-thread time per pan/zoom event 2.03 ms 1.33 ms −34 %
CPU time, 10 shells, 30 s of output (5 × 256 KB/s), whole app 7.05 s 5.56 s −21 %
— renderer only 3.16 s 1.92 s −39 %
CPU time, 420 pan/zoom events, 10 shells 2.20 s 1.86 s −16 %
CPU time, 60 s idle, 10 shells 0.36 s 0.12 s −67 % (small absolute)
First interactive frame, 3 restored terminals, cold (15 launches) 511 ms (p90 529) 434 ms (p90 450) −15 %
Same, warm (30 launches) 509 ms (p90 526) 432 ms (p90 452) −15 %
Hidden native browser tab, timer ticks per second 30 2 −93 % (branch micro-benchmark, estimate)
Even G2 headless parsers kept after showing 10 sessions 10 1 branch micro-benchmark, estimate

Medians; spreads and load are in the full report. Machine load stayed at ~4–13 during the runs, which is why CPU is given as CPU time for a fixed workload and runs were interleaved. Part of the output-flood and idle gain may come from the smaller perf commits in PR 5; the canvas, terminal, size and startup rows are this PR's.

Unchanged: memory per shell card (the xterm WebGL context in the GPU process), pan/zoom CPU with agent cards (−2 % OpenCode, −23 % Claude — within run spread for OpenCode).

What's inside

  • Size: only what main and preload load is packaged; React and xterm move to devDependencies (bundled by Vite); built-in pixel art as AVIF; Even G2 icon 512 px and WOFF2-only Latin + Cyrillic fonts.
  • Canvas: camera in a store outside React; pinch-zoom commits batched per animation frame; snap targets built only for the dragged card; placement candidates reused; edge-pan loop stops on window blur.
  • Terminals: off-screen DOM terminals no longer rebuilt on every scroll; HOME-hidden cards keep their parser current through bounded delivery before ring eviction; their observed xterm screen stops painting, and pending selection/geometry refresh resumes when shown; geometry writes coalesced and pending PTY output capped; unused scrollback dropped from hydration; one surface lifecycle through which terminal cards suspend drawing.
  • Startup: persistence loads overlap; the renderer loads while services start, behind an IPC readiness gate; the canvas mounts from a critical snapshot and heavy surfaces load later; Settings and the launch/link dialogs are lazy; cheap always-on boot marks for main and renderer.
  • Visibility: hidden and idle native browser tabs throttled (off while automation drives them); plugin frames nobody can see are suspended with their state kept; Settings polls its sections only while open; Even G2 and plugin services not polled while Settings is closed; home clock ticks once a minute.
  • Browser / plugins / companion bounds: skip redundant agent-presence evaluations and native view calls; staged upload copies bounded and cleaned; audit disk bounded and verified by streaming, hash-chain root anchored; plugin stdin queue and install previews capped by bytes; showcase manifests per page and reused; Even G2 parsers disposed for sessions no longer shown.
  • Fixes found on the way: git worktree add against another repository judged as a write; a late launch cleanup revokes only its own lease; a late startup list does not restore a card removed meanwhile; window controls accept only the app's renderer.
  • Bench and docs: startup-readiness bench (scripts/bench/startup.mjs) with frame probes; memory snapshot after the canvas settles; docs for the camera store and the cheaper canvas.
  • Docs for the whole series (last commit, cb7835a): README (en/ru/zh-CN) describes launch modes, protection layers, the orchestration tools and the native helper's Go build requirement; getting-started and installing-and-security add Go/bubblewrap requirements, Bypass wording and the git audit; CHANGELOG gets an Unreleased section with the measured A/B numbers (the series' entries previously added under 1.7.0 move there).

Latest review fixes and hidden-terminal measurement

The original parser-state regression is covered by the actual alternate-screen reproduction. Hidden painting now pauses at xterm's observed screen while all output continues to reach its parser. A real-Electron regression checks modes, buffers, cells, history, selection, split escape sequences, hidden resize and full refresh on resume; Linux/macOS CI also run it. Linux CI uses an inactive Xvfb window and bounded real-frame/IntersectionObserver/baseline-render readiness, keeping the positive control and zero-render candidate assertions strict.

Fresh Apple Silicon renderer fixture, five cards, 10 s input, 1024 Ki UTF-16 units/s per card, median of three interleaved pairs:

Metric (total for five cards) Prior hidden painting Painting suspended
Renderer CPU 1.868 s 1.749 s (observed −6.4%)
Hidden render callbacks 225 0

All delivered characters and state hashes match in every pair. This measures installed xterm 6 and the product selection guard in Electron; it excludes the application's main process, PTY/IPC and WebGL. It does not eliminate full-stream parsing, return to the incorrect ring-only replay cost, or establish real-provider/multi-platform speed-ups. The older Node/headless flood table measures parsing and cannot measure this DOM change; its CPU values are totals for all cards.

npm run smoke:terminal-hidden
npm run bench:terminal-hidden-renderer -- 5 10 1024 3

The cumulative review fixes also include environment permission precedence for OpenCode, missing Linux settings protection and first-run directories, .git pointer/common-directory audit, Windows pipe flush/lifecycle handling, bounded orchestration recovery with expired leases revoked, and persistence-fixture cleanup after all owners stop.

Final cumulative local validation: 1608 passed, 3 skipped, 0 failed; production build/typecheck, secret audit and real-Electron hidden-terminal proof passed. All Linux/Windows/macOS jobs passed for the synchronized fixture (be09bd78): https://github.com/BIackFIame/CanvasTTY/actions/runs/36831995308. The final test-launch preflight additionally refuses the known macOS Codex seatbelt context before spawning Electron; its refusal, portable regressions and approved GUI path are verified. Required upstream checks are linked in the latest follow-up comment.

How to verify

npm ci
npm run typecheck
npm test                                             # full suite
npm run build && CSC_IDENTITY_AUTO_DISCOVERY=false npm run package
node scripts/bench/baseline.mjs --size-only release/mac-arm64/CanvasTTY.app
node scripts/bench/baseline.mjs --runs 1 --modes shell     # pan/zoom and flood rows
node --experimental-strip-types scripts/bench/startup.mjs --cold 5 --warm 10

Manual: open 10 terminals, pan and zoom quickly — no card flickers or re-lays out; open Settings, close it, and check in Activity Monitor that CanvasTTY returns to idle; put a plugin card off-screen — its animation stops and resumes with its state.

Risks / follow-ups

  • Packaging now relies on the list of what main and preload load; a new runtime dependency of main must stay in dependencies. The size script reports packages nothing loads.
  • The startup gate changes the order of service start and renderer load; the bench and tests cover restore, but a slow disk may show a short loader longer.
  • Plugin frames: an off-screen card is suspended (hidden smoke: timer 87 → 1 tick/s, animation frames 60 → 0). A card in summary mode or under the HOME editing layer stops drawing (0 animation frames) but its JavaScript timers keep running at ~82–84/s; pausing those is a follow-up.
  • Lightpanda (a lighter headless browser for agent tasks) was evaluated separately and is not part of this PR.
  • Windows and Linux packages were not built or measured locally.

An orchestrator with canvastty_agents did not know which providers it
could spawn, so it searched ~/.local/bin and ~/.config instead of
delegating, and had to poll for results.

- list_providers: every agent provider CanvasTTY can launch as a
  subagent, with the exact spawn_agent.provider id, name, installed and
  available from the provider CLI registry, sign-in state from the last
  usage-limits read (ok, signed_out, expired, unknown; LimitsService.peek
  never starts a read and nothing reads credentials), subagent and
  orchestrator support, and plugin launch options with the plugin tool
  that picks them (such as list_routes). The control CLI answers the
  same list as `providers`.
- spawn_agent lists the known ids (enum, mirrored from providerCatalog
  and kept equal by a test) and refuses an unknown provider with
  INVALID_REQUEST naming list_providers. The stdio helper answers
  malformed core calls itself, since the bridge drops the connection
  over them.
- wait_for_agent({ sessionId, timeoutSeconds <= 600 }): returns when the
  subagent is idle (after its screen settled), needs_approval, exited
  (done/failed), quiet, closed, or on timeout, with status, exit code,
  waited time and the tail masked before it is cut. Own subagents only,
  never the orchestrator itself; it stops at once on cancel or
  disconnect. It reads only metadata and the output offset while it
  waits.
- The MCP instructions, tool descriptions, orchestrator skill, docs and
  both refusal messages give the workflow list_providers -> spawn_agent
  -> wait_for_agent -> get_agent_result and say not to search the
  filesystem for agent CLIs or their configuration.
OpenCode subagents asked "Access external directory ~/Downloads/<name>"
for their own project when its name had a letter with two Unicode
spellings (Cyrillic "й"). Finder stores names decomposed (NFD); the path
an orchestrator types is composed (NFC). macOS opens either, so the CLI
started in the NFC path, read its working folder back from getcwd in
NFD, and OpenCode's string comparison between its project root and the
paths in its prompt made the project external. Codex trust entries had
the same mismatch.

- TerminalManager.create spells the requested cwd as stored on disk
  (onDiskPath: each non-ASCII part replaced with the parent's entry of
  that name, exact match first, else the one entry equal after
  normalization; symlinks are kept, ambiguous or unreadable parts are
  left as given). This covers spawn_agent, the control CLI, plugin
  sessions.create and the launcher.
- Every launch gets PWD set to its folder (the wrapped one for plugin
  environments) instead of inheriting the app's.
- OpenCode gets permission.external_directory "allow" for exactly its
  project folder in its other Unicode spelling (the folder and its
  contents) in the per-run inline config; a person's own external_directory
  word is kept as "*", YOLO and ASCII paths are unchanged.
…effort

The person asked an orchestrator for subagents on "GLM-5.3 Flash"; they
ran on OpenCode's default GLM-5.3 because spawn_agent had no model.

- CreateSessionRequest/SessionMetadata carry model and effort; the
  launch passes them to the CLI's own flags for that run only (--model:
  OpenCode provider/model, Codex, Claude, Qwen, Kimi alias, Grok, OMP,
  Pi, Cursor; effort: Codex -c model_reasoning_effort, Claude --effort,
  Grok --reasoning-effort, as their --help lists them). Hermes, MiniMax,
  Devin and Antigravity take none. Nothing is written to CLI config.
- Checked per provider (shared/launchModel.ts) with the reason in the
  refusal and a pointer to list_providers / the providers command; kept
  on restart and restore (persisted, and a stored value the CLI would
  not take is dropped instead of failing the restore).
- spawn_agent gets model and effort (effort enum mirrored and tested);
  the control CLI gets --model and --effort.
- list_providers shows each provider's model format and effort levels,
  and for OpenCode the models `opencode models` lists: run in the
  background with a 5 s timeout, cached for 10 minutes, never awaited.
- The skill, docs and tool texts say: if the person names a model,
  pass it as model.
- OpenCode stops with only "Unexpected server error" on a model it does
  not know. A model missing from the cached `opencode models` list is
  refused before launch (spawn_agent, control create, TerminalManager
  for the launcher and plugins) with up to five closest ids; when the
  list cannot be read, the model is passed as given.
- A subagent that exits anyway reports the last lines of its screen as
  plain masked text (exitLines) in wait_for_agent, observe_agent and
  get_agent_result.
OpenCode subagents asked the person for every action. OpenCode 1.18 has
no auto flag, so its auto profile is an inline OPENCODE_CONFIG_CONTENT,
merged with the core's own inline config and never written to
~/.config/opencode.

Checked against the installed opencode 1.18.33: Permission.evaluate is
rules.findLast(permission and pattern match), fromConfig turns
{ tool: action } into a "*" rule and { tool: { pattern: action } } into
one rule per pattern in order, and an agent's permission is appended
after the top-level one. So the rules go under agent.build (OpenCode's
default agent): after the person's own rules, and every tool not named
keeps what the person's configuration says.

- read, glob, grep, list: allow (".env" files still ask, as OpenCode's
  default does); edit (edit, write, apply_patch): allow.
- bash: allow only when base protection is on and CanvasTTY's guard is
  installed in that OpenCode (tool.execute.before, hard denies still
  deny before OpenCode's own check); otherwise it asks, so auto never
  grants more than the person chose. A launch contributor's other model
  keeps bash asking, like accept-edits elsewhere.
- external_directory is untouched: outside the project asks as before.
- hasAutoMode includes OpenCode (launcher, card label, control CLI);
  TerminalManager learns base protection through configureBaseProtection.
…ofile

Subagents were always launched in the normal profile, so a person who
ran the orchestrator in auto was still asked about every step of every
subagent.

- spawn_agent takes an optional profile, "normal" or "auto" (auto only
  where the subagent's CLI has one). Without it the subagent gets its
  orchestrator's profile; an auto its CLI lacks becomes normal.
- YOLO is never handed to a subagent: core allows it only in an
  isolated environment the person chose, which spawn_agent cannot pick.
  An explicit "yolo" is refused with the reason, and a YOLO
  orchestrator's subagents run in auto (or normal).
- The answer reports profile and profileInherited, list_agents shows
  each subagent's profile (and model), and the card shows the profile
  it runs in. The control CLI's create --profile stays its equivalent.
- The skill and docs describe it.
…screen tail

How core decided an OpenCode subagent finished: CanvasTTY's OpenCode
plugin reports working on session.status busy and idle on session.idle
through the runtime gateway, so wait_for_agent returned at the right
time. But get_agent_result only returned the raw PTY tail (the last
8,192 characters of a full-screen TUI's redraw sequences), and state
stayed "running" while the CLI was open, so the orchestrator never got
the answer text.

- The OpenCode plugin, when result capture is on for its session, reads
  the session's last assistant message (text parts, not tool, reasoning
  or synthetic parts; v2 and v1 SDK shapes; 2 s timeout) at
  session.idle and reports it with the turn's end, at most 4,096
  characters with the end kept. The runtime gateway accepts a result on
  OpenCode's session.idle as it does on a Stop hook, still only for a
  lease that captures results.
- spawn_agent turns result capture on for Codex and OpenCode subagents
  only (TerminalManager accepts it for both); ordinary cards never keep
  an answer.
- TerminalManager keeps the last answer in memory, masks it when read,
  and clears it when the next turn starts. get_agent_result returns it
  as answer, plus status; wait_for_agent includes it when the turn ended.
- Tests drive the plugin in a separate process with a fake SDK client
  against a real runtime gateway, and cover masking, truncation, the
  v1/v2 client shapes and the no-capture path.
…riables

Base protection let three kinds of calls through unchecked:

- A shell call whose input was over the 40 KB bound reached main as a
  preview only; no command was read, so nothing was denied and, with no
  decision plugin, the call ran. With base protection on, a cut shell
  call, or a cut file write whose target is not at the start of the
  preview, is now denied with advice to split the work.
- A folder named in a variable of the same command (`OUT=/x; rm -rf
  "$OUT"`, `export OUT=/x`) was unresolved. Standalone assignments and
  export/declare/local/readonly/typeset now set the value for the rest of
  the command (a prefix assignment does not, as in the shell), and a value
  that cannot be known clears it.
- `tar -C<dir>` with the folder attached, and `git -c alias.x='!cmd'` or a
  `-c` key whose value git runs (fsmonitor, pager, editor, filters,
  helpers) were not read. Their commands are now analyzed like any other.

A decision-plugin list that cannot be read is no longer an empty list: the
handler fails, so the gateway answers as unavailable (a CLI that can ask asks the
person, a fail-closed gate denies), and a launch still installs the hook.
…s restore, close and cancel

Saved cards:
- A card list written by a newer version is left as it is and never
  saved over for the rest of the run; one that cannot be read at all is
  left alone too. A file that is not a card list is kept beside a fresh
  one (terminal-sessions.json.corrupt-<time>) instead of being replaced
  by an empty state when the manager saves.
- Restore resumes environments only for cards that come back. Cards that
  do not (not restored, or a subagent whose parent is gone) release their
  environment with its data kept, instead of leaking it.
- Subagents whose parents loop back on each other have no owner and do
  not come back.

Owners:
- Closing a card closes its subagents first, so none keeps running
  without an owner.
- An environment a plugin prepared after the launch timed out is still
  heard of (the plugin call keeps the host-call budget) and released.
- A controlled create whose setup failed closes the card it started.

Gateways and cancel:
- The orchestration gateway starts once for concurrent callers, and a
  stop during start leaves nothing listening; the browser agent gateway
  closed while its Unix socket opens closes that socket.
- A Windows pipe host that fails and exits late no longer clears the host
  started after it.
- A late start hook of an earlier turn no longer takes over from the
  newer turn, whose Stop was then dropped as stale.
- A cancelled send_to_agent passes its signal down to the delivery: text
  still waiting for the launch is dropped and the answer is CANCELED.
… keep host replies with their process

- Two installs of the same plugin both passed the "already installed"
  check; the loser's failed move then deleted the winner's directory and
  registry entry. Install, module changes, update and uninstall of one
  plugin now run one after another, the check runs inside that order, and
  a failed install removes only what it created.
- A host call a service made before it crashed was answered into the
  stdin of the restarted process. The reply now goes only to the process
  that asked.
- Uninstall revoked secrets before stopping the plugin, so a secret write
  in flight could recreate the file, and a failed uninstall left the
  plugin without its secrets. Uninstall now stops and removes the plugin
  first, then revokes; secret writes are refused while a revoke is pending
  and re-check permission when they run.
- The companion (glasses) read the terminal screen and captured answers
  without the secret registry. The screen text is now masked as a whole
  before it is cut, so a key the terminal wrapped over two lines is found,
  and captured answers are masked when they arrive.
- An owner holding 64 values dropped the oldest one even when it was still
  in use. Adding a held value again now makes it the newest (the vault adds
  each key it reads), and a drop or a refused owner is logged without the
  value instead of happening silently.
… went in

When focusing or selecting the target opened a JavaScript dialog (a focus
handler's alert), browser_type inserted nothing but answered typed: true.
It now answers DIALOG_OPEN with the dialog type, so the agent handles the
dialog and types again. A dialog raised by the page's input handler comes
after the text went in and still counts as typed. browser_select sets the
selection before it fires its events, so a dialog from those events
leaves it selected and its answer is unchanged.
…v names as the OS does

- Clipboard, settings reads, file pickers, limits, plugin canvas and window
  opening, plugin storage and media, the Hermes HUD, and terminal list,
  resize, bounds, rename, restore and visibility handlers now check that
  the call comes from the main window's own frame, like the other
  privileged channels; the terminal ids they take are checked as text.
- Plugin environment names: __proto__ is refused as a name (assigning it
  was silently dropped), a name such as constructor no longer looks like
  one another plugin already set, and on Windows two spellings of one
  name (Path, PATH) collide with each other and with what CanvasTTY sets.
…plies with what they belong to

- A card whose history snapshot failed also lost its live output: the
  error disposed the only subscription. The failure is still reported, and
  the card now goes on with live output from the oldest event queued.
- Provider keys and API profiles could be saved twice (Enter bypassed the
  busy state), and a save that finished late cleared text typed after it
  was submitted. A second save of the same entry is ignored while the
  first runs, and a draft is cleared only while it still holds what was
  saved. Profile edits no longer merge into a draft from an earlier
  render.
- After the launcher's agent changed, the previous agent's launch plugins
  and environments stayed offered until the new list arrived, so their
  options could be sent with the new agent. Lists are shown only for the
  agent they were loaded for.
- A plugin frame request still running when the frame reloaded or
  navigated was answered into the next document (a secrets.get reply
  included). A reply is now dropped once the frame shows a new document or
  its window is no longer the one that asked.
…cement candidates

- The browser card reports its rectangle on every camera step, resize and
  layout pass. A report that rounds to the placement already applied now
  returns before the native view is re-laid out (hidden reports still run:
  they end gestures). In a simulated slow pan with a duplicate report per
  frame, 1200 reports caused 1200 view syncs before and 300 after; idle
  re-renders go from 1200 to 1. A fast pan is unchanged (every report
  moves the view).
- Placing a new card near Home built and sorted a lattice of about 3000
  candidate positions every time. The sorted list is kept for the same
  Home, card size and options, and the search still stops at the first
  free position: 500 placements beside 20 cards took about 240 ms before
  and under 3 ms after.
…ants and plugin frame replies

- Base protection: rm, rmdir and the other deleters, find -delete, and mv
  operands whose path is a variable or command output that cannot be
  resolved are now denied as an unknown target, with advice to write the
  path out. A variable set in the same command, or a for-loop over plain
  words, still resolves (a loop over a folder outside is the delete
  outside it).
- Uninstall marks the plugin as being removed: a music folder pick that
  finishes while it runs, or after it, stores no grant (the grant is
  checked again right before it is stored).
- A plugin frame reply is tied to the window that asked, the plugin the
  frame served and the frame's document generation, all captured when the
  request arrived. Pointing the frame at another plugin or entry starts a
  new generation at once, so a reply that finishes before the new document
  loads or asks anything is dropped.
…not words in values

A contributed argument was refused when its text anywhere contained words
such as "dangerously" or "approval_policy", so a context plugin whose rule
merely mentioned them blocked the Codex or Grok launch. Arguments are now
read by structure: a flag whose name is core-owned or asks to bypass
approvals, a config override (`-c key=value`, `--config=key=value`,
`-ckey=value`) whose key decides approvals, the sandbox or the hooks, and
the core-owned subcommands are refused as before. Text inside a value is
the plugin's own.
…s reason

Only Claude Code takes "ask" from a hook. For Codex, Qwen, OpenCode and the
rest, an ask from a decision plugin (or a plugin that did not answer in
time) printed nothing, so the tool call simply ran. It is now a deny that
says the person must decide, and the gate turns a gateway-level ask into
the same deny for those CLIs.

Decision services now receive the card's profile and canAsk.
Delegation rules, in the core, for every way an agent starts another one
(spawn_agent, an orchestrator's control connection, restore, restart):
- a subagent never gets more than its orchestrator (plan < normal <
  acceptEdits < auto), never YOLO; a restored record cannot raise it;
- its folder is the orchestrator's project or inside it, compared as real
  paths in both Unicode spellings; /, the home folder and other projects are
  refused with a reason the agent can act on;
- nesting depth (2) and live subagents per orchestration (8) are limits only
  the person sets (Settings -> Agents);
- plugin launch options from an agent reach only plugins that declared
  launch.delegable; OpenCode/Kimi configuration a plugin hands the CLI may not
  set approval keys; cursor's --force/-f and other bypass flags are core-owned;
- YOLO is checked in the main process: acknowledged by the person for that
  CLI, never for a subagent or an orchestrator's connection, and a plugin with
  sessions:launch needs the same acknowledgement;
- an orchestrator gets a control connection of its own (never the app-wide
  descriptor): create makes its subagents under these rules, and theme or
  settings commands are refused.

Agent isolation (Settings -> Agents, on by default) wraps the agent's whole
process tree for every subagent, every plugin-started agent and every agent
not in manual: macOS sandbox-exec with a profile generated per launch, Linux
bubblewrap when installed. Writes only in the project (both spellings), the
launch's own TMPDIR and the CLI's own folders; git hooks, an existing
repository's config and the CLI's permission settings stay unwritable; keys,
other CLIs' credentials and CanvasTTY's tokens are unreadable except what the
launch was handed; no signals to other processes, Launch Services, Apple
events, cfprefsd writes, launchd jobs or foreign Unix sockets. It fails
closed. Without a layer (Windows, Linux without bwrap) a subagent runs in
normal and the card says why; turning it off is the person's opt-in. Inside
the layer Claude Code's own sandbox block is left out (macOS refuses a
sandbox in a sandbox); Codex keeps its flags and escalates through its
reviewer. Plugin environments declare what they keep (keeps.launch,
isolated, confines); an isolated one is not wrapped again, and without
launch only normal runs there. Launches never count as shell-guarded inside
an environment.

Modes: auto is the default launch mode; manual, accept edits, plan and
bypass are offered only where the CLI has them (Codex plan is its read-only
sandbox, OpenCode plan its plan agent, cursor --mode plan). A CLI without an
auto of its own gets its bypass as auto only inside isolation. OpenCode auto
keeps the person's own deny and ask rules after its allow rules. Claude's
sandbox, where it runs, no longer lets commands leave it. A manual card
shows when the CLI's own configuration skips approvals. The card shows the
isolation state; list_providers shows each provider's subagent profiles.
What each layer guarantees and what it does not: the agent's own mode,
base protection, the delegation rules and the OS isolation layer (including
nested sandboxes, network, the macOS login keychain and Windows). The
orchestration guide covers subagent profiles, folders, limits, an
orchestrator's own control connection and environments' keeps; the plugin
guide covers launch.delegable and keeps. The ru and zh-CN security pages get
a short version of the new section.
macOS refuses a sandbox inside another one, so inside CanvasTTY's isolation
layer Codex's seatbelt could not start and every command failed once before
being re-requested. Inside the layer Codex now runs with its own sandbox off
(--sandbox danger-full-access, never the bypass flag) and the same approvals:
on-request, and in auto the reviewer --approve-for-me uses
(approvals_reviewer="auto_review"). Verified with codex debug prompt-input
under a fake HOME, run inside the layer by the tests. Outside the layer Codex
keeps its own sandbox. In plan the layer keeps the project read-only.

Claude Code keeps its sign-in in the macOS login keychain and rewrites that
file from inside the process when it refreshes it. A Claude launch inside
the layer may now write that one file and its temporary siblings (tested
with the security CLI on a temporary keychain in a fake HOME), so a
refreshed sign-in is saved. Other keychain items stay behind their access
lists; the tradeoff is documented.
One static Go binary (standard library only, no cgo) with the subcommands
mcp-browser, mcp-orchestration, permission-gate and hook. Each speaks exactly
the wire protocol of its .mjs original: JSON with V8's parse/stringify
semantics (key order, number form, lone surrogates, invalid UTF-8), the same
NDJSON bounds checked after every append, timeouts, reconnects with the
rotated token, and the gate's fail-closed deny. Windows named pipes use
overlapped I/O through syscall. The tool catalogs are generated from the .mjs
sources by scripts/build-native-helpers.mjs.
The same scenarios run through both implementations: every tool input kind,
every gateway answer and failure for claude/codex/qwen/opencode, lifecycle
hooks with answer grants, MCP handshakes, validation, oversized lines both
ways, malformed gateway lines, cancel, timeouts, gateway down and reconnects,
plus the real Runtime, Agent and Orchestration gateways.
…t one otherwise

The browser and orchestration MCP servers, the decision hook and the
lifecycle hook launch canvastty-helper on macOS and Linux when the binary for
this platform and architecture is present (resources/helpers when packaged,
build/native-helpers in development). Windows keeps the .mjs helpers until the
native named-pipe transport has run on Windows; CANVASTTY_HELPERS=node forces
the JavaScript helpers everywhere and =native opts in anywhere. Hook files
written with either form are recognized as CanvasTTY's own on recovery, and
the native hook commands run inside the isolation layer.
…sources/helpers

`npm run build:helpers` builds every release target (darwin arm64/x64, linux
x64/arm64, windows x64) with -trimpath -ldflags "-s -w", CGO off, no module
proxy; `npm run build` builds this computer's target before bundling, and
electron-builder copies the matching one into resources/helpers. CI and
release jobs set up Go from go.mod, require the helper to build, vet every
target, run the parity tests against it and test the Windows pipe transport
on Windows.
--helpers auto|node|native sets CANVASTTY_HELPERS for the app; the bench app
folder links build/native-helpers so the app finds the binary as it would in
development, and native helper processes and hook runs are attributed as
native:<subcommand>.
With "Agent access" to the browser off, the browser bridge returned no launch
at all, so orchestrators (and sessions a plugin tool applies to) lost their
canvastty_agents MCP server with it. Now such a launch attaches only
canvastty_agents for every provider that takes it (Claude Code, Codex, Qwen
Code, OpenCode, Hermes, Kimi per-run and shared): no browser capability is
issued, no browser entry, environment or allowed browser tools. Hermes and
Kimi's shared temporary configurations record whether they carry the browser
entry and refuse a launch with the other browser access while in use.
…atches the macOS layer

- Other CLIs' credential folders are listed where the launch environment moved them (CODEX_HOME, CLAUDE_CONFIG_DIR, GROK_HOME, HERMES_HOME, KIMI_HOME, OPENCODE_CONFIG_DIR/OPENCODE_CONFIG, QWEN_HOME, XDG_*) as well as by default; only the launched CLI's own variable makes a folder its own (before, any of them was made writable and readable).
- bubblewrap: what a launch was handed (its control grant) is bound read-only, as on macOS; the CLI's own moved home stays writable.
- bubblewrap: a project without .git/hooks gets a throwaway tmpfs there, so a hook written after git init inside the layer never reaches the real repository; the empty mount point is removed afterwards unless a repository was made.
…is released

- restart() clears the ended process's hook count, title state and answered-prompt timer, so a restarted Claude card's idle title counts until the new process's own hooks report.
- EnvironmentRegistry.resume() hears a resume that succeeds after its timeout and releases it again with its data kept (the card still holds it), unless a newer resume of the same session holds it.
- Claude HTTP hook policy: the project walk stops at HOME, stats .claude once per folder instead of reading two missing files, checks .git with one stat and never throws for a missing path (a 12-deep folder under HOME: 520 -> 53 us per launch; 20 deep outside HOME: 749 -> 97 us).
- Terminal mouse adapter: its document-level listeners are attached only while the pointer is over the card or drags from it, not for every card on every mouse move.
- Browser presence expiry ticks only while an agent is present.
- WorkspaceCanvas forgets closed sessions' fullscreen callbacks and master-detail flags.
…mbeds the same bytes

The Windows job reported a generated-versus-embedded catalog mismatch.
It was the checkout, not the catalog: with core.autocrlf=true the
runner's catalog.json had CRLF line endings while the generator writes
LF, and a helper built there would embed different bytes. A
.gitattributes entry pins that one file to LF; the test compares the
content first (a real difference names the regeneration command), then
the bytes, and checks the attribute.
…iled first connection is not final

Both orchestration helpers (JavaScript and canvastty-helper) cached the rejected first connection: one early race with a starting or restarting OrchestrationGateway disabled every agent tool for the rest of the session. A connection lost before authentication is now retried twice, 250 ms apart; after that the call fails and the next call connects again. The parity test runs the same scenarios against both.
…te their platform

The Windows job failed on fixtures, not on the behaviour they check:
- agent-isolation made its short temporary folder under /tmp, which does
  not exist on Windows; it uses the temporary folder there.
- the delegation test used "/" as the file system root, which on Windows
  is the current drive (D: on CI), not the temporary folder's (C:).
- the auto-profile launches did not name the platform although Claude's
  own sandbox exists only on macOS and Linux; they are built for Linux,
  and the Windows launch (auto mode without a sandbox) is asserted too.
  The TerminalManager-level settings test expects the sandbox only where
  the host has one.
- configuredMode builds the paths it reads with the host's rules; the
  fixture now does as well.
- the environment fixture wrapped the agent in "/bin/sh", which is not
  an absolute program path on Windows, so the registry refused it and
  the normal launch never started; it uses node's own path.
scripts/bench/baseline.mjs runs the built app hidden (bench-runtime shims,
isolated HOME and provider homes, no keychain) with plain shells or a stub
opencode/claude CLI that starts the MCP helpers, lifecycle plugin and hooks
CanvasTTY attaches. It reports memory per process kind (RSS and
phys_footprint) for 1/5/10 cards plus an orchestrator, CPU while idle, under
output or status churn and during a canvas pan and zoom, startup time to the
first interactive frame, heaps, and with --size the breakdown of a packaged
app (framework, locales, asar by folder, node_modules that main and preload
never load).
…-separate-git-dir)

The post-session git audit walked only .git folders and skipped every
other entry, so a repository created with `git init --separate-git-dir`
or `git worktree add` inside the agent's project was never reported,
although git recognizes it and runs what its config names.

A .git file is now resolved as git does (`gitdir: <path>`, relative to
the working tree), a .git link to what it names, and each git directory's
`commondir` to the main repository: the shared config, hooks and
info/attributes are audited there (once for all worktrees of one main
repository), the worktree's own config.worktree in its git directory.
Neutralize acts on the same files. The walk itself stays bounded and
still does not follow linked folders; pointer files are read only up to
4 KiB and must name an existing folder. The card shows the working tree.

Tests reproduce the reviewer's case with real git (a separate git
directory inside the project, core.fsmonitor set) and a linked worktree
with a hook and a per-worktree setting.
…ile a launch prepares

With missing protected files now refused on Linux (the host gets a
placeholder first), the moved account home in this pure-arguments test
declares its config.toml; the test also checks that it is read-only
again over the rebound home, and that without it the launch is refused.
… the platform it decides for

ClaudeHttpHookPolicy takes the platform as an option (the host's in the
app) but joined the settings paths with the host's rules, so its macOS
fixture walked backslash paths on the Windows runner. The settings walk
and the managed settings paths now use posix or win32 rules by that
platform; nothing changes in the app.
…tput outgrows the ring

A hidden card was not streamed; when shown it got a replay of at most the
240 000-character ring with a notice for the rest. The notice kept the
history honest but could not restore what the dropped part did to the
terminal: with ESC[?1049h followed by 130 000 CRLF and ESC[HCurrent
approval, a terminal fed the whole stream ends in the alternate screen,
the card fed the trimmed replay in the normal one.

The card's parser is now kept current while its painting stays
suspended: just before the ring would drop output a hidden card has not
received, main hands it the missed stretch as one renderer-only event
(one chunk longer than the ring by itself is handed over whole), so the
replay on show always starts where the card stopped. A quiet hidden card
still gets nothing; a flooding one is parsed in ring-sized pieces. The
ring stays bounded and the renderer still marks any hole it is ever
given.

Cost (scripts/bench/hidden-flood.mjs: real TerminalManager and
attachTerminalOutput, @xterm/headless as the card's parser, Node on an
Apple Silicon Mac, 5 hidden cards, 10 s, median of 3): at 256 KB/s per
card 0.34 s CPU against 0.09 s for the ring-only replay and 0.43 s when
streamed visibly; at 1 MB/s per card 1.71 s, 0.14 s and 2.05 s.
The Windows runner checks sources out with CRLF (core.autocrlf=true), so
the three source assertions this series added that expect a newline right
after a statement (summary mode in TerminalCard, the surfaces gate in
WorkspaceCanvas, the selection-redraw restore before dispose) did not
match there. They now accept \r?\n, like the rest of the suite; checked
against CRLF copies of the sources.
… HOME and the project visible

isolationPaths never hides a folder that holds HOME or the project (a
CLI home variable pointing at HOME or above), and lists the folders an
agent may create below HOME. Both compared with a "/" suffix, which
never matches a Windows path; the Windows runner showed HOME and its
parent listed as unreadable. They now use path.relative, so the rule
holds for either separator.
The audit reports files changed at or after the session's start, taken
with Date.now(). Linux stamps files from its coarse clock, which can
read a tick (up to 10 ms) earlier, so a setting written right after the
start could look older than it and be missed; the Linux CI job showed
exactly that for a config written just after launch. Changes up to
50 ms before the recorded start now count.
The 53 terminal frames and 6 canvas backgrounds were 1536x1024 PNGs, 76 MB of
the 92 MB app.asar. They are already at or below the size they are drawn at
(a 1200x800 card is painted at up to ~1730x1150 device pixels, backgrounds
cover the whole window on Retina), so they keep their resolution and change
format: AVIF 4:4:4 (full-resolution chroma for hard pixel edges), lossless
alpha (the transparent terminal opening is exact), quality 92 for frames and
95 for backgrounds. 76 MB -> 14 MB.

Against the PNG masters (composited on dark and light): frames SSIM >= 0.994
(mean 0.998), PSNR >= 45.6 dB (mean 50.6), alpha identical; backgrounds mean
absolute error < 1 level, 99th percentile <= 3, PSNR >= 46 dB. In the packaged
app, screenshots of the Sakura, Matrix and Gothic Eclipse themes differ from
the PNG build by 0.3 levels on average (PSNR 50-53 dB).

scripts/encode-skin-art.mjs encodes new art with the same settings. Imported
themes stay PNG, so the importer tests build their own PNGs instead of reading
the bundled frames.
… shut down

node:test runs after-hooks in the order they were registered. Four
environment tests registered the removal of their session-store folder
before managerFixture registered the manager's shutdown, so the folder
was removed while the store could still write into it; the Linux job
failed once with ENOTEMPTY. The removal is now registered after the
fixture.
…close

The gateways answer a protocol failure (an unauthenticated command, a
replayed token) with an error message and close the connection. On
Windows both go through the current-user pipe host as a write frame and
a destroy frame; the destroy marked the connection closing at once and
its writer thread stopped without writing what was still queued, so the
client sometimes saw the pipe close without the error. The Windows job
failed orchestration-gateway's "unauthenticated commands and replayed
bootstrap tokens are rejected" this way (a timeout waiting for the
error), intermittently.

A destroy now lets the writer drain what was queued before it, bounded
by two seconds for a client that stops reading, then closes. The relay
self-test (npm run test:windows-pipe-host) sends a last message and the
close in one write and requires the message to arrive.
A Claude Code HTTP lifecycle hook whose body is over the 512 KB bound was
answered (and its connection closed) while Claude was still sending it, so
the sender could see a reset instead of the answer. The state is still
reported at once; the answer now waits until the rest of the body has been
read and dropped.
@BIackFIame

Copy link
Copy Markdown
Contributor Author

Fixed (rebased onto main a0f23b4):

  • P2 terminal state. I chose to keep the card's parser current while painting stays suspended. Just before the ring would drop output a hidden card has not received, main hands it that stretch as one renderer-only event (a single chunk longer than the ring is handed over whole), so the replay on show always starts where the card stopped. A quiet hidden card still gets nothing; the ring stays bounded; the renderer still marks any hole it is ever given. Tests (terminal-hidden-output.test.mjs): the reviewer's case: a TUI enters the alternate screen while hidden and floods; the card ends in the session's state (your ESC[?1049h + 130,000 CRLF + ESC[HCurrent approval, two @xterm/headless terminals, fails on the previous head: card ends in normal), and a hidden stretch longer than the ring reaches the observers and the card whole; the ring stays bounded.

  • Cost, measured honestly (scripts/bench/hidden-flood.mjs: real TerminalManager + attachTerminalOutput, @xterm/headless as the card's parser, Node on an Apple Silicon Mac, 5 hidden cards, 10 s, median of 3; no Electron IPC, DOM or WebGL):

    per card ring-only replay (old) whole stream (now) streamed visibly
    256 KB/s 0.09 s CPU 0.34 s 0.43 s
    1 MB/s 0.14 s 1.71 s 2.05 s

    So a flooding hidden terminal now costs most of its parsing again (the price of exact state); quiet hidden terminals cost nothing. The CHANGELOG and description say so.

  • Windows: the three source assertions added here expected \n right after a statement; the runner checks out with CRLF. They accept \r?\n (checked against CRLF copies). Inherited platform fixtures are fixed in fix: close the gaps found by the code audit (safety, sessions, plugins, IPC, redaction) #108–fix: two bug sweeps across orchestration, isolation, sessions and helpers, plus a git audit after isolated sessions #111.

Verified: fork CI on tip 86e4c33, all three jobs green — https://github.com/BIackFIame/CanvasTTY/actions/runs/36781761157.
Not verified: the description's performance numbers are macOS, stub-agent measurements from before the rebase; they were not re-measured on Linux/Windows or with real provider sessions, and the hidden-flood numbers above are a Node-level measurement, not the Electron renderer.

…owdeploy#109)

Merge OPENCODE_PERMISSION after file and inline rules and verify effective build-agent decisions.

Note: pre-existing failure in the local focused-wheel browser smoke is not addressed by this change.
…#111)

Flush final replies through the Windows pipe buffer, cancel draining on the original deadline or shutdown, and exercise delayed and stalled readers.

Note: pre-existing failure in the local focused-wheel browser smoke is not addressed by this change.
@BIackFIame

Copy link
Copy Markdown
Contributor Author

Rebase onto current main and rerun CI

Carried the remaining OpenCode fix: Auto now reads OPENCODE_PERMISSION after file and inline permissions, preserves the original environment value and the person's build-agent ordering, and tests effective decisions with the environment merge included. Wildcard, tool and command restrictions are covered; malformed JSON follows upstream's skip behavior, while a valid unsupported shape conservatively disables Auto additions.

Also carried the Windows close-race fix: the host drains Windows' pipe buffer before disconnecting, keeps the original two-second bound, and cancels/joins flushing safely on timeout or shutdown. The native self-test exercises delayed readers, duplicate destroy/late writes, a non-reading client, and shutdown during blocked I/O. This addresses the intermittent final-message loss seen on the previous #111 head.

All three upstream checks on 2c1513ee passed: verify, macos-cli-resolution and windows-pipe-host — https://github.com/howdeploy/CanvasTTY/actions/runs/36787425926.

Local validation on this exact cumulative tip: 1581 passed, 3 skipped, no failures; 47 Even G2 tests; typecheck, production build, source/bundle secret audit and MCP-helper smoke passed. Local browser smoke fails at the pre-existing focused-wheel check; Linux CI browser smoke passed. Real OpenCode and isolated provider sessions were not exercised.

Stop Auto and Accept edits before launch for unsupported parsed permission shapes, which OpenCode merges without schema normalization. Preserve malformed JSON skip behavior and the original environment value.
Keep literal __proto__ patterns and preserve numeric pattern decisions without a later generated wildcard. Refuse asking-shell profiles when numeric bash rules cannot be combined safely.
Require real window and page readiness before physical wheel input, verify the actual scroll baseline, and preserve the primary error if renderer cleanup hangs. Exercise the same smoke in macOS CI.
@BIackFIame

Copy link
Copy Markdown
Contributor Author

Keep the terminal parser current while suspending expensive painting.

The lossless bounded delivery and the actual xterm-headless alternate-screen reproduction remain covered. This follow-up now suspends painting at xterm's observed .xterm-screen, rather than relying on visibility:hidden, which leaves its geometry intersecting. Hidden selection updates are deferred and restored on resume; first-hidden measurement and hidden resize are covered too. History, cells, modes, attributes and cursor state are preserved.

A new real-Electron proof runs in Linux/macOS CI. Its Linux/Xvfb fixture is mapped with showInactive and remains unfocused, so window-level pausing cannot make the control and candidate both look idle. Bounded real animation-frame, IntersectionObserver and baseline onRender barriers replace the initial timing assumption; the control must render and the suspended screen must not. It exercises normal/alternate buffers, split escape sequences, selection, hidden resize and one full refresh on resume. Browser input smoke also waits for actual document readiness/focus and has bounded cleanup; it still requires physical wheel input to scroll.

Fresh renderer benchmark on an Apple Silicon Mac: five cards, 10 seconds of input, 1024 Ki UTF-16 units/s per card, three interleaved pairs, installed xterm 6 with the product selection guard. Renderer CPU total for all five cards: old hidden painting 1.868 s, suspended painting 1.749 s (observed median improvement 6.4%); hidden render callbacks 225 → 0. Delivered characters and terminal-state hashes match in every pair. Reproduce with npm run bench:terminal-hidden-renderer -- 5 10 1024 3.

This fixture excludes the app's main process, PTY/IPC and WebGL. The older hidden-flood table measures Node/headless parsing; its CPU values are totals for all cards and cannot measure this DOM change. The parser still consumes the full stream, so the old ring-only CPU saving cannot be retained by this change. No claim of returning to that incorrect replay cost, or of a real-provider/multi-platform speed-up, is made.

The standalone launcher also rejects the known macOS automation CODEX_SANDBOX=seatbelt context before temp setup or Electron spawn. This prevents the reproduced test-launch startup aborts without changing production security settings or automatically escalating; the refusal path, normal GUI launch and portable unit cases were checked.

The cumulative OpenCode permission, Git audit, Windows transport, gateway recovery, lifecycle cleanup and timestamp fixes are included.

Updated head: fbbb023e. All three upstream checks (verify, windows-pipe-host, macos-cli-resolution) passed: https://github.com/howdeploy/CanvasTTY/actions/runs/36833266119.

Independent GPT-6 Luna reviews found no remaining concrete defect in the reported cases. Final cumulative local validation: 1608 passed, 3 skipped, 0 failed; production build/typecheck, secret audit and real-Electron hidden-terminal proof passed. Lower branches were additionally checked with their own gateway/environment/history tests and typecheck.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants