fix(omp,remote): thread remote-omp resume/continue through respawn and reattach - #362
fix(omp,remote): thread remote-omp resume/continue through respawn and reattach#362timkjr wants to merge 3 commits into
Conversation
|
Thanks for this, and sorry it has sat. The branch is conflicting, which means GitHub has run no CI on it at all, and the conflicts are not mechanical: master has moved under two of the three things this PR does. The reconnect-watcher half is already on master, in a form that has since been fixed. #355 (your commit, same title) merged on 2026-09-04. This version's The remote omp branch calls The host-local omp resolver now feeds Smaller: The claude session pin and the trailing-slash normalisation are good and I want them. If you would rather land those quickly, |
… resumes instead of relaunching fresh Two independent defects made ANY clean exit from a remote SSH session (user ctrl-d or ctrl-c, or a dropped pane) relaunch the agent as a NEW conversation: 1. SSH-remote claude was launched as a bare `claude --dangerously-skip-permissions`, so the remote-respawn path (COD-108 reattachRemote re-running the idempotent launch command) started a fresh conversation every time. Pin it to the deterministic Codeman session id, mirroring the docker-claude shape (claudeDockerPaneCommand): `--session-id <id>` to create, with the `|| --resume <id>` fallback so the idempotent re-run resumes instead of erroring with "already in use". A per-host commands.claude override still wins. 2. OMP --resume pinning silently degraded to ambiguous `--continue` whenever a case path ended in a trailing slash (e.g. remote `remotePath` stored verbatim as `/home/user/dotfiles/`): mangleOmpWorkingDir produced `-dotfiles-` while omp persists sessions under `-dotfiles`, readdirSync returned null for an existing dir, and findLatestOmpSessionId/resolveAndClaimOmpSessionId never matched. Normalize the trailing slash before mangling (new exported stripTrailingSlash) and compare the session header cwd against the same normalized value. Both were found live 2026-08-29 on a remote OMP/Claude node: ctrl-c and ctrl-d behaved identically, both relaunching a fresh session.
The COD-108 reconnect watcher treated any dead local pane as a dropped transport and re-ran the pane command — so a normal ctrl-c/ctrl-d on a remote omp/opencode/claude auto-spawned a FRESH agent (claude only looked correct because its '--session-id || --resume' fallback resumed, with a loud 'already in use' error first). Distinguish a transport drop from an intentional exit: only reconnect when the durable remote tmux session (codeman-ssh-*) is verifiably still alive on the remote host. A clean exit tears that session down; the watcher now probes it via ssh has-session and skips (remote-gone) when it is gone OR unknown (fail closed). The probe is cached per-session and fired async so the 5s tick never blocks on ssh. Also thread ompConfig/resumeSessionId into the remote builders so a dead-pane respawn of an omp session resumes (--resume <id>) or continues (--continue) instead of launching bare omp. Tests: 3 new cases pinning remote-gone / unknown / alive decisions; remote omp resume + --continue fallback. Verified live: all three remote CLIs stay dead after exit.
- Remote omp command now renders through buildSpawnCommandFromRegistry (the mode-agnostic engine local/docker spawns use) instead of the buildOmpCommand() the CLI-registry refactor deleted. - Session._pinOmpRespawnId()/_maybeCaptureOmpSessionId() now skip host-local ~/.omp resolution entirely for a remote session and fall back to --continue: that resolver only ever reads THIS host's filesystem, which is meaningless (and could wrongly alias an unrelated local conversation) for a conversation that lives on the remote host. - Remote-claude launch now honors an explicit resumeSessionId distinct from sessionId (mirrors claudeDockerPaneCommand's shape), and validates sessionId the same way that sibling does before interpolating it into the remote shell command. - Add the still-missing header-cwd half of the trailing-slash test, and document respawn/reattach continuation + auto-reconnect-vs- clean-exit in docs/remote-sessions.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
5041bc6 to
797f0d3
Compare
|
Thanks for the detailed review — rebased and addressed all four points. Reconnect-watcher half: dropped entirely, as you said. Confirmed the merge resolved every conflict in that region to master's side ( Remote omp branch: rendered through the registry now. I used Host-local omp resolver on a remote respawn: you're right, and your dotfiles example is exactly the failure case I reproduced. Smaller items: On the cherry-pick-vs-one-PR question: I looked at splitting it back out, but by the time everything's rebased and fixed, the omp threading and the claude pinning share too much of the same code path (both go through CI is green and it's rebased clean on current master. |
Summary
Follow-up to #353 (OMP backend) — this is the remote-continuation threading work that was deliberately split out at the time because it can't apply without
OmpConfigexisting upstream.Two fixes, found live 2026-08-29 on a remote OMP/Claude node, both stemming from the same root cause: a dead/dropped remote pane relaunched the agent as a brand-new conversation instead of resuming.
fix(omp,remote): pin remote conversations on respawnclaudewas launched as bareclaude --dangerously-skip-permissions, so any remote-respawn (COD-108reattachRemotere-running the idempotent launch command) started a fresh conversation every time. Pinned it to the deterministic Codeman session id (--session-id <id>to create,|| --resume <id>fallback so the idempotent re-run resumes instead of erroring "already in use") — mirrors the existing docker-claude shape.--resumepinning silently degraded to ambiguous--continuewhenever a case path ended in a trailing slash (a remote case'sremotePathis stored verbatim, e.g./home/user/dotfiles/).mangleOmpWorkingDirproduced-dotfiles-while omp persists sessions under-dotfiles, so lookup never matched. Normalized the trailing slash before mangling.fix(remote): never auto-revive a remote session after a clean agent exit--session-id || --resumefallback happened to resume it, with a loud "already in use" error first).codeman-ssh-*) is verifiably still alive on the remote host, probed viassh has-session, fails closed (skips) when gone or unknown. The probe is cached per-session and fired async so the 5s watcher tick never blocks on ssh.ompConfig/resumeSessionIdinto the remote command builders so a dead-pane respawn of an omp session resumes (--resume <id>) or continues (--continue) instead of launching bareomp.Test plan
npm run typecheck— cleannode_modules/.bin/prettierin a fresh worktree — resolved bynpm install, confirmed passing after)test/remote-shared-sessions.test.ts(most directly touched) — all passing, including the omp--resume/--continuecommand-building casestest/remote-auto-reconnect.test.ts— new cases for remote-gone / remote-unknown (fail closed) / remote-alive (legitimate reconnect still works)test/omp-session-resolver.test.ts,test/tmux-manager.test.ts— trailing-slash normalization and SSH-remote claude session-id pinning🤖 Generated with Claude Code