Skip to content

Report live view connect failures by stage, without retrying - #417

Open
robertjamesprior wants to merge 3 commits into
mainfrom
hypeship/live-view-connect-failure-reporting
Open

robertjamesprior wants to merge 3 commits into
mainfrom
hypeship/live-view-connect-failure-reporting

Conversation

@robertjamesprior

@robertjamesprior robertjamesprior commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Supersedes #403 and #415 — that stack collapsed into one change, since the retry the base PR added is gone — and folds in #427, which was stacked on this branch. Calibration, reproduction and the hotpatch A/B are in the Linear doc.

Summary

The live view failed silently. onMessage is assigned straight to ws.onmessage, so a throw from createPeer or setRemoteOffer was discarded: no onDisconnected, no parent-frame message, socket still open and healthy. The client could not tell "peer construction failed" from "still connecting."

A single 15s watchdog also covered all three connect stages, so a socket that never opened burned the whole budget and reported only "connection timeout", and a stuck media path was indistinguishable from a stuck socket.

  • Route the throw. onMessage wraps the handler, so a rejection reaches onDisconnected and is reported instead of dropped.

  • Bound each stage. A connect walks transport (socket open) → signaling (signal/provide + offer) → media (ICE starts). Each arms its own bound and reports its own reason.

    stage starts bound
    transport in connect(), before the socket opens 15s
    signaling on ws.onopen 3s
    media after setRemoteOffer, once a peer exists 2s
  • Keep the media bound armed through connected. ICE reaching checking no longer clears the media bound. checking only means the peer started connecting, while onConnected — and KERNEL_CONNECTED — wait for connected, so clearing at checking left a connect that stalled in between with no timer watching and nothing reported. Folds in Keep the live view connect bound armed until ICE connects #427.

  • Report one terminal event with a machine-readable reason. A bound expiring posts KERNEL_CONNECTION_TIMEOUT; a connect that fails outright posts KERNEL_CONNECTION_FAILED. Both carry reason and the ICE, signaling and socket state at the moment it happened.

    reason meaning
    transport the socket never opened, or closed before signaling
    signaling socket open, no offer inside the bound
    media peer exists, ICE never reached connected — covers a peer that stalls after checking
    peer peer construction or setRemoteOffer threw
    unsupported RTCPeerConnection is missing from the browser
    server the server closed the connect
  • No retries. A second attempt against the same peer costs the viewer time the embedder can spend on a new session, and it delays the reason. The client reports the first failure; the embedder decides.

reason replaces the previous prose ("connection timeout", raw error messages) on both events, and attempts is gone with the retry — both are payload changes.

Consequence for embedders

A failed connect now ends within one stage bound instead of a silent 15s wait, with the reason on the parent frame. The neko overlay does not render a terminal state when a connect never succeeded, so an embedder that shows nothing still shows nothing; the parent message is the signal, and it is what lets an embedder choose between remounting and starting a new session.

One result that changes other work: in an injected media-stall (offer stripped, ICE never leaves new), KERNEL_PLAYING still fired while the peer had never connected — the <video> element emits playing without frames flowing. It did not fire in the reproduction (a RTCPeerConnection construction throw). A parent-side gate on playback needs that caveat.

Testing

39 pass, 0 fail in images/chromium-headful/client (bun test tests), including new cases in tests/connect-bound.test.ts covering stage entry, each stage's terminal event, the single-report guarantee, the payload state, the post-connect disconnect path, and the media bound staying armed through connected. tsc --noEmit clean.

The hotpatch A/B numbers in the doc were taken against the stacked build that still retried. Removing the retry changes only the timeout rows, which now post a single event rather than TIMEOUT followed by FAILED. They were not re-run on this revision.

What this does not establish

The affected sessions are a ws close at ~14.9s, which proves the connect watchdog fired and therefore that ICE never reached checking. It does not say why: no peer constructed, a peer that never left new, and signaling that never completed all produce that shape, and no browser-side capture exists from those sessions. The reproduction proves a construction throw produces the shape, not that it is what happened in the field.

The change does not depend on resolving that. reason is what separates the three at the next occurrence, which is the point of the field.

Dependency

The dashboard's live view panel recognises only KERNEL_CONNECTION_TIMEOUT:

if (event.data?.type !== 'KERNEL_CONNECTION_TIMEOUT') return;

so the failures that now report as KERNEL_CONNECTION_FAILED stop showing its failure panel, its Try again button and its GPU failure inspection. kernel/kernel#4421 accepts both and had to land first — accepting both is compatible with either image, so the dashboard goes first and nothing has to be released in lockstep. #4421 merged 2026-10-07, so this dependency is now satisfied.

Not in this PR

Terminal UI with one-click resume, one shared recovery heuristic with the dashboard viewer, a frames-based gate, and awaiting addIceCandidate.


Note

Medium Risk
Changes WebRTC connect lifecycle, parent postMessage payload shape, and failure semantics that embedders depend on; behavior is well-tested but dashboard/other consumers may need updates for KERNEL_CONNECTION_FAILED.

Overview
Replaces the live view client’s single 15s connect watchdog with per-stage bounds (transport 15s → signaling 3s → media 2s) and reports one terminal parent-frame message with a machine-readable reason instead of generic timeout prose.

BaseClient now arms/clears stage timers through connect (socket open → offer handling → ICE connected), wraps onMessage so throws from peer setup reach onDisconnected, and uses giveUp / reportFailure to post KERNEL_CONNECTION_TIMEOUT (stage expired) or KERNEL_CONNECTION_FAILED (unsupported, transport, peer, server, etc.) with ICE/signaling/socket snapshot—no in-client retries after failure (_gaveUp). ICE checking no longer cancels the media timer; only connected does. Server kicks set _failure = 'server' in the neko client subclass.

Adds connect-bound.test.ts with broad coverage of stage transitions, single-report guarantees, and post-connect vs failed-connect paths. Embedders must handle both message types and the new reason values (replacing "connection timeout" / attempts).

Reviewed by Cursor Bugbot for commit 3c72c24. Bugbot is set up for automated code reviews on this repo. Configure here.

A live view that could not start reported nothing to the parent frame until a
15s watchdog fired, and a throw from peer construction was discarded outright,
so an embedder could not tell a failed connect from a slow one.

Bound the connect stages separately -- transport 15s, signaling 3s, media 2s --
and report the failure to the parent frame with a reason an embedder can branch
on: KERNEL_CONNECTION_TIMEOUT when a bound expires, KERNEL_CONNECTION_FAILED when
the connect fails outright. One event per failure, both carrying the ICE,
signaling and socket state at the moment it happened.

The client does not retry. A second attempt against the same peer costs the
viewer time an embedder can spend on a new session, and the reason is what lets
it make that call.
robertjamesprior added a commit to kernel/docs that referenced this pull request Sep 25, 2026
The client reports a failed connect to the parent frame as a single terminal
event, and both terminal events now carry a machine-readable reason instead of
prose.

- add KERNEL_CONNECTION_FAILED to the parent-frame events table
- document the shared reason set -- transport, signaling, media, peer,
  unsupported, server -- so an embedder can branch on it without matching
  message text
- note that a failed connect posts one of the two terminal events, never both,
  and that the client does not retry a connect itself
- treat KERNEL_CONNECTION_TIMEOUT as terminal alongside it, and record the
  older reason string on images that predate the fix

Version-gated on the browser image carrying kernel/kernel-images#417.

@Sayan- Sayan- left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • p1: The client and dashboard use different failure events. A peer-construction error sends only KERNEL_CONNECTION_FAILED, but the dashboard listens only for KERNEL_CONNECTION_TIMEOUT. Its failure panel and “Try again” button never appear; the viewer stays on an empty connect overlay.
  • p2: The field diagnosis rests on 15-second WebSocket closures, which show the watchdog firing but not why. No browser error was captured to confirm that peer construction threw in the affected sessions. The injected failure verifies the code defect, not the claimed field trigger.

@robertjamesprior robertjamesprior left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

p1 - accepted - this is a regression, takes priority. checked against browser-session-client.tsx, fixed in kernel/kernel#4421, which accepts both events and is compatible with either image (goes first)

p2 - accepted - no capture from those sessions, softened the claim in the issue and the PR body, and added a "what this does not establish" section listing the three causes the signature can't separate. the fix doesn't depend on resolving it

@robertjamesprior
robertjamesprior marked this pull request as ready for review September 25, 2026 23:45

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread images/chromium-headful/client/src/neko/base.ts Outdated
Comment thread images/chromium-headful/client/src/neko/base.ts

@Sayan- Sayan- left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • p2: A later signal/offer rearms the two-second media timeout even after the peer has connected. If ICE stays connected without another state-change event, the timer reports a timeout and disconnects. The server can send an offer on renegotiation, but I could not establish that routine operations trigger one. I reproduced the client path with a simulated peer, not a live browser.

A later signal/offer armed the two-second media bound even with ICE already
connected, starting a timer no state change could clear that then reported a
timeout and tore down a live session. Arm it only while the first connect is
still in flight, and reset the per-attempt flag in connect() so a prior
session's success does not suppress the next connect's bound.

disconnect() also left the socket's onopen handler attached, so a socket that
opened as the close ran could re-arm the signaling bound and post a second
failure for a connect that had already given up.
@robertjamesprior

Copy link
Copy Markdown
Contributor Author

p2 — fixed, and it was a real one. The media bound is now armed only while the first connect is still in flight: a later signal/offer on a connected (or checking) peer no longer starts the two-second timer, so a renegotiation can't report a timeout and tear down a live session. connect() also resets the per-attempt connected flag, so a second login on the same client still gets its media bound.

While I was in there I took the two Bugbot findings on the same path, since they're the same class of stale-bound problem:

  • disconnect() now clears the socket's onopen handler too. It was left attached, so a socket that opened as close() ran could re-arm the signaling bound and post a second failure for a connect that had already given up.
  • A renegotiation signal/offer while the connect is still in flight stays bounded — the guard only skips the arm once ICE has reached checking or the peer is connected.

Three tests added in connect-bound.test.ts; reverting the source change fails two of them. bun test tests 38 pass / 0 fail, tsc --noEmit clean.

The p1 (dashboard recognising both events) is kernel/kernel#4421, unchanged and still the one that goes first.

@robertjamesprior

Copy link
Copy Markdown
Contributor Author

Stacked a fix for the one stretch the bounds missed: #427.

oniceconnectionstatechange cleared the bound on checking, but onConnected and KERNEL_CONNECTED wait for connected. A peer that reached checking and then stalled had no timer watching and reported nothing — the silent frame this PR is about, in the only part of the connect still unbounded. Leaving it armed costs nothing healthy: connected follows checking by tens of milliseconds, inside the same 2s budget. reason stays media.

Not included: a stall that begins after playback. That needs a frame-progress signal, and the peer path has no access to the media element, so it lands in the video layer as a separate change.

@robertjamesprior

Copy link
Copy Markdown
Contributor Author

addressed your points. head is 9c222e9e, and all of it landed after 3397bfa, so the diff moved since you last looked.

p1 was real and it's a regression. the client sends KERNEL_CONNECTION_FAILED when peer construction throws, but the dashboard only listened for KERNEL_CONNECTION_TIMEOUT, so the failure panel and the retry button never appeared. that's kernel/kernel#4421 - it accepts both, works against either image, and goes first.

your p2 on the field claim i couldn't argue with. there's no captured browser error, only the 15s closures, so i've softened it in the issue and the PR body and added a section on what the signature can't separate. the fix doesn't depend on settling it.

your second p2 was the real one. a later signal/offer was rearming the two-second media bound even on a connected peer, so a renegotiation could report a timeout and tear down a live session. the bound now arms only while the first connect is in flight, and connect() resets the per-attempt flag. i also took Bugbot's two findings on the same path; disconnect() was leaving onopen attached, the same class of stale bound.

i couldn't figure out whether routine operations actually emit a renegotiation signal/offer. the guard makes it unreachable either way, so i'd rather merge on "no longer reachable" than block on proving it - say the word and i'll instrument it.

what i'd like your eyes on is the added checking->connected arm (#427). the bound used to clear on checking, but onConnected and KERNEL_CONNECTED wait for connected, so a peer that reached checking and stalled had nothing watching. it stays armed now, which is normally tens of ms inside the same budget. same timer, longer arm - that's the change i'm worried could put a false timeout back on a healthy connect.

38 tests pass, tsc clean.

@robertjamesprior

Copy link
Copy Markdown
Contributor Author

#4421 merged, so the dashboard side is in main: the failure panel and retry now listen for both KERNEL_CONNECTION_FAILED and KERNEL_CONNECTION_TIMEOUT. This is the last piece.

robertjamesprior added a commit to kernel/docs that referenced this pull request Oct 8, 2026
The client reports a failed connect to the parent frame as a single terminal
event, and both terminal events now carry a machine-readable reason instead of
prose.

- add KERNEL_CONNECTION_FAILED to the parent-frame events table
- document the shared reason set -- transport, signaling, media, peer,
  unsupported, server -- so an embedder can branch on it without matching
  message text
- note that a failed connect posts one of the two terminal events, never both,
  and that the client does not retry a connect itself
- treat KERNEL_CONNECTION_TIMEOUT as terminal alongside it, and record the
  older reason string on images that predate the fix

Version-gated on the browser image carrying kernel/kernel-images#417.
robertjamesprior added a commit to kernel/docs that referenced this pull request Oct 8, 2026
* Document the KERNEL_CONNECTION_FAILED live view event

The live view client posts KERNEL_CONNECTION_FAILED to the parent frame once it
has given up starting the viewer. Add it to the parent-frame events table,
version-gate it, handle it in the detection sample, and document that the client
retries transient failures internally so an embedder does not remount on top of
a retry the client is already running.

* Document the KERNEL_CONNECTION_FAILED live view event

The client reports a failed connect to the parent frame as a single terminal
event, and both terminal events now carry a machine-readable reason instead of
prose.

- add KERNEL_CONNECTION_FAILED to the parent-frame events table
- document the shared reason set -- transport, signaling, media, peer,
  unsupported, server -- so an embedder can branch on it without matching
  message text
- note that a failed connect posts one of the two terminal events, never both,
  and that the client does not retry a connect itself
- treat KERNEL_CONNECTION_TIMEOUT as terminal alongside it, and record the
  older reason string on images that predate the fix

Version-gated on the browser image carrying kernel/kernel-images#417.
`oniceconnectionstatechange` cleared the connect bound on `checking`, but
`onConnected` — and `KERNEL_CONNECTED` — wait for `connected`. A peer that
reached `checking` and then stalled had no timer watching and reported
nothing to the parent frame, which is the silent black frame this branch
exists to remove.

Leaving the bound armed costs nothing on the healthy path: ICE reaches
`connected` tens of milliseconds after `checking`, inside the same 2s
budget, and `onConnected` clears it either way.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 3c72c24. Configure here.

// `checking` deliberately leaves the media bound armed. It only means the
// peer started connecting, and onConnected — and KERNEL_CONNECTED — wait
// for `connected`, so clearing here left a connect that stalled in
// between with no timer watching and nothing reported.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Media bound ignores ICE completed

High Severity

Leaving the media bound armed through checking without treating completed as success means a peer that jumps checking to completed never clears the bound. ICE can skip connected when gathering is already done, so a live session is torn down two seconds later as a media timeout.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 3c72c24. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants