Skip to content

fix(runtime): prepare only the Session-selected Harness - #431

Open
sam2tom wants to merge 2 commits into
MiniMax-AI:mainfrom
sandbaseai:codex/selected-runtime-harness-upstream
Open

sam2tom wants to merge 2 commits into
MiniMax-AI:mainfrom
sandbaseai:codex/selected-runtime-harness-upstream

Conversation

@sam2tom

@sam2tom sam2tom commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

A Core-managed Codex Session previously waited for Codex, MiniMax Code and Claude SDK discovery before its Runtime connected, even though all three were packaged only to share one template. Carry the owning Session’s immutable Harness through Runtime bootstrap and discover/register only that selection. Unknown or unavailable selections fail without probing another implementation; self-hosted installations retain their installed Harness set.

The bootstrap codec is version 2 with required harness. E2B uses helper version 2, microsandbox helper version 3 and node wire version 5 with exact, unique nested Bootstrap members. E2B’s managed launch file delivers the selection only inside RuntimeBootstrap. Core–Runtime wire, SQL/schema and native dependency pins are unchanged. Publish matching Core/helpers/new Runtime template together; node installations need matching nodes. Existing allocations retain their bootstrap and Runtime.

Codex still validates --version; safe process-spawn/wait timings distinguish its cold probe. Existing production diagnostic evidence showed serial discovery at 16.537 seconds and warm Codex version commands at 49.6/11.0 milliseconds. These are baseline observations, not acceptance or a latency claim for this change.

Validation: selected/unknown/unavailable discovery and strict startup/node codec tests (including race repeats), Provider regressions, Core–Runtime contract checks, Core server package race regressions, 191 E2B Python tests against the pinned SDK, isolated PostgreSQL Session-selection/replay test, vet, docs/name/translation checks and independent review. A fresh-template cloud execution test has not been performed.

This branch is based on upstream main 9fa92df0afb170bd57da8314e2a4280aca1fd4ac and contains no fork deployment assets.

CI passed the aggregate gate, including Core/Runtime/store, official-client, compose/distribution, documentation and Linux/macOS/Windows platform checks.

@sam2tom

sam2tom commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

End-to-end latency breakdown and remediation

This separates three API milestones that should not all be called “Session creation”. The figures below are an anonymized, controlled single cold/warm pair with short successful replies, not a percentile benchmark or a guarantee. Both Turns completed and their replies and usage were persisted. No raw logs, request identifiers, infrastructure configuration, or credentials are included.

Client milestone Cold Reused Session
Session creation response, measured separately 0.534 s —
First input submission returning HTTP 202 9.291 s 1.104 s
Input submission to first visible text 14.162 s 4.900 s
Input submission to completion 15.055 s 5.799 s

Where the time goes

Core measured input admission at 8.838 s cold / 0.490 s warm. Admission to first text was another 4.608 s / 3.748 s. Client first-text time also includes time outside those two Core measurements; the residual must not be labelled model latency or network latency without further evidence.

In this cold sample, provisioning had already started before input reservation, so there was no next-maintenance-tick delay at that boundary. The remaining Provider Create after reservation was approximately 4.147 s, followed by 2.243 s from Create completion to Runtime transport registration and 0.188 s from registration to input selection. Fresh executor readiness took 2.197 s, versus 0.184 s for reuse. Small configuration/admission persistence intervals complete the admission path. These are related measurements with different boundaries, not independently additive top-level spans.

Executor start acknowledgement took 0.447 s cold / 0.195 s warm and is contained in admission-to-first-text time. That remaining interval includes native execution and model response; the current evidence does not isolate it into pure model inference, provider queueing, or network time.

Provider Create is not just sandbox boot

Create component Duration
SDK Sandbox.create 1.687 s
Exact-ID ownership/configuration check 0.221 s
Managed entry-point validation 1.139 s
Bootstrap payload write 0.250 s
Managed initialization command 0.726 s
Final cloud/ready-receipt inspection 0.453 s
Unassigned helper/receipt/control overhead 0.535 s
Total Core Provider Create 5.011 s

The sandbox gateway's corresponding create request took 1.586 s: its upstream creation call took 0.872 s and the following detail lookup 0.230 s. The remaining 0.484 s has no finer confirmed attribution. These are nested inside the SDK/Create measurements, not additional time to add to the table.

The exact Runtime's selected Codex discovery took 1.220 s (process spawn 0.022 s, wait 1.196 s). Its bootstrap took 0.814 s, transport dial 0.751 s, and native executor preparation 2.023 s. Runtime startup overlaps Provider Create, and native preparation is contained in executor readiness: summing all of these with the table would double-count time. The earlier 4–5 s probe observation was a different sample and measured only discovery, not end-to-end cold start.

Remediation and current status

  1. Discover only the selected Harness. This PR addresses unnecessary discovery of other packaged Harnesses. Sharing one image does not require probing every Harness. The controlled downstream sample verified that only Codex was probed; the upstream PR itself is still open. The version probe remains, with timing rather than silently treating an unavailable executable as ready.
  2. Start fresh provisioning after durable Session/input commits. A downstream follow-up, fix: accelerate fresh provisioning and trace E2B create stages sandbaseai/OpenAgentCore#16, adds hints to the existing lifecycle owner and input recovery for missed hints. It has been merged and exercised in the sample above. Ownership, placement, capacity, lease checks, and periodic recovery remain authoritative. Hints are bounded; this is not an immediate-start SLA for every concurrent Session.
  3. Remove expensive remote probe execution. A local, uncommitted follow-up replaces the Python process used only to check the entry point with SDK file-type validation and a streamed open/close. Both checks remain before credential delivery. This uses two file requests, so its benefit must be measured; it cannot be assumed to save the entire previous 1.139 s.
  4. Return the durable bootstrap receipt with the initialization command. The same local follow-up returns the bounded existing receipt on the command stream, eliminating its separate file GET on successful Create. Identity validation and the final cloud ownership/configuration check remain. Unknown responses recover through read-only inspection, never by rerunning Create or initialization. A separate read measured approximately 0.226 s in the baseline, but this is not a measured post-change saving.
  5. Measure the remaining boundaries before changing them. The transport-registration tail, fresh native preparation, gateway's unassigned remainder, and admission-to-first-text interval remain visible costs. Parallelize only independent work; do not remove ownership/readiness checks, assume all registration time is a retry delay, or hide it with a longer timeout. A sustained ready pool would be a separate lifecycle/capacity/cost design, not a claimed result of this patch.

The local helper follow-up has passed generated-contract checks, 11 template tests and 261 helper tests, including SDK transport fixtures and real local subprocess tests for receipt return, failed initialization, and missing/oversized receipts. It has not been published, deployed, or benchmarked against a fresh cloud sandbox. It changes the Core-shipped helper and uses the existing protected template entry point; it requires a Core rebuild, with no new Runtime template or SQL migration.

Acceptance for the next release

Pin the Core/helper revision and Runtime build; repeat controlled cold and reused-Session measurements with the same Harness/model/input and report a distribution rather than selecting one favourable sample. Compare the six Create stages, admission, registration, readiness, first text, and completion. Verify persisted replies/usage, failure recovery, and no duplicate allocation/startup. Keep overlapping stages separate. A healthy deployment or a green fixture suite alone is not evidence of improved cold-start latency.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant