From c56ff745e5fb4e974071c2b85c1e6b623ff2ef29 Mon Sep 17 00:00:00 2001 From: prk-Jr Date: Thu, 17 Sep 2026 11:32:51 +0530 Subject: [PATCH 01/18] Add opt-in reusable sandbox lifecycle for the Fastly adapter Fastly SDK 0.12.1 exposes `fastly::http::serve::Serve`, which lets one Wasm sandbox serve several requests. Adopt it behind a `reusable-sandbox` feature that is off by default, so the shipped entry point is unchanged. Split `main` into a loop owner and `handle_request`. The handler keeps sending its own response and returns `()`, which the SDK's `HandlerResult` impl treats as already sent, so progressive streaming, duplicate `Set-Cookie` headers, response extensions, and post-send pull sync are untouched. No application state is retained yet; that follows in a later change. Three execution modes, only the first identical to today: feature off -> original entry path feature on, effective max_requests <= 1 -> single-request handler path feature on, validated reuse limits -> Serve loop Bounds come from the `edgezero_runtime_env` config store. They cannot be read through `runtime_env_config`, whose `runtime_env_keys` allowlist drops every key outside the adapter, logging, and per-store selectors; routing them there would read as absent and silently disable reuse while passing every test. The adapter opens the store itself and builds the service-scoped key from `service_id()`, since edgezero's own key helper is private. Reads use `try_get`, never `get`: `get` panics on a lookup error, and this runs before the health probe. Any failure resolves to single-request operation. Reuse needs all three bounds. A bare request limit is refused because the SDK reads an omitted lifetime and wait timeout as `Duration::MAX`, and an application-level `0` normalizes to `1` because the SDK reads `with_max_requests(0)` as unlimited. Guard the global logger. `fern`'s `apply()` panics when a logger is already installed, which a reused sandbox would hit on its second request. Startup diagnostics are deferred and flushed once the logger exists, since limit resolution runs before there is anywhere to log. Add measurement so the lifecycle can be evaluated rather than assumed. A `sandbox_metrics_enabled` debug flag attaches instance id, request ordinal, build count, and correlation id to workload responses before headers commit; post-commitment stream failures carry the same context in logs. The flag is skipped while false because `DebugConfig` denies unknown fields and a default blob must stay readable by an older binary. The counters endpoint is feature gated and short-circuits ahead of application construction so polling cannot perturb the build count it reports. Response counters stay feature independent so the feature-off baseline is measurable on the same channel. Add `ts dev sandbox-probe`, which issues a keep-alive sequence and reports reuse only from strictly increasing ordinals under one instance id. Repeated or decreasing ordinals, incomplete attribution, a missing instance identity, and transport errors all report unverified rather than a negative result. `test-fastly` does not pass `--all-features` and CI invokes it directly, so add a `test-fastly-reuse` alias and CI step; clippy compiling the feature is not the same as running its tests. --- .cargo/config.toml | 4 + .github/workflows/test.yml | 3 + AGENTS.md | 3 +- .../trusted-server-adapter-fastly/Cargo.toml | 7 + .../trusted-server-adapter-fastly/src/main.rs | 233 ++++++- .../src/sandbox.rs | 618 ++++++++++++++++++ crates/trusted-server-cli/Cargo.toml | 11 + .../src/commands/dev/mod.rs | 22 +- .../src/commands/dev/sandbox_probe.rs | 533 +++++++++++++++ crates/trusted-server-core/src/settings.rs | 76 +++ ...26-09-17-fastly-reusable-sandbox-design.md | 537 +++++++++++++++ 11 files changed, 2023 insertions(+), 24 deletions(-) create mode 100644 crates/trusted-server-adapter-fastly/src/sandbox.rs create mode 100644 crates/trusted-server-cli/src/commands/dev/sandbox_probe.rs create mode 100644 docs/superpowers/specs/2026-09-17-fastly-reusable-sandbox-design.md diff --git a/.cargo/config.toml b/.cargo/config.toml index 1302091e0..896cd65c8 100644 --- a/.cargo/config.toml +++ b/.cargo/config.toml @@ -30,6 +30,10 @@ build-fastly = "build -p trusted-server-core -p trusted-server-adapter-fastly -p check-fastly = "check -p trusted-server-core -p trusted-server-adapter-fastly -p trusted-server-js -p trusted-server-openrtb --target wasm32-wasip1" clippy-fastly = "clippy -p trusted-server-core -p trusted-server-adapter-fastly -p trusted-server-js -p trusted-server-openrtb --all-targets --all-features --target wasm32-wasip1 -- -D warnings" test-fastly = "test -p trusted-server-core -p trusted-server-adapter-fastly -p trusted-server-js -p trusted-server-openrtb --target wasm32-wasip1" +# Feature-on counterpart of `test-fastly`. `test-fastly` does NOT pass +# --all-features, so without this the reusable-sandbox code is linted by +# `clippy-fastly` but its tests never execute. +test-fastly-reuse = "test -p trusted-server-adapter-fastly --target wasm32-wasip1 --features reusable-sandbox" # --- Axum adapter (native dev server) --- build-axum = "build -p trusted-server-adapter-axum" diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index 97402e6f4..6062e52c5 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -59,6 +59,9 @@ jobs: - name: Run tests run: cargo test-fastly + - name: Run tests (reusable-sandbox feature) + run: cargo test-fastly-reuse + - name: Run template cache ESI local harness run: BID_DELAY=3 ./scripts/template-cache-local-test.sh esi diff --git a/AGENTS.md b/AGENTS.md index 3b7189204..a8809bf59 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -98,6 +98,7 @@ spin up --from crates/trusted-server-adapter-spin # See .cargo/config.toml; default-members = [fastly] so Viceroy can locate # the binary via `cargo run --bin`. cargo test-fastly # Fastly adapter + core (wasm32-wasip1 via Viceroy) +cargo test-fastly-reuse # Fastly adapter with the reusable-sandbox feature on cargo test-axum # Axum dev server adapter (native) cargo test-cloudflare # Cloudflare Workers adapter (native host) cargo test-spin # Spin adapter route tests (native host) @@ -340,7 +341,7 @@ Every PR must pass: 1. `cargo fmt --all -- --check` 2. `cargo clippy-fastly && cargo clippy-axum && cargo clippy-cloudflare && cargo clippy-cloudflare-wasm && cargo clippy-spin-native && cargo clippy-spin-wasm` -3. `cargo test-fastly && cargo test-axum && cargo test-cloudflare && cargo test-spin` +3. `cargo test-fastly && cargo test-fastly-reuse && cargo test-axum && cargo test-cloudflare && cargo test-spin` 4. `cargo test --manifest-path crates/trusted-server-integration-tests/Cargo.toml --test parity` 5. JS build and test (`cd crates/trusted-server-js/lib && npx vitest run`) 6. JS format (`cd crates/trusted-server-js/lib && npm run format`) diff --git a/crates/trusted-server-adapter-fastly/Cargo.toml b/crates/trusted-server-adapter-fastly/Cargo.toml index 65320faa6..710523766 100644 --- a/crates/trusted-server-adapter-fastly/Cargo.toml +++ b/crates/trusted-server-adapter-fastly/Cargo.toml @@ -10,6 +10,13 @@ version = { workspace = true } [lints] workspace = true +[features] +# Opt-in reusable sandbox lifecycle. Off by default, and `fastly.toml`'s build +# command does not pass it, so production builds keep the single-request entry +# point. Enabling the feature alone does not enable reuse: the bounds still +# come from the runtime environment (see `sandbox::resolve_mode`). +reusable-sandbox = [] + [dependencies] async-trait = { workspace = true } base64 = { workspace = true } diff --git a/crates/trusted-server-adapter-fastly/src/main.rs b/crates/trusted-server-adapter-fastly/src/main.rs index 367a3e33f..3466dfee2 100644 --- a/crates/trusted-server-adapter-fastly/src/main.rs +++ b/crates/trusted-server-adapter-fastly/src/main.rs @@ -39,6 +39,7 @@ mod management_api; mod middleware; mod platform; mod rate_limiter; +mod sandbox; mod template_cache; mod tinybird; @@ -49,6 +50,7 @@ use crate::ec_kv::FastlyEcKvStore; use crate::middleware::{HEADER_X_TS_FINALIZED, apply_finalize_headers, resolve_geo_for_response}; use crate::platform::{FastlyPlatformGeo, client_info_from_request}; use crate::rate_limiter::{FastlyRateLimiter, RATE_COUNTER_NAME}; +use crate::sandbox::{Sandbox, SandboxCounters, ServeMode}; /// Opens the Fastly Config Store used by the `EdgeZero` dispatcher. /// @@ -74,9 +76,66 @@ fn health_response(req: &FastlyRequest) -> Option { /// /// Uses an undecorated `main()` with `FastlyRequest::from_client()` instead of /// `#[fastly::main]` so the `EdgeZero` streaming publisher path can call -/// [`fastly::Response::stream_to_client`] explicitly. +/// [`fastly::Response::stream_to_client`] explicitly. It owns the sandbox +/// lifecycle and delegates each request to [`handle_request`]. +/// +/// Without the `reusable-sandbox` feature this takes one request and returns, +/// which is the original behaviour. With the feature, the bounds still have to +/// come from the runtime environment before the SDK serving loop is entered; +/// an unconfigured or partially configured sandbox stays single-request. fn main() { - let req = FastlyRequest::from_client(); + let mut sandbox = Sandbox::default(); + + let (mode, diagnostics) = serve_mode(); + for message in diagnostics { + sandbox.defer_diagnostic(message); + } + + match mode { + ServeMode::Single => handle_request(FastlyRequest::from_client(), &mut sandbox), + ServeMode::Reuse(limits) => serve_loop(limits, &mut sandbox), + } +} + +/// Serves requests from one sandbox under explicit bounds. +/// +/// Only compiled with the `reusable-sandbox` feature; [`serve_mode`] can never +/// return [`ServeMode::Reuse`] without it. +#[cfg(feature = "reusable-sandbox")] +fn serve_loop(limits: crate::sandbox::SandboxLimits, sandbox: &mut Sandbox) { + let summary = fastly::http::serve::Serve::new() + .with_max_requests(limits.max_requests) + .with_max_lifetime(limits.max_lifetime) + .with_timeout(limits.timeout) + .run_with_context(handle_request, sandbox); + + // `handle_request` sends its own response and returns `()`, so the SDK has + // no terminal error to report. The summary is still worth a line: it is the + // only place the sandbox's own view of its request count is visible. + log::info!( + "sandbox retiring after {} request(s), {} build(s)", + summary.requests(), + sandbox.builds() + ); +} + +#[cfg(not(feature = "reusable-sandbox"))] +fn serve_loop(_limits: crate::sandbox::SandboxLimits, _sandbox: &mut Sandbox) { + unreachable!("serve_mode never selects reuse without the reusable-sandbox feature") +} + +/// Handles one request end to end, sending its own response. +/// +/// Returns `()` rather than a response so the streaming publisher path can +/// call [`fastly::Response::stream_to_client`] itself. The SDK's +/// `HandlerResult` impl for `()` treats that as already sent. +/// +/// Every non-panicking path through this function must send exactly once. In a +/// reused sandbox the SDK refuses to wait for the next request until the +/// current one is complete, so a missed send stalls the loop rather than +/// merely dropping one response. +fn handle_request(req: FastlyRequest, sandbox: &mut Sandbox) { + let ordinal = sandbox.begin_request(); // Health probe bypasses logging, settings, and app construction as a cheap liveness signal. if let Some(response) = health_response(&req) { @@ -84,15 +143,133 @@ fn main() { return; } - logging::init_logger(); - edgezero_main(req); + sandbox.ensure_logger(); + + // Correlation is request-local and never retained. `FASTLY_TRACE_ID` names + // the sandbox, not the request, so it is not used here. The id rides on the + // response only when metrics are enabled; the client request is left + // untouched so nothing new reaches origin in the default configuration. + let request_id = req + .get_client_request_id() + .map(str::to_owned) + .unwrap_or_else(|| format!("{}-{ordinal}", instance_id())); + + edgezero_main(req, sandbox, ordinal, &request_id); +} + +/// Builds the counters snapshot response. +/// +/// A snapshot only. It reports the sandbox that served *this* probe, which is +/// not necessarily the sandbox that served any preceding workload request, so +/// reuse is established from the counters attached to workload responses +/// rather than from polling this. +#[cfg(any(feature = "reusable-sandbox", test))] +fn sandbox_metrics_response(sandbox: &Sandbox, ordinal: u64) -> FastlyResponse { + let body = serde_json::json!({ + "instance": instance_id(), + "ordinal": ordinal, + "requests": sandbox.requests(), + "builds": sandbox.builds(), + }); + + FastlyResponse::from_status(fastly::http::StatusCode::OK) + .with_header("cache-control", "private, no-store") + .with_body_json(&body) + .unwrap_or_else(|e| { + log::error!("failed to serialize sandbox metrics: {e}"); + FastlyResponse::from_status(fastly::http::StatusCode::INTERNAL_SERVER_ERROR) + }) +} + +/// Attaches the sandbox counters to a workload response. +/// +/// Called before headers are committed, which on the streaming path means +/// before `stream_to_client`. The counters therefore describe the request up +/// to commitment and cannot report its eventual outcome; a failure after +/// commitment is recorded in logs instead and reconciled during analysis. +fn attach_sandbox_counters(response: &mut HttpResponse, counters: &SandboxCounters) { + let headers = response.headers_mut(); + for (name, value) in [ + (sandbox::HEADER_SANDBOX_INSTANCE, instance_id()), + ( + sandbox::HEADER_SANDBOX_ORDINAL, + counters.ordinal.to_string(), + ), + (sandbox::HEADER_SANDBOX_BUILDS, counters.builds.to_string()), + ( + sandbox::HEADER_SANDBOX_REQUEST_ID, + counters.request_id.clone(), + ), + ] { + match edgezero_core::http::HeaderValue::from_str(&value) { + Ok(value) => { + headers.insert(name, value); + } + Err(e) => log::warn!("sandbox counter `{name}` is not a valid header value: {e}"), + } + } +} + +/// Resolves how this sandbox will serve requests. +/// +/// Without the `reusable-sandbox` feature this is unconditionally +/// [`ServeMode::Single`] and reads nothing, so the default build does no +/// startup work the original entry point did not do. +/// Returns the mode alongside diagnostics that must wait for the logger. +/// +/// This runs before any logger exists, so the reasons reuse was declined are +/// carried out rather than logged here, where they would be discarded. +#[cfg(feature = "reusable-sandbox")] +fn serve_mode() -> (ServeMode, Vec) { + let (raw, diagnostics) = crate::sandbox::read_raw_limits(); + (crate::sandbox::resolve_mode(raw), diagnostics) +} + +#[cfg(not(feature = "reusable-sandbox"))] +fn serve_mode() -> (ServeMode, Vec) { + (ServeMode::Single, Vec::new()) +} + +/// Guest-instance identifier used to attribute requests to a sandbox. +/// +/// `FASTLY_TRACE_ID` describes the sandbox, which is exactly what is wanted +/// here and exactly why it must not be used as a request id. An absent value +/// is reported rather than synthesized, so a measurement run cannot silently +/// claim reuse it never observed. +fn instance_id() -> String { + std::env::var("FASTLY_TRACE_ID") + .ok() + .filter(|value| !value.is_empty()) + .unwrap_or_else(|| sandbox::INSTANCE_ID_UNAVAILABLE.to_owned()) } /// Handles a request through the `EdgeZero` router path. -fn edgezero_main(mut req: FastlyRequest) { +fn edgezero_main(mut req: FastlyRequest, sandbox: &mut Sandbox, ordinal: u64, request_id: &str) { let runtime_env = runtime_env_config(TrustedServerApp::stores()); let runtime_stores = RuntimeStoreConfig::from_env(&runtime_env); + // Short-circuit the sandbox counters probe before app construction. It must + // not build the application: polling it would otherwise increment the very + // build counter it reports. + #[cfg(feature = "reusable-sandbox")] + if req.get_method() == FastlyMethod::GET && req.get_path() == sandbox::SANDBOX_METRICS_PATH { + match load_settings_from_config_store(&runtime_stores) { + Ok(settings) if sandbox::metrics_enabled(&settings) => { + sandbox_metrics_response(sandbox, ordinal).send_to_client(); + } + Ok(_) => { + FastlyResponse::from_status(fastly::http::StatusCode::NOT_FOUND).send_to_client(); + } + Err(e) => { + log::warn!("sandbox metrics endpoint: failed to load settings: {e:?}"); + FastlyResponse::from_status(fastly::http::StatusCode::INTERNAL_SERVER_ERROR) + .with_body_text_plain("Internal Server Error") + .send_to_client(); + } + } + return; + } + // Short-circuit the JA4 debug probe before app construction. Must run here // because TLS/JA4 accessors are only available on FastlyRequest before // conversion to edgezero types. @@ -127,7 +304,10 @@ fn edgezero_main(mut req: FastlyRequest) { }; let (app, app_state) = TrustedServerApp::build_app_with_state(&runtime_stores); + sandbox.record_build(); let settings_snapshot = app_state.as_ref().map(|state| Arc::clone(&state.settings)); + let counters = + SandboxCounters::capture(sandbox, ordinal, request_id, settings_snapshot.as_deref()); let trusted_client_ip = settings_snapshot .as_deref() .and_then(|settings| settings.trusted_client_ip.as_ref()); @@ -220,7 +400,11 @@ fn edgezero_main(mut req: FastlyRequest) { if let Some(settings) = settings_snapshot.as_deref() { match apply_edgezero_ec_finalize(settings, &mut ec_state, &mut response) { Ok(partner_registry) => { - send_edgezero_response(response, request_filter_effects.as_ref()); + send_edgezero_response( + response, + request_filter_effects.as_ref(), + counters.as_ref(), + ); run_edgezero_pull_sync_after_send(settings, &partner_registry, &ec_state); return; } @@ -235,7 +419,11 @@ fn edgezero_main(mut req: FastlyRequest) { Ok(settings) => { match apply_edgezero_ec_finalize(&settings, &mut ec_state, &mut response) { Ok(partner_registry) => { - send_edgezero_response(response, request_filter_effects.as_ref()); + send_edgezero_response( + response, + request_filter_effects.as_ref(), + counters.as_ref(), + ); run_edgezero_pull_sync_after_send( &settings, &partner_registry, @@ -257,7 +445,7 @@ fn edgezero_main(mut req: FastlyRequest) { } } - send_edgezero_response(response, request_filter_effects.as_ref()); + send_edgezero_response(response, request_filter_effects.as_ref(), counters.as_ref()); } fn edge_error_response(error: EdgeError) -> HttpResponse { @@ -338,9 +526,26 @@ fn run_edgezero_pull_sync_after_send( fn send_edgezero_response( mut response: HttpResponse, request_filter_effects: Option<&RequestFilterEffects>, + counters: Option<&SandboxCounters>, ) { apply_terminal_response_effects(&mut response, request_filter_effects); + // Captured before the body is consumed so post-commitment failures can be + // matched back to the response that carried these counters. + let counter_context = counters.map_or_else(String::new, |counters| { + format!( + " [instance={} ordinal={} request={}]", + instance_id(), + counters.ordinal, + counters.request_id + ) + }); + + // Before headers commit, including before `stream_to_client` below. + if let Some(counters) = counters.as_ref() { + attach_sandbox_counters(&mut response, counters); + } + let (parts, body) = response.into_parts(); match body { @@ -353,11 +558,19 @@ fn send_edgezero_response( match futures::executor::block_on(stream_asset_body(body, &mut streaming_body)) { Ok(()) => { if let Err(e) = streaming_body.finish() { - log::error!("failed to finish EdgeZero streaming body: {e}"); + // Also post-commitment: same attribution as the + // streaming failure below. + log::error!( + "failed to finish EdgeZero streaming body{counter_context}: {e}" + ); } } Err(e) => { - log::error!("EdgeZero streaming failed: {e:?}"); + // After commitment: log and stop. Returning an error here + // would let the SDK attempt a second response. Counters + // already went out with the headers, so the failure is + // tagged with the same identity for reconciliation. + log::error!("EdgeZero streaming failed{counter_context}: {e:?}"); drop(streaming_body); } } diff --git a/crates/trusted-server-adapter-fastly/src/sandbox.rs b/crates/trusted-server-adapter-fastly/src/sandbox.rs new file mode 100644 index 000000000..08b6a7e2c --- /dev/null +++ b/crates/trusted-server-adapter-fastly/src/sandbox.rs @@ -0,0 +1,618 @@ +//! Reusable-sandbox lifecycle state and limit resolution. +//! +//! Fastly Compute normally starts a fresh Wasm sandbox per request. SDK 0.12.1 +//! exposes [`fastly::http::serve::Serve`], which lets one sandbox serve +//! several. This module owns the opt-in decision, the bounds, and the +//! per-sandbox bookkeeping the entry point carries across requests. +//! +//! Nothing here retains application state. [`Sandbox`] holds a logger guard and +//! measurement counters only; the retained application arrives in a later +//! change. + +use std::time::Duration; + +use trusted_server_core::settings::Settings; + +/// Header carrying the guest-instance identifier. +pub(crate) const HEADER_SANDBOX_INSTANCE: &str = "x-ts-sandbox-instance"; + +/// Header carrying the request ordinal within the current sandbox, 1-based. +pub(crate) const HEADER_SANDBOX_ORDINAL: &str = "x-ts-sandbox-ordinal"; + +/// Header carrying the number of application builds this sandbox has performed. +pub(crate) const HEADER_SANDBOX_BUILDS: &str = "x-ts-sandbox-builds"; + +/// Header carrying the per-request correlation id. +pub(crate) const HEADER_SANDBOX_REQUEST_ID: &str = "x-ts-sandbox-request-id"; + +/// Path of the counters snapshot endpoint. +#[cfg(any(feature = "reusable-sandbox", test))] +pub(crate) const SANDBOX_METRICS_PATH: &str = "/_ts/debug/sandbox"; + +/// Label used when the runtime reports no usable guest-instance identifier. +/// +/// Recorded rather than papered over: a measurement run that sees this value +/// has no instance identity and cannot claim observed reuse. +pub(crate) const INSTANCE_ID_UNAVAILABLE: &str = "unavailable"; + +/// Per-sandbox state carried across requests by the entry point. +/// +/// Everything here is either a one-time guard or a counter. No request +/// identity, no native handles, no application state. +#[derive(Debug, Default)] +pub(crate) struct Sandbox { + logger_installed: bool, + requests: u64, + builds: u64, + pending_diagnostics: Vec, +} + +impl Sandbox { + /// Installs the global logger on first use. + /// + /// [`crate::logging::init_logger`] panics when a global logger is already + /// installed, which a reused sandbox would otherwise do on its second + /// request. The guard is owned here rather than inside the logger so that + /// ownership of the one-time initialization is visible at the entry point. + pub(crate) fn ensure_logger(&mut self) { + if self.logger_installed { + return; + } + crate::logging::init_logger(); + self.logger_installed = true; + + // Startup limit resolution runs in `main`, before any logger exists, + // so its diagnostics are held here and emitted once there is somewhere + // for them to go. Dropping them would hide the reason reuse is off. + for message in self.pending_diagnostics.drain(..) { + log::info!("{message}"); + } + } + + /// Records a startup diagnostic for emission once the logger is installed. + pub(crate) fn defer_diagnostic(&mut self, message: impl Into) { + self.pending_diagnostics.push(message.into()); + } + + /// Records the start of a request and returns its 1-based ordinal. + pub(crate) fn begin_request(&mut self) -> u64 { + self.requests = self.requests.saturating_add(1); + self.requests + } + + /// Records that the application was constructed in this sandbox. + pub(crate) fn record_build(&mut self) { + self.builds = self.builds.saturating_add(1); + } + + /// Number of requests this sandbox has begun. + #[cfg(any(feature = "reusable-sandbox", test))] + pub(crate) fn requests(&self) -> u64 { + self.requests + } + + /// Number of application builds this sandbox has performed. + pub(crate) fn builds(&self) -> u64 { + self.builds + } +} + +/// Counter values captured for one response. +/// +/// An owned snapshot rather than a borrow of [`Sandbox`], so attaching counters +/// never competes with the mutable borrow the request path holds. +#[derive(Debug, Clone, PartialEq, Eq)] +pub(crate) struct SandboxCounters { + pub(crate) ordinal: u64, + pub(crate) builds: u64, + pub(crate) request_id: String, +} + +impl SandboxCounters { + /// Captures the current counters, or `None` when metrics are disabled. + /// + /// Returning `None` is what keeps the counters off every response in the + /// default configuration, including the correlation id. + pub(crate) fn capture( + sandbox: &Sandbox, + ordinal: u64, + request_id: &str, + settings: Option<&Settings>, + ) -> Option { + settings + .filter(|settings| metrics_enabled(settings)) + .map(|_| Self { + ordinal, + builds: sandbox.builds(), + request_id: request_id.to_owned(), + }) + } +} + +/// Whether the counters endpoint and response counters are enabled. +pub(crate) fn metrics_enabled(settings: &Settings) -> bool { + settings.debug.sandbox_metrics_enabled +} + +/// Resolved sandbox limits, already normalized for the SDK. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) struct SandboxLimits { + pub(crate) max_requests: usize, + pub(crate) max_lifetime: Duration, + pub(crate) timeout: Duration, +} + +/// How this sandbox will serve requests. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum ServeMode { + /// Handle exactly one request, as a non-reusable sandbox does today. + Single, + /// Enter the SDK serving loop under the given bounds. + #[cfg_attr( + not(any(feature = "reusable-sandbox", test)), + expect(dead_code, reason = "constructed only on the reuse path") + )] + Reuse(SandboxLimits), +} + +/// Raw limit values as read from the runtime environment. +#[cfg(any(feature = "reusable-sandbox", test))] +#[derive(Debug, Default, Clone, Copy, PartialEq, Eq)] +pub(crate) struct RawLimits { + pub(crate) max_requests: Option, + pub(crate) max_lifetime_ms: Option, + pub(crate) timeout_ms: Option, +} + +/// Resolves raw configuration into a serving mode. +/// +/// Reuse requires all three bounds. A bare request limit is refused because the +/// SDK's omitted lifetime and wait timeout both default to [`Duration::MAX`], +/// which would leave a sandbox waiting without a bound. An application-level +/// request limit of `0` normalizes to `1`, because the SDK reads +/// `with_max_requests(0)` as unlimited — the opposite of what an operator +/// writing `0` intends. +#[cfg(any(feature = "reusable-sandbox", test))] +pub(crate) fn resolve_mode(raw: RawLimits) -> ServeMode { + let (Some(max_requests), Some(max_lifetime_ms), Some(timeout_ms)) = + (raw.max_requests, raw.max_lifetime_ms, raw.timeout_ms) + else { + return ServeMode::Single; + }; + + let max_requests = max_requests.max(1); + if max_requests <= 1 || max_lifetime_ms == 0 || timeout_ms == 0 { + return ServeMode::Single; + } + + let Ok(max_requests) = usize::try_from(max_requests) else { + return ServeMode::Single; + }; + + ServeMode::Reuse(SandboxLimits { + max_requests, + max_lifetime: Duration::from_millis(max_lifetime_ms), + timeout: Duration::from_millis(timeout_ms), + }) +} + +/// Suffixes of the service-scoped runtime-environment keys holding the bounds. +#[cfg(any(feature = "reusable-sandbox", test))] +const KEY_MAX_REQUESTS: &str = "TS__SANDBOX__MAX_REQUESTS"; +#[cfg(any(feature = "reusable-sandbox", test))] +const KEY_MAX_LIFETIME_MS: &str = "TS__SANDBOX__MAX_LIFETIME_MS"; +#[cfg(any(feature = "reusable-sandbox", test))] +const KEY_TIMEOUT_MS: &str = "TS__SANDBOX__TIMEOUT_MS"; + +/// Builds the service-scoped runtime-environment key for a bound. +/// +/// `EdgeZero`'s own `service_scoped_runtime_env_key` is private, so the shape is +/// reproduced here. It must stay identical to the one `edgezero provision` +/// writes. +#[cfg(any(feature = "reusable-sandbox", test))] +fn scoped_key(service_id: &str, suffix: &str) -> String { + format!("EDGEZERO__SERVICES__{service_id}__{suffix}") +} + +/// Collects raw limits from a fallible key lookup. +/// +/// Separated from the config-store call so the failure modes are testable +/// without a store: a lookup error, an absent key, and an unparseable value +/// must all degrade to "absent" rather than propagate. +/// +/// Diagnostics are returned rather than logged. This runs before the logger +/// exists, so anything logged here would be discarded. +#[cfg(any(feature = "reusable-sandbox", test))] +fn collect_raw_limits(service_id: &str, mut lookup: F) -> (RawLimits, Vec) +where + E: core::fmt::Display, + F: FnMut(&str) -> Result, E>, +{ + let mut diagnostics = Vec::new(); + + let mut read = |suffix: &str| -> Option { + let key = scoped_key(service_id, suffix); + match lookup(&key) { + Ok(Some(raw)) => match raw.trim().parse::() { + Ok(value) => Some(value), + Err(e) => { + diagnostics.push(format!( + "sandbox limit `{key}` is not a non-negative integer ({e}); ignoring" + )); + None + } + }, + Ok(None) => None, + Err(e) => { + diagnostics.push(format!( + "sandbox limit `{key}` lookup failed ({e}); ignoring" + )); + None + } + } + }; + + let limits = RawLimits { + max_requests: read(KEY_MAX_REQUESTS), + max_lifetime_ms: read(KEY_MAX_LIFETIME_MS), + timeout_ms: read(KEY_TIMEOUT_MS), + }; + + (limits, diagnostics) +} + +/// Reads the sandbox bounds from the runtime-environment config store. +/// +/// These keys deliberately bypass `edgezero_adapter_fastly::runtime_env_config`. +/// That helper resolves a closed allowlist — adapter host and port, logging +/// settings, and per-store `__NAME`/`__KEY` selectors — and silently drops +/// everything else, so a sandbox key routed through it would always read as +/// absent and reuse would never engage. +/// +/// Uses [`fastly::ConfigStore::try_get`], never `get`: `get` panics on a +/// lookup error, and this runs in `main` before the health probe, so a panic +/// here would take down liveness rather than merely disabling reuse. +/// +/// Any failure resolves to [`RawLimits::default`], and hence to +/// [`ServeMode::Single`]: an absent or unopenable store, an empty service id, +/// a failed lookup, or a value that does not parse as a `u64`. +#[cfg(feature = "reusable-sandbox")] +pub(crate) fn read_raw_limits() -> (RawLimits, Vec) { + use edgezero_adapter_fastly::RUNTIME_ENV_STORE_NAME; + + let Ok(store) = fastly::ConfigStore::try_open(RUNTIME_ENV_STORE_NAME) else { + return ( + RawLimits::default(), + vec![format!( + "sandbox reuse disabled: config store `{RUNTIME_ENV_STORE_NAME}` unavailable" + )], + ); + }; + + // Viceroy reports a service id of twenty-two zeros, which is a valid + // scope for key construction. Only an empty id is unusable. + let service_id = fastly::compute_runtime::service_id(); + if service_id.is_empty() { + return ( + RawLimits::default(), + vec!["sandbox reuse disabled: no service id available for key scoping".to_owned()], + ); + } + + collect_raw_limits(service_id, |key| store.try_get(key)) +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn scoped_key_matches_the_edgezero_shape() { + // Viceroy reports twenty-two zeros locally; it is a valid scope. + let service_id = "0000000000000000000000"; + + assert_eq!( + scoped_key(service_id, KEY_MAX_REQUESTS), + "EDGEZERO__SERVICES__0000000000000000000000__TS__SANDBOX__MAX_REQUESTS", + "should reproduce the service-scoped key edgezero writes" + ); + assert_eq!( + scoped_key(service_id, KEY_MAX_LIFETIME_MS), + "EDGEZERO__SERVICES__0000000000000000000000__TS__SANDBOX__MAX_LIFETIME_MS", + "should scope the lifetime bound the same way" + ); + assert_eq!( + scoped_key(service_id, KEY_TIMEOUT_MS), + "EDGEZERO__SERVICES__0000000000000000000000__TS__SANDBOX__TIMEOUT_MS", + "should scope the wait timeout the same way" + ); + } + + fn full(max_requests: u64) -> RawLimits { + RawLimits { + max_requests: Some(max_requests), + max_lifetime_ms: Some(30_000), + timeout_ms: Some(500), + } + } + + #[test] + fn absent_configuration_stays_single_request() { + assert_eq!( + resolve_mode(RawLimits::default()), + ServeMode::Single, + "an unconfigured sandbox should behave as it does today" + ); + } + + #[test] + fn zero_request_limit_normalizes_to_single_not_unlimited() { + assert_eq!( + resolve_mode(full(0)), + ServeMode::Single, + "0 should mean one request, never the SDK's unlimited" + ); + } + + #[test] + fn one_request_limit_stays_single_request() { + assert_eq!( + resolve_mode(full(1)), + ServeMode::Single, + "a limit of 1 should not enter the serving loop" + ); + } + + #[test] + fn partial_configuration_refuses_to_reuse() { + let raw = RawLimits { + max_requests: Some(10), + max_lifetime_ms: None, + timeout_ms: Some(500), + }; + + assert_eq!( + resolve_mode(raw), + ServeMode::Single, + "a missing bound should not fall back to the SDK's unbounded default" + ); + } + + #[test] + fn zero_lifetime_or_timeout_refuses_to_reuse() { + assert_eq!( + resolve_mode(RawLimits { + max_lifetime_ms: Some(0), + ..full(10) + }), + ServeMode::Single, + "a zero lifetime should not mean unlimited" + ); + assert_eq!( + resolve_mode(RawLimits { + timeout_ms: Some(0), + ..full(10) + }), + ServeMode::Single, + "a zero wait timeout should not mean unlimited" + ); + } + + #[test] + fn complete_configuration_enables_reuse() { + assert_eq!( + resolve_mode(full(10)), + ServeMode::Reuse(SandboxLimits { + max_requests: 10, + max_lifetime: Duration::from_millis(30_000), + timeout: Duration::from_millis(500), + }), + "all three bounds present should enter the serving loop" + ); + } + + #[test] + fn sandbox_installs_the_logger_once_then_reuses_it() { + let mut sandbox = Sandbox::default(); + sandbox.defer_diagnostic("startup note"); + + // First request: really installs the global logger. + sandbox.ensure_logger(); + assert!( + sandbox.logger_installed, + "the first request should install the logger" + ); + assert!( + sandbox.pending_diagnostics.is_empty(), + "installation should flush the deferred startup diagnostics" + ); + + // Second request in the same sandbox. Without the guard this reaches + // `fern`'s `apply()` a second time and panics, which is the failure + // this test exists to catch. + sandbox.ensure_logger(); + sandbox.ensure_logger(); + + assert!( + sandbox.logger_installed, + "the guard should stay set across requests" + ); + } + + /// Stand-in for `fastly::config_store::LookupError`, which cannot be + /// constructed outside the SDK. + #[derive(Debug, derive_more::Display)] + #[display("simulated lookup failure")] + struct LookupFailed; + + #[test] + fn a_failed_lookup_degrades_to_single_request_instead_of_panicking() { + // `ConfigStore::get` panics on a lookup error and runs before the + // health probe, so the fallible path must absorb the error. + let (limits, diagnostics) = + collect_raw_limits::("svc", |_key| Err(LookupFailed)); + + assert_eq!( + limits, + RawLimits::default(), + "a failed lookup should read as absent" + ); + assert_eq!( + resolve_mode(limits), + ServeMode::Single, + "a failed lookup should leave the sandbox single-request" + ); + assert_eq!( + diagnostics.len(), + 3, + "each failed key should report why reuse was declined" + ); + assert!( + diagnostics[0].contains("lookup failed"), + "the diagnostic should name the failure, got: {}", + diagnostics[0] + ); + } + + #[test] + fn an_unparseable_value_degrades_to_single_request() { + let (limits, diagnostics) = collect_raw_limits::("svc", |key| { + Ok(Some(if key.ends_with(KEY_MAX_REQUESTS) { + "ten".to_owned() + } else { + "500".to_owned() + })) + }); + + assert_eq!( + limits.max_requests, None, + "an unparseable limit should read as absent" + ); + assert_eq!( + resolve_mode(limits), + ServeMode::Single, + "an unparseable request limit should not enable reuse" + ); + assert!( + diagnostics + .iter() + .any(|d| d.contains("not a non-negative integer")), + "should explain the rejected value, got: {diagnostics:?}" + ); + } + + #[test] + fn absent_keys_report_nothing_and_stay_single_request() { + let (limits, diagnostics) = collect_raw_limits::("svc", |_key| Ok(None)); + + assert_eq!( + limits, + RawLimits::default(), + "absent keys should read absent" + ); + assert!( + diagnostics.is_empty(), + "an unconfigured sandbox is the normal case, not a diagnostic" + ); + } + + #[test] + fn a_complete_store_enables_reuse_end_to_end() { + let (limits, diagnostics) = collect_raw_limits::("svc", |key| { + Ok(Some( + if key.ends_with(KEY_MAX_REQUESTS) { + "10" + } else if key.ends_with(KEY_MAX_LIFETIME_MS) { + "30000" + } else { + "500" + } + .to_owned(), + )) + }); + + assert!(diagnostics.is_empty(), "a valid store should be quiet"); + assert_eq!( + resolve_mode(limits), + ServeMode::Reuse(SandboxLimits { + max_requests: 10, + max_lifetime: Duration::from_millis(30_000), + timeout: Duration::from_millis(500), + }), + "a fully configured store should enter the serving loop" + ); + } + + #[test] + fn deferred_diagnostics_survive_until_the_logger_exists() { + let mut sandbox = Sandbox::default(); + sandbox.defer_diagnostic("reuse declined"); + + assert_eq!( + sandbox.pending_diagnostics.len(), + 1, + "startup runs before the logger, so the message must be held" + ); + + // Simulates `ensure_logger` past the point of installation; calling it + // for real would install a global logger and break sibling tests. + sandbox.logger_installed = true; + let drained: Vec = sandbox.pending_diagnostics.drain(..).collect(); + + assert_eq!( + drained, + vec!["reuse declined".to_owned()], + "the held diagnostic should be emitted, not dropped" + ); + } + + #[test] + fn counters_are_captured_only_when_metrics_are_enabled() { + let sandbox = Sandbox::default(); + let mut settings = Settings::default(); + + assert_eq!( + SandboxCounters::capture(&sandbox, 1, "req-1", None), + None, + "no settings should mean no counters" + ); + + settings.debug.sandbox_metrics_enabled = false; + assert_eq!( + SandboxCounters::capture(&sandbox, 1, "req-1", Some(&settings)), + None, + "counters must stay off in the default configuration" + ); + + settings.debug.sandbox_metrics_enabled = true; + assert_eq!( + SandboxCounters::capture(&sandbox, 7, "req-7", Some(&settings)), + Some(SandboxCounters { + ordinal: 7, + builds: 0, + request_id: "req-7".to_owned(), + }), + "enabling the flag should capture the current counters" + ); + } + + #[test] + fn ordinals_increase_and_builds_count_separately() { + let mut sandbox = Sandbox::default(); + + assert_eq!( + sandbox.begin_request(), + 1, + "first request should be ordinal 1" + ); + sandbox.record_build(); + assert_eq!(sandbox.begin_request(), 2, "ordinals should increase"); + + assert_eq!(sandbox.requests(), 2, "should have begun two requests"); + assert_eq!( + sandbox.builds(), + 1, + "a reused sandbox should report fewer builds than requests" + ); + } +} diff --git a/crates/trusted-server-cli/Cargo.toml b/crates/trusted-server-cli/Cargo.toml index 44ad1d443..7caac2dff 100644 --- a/crates/trusted-server-cli/Cargo.toml +++ b/crates/trusted-server-cli/Cargo.toml @@ -33,6 +33,17 @@ trusted-server-core = { workspace = true } url = { workspace = true } which = { workspace = true } +# `ts dev sandbox-probe` needs a real HTTP/1.1 client with keep-alive. These +# are shared with the macOS-only proxy but scoped to every non-wasm host, so +# the probe is built and linted on Linux CI too. The exclusion still keeps +# `tokio`/`ring` off the repo-default `wasm32-wasip1` target. +[target.'cfg(not(target_family = "wasm"))'.dependencies] +bytes = { workspace = true } +http-body-util = { workspace = true } +hyper = { workspace = true, features = ["http1", "client"] } +hyper-util = { workspace = true, features = ["tokio"] } +tokio = { workspace = true, features = ["net", "rt-multi-thread"] } + # `ts dev proxy` is macOS-only — CA trust via the login keychain, Safari # automation via `networksetup`, and a native TLS / networking stack. Scoping # these dependencies to macOS keeps unsupported targets (notably the diff --git a/crates/trusted-server-cli/src/commands/dev/mod.rs b/crates/trusted-server-cli/src/commands/dev/mod.rs index 7a61d769b..702a5be32 100644 --- a/crates/trusted-server-cli/src/commands/dev/mod.rs +++ b/crates/trusted-server-cli/src/commands/dev/mod.rs @@ -3,6 +3,9 @@ // other host targets `ts dev` parses but exposes no subcommands. #[cfg(target_os = "macos")] pub mod proxy; +// Needs a real HTTP client; its dependencies are scoped to non-wasm hosts. +#[cfg(not(target_family = "wasm"))] +pub mod sandbox_probe; /// The `ts dev …` command group. #[derive(Debug, clap::Subcommand)] @@ -10,27 +13,20 @@ pub enum DevCommand { /// Run the local production-hostname dev proxy (macOS only). #[cfg(target_os = "macos")] Proxy(proxy::ProxyArgs), + /// Measure sandbox reuse from the counters on workload responses. + #[cfg(not(target_family = "wasm"))] + SandboxProbe(sandbox_probe::SandboxProbeArgs), } /// Dispatches a `dev` subcommand. /// /// # Errors -/// Returns the subcommand's failure rendered as a message. On non-macOS targets -/// `DevCommand` has no variants, so this never returns an error there. -// On non-macOS targets `DevCommand` is an empty enum: the by-value parameter is -// consumed by an empty `match`, which clippy reads as a needless by-value pass. -// Taking `&DevCommand` is not an option — a zero-arm `match` is not exhaustive -// over a reference type — so the owned parameter is required. -#[cfg_attr( - not(target_os = "macos"), - allow( - clippy::needless_pass_by_value, - reason = "empty enum requires owned value for exhaustive match" - ) -)] +/// Returns the subcommand's failure rendered as a message. pub fn run(command: DevCommand) -> Result<(), String> { match command { #[cfg(target_os = "macos")] DevCommand::Proxy(args) => proxy::run(&args).map_err(|report| format!("{report:?}")), + #[cfg(not(target_family = "wasm"))] + DevCommand::SandboxProbe(args) => sandbox_probe::run(&args), } } diff --git a/crates/trusted-server-cli/src/commands/dev/sandbox_probe.rs b/crates/trusted-server-cli/src/commands/dev/sandbox_probe.rs new file mode 100644 index 000000000..d29c1de53 --- /dev/null +++ b/crates/trusted-server-cli/src/commands/dev/sandbox_probe.rs @@ -0,0 +1,533 @@ +//! `ts dev sandbox-probe` — measures sandbox reuse from workload responses. +//! +//! Issues a keep-alive sequence against a locally served Trusted Server over a +//! single connection and reads the counters the Fastly adapter attaches to +//! each response. Reuse is reported only when one instance id serves several +//! *workload* requests with strictly increasing ordinals. +//! +//! The counters endpoint is deliberately not used to establish reuse: a +//! snapshot identifies the sandbox that served the probe, which need not be +//! the one that served the preceding request. + +use std::collections::BTreeMap; +use std::fmt::Write as _; +use std::time::{Duration, Instant}; + +use http_body_util::BodyExt as _; +use hyper_util::rt::TokioIo; + +/// Response headers the Fastly adapter attaches when sandbox metrics are on. +const HEADER_INSTANCE: &str = "x-ts-sandbox-instance"; +const HEADER_ORDINAL: &str = "x-ts-sandbox-ordinal"; +const HEADER_BUILDS: &str = "x-ts-sandbox-builds"; +const HEADER_REQUEST_ID: &str = "x-ts-sandbox-request-id"; + +/// Value the adapter reports when the runtime exposes no instance identity. +const INSTANCE_UNAVAILABLE: &str = "unavailable"; + +/// Paths that never carry counters because they short-circuit before dispatch. +const NON_WORKLOAD_PATHS: &[&str] = &["/health", "/_ts/debug/sandbox", "/_ts/debug/ja4"]; + +/// Arguments for `ts dev sandbox-probe`. +#[derive(Debug, clap::Args)] +pub struct SandboxProbeArgs { + /// Host and port of the locally served instance. + #[arg(long, default_value = "127.0.0.1:7676")] + pub authority: String, + + /// Workload path to exercise. + /// + /// Required, and it must reach the router. `/health` and the debug probes + /// short-circuit ahead of the point where counters are attached, so + /// probing one of those reports nothing. + #[arg(long)] + pub path: String, + + /// Number of requests to issue over one connection. + #[arg(long, default_value_t = 6)] + pub requests: u32, + + /// Per-request timeout in milliseconds. + #[arg(long, default_value_t = 10_000)] + pub timeout_ms: u64, +} + +/// One observation taken from a workload response. +#[derive(Debug, Clone)] +struct Observation { + instance: Option, + ordinal: Option, + builds: Option, + request_id: Option, + status: u16, + elapsed: Duration, +} + +/// Why the request sequence stopped, if it stopped early. +#[derive(Debug, Clone, PartialEq, Eq)] +enum Ending { + /// Every requested iteration completed. + Completed, + /// The connection closed. Expected when a bounded sandbox retires, but a + /// client cannot distinguish that from any other close. + ConnectionClosed(String), + /// A transport or body error, which invalidates the run. + TransportError(String), +} + +/// Runs the probe and prints a report. +/// +/// # Errors +/// Returns a message when the path is not a workload route, the async runtime +/// cannot start, or the connection cannot be established. +pub fn run(args: &SandboxProbeArgs) -> Result<(), String> { + if NON_WORKLOAD_PATHS.contains(&args.path.as_str()) { + return Err(format!( + "`{}` short-circuits before counters are attached; choose a workload route", + args.path + )); + } + + let runtime = tokio::runtime::Builder::new_current_thread() + .enable_all() + .build() + .map_err(|e| format!("failed to start async runtime: {e}"))?; + + let (observations, ending) = runtime.block_on(collect(args))?; + crate::output::info(report(&observations, &ending).trim_end()); + Ok(()) +} + +/// Issues the keep-alive sequence over a single HTTP/1.1 connection. +async fn collect(args: &SandboxProbeArgs) -> Result<(Vec, Ending), String> { + let stream = tokio::net::TcpStream::connect(&args.authority) + .await + .map_err(|e| format!("failed to connect to {}: {e}", args.authority))?; + + let (mut sender, connection) = hyper::client::conn::http1::handshake(TokioIo::new(stream)) + .await + .map_err(|e| format!("HTTP/1.1 handshake with {} failed: {e}", args.authority))?; + + // The connection task drives the socket; it resolves when the peer closes. + let connection = tokio::spawn(connection); + + let timeout = Duration::from_millis(args.timeout_ms); + let mut observations = Vec::with_capacity(args.requests as usize); + let mut ending = Ending::Completed; + + for _ in 0..args.requests { + let request = hyper::Request::builder() + .method(hyper::Method::GET) + .uri(&args.path) + .header(hyper::header::HOST, &args.authority) + .body(String::new()) + .map_err(|e| format!("failed to build request: {e}"))?; + + let started = Instant::now(); + + let response = match tokio::time::timeout(timeout, sender.send_request(request)).await { + Ok(Ok(response)) => response, + Ok(Err(e)) if e.is_closed() || e.is_incomplete_message() => { + ending = Ending::ConnectionClosed(e.to_string()); + break; + } + Ok(Err(e)) => { + ending = Ending::TransportError(e.to_string()); + break; + } + Err(_) => { + ending = Ending::TransportError(format!("request timed out after {timeout:?}")); + break; + } + }; + + let (parts, body) = response.into_parts(); + + // Hyper handles chunk extensions, trailers, and close-delimited + // bodies. The body must be drained before the next request is sent. + match tokio::time::timeout(timeout, body.collect()).await { + Ok(Ok(collected)) => drop(collected.to_bytes()), + Ok(Err(e)) => { + ending = Ending::TransportError(format!("body read failed: {e}")); + break; + } + Err(_) => { + ending = Ending::TransportError(format!("body read timed out after {timeout:?}")); + break; + } + } + + observations.push(observation_from(&parts, started.elapsed())); + } + + drop(sender); + if let Ok(Err(e)) = connection.await + && ending == Ending::Completed + { + ending = Ending::ConnectionClosed(e.to_string()); + } + + Ok((observations, ending)) +} + +/// Extracts the counters from one response. +fn observation_from(parts: &hyper::http::response::Parts, elapsed: Duration) -> Observation { + let header = |name: &str| { + parts + .headers + .get(name) + .and_then(|value| value.to_str().ok()) + .map(str::to_owned) + }; + + Observation { + instance: header(HEADER_INSTANCE), + ordinal: header(HEADER_ORDINAL).and_then(|v| v.parse().ok()), + builds: header(HEADER_BUILDS).and_then(|v| v.parse().ok()), + request_id: header(HEADER_REQUEST_ID), + status: parts.status.as_u16(), + elapsed, + } +} + +/// Renders the report, stating explicitly what was and was not established. +fn report(observations: &[Observation], ending: &Ending) -> String { + let mut out = String::new(); + let _ = writeln!(out, "responses: {}", observations.len()); + + match ending { + Ending::Completed => {} + Ending::ConnectionClosed(reason) => { + let _ = writeln!( + out, + "ended: connection closed ({reason}); a client cannot tell sandbox \ + retirement from any other close" + ); + } + Ending::TransportError(reason) => { + let _ = writeln!(out, "ended: transport error ({reason})"); + } + } + + if observations.is_empty() { + let _ = writeln!(out, "reuse: unverified (no responses)"); + return out; + } + + for (index, observation) in observations.iter().enumerate() { + let _ = writeln!( + out, + " {:>3}. status={} instance={} ordinal={} builds={} request={} elapsed={:?}", + index + 1, + observation.status, + observation.instance.as_deref().unwrap_or("-"), + observation + .ordinal + .map_or_else(|| "-".to_owned(), |v| v.to_string()), + observation + .builds + .map_or_else(|| "-".to_owned(), |v| v.to_string()), + observation.request_id.as_deref().unwrap_or("-"), + observation.elapsed, + ); + } + + out.push_str(&verdict(observations, ending)); + out +} + +/// Decides whether the observations establish reuse. +/// +/// Deliberately conservative: only strictly increasing ordinals on one +/// instance count. A repeated or decreasing ordinal under one instance id +/// means the attribution is unreliable — the id is not identifying what it +/// claims to — so no conclusion is drawn either way. Counting bare repeats +/// would instead report those runs as reuse. +fn verdict(observations: &[Observation], ending: &Ending) -> String { + if let Ending::TransportError(reason) = ending { + return format!("reuse: unverified (transport error: {reason})\n"); + } + + if observations.iter().all(|o| o.instance.is_none()) { + return "reuse: unverified (no counters; enable debug.sandbox_metrics_enabled)\n" + .to_owned(); + } + + if observations + .iter() + .any(|o| o.instance.as_deref() == Some(INSTANCE_UNAVAILABLE)) + { + return "reuse: unverified (runtime reported no instance identity)\n".to_owned(); + } + + // Partial attribution means some responses cannot be placed in a sandbox, + // so neither reuse nor its absence can be concluded. + if observations + .iter() + .any(|o| o.instance.is_none() || o.ordinal.is_none()) + { + return "reuse: unverified (incomplete attribution: some responses carried no \ + instance id or ordinal)\n" + .to_owned(); + } + + let mut by_instance: BTreeMap<&str, Vec> = BTreeMap::new(); + for observation in observations { + if let (Some(instance), Some(ordinal)) = + (observation.instance.as_deref(), observation.ordinal) + { + by_instance.entry(instance).or_default().push(ordinal); + } + } + + let mut out = String::new(); + let _ = writeln!(out, "distinct instances: {}", by_instance.len()); + + // Ordinals from one sandbox must arrive strictly increasing. Anything else + // means the ids are not the identities they claim to be. + let anomalous: Vec<&str> = by_instance + .iter() + .filter(|(_, ordinals)| !ordinals.windows(2).all(|pair| pair[0] < pair[1])) + .map(|(instance, _)| *instance) + .collect(); + if !anomalous.is_empty() { + let _ = writeln!( + out, + "reuse: unverified (instance(s) {} reported repeated or decreasing ordinals)", + anomalous.join(", ") + ); + return out; + } + + let reused = by_instance + .values() + .filter(|ordinals| ordinals.len() > 1) + .count(); + let max_per_instance = by_instance.values().map(Vec::len).max().unwrap_or_default(); + let max_builds = observations + .iter() + .filter_map(|o| o.builds) + .max() + .unwrap_or_default(); + + let _ = writeln!(out, "max requests per instance: {max_per_instance}"); + let _ = writeln!(out, "max builds per instance: {max_builds}"); + + if reused == 0 { + let _ = writeln!( + out, + "reuse: not observed (every response came from a distinct instance)" + ); + } else { + let _ = writeln!( + out, + "reuse: observed ({reused} instance(s) served more than one workload request)" + ); + } + out +} + +#[cfg(test)] +mod tests { + use super::*; + + fn observation( + instance: Option<&str>, + ordinal: Option, + builds: Option, + ) -> Observation { + Observation { + instance: instance.map(str::to_owned), + ordinal, + builds, + request_id: Some("req".to_owned()), + status: 200, + elapsed: Duration::from_millis(1), + } + } + + #[test] + fn absent_counters_report_unverified_not_negative() { + let observations = vec![observation(None, None, None), observation(None, None, None)]; + + let verdict = verdict(&observations, &Ending::Completed); + + assert!( + verdict.contains("unverified"), + "missing counters are not evidence that reuse failed" + ); + assert!( + !verdict.contains("reusable-sandbox"), + "counters are feature-independent, so the hint must not name the feature: {verdict}" + ); + } + + #[test] + fn an_unavailable_instance_id_reports_unverified() { + let observations = vec![ + observation(Some(INSTANCE_UNAVAILABLE), Some(1), Some(1)), + observation(Some(INSTANCE_UNAVAILABLE), Some(2), Some(1)), + ]; + + assert!( + verdict(&observations, &Ending::Completed).contains("unverified"), + "a runtime with no identity cannot establish reuse" + ); + } + + #[test] + fn repeated_ordinals_on_one_instance_are_not_reuse() { + // Unreliable attribution: one id reporting ordinal 1 twice. A bare + // repeat count would call this reuse. + let observations = vec![ + observation(Some("a"), Some(1), Some(1)), + observation(Some("a"), Some(1), Some(1)), + ]; + + let verdict = verdict(&observations, &Ending::Completed); + + assert!( + verdict.contains("unverified"), + "duplicate ordinals must not read as reuse, got: {verdict}" + ); + assert!( + verdict.contains("repeated or decreasing"), + "should name the anomaly, got: {verdict}" + ); + } + + #[test] + fn decreasing_ordinals_are_not_reuse() { + let observations = vec![ + observation(Some("a"), Some(3), Some(1)), + observation(Some("a"), Some(2), Some(1)), + ]; + + assert!( + verdict(&observations, &Ending::Completed).contains("unverified"), + "ordinals going backwards mean the identity is untrustworthy" + ); + } + + #[test] + fn incomplete_attribution_reports_unverified() { + let observations = vec![ + observation(Some("a"), Some(1), Some(1)), + observation(None, None, None), + ]; + + let verdict = verdict(&observations, &Ending::Completed); + + assert!( + verdict.contains("incomplete attribution"), + "a response that cannot be placed in a sandbox blocks a verdict, got: {verdict}" + ); + } + + #[test] + fn a_transport_error_invalidates_the_run() { + let observations = vec![ + observation(Some("a"), Some(1), Some(1)), + observation(Some("a"), Some(2), Some(1)), + ]; + + let verdict = verdict(&observations, &Ending::TransportError("reset".to_owned())); + + assert!( + verdict.contains("unverified"), + "a transport error makes the run unreliable, got: {verdict}" + ); + } + + #[test] + fn distinct_instances_report_reuse_not_observed() { + let observations = vec![ + observation(Some("a"), Some(1), Some(1)), + observation(Some("b"), Some(1), Some(1)), + ]; + + let verdict = verdict(&observations, &Ending::Completed); + + assert!( + verdict.contains("not observed"), + "one request per instance is arm A, got: {verdict}" + ); + assert!( + verdict.contains("distinct instances: 2"), + "should report the instance count, got: {verdict}" + ); + } + + #[test] + fn increasing_ordinals_on_one_instance_report_reuse() { + let observations = vec![ + observation(Some("a"), Some(1), Some(1)), + observation(Some("a"), Some(2), Some(1)), + observation(Some("a"), Some(3), Some(1)), + ]; + + let verdict = verdict(&observations, &Ending::Completed); + + assert!( + verdict.contains("reuse: observed"), + "three increasing ordinals on one instance is reuse, got: {verdict}" + ); + assert!( + verdict.contains("max requests per instance: 3"), + "should report the depth reached, got: {verdict}" + ); + assert!( + verdict.contains("max builds per instance: 1"), + "one build across three requests is the arm-C signal, got: {verdict}" + ); + } + + #[test] + fn a_closed_connection_is_reported_without_being_read_as_retirement() { + let observations = vec![ + observation(Some("a"), Some(1), Some(1)), + observation(Some("a"), Some(2), Some(1)), + ]; + + let report = report( + &observations, + &Ending::ConnectionClosed("closed".to_owned()), + ); + + assert!( + report.contains("cannot tell sandbox"), + "a close is ambiguous and must be reported as such, got: {report}" + ); + assert!( + report.contains("reuse: observed"), + "a close does not invalidate observations already collected, got: {report}" + ); + } + + #[test] + fn non_workload_paths_are_rejected_before_connecting() { + for path in NON_WORKLOAD_PATHS { + let args = SandboxProbeArgs { + authority: "127.0.0.1:7676".to_owned(), + path: (*path).to_owned(), + requests: 2, + timeout_ms: 100, + }; + + let error = run(&args).expect_err("should refuse a short-circuiting path"); + + assert!( + error.contains("short-circuits"), + "should explain why `{path}` cannot measure reuse, got: {error}" + ); + } + } + + #[test] + fn an_empty_run_is_unverified() { + assert!( + report(&[], &Ending::Completed).contains("unverified"), + "no responses cannot establish anything" + ); + } +} diff --git a/crates/trusted-server-core/src/settings.rs b/crates/trusted-server-core/src/settings.rs index f57714dcc..a70ddd1cd 100644 --- a/crates/trusted-server-core/src/settings.rs +++ b/crates/trusted-server-core/src/settings.rs @@ -2548,6 +2548,31 @@ pub struct DebugConfig { /// un-sanitized creative for diagnostics, so never enable in production. #[serde(default)] pub inject_adm_for_testing: bool, + + /// Expose the reusable-sandbox counters endpoint at `GET /_ts/debug/sandbox` + /// and attach the same counters to workload responses. + /// + /// The counters are the guest-instance identifier, the request ordinal + /// within that instance, the application build count, and the request + /// correlation id. They carry no settings, secrets, or request content. + /// + /// Independent of the adapter's `reusable-sandbox` Cargo feature by design: + /// the feature decides whether a `Serve` loop exists, this flag decides + /// whether counters are emitted. Keeping them separate is what lets the + /// feature-off baseline be measured on the same channel as the reuse arms. + /// + /// Skipped from serialization while false: [`DebugConfig`] denies unknown + /// fields, so a default blob must stay readable by a binary built before + /// this field existed. A blob with it enabled requires restoring a + /// compatible blob before rolling back, the same trade + /// [`DebugConfig::auction_html_comment_options`] makes. + #[serde(default, skip_serializing_if = "is_false")] + pub sandbox_metrics_enabled: bool, +} + +/// Serde predicate for omitting `false` flags from serialized config blobs. +fn is_false(value: &bool) -> bool { + !*value } /// Metadata keys safe to surface in the `ts-debug` auction comment. @@ -3685,6 +3710,57 @@ mod tests { use std::collections::HashSet; use std::sync::Arc; + /// `DebugConfig` denies unknown fields, so a binary built before + /// `sandbox_metrics_enabled` existed must still accept a default blob. + /// That only holds while the flag is skipped during serialization. + #[test] + fn default_debug_config_omits_sandbox_metrics_for_rollback() { + let serialized = + serde_json::to_value(DebugConfig::default()).expect("should serialize debug config"); + + assert!( + serialized.get("sandbox_metrics_enabled").is_none(), + "a default blob must not carry the field, or an older binary rejects it: {serialized}" + ); + } + + #[test] + fn enabled_sandbox_metrics_serializes_and_round_trips() { + let config = DebugConfig { + sandbox_metrics_enabled: true, + ..DebugConfig::default() + }; + + let serialized = serde_json::to_value(&config).expect("should serialize debug config"); + assert_eq!( + serialized.get("sandbox_metrics_enabled"), + Some(&json!(true)), + "an enabled flag must be written so the setting survives a round trip" + ); + + let restored: DebugConfig = + serde_json::from_value(serialized).expect("should deserialize debug config"); + assert!( + restored.sandbox_metrics_enabled, + "the flag should survive a round trip" + ); + } + + #[test] + fn debug_config_accepts_a_blob_without_the_sandbox_field() { + let restored: DebugConfig = serde_json::from_value(json!({"ja4_endpoint_enabled": true})) + .expect("should deserialize a blob written before the field existed"); + + assert!( + restored.ja4_endpoint_enabled, + "existing fields should still load" + ); + assert!( + !restored.sandbox_metrics_enabled, + "an absent flag should default to off" + ); + } + use crate::auction::build_orchestrator; use crate::integrations::{ IntegrationRegistry, gpt::GptConfig, nextjs::NextJsIntegrationConfig, diff --git a/docs/superpowers/specs/2026-09-17-fastly-reusable-sandbox-design.md b/docs/superpowers/specs/2026-09-17-fastly-reusable-sandbox-design.md new file mode 100644 index 000000000..6491d4a08 --- /dev/null +++ b/docs/superpowers/specs/2026-09-17-fastly-reusable-sandbox-design.md @@ -0,0 +1,537 @@ +# Fastly reusable sandbox adoption + +Issue: [#856](https://github.com/IABTechLab/trusted-server/issues/856) + +## Problem + +Every Fastly Compute request currently starts a fresh Wasm sandbox, so the +entry point repeats the whole initialization sequence: read the runtime env +config store, open the application config store, load and parse settings, +resolve secrets, compile the auction plan, build the orchestrator, build the +integration registry, and construct the telemetry sink. The `OnceLock` regex +caches in `settings.rs` are initialized and discarded within a single request. + +Fastly SDK 0.12.1 exposes `fastly::http::serve::Serve`, which lets one sandbox +handle several requests. This design adopts it so that initialization can be +amortized, while keeping reuse opt-in and leaving the default deployment +behaviour unchanged. + +## Scope + +In scope: the Fastly adapter entry point and the state it owns, plus one +`trusted-server-core` change — moving the script rewriters' accumulation +buffers out of registry-lifetime objects, without which retention is unsafe. +See [Retained rewrite buffers](#retained-rewrite-buffers-blocking). + +Out of scope: Spin, Cloudflare, and Axum adapters; any other change to routing, +auction, EC, or integration behaviour; memory-bounded retirement (see +[Deferred](#deferred)). + +## Current entry point + +The adapter does not use `#[fastly::main]`. It owns its own `main`, converts +the raw request itself, dispatches straight into the router, and sends the +response explicitly. That shape is load-bearing and must survive the change. + +| Concern | Location | +| ----------------------------------------------------- | ----------------------------------------------------------------- | +| Raw `FastlyRequest::from_client()` | `main.rs:78-89` | +| Health probe ahead of all initialization | `main.rs:65-72`, called at `main.rs:81` | +| Global logger install | `main.rs:87` → `logging.rs:82-106` | +| Runtime env read | `main.rs:93-94` | +| Native config-store handle → request extension | opened `main.rs:58-63`, `main.rs:117-127`; inserted `main.rs:178` | +| Application build | `main.rs:129` → `app.rs:1278-1297` | +| Direct router dispatch (keeps duplicate `Set-Cookie`) | `main.rs:170-180` | +| Progressive streaming via `stream_to_client` | `main.rs:338-368` | +| Synchronous post-send pull sync | `main.rs:450-466` | + +Two properties of this path block naive reuse. + +1. `logging.rs:105` ends in `.apply().expect("should initialize logger")`. + A second call panics, because a global logger is already installed. +2. `app.rs:1288-1296` returns `startup_error_router` with `state: None` when + settings fail to load. Retaining that result would pin a sandbox into + permanent error mode for every subsequent request it serves. + +## Dependencies + +No EdgeZero version change. The workspace pin stays at `v0.0.8`. + +`Serve`, `ServeSummary`, and `HandlerResult` all come from `fastly 0.12.1`, +which is already resolved in `Cargo.lock`. EdgeZero PR #379 re-exports only +`Serve` and `ServeSummary` (not `HandlerResult`) and adds `serve_app` / +`serve_app_with_request_extensions`. + +Those helpers are unusable here, but not because of header handling: EdgeZero's +`to_fastly_response` uses `append_header` (v0.0.8 +`crates/edgezero-adapter-fastly/src/response.rs:28`), so duplicate `Set-Cookie` +values survive that conversion. The disqualifying part is that the same +function drains `Body::Stream` into a buffered `fastly::Body` before sending, +which ends progressive streaming, and that the helper path leaves no place for +the response-extension finalization and post-send work this adapter performs. +That holds regardless of whether PR #379 merges, so this work takes no +dependency on it. + +`HandlerResult` is implemented for `()`, for handlers that have already called +`send_to_client` or `stream_to_client`. That is exactly the current handler +shape, so the existing send logic is reused unchanged. + +## Design + +### Three execution modes + +Compatibility is layered, and the layers are not equivalent. Only the first +reproduces today's execution. + +```text +feature off → original entry path +feature on, effective max_requests <= 1 → single-request handler path +feature on, validated reuse limits → Serve loop +``` + +The feature-on single-request path is **not** byte-identical to the feature-off +path: it performs the startup mode lookup, which opens a config store and reads +three keys before the first request. It does not construct a `Serve` or enter +its loop — it calls the handler once directly, as the code below shows. It is +described as _single-request operation_, not as identical execution. + +The startup mode lookup must not be able to break the health probe. If the +lookup fails for any reason, the adapter falls back to single-request operation +and the health probe continues to answer. + +### Commit 1 — lifecycle, no retention + +Add a `reusable-sandbox` Cargo feature to `trusted-server-adapter-fastly`, off +by default. `fastly.toml`'s build command does not pass it, so production +builds are unaffected. + +Split `main` into a loop owner and a handler: + +```rust +fn main() { + let mut sandbox = Sandbox::default(); + match serve_mode() { + ServeMode::Single => { handle_request(FastlyRequest::from_client(), &mut sandbox); } + ServeMode::Reuse(serve) => { let _ = serve.run_with_context(handle_request, &mut sandbox); } + } +} + +fn handle_request(req: FastlyRequest, sandbox: &mut Sandbox) { + // today's `main` body: health probe, logger, edgezero_main +} +``` + +`handle_request` returns `()`. It keeps sending its own response, so streaming, +duplicate `Set-Cookie`, response extensions, and post-send pull sync are +untouched. + +In this commit `Sandbox` holds no application state. It holds: + +- `logger_installed: bool`, which fixes the `logging.rs:105` double-`apply()` + panic. +- The measurement bookkeeping from + [Ownership of the measurement surface](#ownership-of-the-measurement-surface): + the validated guest-instance identifier, the request ordinal within that + instance, and the application build counter. + +The build counter lives here from commit 1 even though nothing increments it +past one until commit 3 — that is what makes arm B's "one build per request" +baseline measurable rather than assumed. + +Counters are attached to the response **before headers are committed**. On the +streaming path that is before `stream_to_client`, which means they describe the +request up to commitment and cannot report its eventual outcome. A failure +after commitment is therefore recorded separately, in logs, and is reconciled +with the counters during analysis rather than being expected to appear in +them. + +Every non-panicking path through `handle_request` must send exactly once. This +is already true of the current code (each early return sends before returning), +but it becomes a correctness requirement of the loop rather than an incidental +property of a process that is about to exit, so it is asserted by test. + +The loop owner introduces one new failure mode. `run_with_context` panics +(`serve.rs:320`) if `RequestPromise::new` fails mid-loop, which cannot happen in +today's one-shot `main`. This is accepted rather than worked around: it fires +only between requests, after the current response has been sent, and the SDK +treats it as a sandbox failure — the same outcome as any other guest panic. It +is recorded here so it is not mistaken for an application fault during +validation. + +### Commit 2 — per-document rewrite buffers + +A `trusted-server-core` prerequisite for retention, specified under +[Retained rewrite buffers](#retained-rewrite-buffers-blocking). No adapter +change and no behavioural change under today's one-request-per-sandbox model. + +### Commit 3 — retention + +`Sandbox` gains `app: Option`, where `RetainedApp` owns the built +`App` and its `Arc`. It is constructed lazily on the first request +that is not one of the three pre-build short-circuits — the health probe, the +JA4 debug probe, and the counters endpoint from +[Ownership of the measurement surface](#ownership-of-the-measurement-surface) — +so none of those paths pays for construction. + +A failed build is never stored. `router_with_state` already returns +`(startup_error_router, None)` on failure; that result serves the current +request and is dropped, so a transient config-store failure cannot poison the +sandbox. + +### Ownership + +| Retained in `Sandbox` | Rebuilt every request | +| --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------- | +| `logger_installed` | `ConfigStoreHandle` (native handle) | +| `App` + `Arc`: settings, auction plan, orchestrator, integration registry, telemetry sink, default KV store | `EnvConfig` / `RuntimeStoreConfig` | +| | `ClientInfo`, `DeviceSignals`, TLS/JA4 metadata, resolved client IP | +| | `RuntimeServices` (`app.rs:285`) | +| | Correlation id from `get_client_request_id()` | +| | `EcFinalizeState`, `RequestFilterEffects`, auth results, bodies | + +Native store handles are never retained alongside the app. + +Correlation is request-local. `get_client_request_id()` supplies it; +`FASTLY_TRACE_ID` identifies the sandbox, not the request, and is never used as +a request id. Correlation values are never written into app state or into +global logger configuration. + +### Nested-state audit + +Retaining `AppState` is only safe if everything reachable from it is +config-derived and free of request state and native handles. + +- `AuctionOrchestrator` (`orchestrator.rs:307-317`): `enabled`, `plan_backed`, + `plan`, `planned_providers`, `mediator`. All config-derived. The per-auction + `PlannedLaunchState` is a local, not a field. +- `IntegrationRegistry` (`registry.rs:780-783`): `Arc` + plus an optional plan. **Not fully config-derived — see + [Retained rewrite buffers](#retained-rewrite-buffers-blocking).** +- `FastlyTinybirdAuctionTelemetrySink` (`tinybird.rs:36-48`): owned strings and + a backend spec. No handles. +- `default_kv_store` is `UnavailableKvStore`, inert. + +Two process-wide statics outlive a request for the first time under reuse. +Neither is reachable from `AppState`, but both change behaviour: + +- `IP_CIDR_SOURCE_CACHE` (`protection_scope.rs:200`), covered under + [Behavioural tests](#behavioural-tests). +- `MISSING_GEO_WARNING_LOGGED` (`consent/mod.rs:64`), an `AtomicBool` that + becomes log-once-per-sandbox instead of log-once-per-request. Benign and + arguably the intent, but recorded here so a reduced warning count during + validation is not mistaken for a dropped warning. + +A sweep of the core and adapter crates for `static` interior mutability finds +only these two; everything else is immutable `LazyLock` regexes and sets. + +### Retained rewrite buffers (blocking) + +The registry is **not** safe to retain as it stands. Two script rewriters hold +request content in interior-mutable state on the rewriter object itself: + +- `GoogleTagManagerIntegration.accumulated_text: Mutex` + (`google_tag_manager.rs:366`, used at `:1031`) +- `NextJsNextDataRewriter.accumulated_text: Mutex` + (`nextjs/script_rewriter.rs:24`, used at `:77`) + +Both accumulate fragments of an inline `