Skip to content

v0.9.17: tables ledger, oracle integrations, mship improvements - #8853

Open
waleedlatif1 wants to merge 79 commits into
mainfrom
staging
Open

waleedlatif1 wants to merge 79 commits into
mainfrom
staging

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Companion: simstudioai/mothership#656 (merge this first; then re-run sim-sync on #656)

waleedlatif1 and others added 30 commits October 8, 2026 07:51
* feat(integrations): add saved Ramp and Vanta credentials

* fix(integrations): harden credential renewal and connection setup

* fix(integrations): make credential renewal waits abortable

* docs(integrations): document token credential availability registration

* improvement(integrations): refresh Ramp branding and simplify credential forms
Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Bill Leoutsakos <bill@sim.ai>
…rface lint on deploy (#8803)

* fix(workflows): lint unquoted string references in JSON fields and surface lint on deploy

* fix(workflows): lint only the JSON fields the operation sends, and lint deploys as the acting user

* fix(workflows): check JSON editors sent under a canonical param
…th Sim? (#8804)

* feat(library): How Do You Build a Salesforce AI Lead-Scoring Agent With Sim?

* Pi Babysit: address PR #8804 feedback

---------

Co-authored-by: Sim Pi Agent <pi@sim.ai>
* docs(library): update best-ai-agents-support-ticket-triage

* Pi Babysit: address PR #8805 feedback

---------

Co-authored-by: Sim Pi Agent <pi@sim.ai>
… 2026 (#8806)

Co-authored-by: Sim Pi Agent <pi@sim.ai>
…ags (#8808)

* chore(config): retire unused env vars and fully-rolled-out feature flags

* chore(helm): bump chart version for removed TABLE_ROW_TTL value
* feat(oracledb): add Oracle Database integration

* test(oracledb): remove retired block secret fields from masking inventory

---------

Co-authored-by: Waleed Latif <walif6@gmail.com>
* fix(files): support arbitrary uploads and in-document links

* fix(files): normalize previews and avoid oversized retries

* fix(files): bound decoded text for encoded responses
…es a missing organization (#8809)

* fix(billing): acknowledge Stripe webhooks whose subscription references a missing organization instead of retrying for days

* test(billing): prove the not-yet-matching issuance branch by redelivery; drop the logging-only update catch

* test(testing): a swapped price in the Stripe fake no longer reports the previous amount
* fix(db): retire unused keyword search projections

* fix(db): retire keyword writers before schema push

* fix(db): require approval before retiring push writers
* fix(google-ads): align reporting with current API access

* fix(google-ads): preserve nullable inputs and redact retired credentials
* feat(dashboards): add empty-state graphic for the dashboard page

Claude-Session: https://claude.ai/code/session_01J7A6CWpREr1jiPAdWyTQbA

* improvement(dashboards): mark empty-state chart points as const

Claude-Session: https://claude.ai/code/session_01J7A6CWpREr1jiPAdWyTQbA
…table definition row (#8765)

* fix(tables): log row writes to the change log instead of locking the table definition row

* fix(tables): lock the change log before the dev count reconcile, and tighten the held-row test

* test(tables): name the renumbered change-log migration
* fix(tables): drop the per-table row-order lock from inserts

Appends mint keys in a random slot after the last key, so concurrent appends
never share a key and a batch stays contiguous. Positions are left best-effort
and may repeat; the run dispatcher finishes a tied position before advancing
its cursor. Positional inserts skip tied keys. Concurrent replaces stay
serialized by the table's unique lock, which every replace already takes.

* test(tables): order rows bytewise in the lockless-insert tests and stub the job queue locally
…ition source (#8815)

* feat(analytics): attribute sign-ups and demo requests to their acquisition source

- Record first and last marketing touch (campaign params, referring domain, landing path) in consent-gated first-party cookies on marketing and sign-in pages; attach them to user_created and the PostHog person
- Capture $pageview on marketing routes, landing_demo_request_submitted, landing_demo_booked, external_sign_in_started, and email_type on identify
- Add useCaptureWhenReady so view events captured on mount are no longer dropped before PostHog initializes
- Include attribution in the demo-request sales notification
- List the attribution cookies in the cookie policy

* fix(analytics): cap attribution cookies at the consent grant, enforce stored field shapes, and lock sign-in buttons while pending
* feat(changelog): publish curated product updates

* fix(changelog): harden feed rendering and media validation

* improvement(changelog): present updates in a crawlable list

* improvement(changelog): define the Sim editorial workflow

Document scheduled discovery, durable candidate state, draft PR preparation, real-media review, and operational checks. Reuse the existing cached checkout for trusted Helm PR checks while preserving full history.

* fix(changelog): preserve editorial decisions in the workflow plan

* docs(changelog): refine product stories and define editorial quality

* refactor(changelog): simplify presentation and validate model links
… leave-site prompt is answered (#8432)

* fix(desktop-browser): ask the user before a page's alert, confirm, or leave-site prompt is answered

* fix(desktop-browser): ask only about the user's own on-screen page, never the agent's

* fix(desktop-browser): label frame dialogs with their own origin and cover reload leave prompts

* fix(desktop-browser): preserve native dialog ownership and keyboard focus
…IPAA, GDPR, Data Residency, On-Prem, and VPC (#8822)

Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
* fix(mcp): add directory verification compatibility

* test(mcp): verify ownership challenge over HTTP
* fix(chat): send entitlements and agent mode from the public chat API

`/api/v2/chat` (what `sim chat` uses) stopped sending `entitlements` and `mode`
in #8208, so CLI and API chats got no entitlement-gated tools and every CLI
service refused them with "CLI services require agent mode". The route now
computes entitlements per turn like the workspace chat and sends
`mode: 'agent'`. Its test still mocked the old entitlements function and
listed `mode` as a forbidden legacy field; both are updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* feat(tests): workflow tests as a workspace resource

Workflow tests are a workspace resource whose source is a plain vitest file,
`tests/<name>.test.js`, owned by the test (workspace_files.context = 'test').
Sim creates a test's metadata with the tests tool and writes its cases with
the file tools; every write is collected in the sandbox and refused if the
file does not load.

- Runner: test files run in the isolated-vm sandbox against draft or
  deployed workflows. `runWorkflow` executes real runs; `mockBlock`,
  `mockTool` (Agent tool calls) and `spyOnBlock` reach blocks in the tested
  workflow and in every child workflow it runs, matched by name as each
  workflow starts. `.mockSampleOutput()` builds outputs shaped like the real
  block or tool. `toMatchRubric` asks a model judge for pass or fail.
- Runs record live per-case progress, the source hash, and the deployment of
  every workflow they ran, so results show as out of date once the test or
  a workflow changes.
- UI: Tests page and test page (Edit / Split / Preview over the file, the
  preview a dashboard of the selected run), a test resource type in chat,
  and a Tests sidebar entry behind the `workflow-tests` flag.
- Owned files never open as file tabs in chat: only workspace files and
  chat uploads do.
- Migration 0400 adds workflow_test and workflow_test_run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* feat(tests): pick a run from a dropdown and open what each run ran against

The test page shows one run at a time, chosen from a run picker with status
dots and Draft / Out of date chips. Case statuses use the Badge status chip.
Each ran-against entry records one execution, so a draft row opens the
workflow snapshot from that run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* fix(tests): type errors and tests broken by workflow tests

- Narrow the test principal to the kinds workflow_tests.run admits before
  handing it to executeWorkflow.
- Select progress with the latest-run rows, guard file upsert ids in the
  tab filter, and set the sandbox Event polyfills through Reflect.
- Cover the tests tool in the management tool contract, expect content
  writes to reach test files, and stub test availability in the payload test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* fix(tests): address review on redaction, staleness, and the harness

- Redact each run's resolved secrets from what returns to the sandbox
  (output, errors, mocked tool inputs); mocked tools get only declared params.
- Custom blocks no longer receive the consumer's test hooks.
- Draft runs go stale when the draft changes; children a run calls are
  recorded in ran-against.
- Test cases commit in the same transaction as the source file write.
- Harness: runWorkflow is rejected in suite hooks, a timed-out case stops
  the file, and expect.assertions/hasAssertions are supported.
- Insert run rows in one statement and start each run's clock with its file;
  check bans before each workflow run; restrict owned-file access to Copilot
  delegation; validate names in the tool contract.
- Delete soft-deletes the test file and removes the chat tab; a finished run
  shows its own cases; polling at 3s on a separate read bucket; list error
  state; store reset; tests stay in the org Add Resource picker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* fix(tests): reserve run slots, await pending assertions, keep dynamic tool args

- Each test workflow run reserves and releases an execution slot.
- A case waits for assertions it did not await and fails if one fails.
- Mocked MCP and custom tools keep the arguments their schema declares.
- A closed session refuses starts still awaiting their lookups.
- Stable refresh callback; scroll fade on the results pane.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* fix(tests): keep test sources out of file tabs, refresh after runs, reopen tests

- File-edit tool results mark a non-tab file `fileTab: false`, and the browser
  skips promoting it.
- Idle test pages poll every 15s so runs started elsewhere appear; a Mothership
  run returns its tests as resource changes.
- open_resource accepts test resources through an authorized read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* fix(tests): refresh test tabs after a run, keep saved edits successful

- A finished Mothership run refreshes its tests instead of upserting tabs, so a
  test deleted mid-run does not come back.
- A failed file-tab lookup after a saved edit opens no tab instead of
  reporting the edit as failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* fix(tests): starter source imports every test helper

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* build(tests): add @vitest/expect and @vitest/spy for the sandbox bundle

The vitest-expect sandbox bundle builds from these packages; rebuilt with the
Reflect-based event polyfills.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* build(tests): tell knip the sandbox bundle uses @vitest/expect and @vitest/spy

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* feat(tests): name MCP and custom tool mocks by server and title

MCP tool ids embed the server's database id, which changes when a server is
re-added or a workspace is forked, so a stale mock silently stopped matching
and the real server was called. Tests now name workspace tools the way the
workspace does: mockTool({ mcp: 'Server', tool: 'name' }) resolved per run
(failing on an unknown or ambiguous server), and mockTool({ customTool:
'Title' }) matched case- and space-insensitively. Raw mcp- and custom_ ids
are rejected; built-in catalog ids are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* fix(tests): reserve the run name, hold Run for unsaved edits, cancel judges on close

- A test named "run" collided with the static run endpoint, so its detail
  page got a 405; the name is now reserved.
- Run is disabled while the open editor holds edits the server has not
  saved (including a refused save), so a run never uses the previous source.
- toMatchRubric model calls are aborted when the sandbox run ends, so a
  stopped test no longer keeps calling or billing the judge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

* fix(tests): fail a run whose selected case names no test in the file

A renamed or misspelled `only` path skipped every case, and the run was then
saved as passing. The harness now rejects unknown names, so the run is
recorded as an error with the names it could not find.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(providers): remove default agent tool-call iteration cap

* docs(providers): note no execution timeout when billing is disabled
* fix(desktop): ask before accessing local files

* fix(desktop): remember folder permissions across chats

* fix(desktop): cancel pending file consent and recheck access

* fix(desktop): enumerate approved directory descriptors

* fix(desktop): reject replaced directory listings

* fix(desktop): share pending folder consent decisions

* fix(desktop): complete file consent and add full file access

* fix(desktop): revalidate folder consent and preserve exact identities
…ords (#8864)

* fix(files): refuse Chat current-file reads no provenance observer records

* test(files): run Chat file reads through the production CLI transport composition

* test(sandbox): assert the observed file read is admitted, not the recorder call
…and runs (#8874)

The inventory mapped paths back to operations by re-deriving every route, so an operation whose derived path equals a sibling's (undeployWorkflow derives `workflows deploy`) overwrote it. Generated commands now record the operation they run and the inventory reads it.
* fix(pii): stop one half-emoji from failing log redaction

A lone UTF-16 surrogate (e.g. an emoji cut by display truncation) made
Presidio fail encoding its /redact_batch response, so mask-batch answered
500, each chunk was retried 8 times, and the default scrub mode replaced
every string in the execution log with [REDACTION_FAILED].

- validate_pii: every string sent to Presidio goes out well-formed
  (toWellFormed, length-preserving U+FFFD); Presidio 4xx (other than
  408/429) becomes PiiServiceRejectedError, unreachable / 502-504 becomes
  PiiServiceUnavailableError
- mask-batch route: rejected -> 422 (never resent), unavailable -> 503
- mask client: optional failedChunkPlaceholder scrubs only the failed
  chunk and reports scrubbedCount; an outage (connection error or
  502/503/504) stops sending queued chunks, a 500 does not
- pii-redaction: scrub mode uses it and logs scrubbedCount, so one bad
  chunk no longer erases the whole payload; throw mode is unchanged
- display-filters: truncate on a code-point boundary with an exact count
- Presidio: any unencodable string in a request body is a 422, and 422
  bodies no longer echo request input

* test(pii): assert the outage breaker through masked output, not request counts

* refactor(pii): drop unused error status, exports, and dead 422 escaping
…paging old ones (#8869)

A billing Stripe sync that exhausted its retries stayed dead-lettered forever:
requeue only sends it back through the same failure, retention never prunes a
dead letter, and the hourly billing reconcile cron logged every one at ERROR
on every run.

- POST /api/v1/admin/outbox/[id]/resolve closes a dead-lettered billing Stripe
  event (dead_letter -> completed). last_error keeps who resolved it, why, and
  the last failure; the action is audited. Other domains count completed rows
  as delivered, so only billing Stripe events can be resolved.
- The reconcile cron logs dead letters from the last two hourly runs at ERROR
  and older unresolved ones at WARN with the same identifiers; the dead-letter
  scan returns newest first so its cap never hides a new one.
- The admin dashboard reports a resolved cancellation as resolved, not
  applied.
…s behind it (#8871)

* fix(serializer): skip tool selection for trigger-mode blocks

Trigger-mode blocks run through TriggerBlockHandler and never read a tool id,
while their tool-mode sub-blocks (e.g. operation) are not serialized, so
selecting a tool always threw and logged a fallback warning.

* improvement(executor): memoize block output schemas per serialized block

collectBlockData re-derived every block's output schema on each condition,
function, agent, and reference evaluation, re-parsing an invalid agent
response format and logging its full stack every time. Memoize the schema per
immutable serialized block, attribute parse failures to the real block id, and
log them at debug without the stack.

* fix(webhooks): stop warning on accountless webhook credentials

Service-account and managed-OAuth credentials (e.g. a custom Slack bot) have no
OAuth account row, so the account-owner lookup always missed and warned on
every webhook run. Skip the lookup for them and warn only when the credential
or its account is genuinely missing.

* improvement(webhooks): log Slack reactions.get missing_scope once per webhook

A Slack bot without reactions:read fails reactions.get identically for every
reaction event until it is reinstalled. Record that configuration state once
per webhook per hour at info and keep the warning for other failures.

* improvement(workflows): repair drifted subBlock types quietly

Deployed snapshots are immutable and re-sanitized on every load, so a field
whose declared type later changed (e.g. dropdown to combobox) warned on every
materialization forever. Log repairs at debug when the stored type is one some
registered block declares; keep warning for missing or unknown types.

* improvement(catalog): skip custom block input derivation for tool reads

Listing, reading, and executing a catalog tool resolved the catalog gate with
every custom block's input fields, which loads and sanitizes each custom
block's deployed workflow per request. Tool scope only needs the block types
(every custom block exposes workflow_executor), so tool reads use the
lightweight overlay rows; block reads still derive the inputs.

* improvement(webhooks): debug-log missing idempotency ids for providers without one

A generic webhook without an idempotency header or configured field has no
delivery identifier by design, so warning on every delivery was noise. Warn
only when the provider normally extracts an id from the body.

* fix(flint): declare generate_pages items as an array param

items was typed json, so the LLM schema advertised it as an object and
dropped its item schema with a warning on every schema build. Declare it as an
array of 1-10 page objects; block-provided JSON strings still parse through
parseItems. Regenerate tool metadata and docs.

* improvement(executor): skip JSON parsing for json inputs holding plain strings

json-typed block inputs also accept plain strings (a file id, URL, or
reference) that are kept as-is, so parsing them could only fail and warn.
Parse only text that can be JSON; every valid JSON text still parses.

* improvement(mothership): classify catalog routes and dedupe missing-provenance warnings

Catalog responses (blocks, tools, connector types) carry no workspace data, so
no producer records provenance for them and the per-request warning was
noise. Skip them statically, and report a data-bearing route without a
producer once per route per hour instead of on every request. Admission and
recording behavior is unchanged.

* fix(tokenization): resolve tiktoken encodings by model family

encodingForModel only knows exact OpenAI names, so newer or provider-prefixed
OpenAI ids fell back to cl100k_base (miscounting o200k models) and every
unknown model built and cached its own copy of the rank table. Resolve exact
names through js-tiktoken, newer OpenAI ids by family, everything else to
cl100k_base, and cache one instance per encoding.

* improvement(proxy): log blocked scanner requests at debug

Every blocked request is an empty or tool user agent probing paths like
/admin.php; the 403 is the whole response, so the per-request warning was
pure ingest cost.

* improvement(telemetry): downgrade best-effort collector failures to warn

Forwarding to the external collector is best effort and the route still
returns success, so a collector timeout or error is not an application error.

* fix(admin): log mothership admin proxy failures and map upstream 5xx to 502

The proxy logged nothing, passed an upstream 5xx straight through as its own
500, threw on a non-JSON upstream body, and returned a silent 500 when the
admin key was missing. Log the method, environment, endpoint, status, and a
truncated upstream body; answer upstream 5xx and unparseable bodies with 502;
and log the missing-key misconfiguration.

* improvement(logs): move per-call success chatter to debug

PII batch masking, function execution request/success, sandbox mount
resolution, table row queries and limits, usage-limit statistics, DAG builds,
and mothership tool routing logged an info line on every call. They are
diagnostic detail, not events, so log them at debug.

* chore(audits): shrink explicit-any baseline after tokenization cleanup

* improvement(mothership): keep tool call routing decision at info

It is the diagnostic for stale-catalog 'No handler for tool' dispatches and
in-band double-dispatch races, so it stays visible in production.

* fix(admin): keep upstream message and empty 2xx bodies in the mothership proxy

A gateway 502 now carries the upstream's error or message field, which the
admin UI shows, and an empty 2xx body (e.g. 204) passes through as an empty
response instead of failing JSON serialization into a misleading 502.

* improvement(tokenization): memoize model to encoding name resolution

A non-OpenAI model threw and caught inside getEncodingNameForModel on every
count. Cache the resolved encoding name per model id in a bounded LRU.

* improvement(mothership): resolve catalog route patterns through the route table

Derive each catalog pattern by matching its contract path against the
generated v2 route table, so a renamed path parameter cannot drift from the
pattern matchV2Route reports and silently restore the warning.

* fix(webhooks): warn on missing idempotency ids unless the provider declares them optional

Header-based providers (GitHub, GitLab, Shopify, Linear, Svix) have no body
extractor, so keying the warning on the extractor silenced exactly the
anomalous case of their delivery header going missing. Providers now declare
deliveryIdOptional; only the generic webhook does, since its deliveries carry
no identifier unless an idempotency field is configured.

* refactor(logs): simplify log-noise fixes after review

- Skip reactions.get entirely while a webhook's bot is known to lack the
  scope, instead of repeating a call known to fail.
- Treat any intact sub-block whose stored type differs from the registry as
  drift, dropping the registry-wide type scan.
- Move the JSON-candidate check next to isJSONString and trim once.
- Make the catalog gate's cheap overlay rows the default; only block reads
  derive custom block inputs.
- Collapse the admin mothership proxy's three copied handlers into one.

* chore(logs): tighten comments and rename trigger-mode serializer test

* test(mothership): guard sandbox catalog route classification against the route table

Extract the catalog route set into isCatalogRoute, built lazily from the catalog
contracts resolved through the generated v2 route table, and test it against
real request paths so a route or parameter rename cannot silently restore the
missing-provenance warning.

* fix(webhooks): keep calling reactions.get after a missing-scope report

Skipping the call for the cooldown window left a reinstalled Slack bot with
empty reaction text for up to an hour. Keep the once-per-window log but always
make the call. Block listing also uses the lightweight custom-block rows, since
block summaries never read custom block inputs.

* chore(logs): tighten log-noise changes after a line-by-line audit

- Drop the blockId option threaded through getEffectiveBlockOutputs; it only
  labeled a debug log.
- Name the log-dedupe windows and cache ceilings.
- Build the sandbox catalog route set by converting contract paths to the
  route table's pattern syntax; the route-table test guards drift.
…8866)

* improvement(changelog): open recordings in the shared lightbox

* improvement(changelog): refresh product captures and released updates

* improvement(changelog): write concise product announcements

* docs(changelog): align the editorial body limit
… reconnect (#8868)

* fix(oauth): record revoked refresh tokens durably and surface them as reconnect

A refresh token the provider revoked (invalid_grant and kin) was only
flagged in Redis for an hour, so every scheduled run kept failing with a
generic error logged at ERROR three times, and the provider was asked
again each hour, indefinitely.

- account gains refresh_revoked_at/_code/_token_hash (expand-only,
  nullable). The record holds only while its hash fingerprints the stored
  refresh token, so a writer storing a new chain supersedes it; every
  reconnect path also clears it explicitly, since a provider can
  reauthorize without issuing a new refresh token.
- The record is written only on rows still holding the rejected token,
  so a refresh that lost a race to a newer chain cannot mark the live
  credential revoked. Revocations no longer use the Redis dead flag,
  which now holds only app-registration faults.
- A recorded revocation answers without the provider; one re-probe a day
  lets rejections that clear without a reconnect recover. A successful
  refresh clears it; a re-probe that fails for a passing reason keeps it.
- refreshTokenIfNeeded throws CredentialRevokedError; token resolution
  returns code OAUTH_CREDENTIAL_REVOKED (POST and GET) and logs it once at
  WARN; the executor, tool and connector sites no longer log it at ERROR.

* fix(oauth): guard revocation records against in-flight reconnects

- Record a revocation only when the row's updated_at has not moved since
  the leader read it, so a reconnect that keeps the same refresh token
  while a rejected refresh is in flight is not undone.
- Reread the row under the refresh lock, so a revocation another process
  just recorded is honored instead of calling the provider again.
- A re-probe whose transient failure meets a moved chain reports no
  revocation.
- Slack fan-out clears the record only for a chain the provider just
  issued, not when re-spreading a stored one.
- A reconnect whose cleanup fails now fails the callback.
- Revocation rejections log at WARN in oauth.ts too; the pure refresh
  error classifiers move to refresh-error-codes.ts so that client-reachable
  module can use them.
- The connector path leaves the revocation log to executeSync.

* fix(oauth): keep Slack connect fan-out unchanged and skip the cross-process check without Redis

* test(oauth): clear only the fixture's Redis flags in the revocation suite

* fix(oauth): clear only the recorded revocation on in-place relinks, with no Redis on the sign-in path
…rites (#8870)

* fix(execution): tag deterministic admission rejections and throttle blocked-run logs

Usage-limit, suspended-account, and missing-billing-account refusals now carry
stable codes, so surfaces can tell a refusal that holds until a person acts
from a transient one.

Unattended surfaces can opt into throttleErrorLogs: each refusal of a workflow
by the same gate records at most one execution error row per 15-minute window
(Redis SET NX, in-process LRU without Redis, fail-open on Redis errors). A sender
resending a refused delivery no longer writes an execution row, trace archive,
and workspace_files row per attempt. A caller-supplied logging session is
always completed.

* fix(webhooks): acknowledge deterministic admission rejections for Telegram and Slack

Telegram resends a non-2xx update until it succeeds or 24 hours pass, and Slack
disables an app's event subscription once most deliveries fail, so answering a
usage-limit refusal with 402 only loops. Providers opt in with
acknowledgeAdmissionRejections and get an empty 200 with an ignored outcome;
transient refusals (rate limit, concurrency, reservation outage) still fail so
the sender retries. Generic webhooks keep the 402. Polling receives the raw
refusal with its code. Webhook preprocessing throttles its error rows.

* fix(webhooks): skip polls for over-limit payers and back off failing sources

The poll orchestrator checks each workspace's payer once per tick and skips its
webhooks while the payer is over its usage limit: nothing is fetched, marked
seen, or counted as a failure, so items deliver once the payer is back under
the limit and triggers are no longer auto-disabled over billing state.

A webhook with consecutive poll failures waits 2^(n-1) minutes (capped at an
hour) before the next fetch, and a source's Retry-After or FLOOD_WAIT_<n> is
persisted and honored. RSS logs a source's 4xx once at warn, and an admission
refusal mid-poll leaves the remaining items unseen.

* fix(webhooks): stop Slack redelivering to deleted trigger paths and log them once

A POST to a path with no webhook keeps its 404 but now carries
x-slack-no-retry, sent unconditionally so it reveals nothing the 404 does not.
The processor and route lines for an unknown path drop to debug, leaving the
route handler's single client-error line.

* fix(telegram): verify webhook deliveries with a per-webhook secret token

Register a generated secret_token with setWebhook, store it as
providerConfig.secretToken, and reject deliveries whose
X-Telegram-Bot-Api-Secret-Token header does not match (401). Webhooks
registered before this change have no stored secret and stay accepted
until their next deploy registers one. A new registration reuses the
active deployment's secret for the same bot so deliveries keep verifying
during cutover. Drops the stale empty User-Agent warning: the proxy
exempts webhook trigger routes from the empty-UA block.

* refactor(webhooks): declare the Telegram admission opt-in on its handler

* fix(execution): only treat billing and account refusals as deterministic

An unreadable usage ledger fails closed as exceeded; it said nothing about the
payer, yet it was tagged USAGE_LIMIT_EXCEEDED and acknowledged-and-dropped for
Telegram and Slack. It now stays untagged and retryable. Reservation headroom
denials clear as in-flight runs settle, so they leave the deterministic set too.

The blocked-run log claim drops its in-process fallback: usage refusals only
happen on hosted billing deployments, which run Redis, and without Redis every
refusal records its row. A Redis failure logs at debug.

* fix(webhooks): keep a fan-out target's retryable failure visible past a dropped refusal

An acknowledged admission refusal counted as an acknowledgment, so a Slack or
path fan-out answered 200 even when another target failed and needed the sender
to retry. Dropped targets (block missing, acknowledged refusal) now answer 200
only when no other target failed.

* fix(telegram): match the active bot through env-var token references

Active rows store the bot token as authored, often a {{VAR}} reference, while
subscription calls receive it resolved, so the comparison never matched: every
deploy minted a fresh secret (a candidate that never activated left the bot
rejected by the active row), and retiring an old version could delete the
webhook the active version still used. Stored tokens are now resolved against
the background webhook env before comparing.

* fix(webhooks): skip polls only after a recorded refusal and back off source failures

The per-tick payer pre-check read billing attribution and the usage ledger for
every polled workspace, including healthy idle ones. A deterministic admission
refusal of a polled event now records the workspace in Redis for five minutes,
and the orchestrator skips only those workspaces in one MGET; healthy payers
cost no billing reads, and billing-disabled deployments skip the mechanism.

Backoff now follows source fetch failures only, tracked in providerConfig and
stamped from the failed poll's start, so transient item-processing refusals
never back a webhook off. Poll state keys are system-managed so deploy change
detection ignores them.

* refactor(webhooks): route every poller's source failures through one backoff

Every poller's outer catch now records a source failure, so a failing Gmail,
Outlook, IMAP, Drive, Sheets, Calendar, or HubSpot source backs off like RSS
instead of only RSS. markWebhookSuccess clears the backoff in its existing
reset write, the window uses the shared jittered backoff, and the orchestrator
asks a boolean isPollBackedOff.

Smaller cleanups: one isDroppedDispatch predicate for the Slack and path
fan-outs, explicit precedence for a polled refusal's code, a typed RSS refusal
error instead of a flag, Telegram resolves only the stored bot token, and
PollOutcome lives with the polling types.

* chore(webhooks): tighten poll comments and backoff tests

Poll outcome docs describe what a skipped poll actually does, source failures
keep logging the full error object as the pollers did before, the backoff table
pins the clock past the poll start so it proves the window is anchored there,
and the RSS rate-limit test drives a Retry-After header.

* fix(webhooks): stop every poller's batch on a deterministic admission refusal

Only RSS stopped at a refused item; Gmail, Outlook, and IMAP advanced their
cursors past refused emails, and every poller counted the refusals toward
auto-disable. A shared PollAdmissionRefusedError now leaves the idempotency
callback, stops the batch, and returns skipped before any cursor update or
failure count; items that already ran replay as idempotent no-ops.

Source backoff goes back to RSS only, where the rate-limited feed was: the other
pollers' fetch helpers do not carry status or Retry-After, so routing their
failures through it would back off on a guess. The block-missing 404 also tells
Slack not to redeliver.

* fix(webhooks): never replay completed poll work or mask a retryable fan-out failure

A poller stops its batch on a deterministic refusal only while nothing in the
batch has completed; once an item has run, the refusal is an ordinary item
failure, so the poller saves its completed work exactly as before and no
completed event can replay after the idempotency window.

In a multi-target delivery a missing block's no-retry 404 no longer stands in
for a target that failed and needs the sender to retry. A source's Retry-After
is counted from its answer rather than the poll's start. The RSS backoff keys
are cleared by RSS's own state write, so other pollers' success path is
unchanged, and two fields nothing reads are dropped.

* fix(execution): drop the unreachable billing-account admission code

A workspace without a billing account fails inside payer resolution and takes
the retryable attribution-error path; the branch that tagged
BILLING_ACCOUNT_REQUIRED only ran for an attribution with no actor, which
system attribution never produces. The branch goes back to its staging form
and the deterministic set keeps the usage limit and suspended accounts.

* chore(webhooks): key blocked-run claims by gate and centralize the polling utils mock

Two gates that fail without a code (a ban lookup error and a usage lookup
error) no longer share one throttle claim, so neither hides the other's row.
The polling utils module gets one central mock in @sim/testing, replacing the
partial importOriginal mocks and the hand-rolled factory in the table trigger
test.

* fix(telegram): match the active bot with the env the caller resolved its token with

Subscription creation resolves the incoming bot token with the deployer's env,
while cleanup resolves with the background env; the active-row matcher now
uses the same env as its caller, so a bot referenced through a personal
variable still reuses the active secret and is not deleted from under the
active deployment.

Tests: the idempotency service gets one central mock in @sim/testing, used by
every test that mocked it locally, and the new tests import single factory and
mock files instead of the @sim/testing barrel.

* chore(testing): stub IdempotencyService.createWebhookIdempotencyKey in the central mock
…ons they declare (#8877)

* improvement(sim-cli): commands reach the API only through the operations they declare

Hand-written commands called SimClient by path and declared nothing, so the command inventory could not mark one Mothership-unavailable when it calls a refused route. Every command now gets its client from callsOperations: calls name an operation, route and method come from the operation table, and the client is typed to the declared operations.

* test(sim-cli): assert declared-operation routing at the transport

* test(sim-cli): assert directory requests at the transport
…dary (#8872)

* improvement(logs): log each execution failure once at its owning boundary

A single external tool failure was logged at ERROR up to eight times as it
propagated from the tool layer through the block executor, the engine,
execution-core, and the trigger surfaces. Each boundary now calls
logFailureOnce, which skips a failure an inner boundary already logged and
sets severity by attribution: the author's input, configuration, or code at
info, a third-party 4xx or 5xx at warn, and Sim's own faults (database,
retryable setup, hosted-key rejection, anything unattributed) at error.

The logged mark and attribution live in WeakMap/WeakSet side tables keyed by
the thrown value, walked through the cause chain, and carried on the
flattened tool failure's output so they survive a handler rebuilding a failed
tool result as a new error. Run logs and trace spans are unchanged.

* test(logs): pin the logged mark and attribution at the block executor

Asserts the block executor's thrown error is already logged and attributed
(database failure internal, user Stop user), switches the tool-boundary test
to the central jsonResponse helper, and updates exact-argument log
assertions for the added failureKind and executionId fields.

* fix(logs): scope the logged mark to one propagation and attribute in-process failures

Never mark a raw thrown value as logged: a persistent fault that rethrows one
object (a rejected dynamic import, a memoized rejected promise) was logged once
per process and then silenced everywhere. Only carriers created during the
propagation (the flattened tool failure output, the block error, a handler's
rebuilt error) carry the mark, and execution-level boundaries scope a raw
value's mark to their execution id.

Handlers that rebuild a failed tool result now call adoptToolFailure instead of
relying on the result output being copied by reference. The child workflow and
agent handlers no longer log failures the block executor logs with the block's
and run's identity. In-process operations are attributed to Sim (4xx user, 5xx
internal) rather than to a third party, only external HTTP failures log their
response body, a missing required field is the only serializer refusal at info,
and the size-limit and JSON-parse paths no longer log twice.

* refactor(logs): build failure log metadata lazily and attribute author failures by type

logFailureOnce takes its metadata as a thunk so a skipped boundary does no
secret projection, adds the execution id itself, and returns the attribution
it logged with. UserFailure attributes author-caused failures by type
(missing required fields, boundary-safe custom block refusals, no start
block) wherever they surface, replacing a try/catch at one serializer caller.
One inheritFailureMarks carries both marks across a boundary that drops
cause, the hosted-key check reuses classifyHostedKeyFailure, the Pi backends
share one toolResultError, and the sandbox and Function route no longer log
author code failures the tool boundary already logs.

* test(logs): attribute the serializer refusal at its owner and pin block log redaction

Moves the missing-required-fields attribution check to the serializer that
throws it, drops an execution-core case whose internal row passed by default,
and covers the block failure line's secret projection now that it, not the
Agent handler, logs provider errors. Tightens comments the diff added.

* improvement(logs): drop casts and test-only exports from the failure log

Reads status and cause without casts, keeps the thrown missing-required-fields
error's name and stack unchanged, ignores an empty execution id when scoping a
logged mark, unexports wasFailureLogged and FailureKind (tests observe the
outer boundary through logFailureOnce instead), and shrinks the explicit-any
baseline for the Agent handler.

* fix(logs): close the remaining double logs found in review

Carries failure marks through the scrubbed Pi error and normalizes a
non-Error node failure before logging so the engine sees its mark; drops the
condition handler's and response-size handler's own error lines in favor of
the owning boundary; logs MCP failures once and marks their output; passes
the block id to the API and Function tool calls so the tool line carries it;
attributes custom tool parameter validation and condition expression errors
to the author; checks the whole cause chain for a retryable setup failure
before honoring a mark; and restores the sandbox failure logs, whose adapters
cannot yet tell a provider exception from a non-zero exit.

* fix(logs): keep run identity on every remaining failure line

Fixes the type error on the external failure body, attributes a provider
error payload on a 2xx to the provider, and adds workflow, execution, and
block ids to the MCP failure lines that replace the block executor's. The
condition batch carries its failed result's marks and passes its block id.
Pi GitHub calls carry no run identity, so their errors now carry only the
tool layer's attribution and the block executor still logs them with ids.
Restores the HITL notification warning, which covers soft failures the tool
layer never logs, and attributes refused tool and proxy URLs to the author.

* fix(logs): attribute a revoked OAuth credential to its owner

A CredentialRevokedError anywhere in the cause chain classifies as a user
failure, so the boundaries that log it after token resolution's WARN write
INFO rather than ERROR.

* fix(logs): keep one logged mark per execution and attribute unparseable responses

A raw value logged at an execution boundary now remembers every execution
that logged it (bounded), so overlapping runs sharing one persistent
rejection each log it once. A 2xx body that is not JSON is the endpoint's
failure, not Sim's.
* improvement(app): show branded wordmark while loading

* fix(app): preload workspace data alongside branding
…data-typed next/script (#8875)

* fix(docs): render the FAQ JSON-LD as a native script tag, and forbid data-typed next/script

The FAQ component rendered its FAQPage JSON-LD through next/script with the
default afterInteractive strategy, which returns null on the server and injects
the tag from a client effect, so the served HTML never carried it. Render a
native <script> the way #6763 fixed the other docs JSON-LD blocks;
serializeJsonLd already escapes '<'.

check:source-text now also fails on a next/script element in apps/** whose type
is not JavaScript, naming the native-script fix.

* fix(audits): match only the real type prop, fail closed on expression types, and scan MDX pages

* fix(audits): recognize default-as next/script imports and skip MDX code examples

* fix(audits): keep JSX template literals intact when skipping MDX inline code
…ode; fix the PPTX blend-group call they hid (#8876)

* chore(audits): drop dead ESLint directives and the redundant check:dead-code script

- Remove all 89 `eslint-disable` comments. Nothing runs ESLint (no eslint dependency or
  config in any workspace; Next 16 `next build` has no lint step), so they suppressed
  nothing. Reasons they carried are kept as plain `//` whys.
- `check:comment-hygiene` now rejects `eslint-disable`/`eslint-enable` comments so they
  cannot return.
- Remove `check:dead-code`: `check:unused-exports` already runs knip with the same config
  and a superset of its issue types, and `check:audits` skipped the alias.
- Ignore unused exports in the generated `apps/docs/components/icons.tsx` (a verbatim copy
  of the sim icon set, whose export surface is ratcheted at the source) and shrink the
  unused-exports baseline by its 66 entries.

* fix(pptx): stop calling the appended rect blend group as a function

* docs(skills): name both ESLint directives the comment audit rejects

* docs: name both ESLint directives in the comment rule

* test(pptx): cover the rect gradient blend-group path
…adline (#8878)

* fix(tools): stop provider timeout params from becoming the request deadline

The request transport and the internal-operation path read params.timeout as a
millisecond deadline for every tool. Twilio make_call, New Relic NRQL, Apify
(3 tools), Daytona (2 tools), and Trigger.dev waitpoint tokens declare their own
timeout param in seconds or as a duration, so a 60-second setting aborted the
call after 60 ms. A declared timeout param is now the deadline only when the tool
sets timeoutParamIsDeadline (http_request, firecrawl_map, firecrawl_parse);
callers can still bound tools that declare none. No param ids change.

Redis and Upstash coerced params with Number() inside tools.config.tool, which
runs at serialization on the serialized params object, turning <Block.output>
references into NaN. The coercions now run in tools.config.params.

Guardrails: check-block-registry rejects coerced assignments to params inside an
inline tools.config.tool; check-tool-param-reachability rejects a method param on
a fixed-verb external tool, which the transport would send as the HTTP verb.

* fix(audits): catch coercions under fallbacks in tools.config.tool, and clarify declared timeout params

* fix(audits): treat every value-deriving expression as a selector coercion

* fix(audits): reject compound assignments and increments on params in selectors
…through wrappers (#8885)

- v1 admin: unauthorized/forbidden/notFound/badRequest/internalErrorResponse -> admin*Response; errorResponse module-local; conflict/notConfiguredResponse also admin-prefixed
- files createErrorResponse -> createFileErrorResponse; workflows createErrorResponse -> createCodedErrorResponse (+ central mock)
- import VFS segment codecs from @/lib/vfs/path; delete the mothership pass-throughs and the dead canonicalizeVfsPath
- search-replace resource group key uses sortObjectKeysDeep instead of a localeCompare serializer
… discovers them (#8882)

* chore(audits): name generated-artifact checks check:<x> so run-audits discovers them

Rename tool-metadata, deployment-config, integration-catalog, docs, agent-stream-docs
and docs-manifest checks to check:<x> (generators to generate:<x>, matching the
existing generate:openapi/check:openapi pairs). run-audits now derives every audit
from the check:* prefix, so the hand-kept EXTRA_AUDITS list is gone, and
check:docs-manifest (local fs only, ~1s) joins check:audits instead of its own CI
step and CLAUDE.md gate line. The <x>:check suffix stays only for checks this job
cannot run (sibling copilot repo, Helm, per-machine covers).

Add module headers to check-monorepo-boundaries, check-realtime-prune-graph and
check-openapi saying what each protects and how to fix a finding.

* docs(ship): regenerate tool metadata in Phase A; drop the actionlint comment with no command

This branch was previously deployed

1 inactive deployment
Preview — 9d8e4d81 Deployed Oct 10, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants