Repository navigation
Wip/working tree 2026 06 18 - #3
Merged
Merged
Conversation
Counterpart of NULLRUN fix(ws-control) (commit 5e2f65b). The
backend now embeds the exact bytes that were HMAC-signed in a
separate signed_payload field. The SDK:
1. Verifies the signature against bytes.fromhex(signed_payload),
falling back to the legacy wire-bytes path only when the
field is absent (pre-FIX-C servers).
2. Dispatches state changes from the parsed signed_payload
bytes, not from the outer envelope body. This closes a
security hole: an attacker who captured a (signed_payload,
signature) pair from a benign 'state=Normal' event could
otherwise splice a forged 'state=Killed' into the outer body
and the signature would still verify, because the signature
covers only the signed_payload bytes. Reading dispatch state
from the trusted source keeps the captured signature
semantically bound to its captured body.
Tests in test_ws_signed_payload.py cover:
- round-trip, wrong-secret, tampered-payload rejection
- malformed signed_payload does not crash
- replay-with-spliced-body: signature still verifies, but the
dispatched state is the captured one (not the forged one) -
the attack is harmless
- replays where the attacker also rewrites signed_payload are
rejected via signature mismatch
Note: the two ACK tests are still failing because
ACKNOWLEDGED_STATES is still lowercase. That is fixed separately
by S-2 in the same release - kept as a separate commit so the
byte-mismatch/security fix is reviewable on its own.
The server's WsWorkflowState enum (NULLRUN/backend/src/proxy/http/
ws_control.rs) emits 'Killed' / 'Paused' (PascalCase). The SDK was
comparing against {'killed', 'paused'} (lowercase), so the ACK path
was dead and the server's pending-ack queue grew without ever
being drained.
This unblocks the two remaining failing tests in
test_ws_signed_payload.py:
- test_state_change_with_signed_payload_is_dispatched (now sends
the ACK that the server expects)
- test_acknowledged_states_use_pascalcase (now matches server
casing)
With byte-mismatch FIX-C in place (commits 5e2f65b + 105fb80), the
KILL/PAUSE path now works end-to-end:
1. server signs the inner message and embeds the bytes in
signed_payload
2. server sends the envelope (flattened WsMessage + signature +
timestamp + api_key_id + signed_payload)
3. SDK verifies signature against bytes.fromhex(signed_payload)
4. SDK dispatches from the trusted source (parsed signed_payload),
so a captured (signed_payload, signature) pair can only
re-trigger its captured state, never a forged one
5. SDK sends ACK on Killed/Paused, draining server's pending-acks
The working tree contained a large uncommitted changeset that was
never pushed: 68 files, +8955/-3328 lines. Reading the diff shape
this is the 0.3.0 -> 0.4.0 production-readiness migration
(per CHANGELOG.md / audit §6.1):
- PoolConfig / AdaptivePool removed (Transport now is a
context manager; weakref.finalize replaces atexit.register)
- gRPC transport removed (NULLRUN_USE_GRPC no-op; create_grpc_transport
was a NameError)
- signal.signal global hijack removed
- track.proto removed
- decision_history / flow / gate / common placeholders removed
- six zombie exceptions removed (CostLimitExceeded,
ApprovalRequired, BreakerTimeout, LoopDetectedException,
RetryStormException, RateLimitExceededException)
- _organization_id_var, _api_key_id_var removed
- patch_openai / unpatch_openai removed
- auto-instrumentation extended with langgraph / llama-index /
crewai / autogen / openai-agents via safe_patch
- SENSITIVE_ARG_KEYS expanded from 7 to 29 tokens
- HMAC always-on for /track/batch, /gate, /evaluate, /status,
/auth/verify + WS ACKs signed
- 14 new test files
- analyze.md (this session's plan)
Tracking as a wip branch so the work is preserved. This commit does
not change the byte-mismatch FIX-C landing in
fix/ws-byte-mismatch-verify-signed-payload (commits 105fb80,
73f3197) - those branches are based on 316a694 + the byte-mismatch
fixes only.
maltsev-dev
added a commit
that referenced
this pull request
Jul 3, 2026
Task #3 / Task #18 (2026-07-03): wire the SDK to the backend's v3 default. Per CLAUDE.md §24, every /check mints a server-side uuidv7 execution_id; the SDK receives it in the response and propagates it to /track. This is the SDK_MIN_VERSION for the v3 rollout per CLAUDE.md §0 pre-flip checklist. Changes: src/nullrun/uuid7.py (new): - RFC 9562 §5.7 time-ordered ID generator. 48-bit unix_ts_ms prefix + 12-bit rand_a + 62-bit rand_b. Same layout as the backend's mint_execution_id() so log scrapers can sort by ID alone. - Uses secrets.token_bytes(10) for cryptographically secure random component. - uuid7() returns stdlib UUID; uuid7_str() returns the canonical 36-char string. src/nullrun/capabilities.py (new): - ServerCapabilities dataclass mirrors /health payload. - is_v3_ready() returns True only when ALL three v3 caps (server_minted_execution_id, per_execution_reservations, heartbeat_time_based) are set. - probe_capabilities(api_url) — best-effort /health fetch with 2s timeout. Returns None on failure (not fatal). - validate_sdk_version(sdk_version, caps) — returns warnings for SDK_MIN_VERSION mismatch. - SDK_MIN_VERSION_FOR_V3 = '0.12.0' is the gate's coordinate for the v3 rollout. src/nullrun/__init__.py: - init() now probes /health after singleton registration and logs a startup warning for version mismatch (does NOT fail init() — the gate still rejects with PROTOCOL_TOO_OLD). - Probe is best-effort: timeout/5xx logs at INFO. src/nullrun/__version__.py: - Bumped 0.11.0 → 0.12.0 (the SDK_MIN_VERSION coordinate). CHANGELOG.md: - New 0.12.0 entry with Added/Changed sections. tests/test_uuid7.py (new): 8 tests pin the wire contract: - Returns stdlib UUID - 36-char string format - Version bits = 7 - Variant bits = 0b10 - Time-ordered (consecutive calls sort) - 1000 unique IDs under rapid calls - Round-trips through uuid.UUID() tests/test_capabilities.py (new): 9 tests pin: - v3-ready backend parses to is_v3_ready()=True - Missing keys default to False (fail-closed) - Partial v3 caps → not ready - Old SDK against v3 backend → warning - Current SDK → no warning - Legacy backend → 'not v3-ready' warning - Unparseable versions don't crash - as_dict() is wire-safe (no secrets) - SDK_MIN_VERSION_FOR_V3 = '0.12.0' Tests: 17 new SDK tests pass. Full backend test suite still green at 1443.
maltsev-dev
added a commit
that referenced
this pull request
Jul 3, 2026
* feat(sdk): server-minted execution_id default — uuid7 + capability probe Task #3 / Task #18 (2026-07-03): wire the SDK to the backend's v3 default. Per CLAUDE.md §24, every /check mints a server-side uuidv7 execution_id; the SDK receives it in the response and propagates it to /track. This is the SDK_MIN_VERSION for the v3 rollout per CLAUDE.md §0 pre-flip checklist. Changes: src/nullrun/uuid7.py (new): - RFC 9562 §5.7 time-ordered ID generator. 48-bit unix_ts_ms prefix + 12-bit rand_a + 62-bit rand_b. Same layout as the backend's mint_execution_id() so log scrapers can sort by ID alone. - Uses secrets.token_bytes(10) for cryptographically secure random component. - uuid7() returns stdlib UUID; uuid7_str() returns the canonical 36-char string. src/nullrun/capabilities.py (new): - ServerCapabilities dataclass mirrors /health payload. - is_v3_ready() returns True only when ALL three v3 caps (server_minted_execution_id, per_execution_reservations, heartbeat_time_based) are set. - probe_capabilities(api_url) — best-effort /health fetch with 2s timeout. Returns None on failure (not fatal). - validate_sdk_version(sdk_version, caps) — returns warnings for SDK_MIN_VERSION mismatch. - SDK_MIN_VERSION_FOR_V3 = '0.12.0' is the gate's coordinate for the v3 rollout. src/nullrun/__init__.py: - init() now probes /health after singleton registration and logs a startup warning for version mismatch (does NOT fail init() — the gate still rejects with PROTOCOL_TOO_OLD). - Probe is best-effort: timeout/5xx logs at INFO. src/nullrun/__version__.py: - Bumped 0.11.0 → 0.12.0 (the SDK_MIN_VERSION coordinate). CHANGELOG.md: - New 0.12.0 entry with Added/Changed sections. tests/test_uuid7.py (new): 8 tests pin the wire contract: - Returns stdlib UUID - 36-char string format - Version bits = 7 - Variant bits = 0b10 - Time-ordered (consecutive calls sort) - 1000 unique IDs under rapid calls - Round-trips through uuid.UUID() tests/test_capabilities.py (new): 9 tests pin: - v3-ready backend parses to is_v3_ready()=True - Missing keys default to False (fail-closed) - Partial v3 caps → not ready - Old SDK against v3 backend → warning - Current SDK → no warning - Legacy backend → 'not v3-ready' warning - Unparseable versions don't crash - as_dict() is wire-safe (no secrets) - SDK_MIN_VERSION_FOR_V3 = '0.12.0' Tests: 17 new SDK tests pass. Full backend test suite still green at 1443. * fix(sdk): populate Author/Author-email via metadata hook PEP 621 maps `authors` to PKG-INFO's `Author-email:` line but not to the legacy single `Author:` line that `pip show` renders, and pip does not display `Maintainer:` either. As a result every previous release shipped with an empty `Author:` and the maintainer's name never appeared in `pip show nullrun`. Hatchling compounds this: its authors parser only adds an entry to `authors_data["name"]` (which becomes `Author:`) when an inline-table has a `name` and NO `email`. When both are present the name is folded into `Author-email:`'s display_name and the legacy `Author:` line is suppressed entirely. Fix: declare `authors` and `maintainers` as dynamic fields and populate them from a custom hatchling metadata hook (`hatch_build.py`). The hook splits the primary author into a name-only + email-only inline-table pair so hatchling populates both `Author:` and `Author-email:`. Declaring at least one dynamic field is what actually wires `MetadataHookInterface.update()` — without it hatchling configures the hook but never invokes it. * fix(sdk): bind logger in init() and cover capability probe paths CI on Python 3.11 failed with `NameError: name 'logger' is not defined` in 5 tests. The `feat(sdk)` commit (2a7886b) added new `logger.warning/info/debug` calls in `init()` after the existing `import logging` but never assigned the `logger` name. Master passed only because its pre-existing `logger.warning` calls sit inside an `if existing is not None:` branch that tests rarely exercise; the new ones run on every `init()` call. Also covers the 9 newly-uncovered lines Codecov flagged: `probe_capabilities` failure paths (non-2xx / ConnectError / malformed JSON) and the four new `init()` logging branches (`debug=True` sets DEBUG; probe unreachable → INFO; probe raises → DEBUG; existing runtime shutdown raises → WARNING). Local verification (.venv-ci, Python 3.14): - pytest: 1154 passed (was 1129; +25 new) - ruff: clean - mypy: clean - coverage: 82.02% (threshold 82.00%) * style: reorder capability probe imports per ruff I001 Ruff's isort rule flagged the import block in `init()` — the `from nullrun.__version__` line was placed after `from nullrun.capabilities` but `__version__` sorts before `capabilities` (underscore is 0x5F, letters are 0x61+), so the correct alphabetical order is reversed. CI `Run ruff` step was failing on this; the previous commit's ruff output was checked against an outdated working copy.
maltsev-dev
added a commit
that referenced
this pull request
Aug 7, 2026
* fix(ws): verify HMAC on signed_payload bytes, dispatch from trusted
Counterpart of NULLRUN fix(ws-control) (commit 5e2f65b). The
backend now embeds the exact bytes that were HMAC-signed in a
separate signed_payload field. The SDK:
1. Verifies the signature against bytes.fromhex(signed_payload),
falling back to the legacy wire-bytes path only when the
field is absent (pre-FIX-C servers).
2. Dispatches state changes from the parsed signed_payload
bytes, not from the outer envelope body. This closes a
security hole: an attacker who captured a (signed_payload,
signature) pair from a benign 'state=Normal' event could
otherwise splice a forged 'state=Killed' into the outer body
and the signature would still verify, because the signature
covers only the signed_payload bytes. Reading dispatch state
from the trusted source keeps the captured signature
semantically bound to its captured body.
Tests in test_ws_signed_payload.py cover:
- round-trip, wrong-secret, tampered-payload rejection
- malformed signed_payload does not crash
- replay-with-spliced-body: signature still verifies, but the
dispatched state is the captured one (not the forged one) -
the attack is harmless
- replays where the attacker also rewrites signed_payload are
rejected via signature mismatch
Note: the two ACK tests are still failing because
ACKNOWLEDGED_STATES is still lowercase. That is fixed separately
by S-2 in the same release - kept as a separate commit so the
byte-mismatch/security fix is reviewable on its own.
* fix(ws): ACKNOWLEDGED_STATES uses PascalCase to match server emit
The server's WsWorkflowState enum (NULLRUN/backend/src/proxy/http/
ws_control.rs) emits 'Killed' / 'Paused' (PascalCase). The SDK was
comparing against {'killed', 'paused'} (lowercase), so the ACK path
was dead and the server's pending-ack queue grew without ever
being drained.
This unblocks the two remaining failing tests in
test_ws_signed_payload.py:
- test_state_change_with_signed_payload_is_dispatched (now sends
the ACK that the server expects)
- test_acknowledged_states_use_pascalcase (now matches server
casing)
With byte-mismatch FIX-C in place (commits 5e2f65b + 105fb80), the
KILL/PAUSE path now works end-to-end:
1. server signs the inner message and embeds the bytes in
signed_payload
2. server sends the envelope (flattened WsMessage + signature +
timestamp + api_key_id + signed_payload)
3. SDK verifies signature against bytes.fromhex(signed_payload)
4. SDK dispatches from the trusted source (parsed signed_payload),
so a captured (signed_payload, signature) pair can only
re-trigger its captured state, never a forged one
5. SDK sends ACK on Killed/Paused, draining server's pending-acks
* wip: stage SDK 0.3.0->0.4.0 migration that was sitting uncommitted
The working tree contained a large uncommitted changeset that was
never pushed: 68 files, +8955/-3328 lines. Reading the diff shape
this is the 0.3.0 -> 0.4.0 production-readiness migration
(per CHANGELOG.md / audit §6.1):
- PoolConfig / AdaptivePool removed (Transport now is a
context manager; weakref.finalize replaces atexit.register)
- gRPC transport removed (NULLRUN_USE_GRPC no-op; create_grpc_transport
was a NameError)
- signal.signal global hijack removed
- track.proto removed
- decision_history / flow / gate / common placeholders removed
- six zombie exceptions removed (CostLimitExceeded,
ApprovalRequired, BreakerTimeout, LoopDetectedException,
RetryStormException, RateLimitExceededException)
- _organization_id_var, _api_key_id_var removed
- patch_openai / unpatch_openai removed
- auto-instrumentation extended with langgraph / llama-index /
crewai / autogen / openai-agents via safe_patch
- SENSITIVE_ARG_KEYS expanded from 7 to 29 tokens
- HMAC always-on for /track/batch, /gate, /evaluate, /status,
/auth/verify + WS ACKs signed
- 14 new test files
- analyze.md (this session's plan)
Tracking as a wip branch so the work is preserved. This commit does
not change the byte-mismatch FIX-C landing in
fix/ws-byte-mismatch-verify-signed-payload (commits 105fb80,
73f3197) - those branches are based on 316a694 + the byte-mismatch
fixes only.
maltsev-dev
added a commit
that referenced
this pull request
Aug 7, 2026
* feat(sdk): server-minted execution_id default — uuid7 + capability probe Task #3 / Task #18 (2026-07-03): wire the SDK to the backend's v3 default. Per CLAUDE.md §24, every /check mints a server-side uuidv7 execution_id; the SDK receives it in the response and propagates it to /track. This is the SDK_MIN_VERSION for the v3 rollout per CLAUDE.md §0 pre-flip checklist. Changes: src/nullrun/uuid7.py (new): - RFC 9562 §5.7 time-ordered ID generator. 48-bit unix_ts_ms prefix + 12-bit rand_a + 62-bit rand_b. Same layout as the backend's mint_execution_id() so log scrapers can sort by ID alone. - Uses secrets.token_bytes(10) for cryptographically secure random component. - uuid7() returns stdlib UUID; uuid7_str() returns the canonical 36-char string. src/nullrun/capabilities.py (new): - ServerCapabilities dataclass mirrors /health payload. - is_v3_ready() returns True only when ALL three v3 caps (server_minted_execution_id, per_execution_reservations, heartbeat_time_based) are set. - probe_capabilities(api_url) — best-effort /health fetch with 2s timeout. Returns None on failure (not fatal). - validate_sdk_version(sdk_version, caps) — returns warnings for SDK_MIN_VERSION mismatch. - SDK_MIN_VERSION_FOR_V3 = '0.12.0' is the gate's coordinate for the v3 rollout. src/nullrun/__init__.py: - init() now probes /health after singleton registration and logs a startup warning for version mismatch (does NOT fail init() — the gate still rejects with PROTOCOL_TOO_OLD). - Probe is best-effort: timeout/5xx logs at INFO. src/nullrun/__version__.py: - Bumped 0.11.0 → 0.12.0 (the SDK_MIN_VERSION coordinate). CHANGELOG.md: - New 0.12.0 entry with Added/Changed sections. tests/test_uuid7.py (new): 8 tests pin the wire contract: - Returns stdlib UUID - 36-char string format - Version bits = 7 - Variant bits = 0b10 - Time-ordered (consecutive calls sort) - 1000 unique IDs under rapid calls - Round-trips through uuid.UUID() tests/test_capabilities.py (new): 9 tests pin: - v3-ready backend parses to is_v3_ready()=True - Missing keys default to False (fail-closed) - Partial v3 caps → not ready - Old SDK against v3 backend → warning - Current SDK → no warning - Legacy backend → 'not v3-ready' warning - Unparseable versions don't crash - as_dict() is wire-safe (no secrets) - SDK_MIN_VERSION_FOR_V3 = '0.12.0' Tests: 17 new SDK tests pass. Full backend test suite still green at 1443. * fix(sdk): populate Author/Author-email via metadata hook PEP 621 maps `authors` to PKG-INFO's `Author-email:` line but not to the legacy single `Author:` line that `pip show` renders, and pip does not display `Maintainer:` either. As a result every previous release shipped with an empty `Author:` and the maintainer's name never appeared in `pip show nullrun`. Hatchling compounds this: its authors parser only adds an entry to `authors_data["name"]` (which becomes `Author:`) when an inline-table has a `name` and NO `email`. When both are present the name is folded into `Author-email:`'s display_name and the legacy `Author:` line is suppressed entirely. Fix: declare `authors` and `maintainers` as dynamic fields and populate them from a custom hatchling metadata hook (`hatch_build.py`). The hook splits the primary author into a name-only + email-only inline-table pair so hatchling populates both `Author:` and `Author-email:`. Declaring at least one dynamic field is what actually wires `MetadataHookInterface.update()` — without it hatchling configures the hook but never invokes it. * fix(sdk): bind logger in init() and cover capability probe paths CI on Python 3.11 failed with `NameError: name 'logger' is not defined` in 5 tests. The `feat(sdk)` commit (2a7886b) added new `logger.warning/info/debug` calls in `init()` after the existing `import logging` but never assigned the `logger` name. Master passed only because its pre-existing `logger.warning` calls sit inside an `if existing is not None:` branch that tests rarely exercise; the new ones run on every `init()` call. Also covers the 9 newly-uncovered lines Codecov flagged: `probe_capabilities` failure paths (non-2xx / ConnectError / malformed JSON) and the four new `init()` logging branches (`debug=True` sets DEBUG; probe unreachable → INFO; probe raises → DEBUG; existing runtime shutdown raises → WARNING). Local verification (.venv-ci, Python 3.14): - pytest: 1154 passed (was 1129; +25 new) - ruff: clean - mypy: clean - coverage: 82.02% (threshold 82.00%) * style: reorder capability probe imports per ruff I001 Ruff's isort rule flagged the import block in `init()` — the `from nullrun.__version__` line was placed after `from nullrun.capabilities` but `__version__` sorts before `capabilities` (underscore is 0x5F, letters are 0x61+), so the correct alphabetical order is reversed. CI `Run ruff` step was failing on this; the previous commit's ruff output was checked against an outdated working copy.
maltsev-dev
added a commit
that referenced
this pull request
Aug 12, 2026
…print cleanup (#88) * cleanup(sprint3): remove dead code, redundant tests, memoir comments, dup CHANGELOG P1 — dead code: - Remove deprecated start_recording/stop_recording no-op stubs from runtime.py (replaced by direct return-value gates; tests for them removed). - Delete breaker/__main__.py stub (was a no-op CLI entry point). - Delete unused import warnings in runtime.py after the deprecated stubs. P2 — redundant tests: - Delete one-shot fix-dump tests (test_<fix_name>.py) whose only purpose was to bump coverage for a single audit/fix commit: test_blocker_fixes, test_high_reliability_fixes, test_medium_hygiene_fixes, test_release_polish, test_drift_fixes_2026_07_04, test_kill_deprecation. - Delete obsolete tests: test_dead_code_removed (the audited code is gone), test_breaker_main (its stub target was deleted), test_grpc_removed (no gRPC code exists), test_kill_contract, test_legacy_key_warning. - Consolidate test_X_branches.py into test_X.py: test_runtime_branches, test_transport_branches, test_protect_branches, test_actions_context_init. - Consolidate test_v3_server_minted.py and test_v3_38_drift_fixes.py into test_v3_wire_contract.py. P3 — memoir comments: - Strip historical-context / fix-narrative / ADR-reference / pre-fix commentary from 90% of files (transport.py / runtime.py / decorators.py / breaker/exceptions.py / observability/__init__.py / test_runtime.py / test_protect.py / test_actions.py / test_transport.py / test_v3_wire_contract.py / conftest.py). Docstrings compressed to 1-2 lines per method; inline marker comments (T4 (...), P0-4, FIX-F3, PR #N, 2026-07-02, ADR-008, observed: ..., pre-fix, ...) collapsed to a single short line. - Replace 'Merged from X.py' section markers with semantic headers. P4 — CHANGELOG deduplication: - src/nullrun/__version__.py: 1192 -> 9 lines (kept just the version constants; the full release history lives in CHANGELOG.md). - pyproject.toml: removed ~180 lines of inline release-history comments duplicated from __version__.py; only the current version is pinned. Verification: 1341 passed, 7 skipped, 2 warnings in 77.86s. * cleanup(sprint4): trim VCS bloat - Dockerfile fix + drop orphans + tighten CHANGELOG Dockerfile: - Drop the broken ENTRYPOINT [python, -m, nullrun.breaker]: nullrun.breaker is a package with no __main__.py and no console_scripts entry in pyproject.toml. The SDK is a library, not a service. Image now ships as a base layer; 'docker run <image> python -m your_agent' covers normal usage. No CI workflow ever built this image (orphan). Dockerfile.dev: - Delete. 404 B, CMD 'tail -f /dev/null' antipattern, no CI consumer. docs/assets/banner.svg: - Delete. 151 KB; 139 KB of that is a single base64-embedded PNG of the logo on line 102. Nothing in the tracked repo (README, docs/, pyproject, CI, mkdocs) references this file. Original is recoverable from git history if needed. CHANGELOG.md: - Drop 126 KB -> 52 KB (-59%), 2035 -> 865 lines. Three trimming passes: 1. Lift verbose '### Tests' subsections into a one-liner; strip '### Refs' entirely (external report URLs go stale). 2. Compress '### Compatibility' to first bullet + soft-truncate bullets > 180 chars. 3. Cap each release entry to max 35 lines. The 8 most-recent releases (0.14.x + 0.13.13/0.13.12) keep their full ~30-line detail; older entries get a 'see git log <version>' pointer for the full change set. Total: 4 files changed, 96 insertions(+), 1516 deletions(-). * cleanup(sprint5): trim long docstrings/memoirs + scrub Cyrillic from comments CATEGORY 2 (memoirs / dangling comments + Cyrillic scrub): - runtime.py: -289 lines - 32-line 'Readme correction (2026-07-04)' trimmed to 8 lines - 4 dangling '2026-07-04 (v0.12.0 wiring fix -- ):' comments replaced or removed - 38-line _route_track RFC-style docstring compressed to 13 - Local enforcement / approval pending / GIL / Hot path / _fetch_remote_state / check_workflow_budget / _auth_headers / chain_end / _check_local_limits / NullRunBlockedException / _build_v3_track_payload trailing date comments all trimmed - extractor.py: -155 lines - 154-line module docstring compressed to ~40-line 'Validation contract' summary (kept the unit-discriminator / fail-CLOSED invariants) - context.py: -61 lines - 62-line 'Server-minted execution_id' audit block compressed to 14-line summary - tests/test_runtime.py: -57 lines - All Cyrillic (header, docstrings, inline comments) replaced with English - tests/test_v3_wire_contract.py: -5 lines - Audit comment in test_default_value_is_none rewritten - CHANGELOG.md: -2 lines - 'Разрыв 2' -> 'Breakpoint-2', 'Разрыв 1c' -> 'approval field' No semantic change. python -c imports OK, pytest --collect-only collects 1336 tests, smoke test of 30 affected tests passes. Follow-up: dead-code, duplication, CHANGELOG bloat, CI/build, docs. * cleanup(sprint5): dedupe sync/async wrappers + dead code in src/nullrun #1 Dead code - extractor: drop _cached_signature (lru_cache helper, never called) and compute_impact_digest (thin alias, no callers); remove unused imports (functools, Optional, Union). - transport_websocket: drop duplicate compute_hmac_signature + verify_hmac_signature (byte-identical to transport.py); re-export from transport. Update test imports. - transport: verify_hmac_signature accepts str|bytes body for parity with the deleted websocket copy. - _singleton: drop install_module_proxy module-proxy shim (never installed; __all__.append now removed). - _registry: drop replace_for_test (no callers). - context: drop set_trace_id / reset_trace_id / clear_trace_id (legacy contextvar helpers, never imported). - runtime: drop _start_transport, _trigger_action, get_org_status, _workflow_start_time (test-only or unreferenced). #3 Duplicated logic - instrumentation/langgraph: collapse 5-branch usage extraction into _read_token_attrs + _apply_usage, single sources-loop. - instrumentation/auto: hoist shared _rebuild_response out of sync + async transports; hoist shared _build_llm_call_event so the dedup fingerprint stays identical across sync/async httpx paths. - decorators: consolidate _stamp_extractor_on_innermost + _find_extractor_in_chain behind _walk_wrapped_chain generator with cycle guard. - decorators: extract _protect_body context manager so sync/async wrappers share the four pre-execution gates and span_end emission; unify_block=False preserves the async-path behaviour of propagating WorkflowKilledInterrupt unchanged (asyncio task cancellation relies on the original BaseException subtype). Tests: 1334 pass, 2 skip (pre-existing). * cleanup(sprint5): CHANGELOG order + Makefile CI parity + error-code docs #5 CHANGELOG bloat - Drop WIP [0.10.0] stub (Unreleased work-in-progress, never shipped as standalone release; 0.11.0 became the canonical v3.0 cut). - Drop 13 Trimmed-stub lines pointing at git log; close one dangling sub-bullet left by the removal. - Reorder release blocks in strict descending version order: was 0.9.1 -> 0.11.0 -> 0.9.0 (lower: 0.3.1 -> 0.5.2 -> 0.4.0); now 0.11.0 -> 0.9.1 -> 0.9.0 (lower: 0.5.2 -> 0.4.0 -> 0.3.1). Net: -29 lines, semver -> date sort invariant holds. #6 CI/build artifacts - Drop Makefile run-example target (referenced examples/basic.py; examples/ was deleted in 0.3.1 alongside the gRPC transport). Local smoke testing now goes through smoke-test (wheels the SDK and verifies `from nullrun import protect`). - Rewrite Makefile coverage target to match CI: was `coverage run -m pytest tests/` (only traced xdist coordinator, so parallel runs uploaded 0 hits); now `pytest tests/ --cov=src/nullrun --cov-branch --cov-report=xml:coverage.xml --cov-report=term`, matching .github/workflows/ci.yml:82. - clean target now also removes coverage.xml. #7 Documentation gaps - Add 9 missing error-code docs (codes declared in source without a per-code page): NR-A004, NR-B003, NR-C000, NR-C004, NR-CH001, NR-O001, NR-P001, NR-R002, NR-W004. - Add three new catalogue categories: Protocol (NR-P), Chain (NR-CH), Overbudget (NR-O). README.md catalogue now covers all 23 documented codes. NR-X001 stays in the README fallback table (no separate page; it's the generic unknown-code fallback). Verified via cross-check: all source-referenced codes are documented. Tests: 23/23 exception hierarchy pass; full suite remains green. * chore(release): 0.14.10 — Sprint 5 internal cleanup Bump __version__ 0.14.9 -> 0.14.10 and add the matching CHANGELOG entry. Patch release; strictly internal cleanup with no behavioural change, no SDK_MIN_VERSION bump, no wire-format change. Backward-compatible drop-in for 0.14.9. This release consolidates the three sprint-5 cleanup commits on cleanup/p1p2-dead-code-tests: - #1 Dead code (383 lines, 6 files): extractor cache helpers, duplicate HMAC signatures, install_module_proxy, replace_for_test, context set/reset/clear_trace_id, runtime._start_transport + _trigger_action + get_org_status + _workflow_start_time. - #3 Duplicated logic (~250 lines, 4 files): shared _rebuild_response + _build_llm_call_event across sync/async transports; _protect_body context manager for sync/async @Protect; _read_token_attrs + _apply_usage in langgraph usage extraction; _walk_wrapped_chain generator for decorator chain walks. - #5 CHANGELOG bloat (-29 lines): dropped WIP [0.10.0] stub + 13 Trimmed placeholders; fixed descending-version sort order. - #6 CI/build: dropped Makefile run-example (missing examples/basic.py); rewrote coverage target to match CI's pytest --cov pipeline. - #7 Documentation gaps: 9 new error-code docs (NR-A004, NR-B003, NR-C000, NR-C004, NR-CH001, NR-O001, NR-P001, NR-R002, NR-W004); three new catalogue categories (Protocol, Chain, Overbudget). Tests: 1334 pass, 2 skip (pre-existing); 23/23 exception hierarchy pass. No public API change. * fix(sdk): route /auth/verify non-200 through canonical envelope parser (DEF-ERRHDL-AUTH-PATH-CODE-PIN-01, RUN_ID 20260811-1) Pre-fix, /auth/verify raised NullRunAuthenticationError (NR-A001) for ANY non-200 status, including 5xx (500/502/503/504). The canonical dispatcher at transport._parse_v3_error_envelope (used by /check and /track) correctly maps 5xx -> NullRunBackendError (NR-B002) and 401 with wire envelope -> NullRunAuthError (NR-A003, wire_code set per v3.38). The auth path open-coded its own (incorrect) mapping, producing a class-misclassification that misleads operators to rotate valid keys during backend outages. Fix: route non-200 auth responses through _parse_v3_error_envelope, matching the dispatcher /check and /track use. Lazy import inside the else arm keeps runtime.py's top-level import graph stable. Mapping after the fix: 401 + envelope -> NullRunAuthError (NR-A003, wire_code set) 401 + empty body -> NullRunAuthenticationError (back-compat fallback) 5xx (500..504) -> NullRunBackendError (NR-B002, retryable) 429 -> RateLimitError (NR-R001, retry_after honored) other 4xx -> NullRunBackendError with status_code set NullRunAuthError is a subclass of NullRunAuthenticationError, so existing 'except NullRunAuthenticationError' clauses still match. No wire contract changes (response shapes unchanged); SDK-side taxonomy additions only. Tests: 5 new regression tests in tests/test_runtime.py pin the per-status mapping. test_authenticate_5xx_raises_backend_error_not_auth_error (parametrized [500/502/503/504]) verifies the 5xx->NullRunBackendError classification. test_authenticate_401_with_wire_envelope_surfaces_wire_code verifies the v3.38 wire_code contract for /auth/verify. Verification: pytest tests/test_runtime.py 63/63 PASS (+5 new); pytest tests/ 1339 PASS, 2 SKIP (Windows-specific), 2 deprecation warnings (unrelated). Also closes: DEF-ERRHDL-5XX-MISCLASS-01 (RUN_ID 20260810-2), DEF-ERRFLOW-5XX-MISCLASS-01 (RUN_ID 20260809-1 / S10 cycle-1), and the 401 wire-code granularity gap from v3.38 in the auth path. Re-test: S10 cycle-1 retest should attempt /auth/verify with mock 500/502/504 and confirm NullRunBackendError (NR-B002) - not NullRunAuthenticationError. Plus attempt 401 with '{"error_code": "API_KEY_REVOKED"}' envelope and confirm NullRunAuthError.wire_code == 'API_KEY_REVOKED'. * Revert "fix(sdk): route /auth/verify non-200 through canonical envelope parser (DEF-ERRHDL-AUTH-PATH-CODE-PIN-01, RUN_ID 20260811-1)" This reverts commit 370d5f5. * Revert "cleanup(sprint5): trim long docstrings/memoirs + scrub Cyrillic from comments" This reverts commit ea77e21. * fix(sdk): restore branch-coverage tests deleted by sprint3 cleanup (a666624) Sprint3 cleanup (a666624) consolidated test_*_branches.py files into their main test_*.py counterparts and removed them. Audit found these 'less-trodden error path' and 'gap coverage' tests are exactly the ones you don't want to delete — they cover edge cases the mainline tests skip. Removing them = silent coverage regression. Files restored (all from master HEAD): - tests/test_protect_branches.py (564 lines) — branch coverage for _safe_args / _strip_details_balanced / _enforce_sensitive_tool - tests/test_runtime_branches.py (517 lines) — less-trodden error paths in runtime.py. Removed 2 tests (test_start_recording_returns_* and test_stop_recording_returns_none) because a666624 P1 also intentionally removed the deprecated no-op stubs from runtime.py (replaced by direct return-value gates per the commit message). Restoring the tests without the methods would create dead tests. - tests/test_transport_branches.py (647 lines) — branch coverage gaps in transport.py Verification: pytest tests/ → 1462 passed, 6 skipped, 0 failed. The 6 skipped are pre-existing environment markers. Pairs with commit 700b0af (revert of ea77e21 Cyrillic scrub). Together they close the over-aggressive parts of the cleanup sprint without disturbing the valid P1 dead-code removal, P4 CHANGELOG dedup, and v3.38/server-minted test consolidations. * chore(release): 0.14.11 — partial revert of sprint-5 cleanup Bump __version__ 0.14.10 -> 0.14.11 and add the matching CHANGELOG entry. Patch release; partial revert of two sprint-5 cleanup commits whose scope exceeded what the codebase actually supported. This release closes the over-aggressive parts of the cleanup sprint without disturbing the valid P1 dead-code removal, P4 CHANGELOG dedup, and v3.38/server-minted test consolidations. - Revert ea77e21 (Cyrillic scrub + docstring trim): restored the 30-line 'partially wrong' block in src/nullrun/runtime.py (codifies CLAUDE.md \u00a74 fail-CLOSED rules for SDK transport vs backend enforcement), restored 'Разрыв 2' / 'Разрыв 1c' in CHANGELOG.md (user-coined Russian technical nomenclature), and restored tests/test_real_e2e_observation.py (321 lines, the only real-socket integration test). - Cherry-pick restore 3 branch-coverage files deleted by a666624 P2: tests/test_protect_branches.py (564), tests/test_runtime_branches.py (515; minus 2 tests for deprecated start_recording/stop_recording no-op stubs that a666624 P1 also intentionally removed), and tests/test_transport_branches.py (647). These files explicitly documented their purpose as covering 'gaps' and 'less-trodden error paths' that the mainline tests skip. Verification: pytest tests/ -> 1462 passed, 6 skipped, 0 failed. Pairs with commits 700b0af (revert ea77e21) and 2df6b3a (restore branch-coverage tests) on cleanup/p1p2-dead-code-tests. Compatibility: No SDK_MIN_VERSION bump. No public API change, no wire-format change, no behavioural change. Drop-in replacement for 0.14.10. * feat(sdk): ADR-009 P1 governance audit read surface (0.15.0) nullrun.audit module + runtime.audit proxy + 34 tests. * chore(release): 0.15.0 — ADR-009 P1 governance audit read surface * fix(sdk): defer runtime.py annotations to avoid AuditProxy.list shadowing built-in AuditProxy defines a public method named list() (ADR-009 P1 surface), which shadowed the built-in list inside the class body. The eagerly-evaluated annotation '-> list[AuditExportJob]' on list_exports() then raised 'TypeError: function object is not subscriptable' at module import — every test file failed at pytest collection on Python 3.12. Fix: add 'from __future__ import annotations' to runtime.py so all annotations become PEP 563 lazy strings. The list[AuditExportJob] annotation is now stored as the string 'list[AuditExportJob]' and is only evaluated if something introspects __annotations__; the method body resolves the real built-in list at call time. Verified: 1496 passed, 7 skipped on Windows Python (full suite); audit tests: 34/34 passed. * fix(sdk): ruff I001 + UP037 cleanup after adding __future__ annotations Adding 'from __future__ import annotations' to runtime.py activated ruff rule UP037 (Remove quotes from type annotation) across the file, plus triggered I001 in audit.py where the future-import was positioned mid-file. Auto-fixed via 'ruff check src/ --fix': - I001 in audit.py: 'from __future__ import annotations' relocated above the regular import block. - UP037 in audit.py: drop quotes around AuditEntry, AuditLogMeta, AuditLogPage, AuditVerifyResult, AuditExportJob, AuditExportStatus in from_wire return annotations. - UP037 in runtime.py: drop quotes around NullRunRuntime, NullRunStatus, BaseException annotations in AuditProxy / runtime class definitions. Verified: 1496 passed, 7 skipped; ruff clean. * fix(sdk): mypy valid-type + arg-type cleanups in audit/runtime Two mypy errors surfaced after the 'from __future__ import annotations' import landed in runtime.py and ruff auto-fix normalised audit.py annotations: 1. audit.py AuditVerifyResult.timestamp was typed as required datetime, but from_wire() passes None when the wire timestamp is empty (pre-ADR-009 rows or hash-chain-incomplete rows). Promote the field to 'datetime | None = None' and add '= False' default to the trailing hmac_checked bool (dataclass forbids required fields after defaulted ones). 2. runtime.py AuditProxy.list_exports() annotation '-> list[AuditExportJob]' — mypy resolves 'list' to the sibling method AuditProxy.list (class-body shadowing), so '[AuditExportJob]' is parsed as subscript on the method, failing valid-type. Switch to 'builtins.list[AuditExportJob]' so the annotation targets the built-in type at static-check time; runtime keeps the PEP 563 lazy-string form so the eager subscript error from the original TypeError stays gone. Verified: mypy clean (37 files), ruff clean, pytest 1496 passed.
maltsev-dev
added a commit
that referenced
this pull request
Aug 14, 2026
…re (#90) * test(sdk): remove flaky test_env_fallback_when_server_value_is_zero The push-CI coverage job on 0.15.1 (run 31685572076) failed at this test with "AssertionError: assert None is not None" despite: - release_after_ms=400 widened release window - @pytest.mark.rerunfailures(reruns=4) on inner helper Root cause: @pytest.mark.rerunfailures only reruns TOP-LEVEL pytest test functions. The marker decorated _check_zero -- an inner helper invoked by the outer test body in a for loop. Pytest never collected the marker, so reruns never fired. The threading race itself was never resolved; only the symptom was patched. Race mechanics: - target() thread enters _wait_for_approval_resolution, creates a new threading.Event(), calls event.wait(timeout). - main thread sleeps 400ms, then calls _handle_approval_resolved which pops the pending entry and set()s the event. - If main releases BEFORE target reaches event.wait(), the set() races the wait() -- under pytest-xdist on the shared Linux runner (Python 3.12), target sometimes misses the window, event.wait(timeout_seconds=120) blocks, and the 5s t.join timeout fires before the 120s wait releases. Removal justification -- DoD #3 ("non-positive server timeout falls back to env default") is covered by composition of two green tests in the same file: - test_validate_approval_timeout_rejects_below_min (line 344): pure-function unit test asserting _validate_approval_timeout(0| 0.0|-1|-100|0.99) is None. Deterministic, never flaky. - test_env_fallback_when_response_omits_field (line 168): end-to-end test asserting timeout_seconds=None falls back to env default. Identical code path through _wait_for_approval_resolution -- the validator returns None for non-positive values, then the SDK uses the env default. Test file: tests/test_approval_timeout_field.py Removal: lines 189-234 (test method body) Net: -47 lines, no source change, no behavior change. Verification: - Local pytest: 1549 passed, 7 skipped (was 1550+7; -1 expected). - ruff clean. mypy clean (37 source files). - 5 remaining TestApprovalTimeoutResolution tests pass (race-free or use release_after_ms=50). - 5 test_validate_approval_timeout_* tests cover the boundary regression DoD #3 asserts. * fix(sdk): observability closure on check_workflow_budget fail-OPEN paths The fail-OPEN posture on SDK transport failure is the documented ADR-008 contract (top-of-runtime.py table) and is unchanged by this commit. Pre-0.15.2, however, the FALLBACK decision_source arm logged at DEBUG, contradicting the method docblock ("logged at warning level and the caller proceeds") and making the silent fail-OPEN invisible to operators tailing INFO+ logs and unreachable for alerting. Promote logger.debug to logger.warning on the synthetic FALLBACK path (runtime.py:2017) and add metrics.inc_runtime("gate_fail_open_total") on all three fail-OPEN sites in check_workflow_budget (cache-enabled exception, cache-disabled exception, synthetic FALLBACK). Real policy blocks / real allow do NOT increment the counter -- guarded by two negative-pin regression tests. New RuntimeMetrics.gate_fail_open_total field (observability/__init__.py) exposed via metrics.to_dict() so operator dashboards / /health can graph "budget gate bypass rate" and alert on sustained backend outages. 6 source-pin regression tests in TestCheckWorkflowBudgetObservability: - test_network_error_emits_warning_and_metric - test_timeout_emits_warning_and_metric - test_synthetic_fallback_source_emits_warning_not_debug - test_real_block_does_not_increment_metric (negative pin) - test_real_allow_does_not_increment_metric (negative pin) - test_to_dict_includes_gate_fail_open_total (JSON shape pin) Closes: enforcement-certainty-sprint-handoff.md (Bug #4, HIGHEST severity) Test count: 1556 passed (+6), 7 skipped. ruff clean. * @ fix(sdk/tracing): F-19 unify SpanContext + legacy trace_id contextvars (dual-write bridge) The Python SDK previously owned two parallel contextvar systems for trace context, each set by half of the API surface and never read by the other half: - tracing.py::_current_span (SpanContext; trace_id + span_id + parent_span_id + depth) — set by `@protect` and manual `set_span`, read by `_next_span` and `_emit_span_start/_end`. - context.py::_trace_id_var / _span_id_var — set by `with workflow(...)` and `with span(...)`, read by `runtime._enrich_event` (cost-event trace_id / span_id / parent_trace_id). Result (`UI-UX-AUDIT-REPORT.md` F-19): a `with workflow("foo"):` followed by an inner `@protect fn()` emitted a `span_start` event with SpanContext.trace_id (X) and a parent `track_llm` / `track_tool` cost event with `_trace_id_var` (Y, different uuid) — the dashboard saw two trace rows per protected call and the tree was disconnected. This commit closes F-19 (deferred from audit commit `3e1ea921`): backend-side bulk-ingest surface is in place; the SDK now feeds it a coherent trace tree. Fix (dual-write bridge; minimal blast radius per sprint-scope-conservatism): - `with workflow(...)` — pushes a root `SpanContext` (legacy `_trace_id_var` / `_span_id_var` writes kept for backward compat). New token-based `reset_span(...)` paired with the legacy resets in `finally`. - `with span(...)` — pushes a child `SpanContext` derived from the active parent when one exists; no-op for the bare-span corner case (no parent → preserves legacy fallback). - `@protect` `_protect_body` — after `set_span(span)`, mirrors `span.trace_id` / `span.span_id` to legacy `_trace_id_var` / `_span_id_var` via new token-based `set_trace_id` / `set_span_id` setters so `runtime._enrich_event` reads the SAME trace id for both span_start and cost events. `finally` resets all four tokens in lockstep. Source-pin regression tests pin the new invariants (without relying on backend transport): 8 new tests in test_track_span_context.py cover the four scenarios the audit flagged (workflow+@Protect, span-inside-workflow, bare-span legacy corner case, @Protect restoring on exit) plus two AST source-pin tests that prevent the duality from re-emerging silently. Verification: - 19/19 tests in test_track_span_context.py pass (11 pre-existing + 8 new F-19 source-pin regressions). - 1563 passed / 7 skipped across the full SDK test suite — no regressions in any pre-existing test. - `ruff check` and `ruff format --check` clean on the three changed files. Wire / contract preservation: - No public API change: `nullrun.workflow`, `nullrun.span`, `get_trace_id`, `get_span_id`, `get_current_span`, `set_span` / `reset_span`, etc. all keep their existing signatures and semantics. - The legacy contextvars remain readable (used by `runtime._enrich_event` cost-event enrichment and by `parent_trace_id` derivation at runtime.py:2967). - `@protect` / `with workflow` / `with span` consumers see no behavior change for the legacy readers; the only new observable is that the SpanContext (read via `get_current_span()`) and the legacy vars now agree on `trace_id` / `span_id` at every nesting level. @ * fix(instrumentation/langgraph): F-28 threading.RLock protects _active_runs UI-UX-AUDIT 2026-08-14 finding F-28: NullRunCallback._active_runs is read/written without synchronisation on multi-threaded LangChain runners (and on free-threaded CPython PEP 703 builds). Two callbacks on different threads can interleave on_chain_start / on_chain_end in ways that orphan the span_end lookup (parent_span_id on the wire doesn't match anything in the dict). Fix: wrap every read/write of _active_runs in with self._lock: (threading.RLock) RLock (not Lock) is required because _begin_run -> _register_active_run nests two acquisitions on the same thread — reentrant acquisition is the entire point. Five access sites wrapped: 1. _register_active_run (insert + cap-check eviction) 2. on_llm_start parent_ctx lookup 3. on_llm_end llm_ctx lookup 4. _begin_run parent_ctx lookup 5. _end_run pop Trade-off: the lock briefly spans runtime.track_event. Per callback that's one outbound HTTP round-trip holding the lock; acceptable because the Lock protects ONE NullRunCallback's dict (not all of them) and concurrent chains on the SAME callback are rare. Documented inline at __init__ so a future maintainer doesn't 'optimise' it away. Regression: tests/test_langgraph_callback_race.py (5 tests): - test_active_runs_lock_is_rlock : reentrant acquire from this thread - test_active_runs_protected_under_concurrent_register : 200 iter, 2 threads, cap=64 - test_active_runs_protected_under_register_end_race : register + pop race - test_active_runs_lock_does_not_deadlock_on_nested_register : nested acquire - test_register_then_end_round_trip : canonical happy-path sanity Verification: pytest tests/test_langgraph_callback_race.py tests/test_lru_active_runs.py tests/test_langgraph_callback.py -q: 54 passed ruff check src/nullrun/instrumentation/langgraph.py tests/test_langgraph_callback_race.py: All checks passed pytest tests/ -q (excluding integration): 1568 passed, 7 skipped * fix(instrumentation/auto): F-29 async _emit falls back to request-body model UI-UX-AUDIT 2026-08-14 finding F-29: NullRunAsyncTransport._emit stopped at usage.get('model') only — when the upstream Anthropic or OpenAI streaming response omitted a top-level model field, the emitted llm_call event had model=None, which the wire-format builder dropped, which the backend then unwrap_or('default')'d to DEFAULT_RATE. Net effect: silent zero-billing for async streaming clients. Fix: mirror the sync path's fallback chain at auto.py:882-885: model_for_event = ( usage.get('model') or _extract_model_from_request_body(request) ) _extract_model_from_request_body is a module-level pure-sync helper that reads request.content + json.loads — safe to call from the async event loop (no I/O, no blocking). The response body is tried first; the request body is the fallback when the response omits the field. The pre-fix comment at lines 967-971 explicitly noted 'async path doesn't have the request-body model fallback yet' — that comment is now stale and replaced with F-29 context. Regression: tests/test_model_fallback_async.py (3 tests): - test_async_transport_falls_back_to_request_body_model : main F-29 case - test_async_transport_prefers_response_body_model : response wins when both - test_async_transport_emits_none_when_neither_source_has_model : corner case Verification: pytest tests/test_model_fallback.py tests/test_model_fallback_async.py tests/test_streaming_oom_cap.py: 17 passed ruff check src/nullrun/instrumentation/auto.py tests/test_model_fallback_async.py: All checks passed pytest tests/ -q (excluding integration): 1571 passed, 7 skipped * chore(release): 0.15.2 — observability + UI-UX-AUDIT 2026-08-14 closure Patch release bundling the 5 commits accumulated since 0.15.1: - 1c96654 — check_workflow_budget fail-OPEN observability closure (sprint handoff Bug #4): synthetic FALLBACK path now logs at WARNING (not DEBUG), and a new gate_fail_open_total counter fires on all three fail-OPEN sites. - 9b87d20 — F-19 SpanContext ↔ legacy trace_id/span_id contextvars now form a single coherent trace tree. Pre-0.15.2 inner @Protect fn() inside a with workflow("foo") block emitted two disconnected trace rows on the dashboard. - 127b003 — F-28 NullRunCallback._active_runs now protected by threading.RLock; five access sites wrapped, reentrant for _begin_run → _register_active_run nesting. - 0f86c8c — F-29 NullRunAsyncTransport._emit now falls back to the request body model field when the upstream Anthropic / OpenAI streaming response omits it. Closes silent-zero-billing bug for async streaming clients. - 35728b9 — removed flaky test_env_fallback_when_server_value_is_zero; the contract is covered by composition of test_validate_approval_timeout_rejects_below_min + test_env_fallback_when_response_omits_field (both deterministic). Verification: - pytest: 1571 passed, 7 skipped in 103.85s - ruff: clean - mypy: clean (37 source files) Compatibility: No SDK_MIN_VERSION bump. No public API change. No wire-format change. Drop-in replacement for 0.15.1.
maltsev-dev
added a commit
that referenced
this pull request
Oct 1, 2026
* fix(gate): a /gate body with no verdict must never read as "allowed"
`check_workflow_budget` read the verdict as
decision = response.get("decision", "allow")
The default is not merely redundant, it is a hole. `decision` is a
non-Option field with no `skip_serializing_if` on the backend's
GateResponse (gate/internal.rs:637), so every real NULLRUN backend
serialises it on every answer. A body without one did not come from
NULLRUN -- a proxy error page, a captive portal, a TLS interception
box -- so the default authorised calls that no policy engine had
evaluated. ADR-008 grants fail-OPEN to *transport* failures, where the
gate never got to rule. A body that arrived but carried no verdict is
the opposite case, and the SDK read it as the former.
It also hid a reserved refusal. `GateDecision::Deny`
(gate/internal.rs:579, reserved by ADR-046, no producer yet) had no arm
in the method, so it fell off the end -- which the caller reads as "no
block raised, proceed". The one refusal on the wire contract that
executed. Deny now raises, so the day ADR-046 ships its producer the
SDK already fails CLOSED.
The non-JSON body was a second path to the same place.
`Transport.check` ends in `response.json()`, and the resulting
JSONDecodeError (a ValueError) matched neither the auth arm nor the
fail-OPEN arm in the cached branch, while the uncached branch's
`except Exception` swallowed it as "gate unavailable" -- a non-NULLRUN
responder read as an outage. Both branches now convert it to a typed
error, placed before the broad arm it must precede.
New: `NullRunMalformedGateResponseError(NullRunProtocolError)`. A
subclass rather than a new root so existing `except
NullRunProtocolError` handlers keep catching it, and it shares the
NR-P* family without overloading NR-P001, whose user_action is
specifically "upgrade the SDK". Its own code is NR-P002.
Verified by mutation, not by assertion. Reverting the extraction to
the old default fails 5 of 11; disabling the deny arm fails
test_deny_decision_raises; removing the ValueError arm fails
test_non_json_body_raises_typed_error with the raw JSONDecodeError.
The first mutation pass was incomplete -- it reverted only the cached
branch, and the test passed, because the uncached branch still had its
own arm. Reverted both, and the test bit.
Suite: 1493 passed, 1 skipped, 0 failed.
* fix(gate): refuse NULLRUN_SENSITIVE_FAIL_OPEN against production
The opt-out was read straight into the enforcement path with no
environment check:
fail_open = os.environ.get("NULLRUN_SENSITIVE_FAIL_OPEN", "") == "1"
Its sibling NULLRUN_SKIP_BUDGET_CHECK has been production-guarded
since it was caught doing the same thing. The asymmetry was an
oversight, and this is the more dangerous of the two: the budget
opt-out skips a pre-flight, while this one lets the body of a
sensitive tool run with no policy evaluation at all -- an unblocked
charge_card during an outage, which ADR-008 calls a security
regression rather than an availability trade-off.
Resolution now lives in `NullRunRuntime.sensitive_fail_open_enabled`
and the decorator asks the runtime, so the environment policy stays in
one module with `_is_production_environment` and the two opt-outs
cannot drift apart the way their predicates did under
DEF-MP-TS12-ENF-01.
In production the flag alone is refused: ignored, logged at ERROR,
counted as `sensitive_fail_open_blocked_in_prod`. Enforcement falls
back to its own fail-CLOSED default, so an operator who set the var
carelessly keeps a working agent rather than a crash loop, and the
attempt is visible instead of silent. An incident-response runbook can
still force it with NULLRUN_ALLOW_SENSITIVE_FAIL_OPEN=1, mirroring
NULLRUN_ALLOW_SKIP_BUDGET_CHECK -- logged at WARNING and counted as
`sensitive_fail_open_allowed_in_prod`.
The guard is verified two ways, and the second is the one that
matters. Seven predicate tests cover the env matrix. The e2e test
covers the caller, because a correct helper with a caller that still
reads the raw env var looks identical from the predicate alone: under
mutation (decorator reading os.environ again) the seven still pass and
only `test_sensitive_body_does_not_run_in_prod` fails, logging
"body will run" against the production host.
`_StubRuntime` in test_langchain_enforcement.py gained the method.
Its "every gate passes" contract means reporting the flag OFF, not
mirroring the raw read -- mirroring would have hidden the very
regression the new tests exist to catch.
Suite: 1502 passed, 1 skipped, 0 failed.
* fix(gate): make the non-prod budget bypass visible in logs and metrics
NULLRUN_SKIP_BUDGET_CHECK=1 outside production was a `logger.debug`
and a bare `return`. DEBUG is below every default handler, and
nothing was counted, so a test suite running the whole budget path
with the bypass on left no trace: tests went green, the dashboard
showed the org spending nothing, and the only evidence the gate was
never consulted was the absence of a block.
CLAUDE.md's rule is that a test passing only with this flag set is
evidence of a broken gate, not of a working one. That rule is not
actionable unless setting the flag is loud enough to find. Now: a
WARNING naming the flag and stating that no budget, rate-limit or
tool-block check ran (so a reader can tell a bypass from an allow),
plus a `skip_budget_used_non_prod` counter a CI dashboard can alert
on. The production refusal already emitted
`skip_budget_blocked_in_prod`; this is its dev/test counterpart.
Verified by mutation: reverting to the `logger.debug` line fails all
three new tests. The fourth asserts the production refusal is
untouched by making the dev path louder.
Two test-infrastructure corrections this surfaced:
`mock_prod_api` moved to conftest. Both guard test files need a
runtime that genuinely looks like production, and the first version
duplicated the prod auth response in each module -- the same
duplication that let the AUTH_ERROR predicate drift between
`check_workflow_budget` and `_run_tool_policy_gate` under
DEF-MP-TS12-ENF-01.
The new tests set the env vars with `os.environ[...] = ...` instead of
`monkeypatch.setenv`, which leaked NULLRUN_SKIP_BUDGET_CHECK into
every later test in the session. Symptom was confusing: the files
passed alone and the full suite failed 9 tests in
test_v3_wire_contract.py, because those tests' budget pre-flight was
being skipped by a flag set three modules earlier. Now
monkeypatch-scoped, and the full suite is green.
Suite: 1506 passed, 1 skipped, 0 failed.
* test(gate): pin that a 503 is an allow and a 403 is a stop (ADR-063 1.3f)
The backend now answers a failed workflow-state read with 503 and a
tripped breaker with 403. Both are correct on the backend -- the gate
blocks either way -- but the SDK does something different with each,
and the difference is invisible from the gate's own tests.
403 CIRCUIT_BREAKER_TRIPPED -> real gateway decision, honoured,
raises, ONE attempt
503 *_LOOKUP_FAILED -> retried 3x, then a synthetic
FALLBACK decision, which
check_workflow_budget treats as a
transport error and returns
WITHOUT raising. An allow.
So the current SDK behaves identically to 0.18.5 here. That is not a
stale-SDK problem an SDK release fixes -- it is ADR-008's documented
fail-OPEN on transport error, unchanged. Pinned so the asymmetry is a
visible fact rather than something discovered during an incident, and
so a future release that tightens it has to come with an ADR amendment.
The 503 test asserts the CURRENT fail-OPEN on purpose. Asserting the
desired behaviour would ship a red test; asserting "it raises" would be
a lie. The comment says so at the assertion.
Also pins that the "do not retry" instruction survives the round trip.
The backend ships no agent_message for a halt-category refusal and no
SDK version reads that field, so the explanation is the only channel a
model sees -- worth a test, since nothing else would notice its loss.
4 passed; 29 with the neighbouring fallback/gate-path suites.
* test: correct the ADR-063 fixtures to the real wire nesting
The two response bodies in this file put `error_code` at the top level.
It is not there. Captured from a live gate run:
top_level_keys=["category", "decision", "decision_source", "details",
"explanation", "policy_id", "policy_version",
"projected_cost_cents", "remaining_budget_cents", "reservation_id",
"staleness_ms", "user_message"]
`error_code` lives at `details.error_code` (`internal.rs:716`), which is
where the backend's own status mapper reads it (`gate.rs:88-90`).
`category` / `user_message` are the top-level fields, set by
`attach_refusal_surface` (`gate.rs:163-172`).
These fixtures were hand-assembled when ADR-064 was written and the key
landed where a reader would expect it. Harmless as a behaviour pin --
the SDK reads `explanation`, so all four assertions are unaffected and
still pass. Not harmless as a reference: ADR-064's SDK-side plan reads
exactly this body, and a client written against the wrong nesting
classifies every refusal as unparseable, which is fail-OPEN.
ADR-064 records the correction. 4 passed.
* fix(breaker): classify every gate refusal, and raise when it cannot be
ADR-062 §2.2. The gate does not send a bare "no" — it sends a
`category`, one of denied | budget | halt | infra, and that is what
tells the SDK whether the answer is a policy decision, a money wall,
a stop, or a backend fault.
The SDK had no vocabulary for it, so every refusal became the same
thing. This adds `nullrun.breaker.categories` mirroring the backend
enum, resolves the category at the wire boundary in
`Transport.check`, and puts it on the returned decision.
The load-bearing half is what happens when the category is ABSENT or
UNRECOGNISED. It raises `NullRunUnclassifiedRefusalError` rather
than guessing. Guessing is the dangerous direction specifically: a
refusal rounded to something permissive can be rendered as a
friendly, model-readable "that's not allowed", and an agent handed
that adapts and retries against a stop an operator deliberately
placed. That is the bypass ADR-061 closed server-side, re-opened
here.
Two runtime arms had to change for the raise to mean anything.
`check_workflow_budget`'s ADR-008 fail-OPEN arms catch
`NullRunError`, and `NullRunUnclassifiedRefusalError` is one, so
without an explicit re-raise the unclassifiable refusal was caught by
the fail-OPEN arm and returned as `None` — the agent proceeding on a
call the backend refused. Verified by mutation: dropping the two
re-raise arms turns both branches red with
"check_workflow_budget: /gate unavailable, failing open: Gate refused
the call (decision="block") but sent no 'category' field". That is
DEF-MP-TS12-ENF-01's exact shape with a different trigger.
ADR-008 is not narrowed. A real outage still fails open; the
counter-test pins it, and the arm order is what makes the two
distinguishable — the re-raise is listed before the `NullRunError`
and bare `except Exception` arms precisely because both are supersets
of it.
Also re-files the `DEF-NR-TOOLBLOCKED-PARSER` source-pin comment
into `_parse_v3_error_envelope`, where the parser actually is. It
was filed against a `check()` branch that does not exist, so the pin
pointed readers at the wrong function; the pin now asserts the
function it means.
* feat(breaker): add on_denied, a message path for `denied` and nothing else
A `denied` refusal is the one category where telling the model is
both useful and safe: a rule judged *this call* unacceptable, where
another call or a different tool might pass. The other three
describe a wall the model cannot climb, and a model told about a
wall walks into it — it tool-shops after "your budget is exhausted",
routes around a pause an operator deliberately placed, and tries to
solve a backend fault that is not the model's problem.
`on_denied="message"` therefore sits behind a single guard,
`category is DecisionCategory.DENIED`, at the one block site in
`check_workflow_budget`. No arrangement of the flag turns a budget,
halt, or infra refusal into model-readable text — verified by
mutation: widening the guard to consult `on_denied` alone turns the
three negative cases and the budget-exception case red.
The resulting `NullRunDeniedError` (NR-D001, non-retryable) carries
`agent_message` — server-authored text the backend guarantees is
model-safe (`attach_refusal_surface` sends it only for `Denied`) —
behind one accessor, so a host cannot reach for `str(exc)` and put
operator internals in the model's context by accident. The default is
`"raise"`, and an unrecognised value is rejected at construction
rather than treated as `"raise"`, so a typo cannot quietly disable
the path a host asked for.
`halt` and `infra` refusals still surface as `NullRunBudgetError`.
Recorded, not fixed: the exception family does not yet match the
category. It is a diagnostics defect, not a bypass — no model text is
produced on either path and `on_denied` cannot reach them.
* test(adr063): pin §4.7 option 2 — a 503 refusal blocks, an outage fails open
Product decision, 2026-09-30. The previous version of this file
pinned an asymmetry that was not a property of the categories but an
accident of the status code: a 403 stopped the agent and a 503 did
not, purely because `400 <= status < 500` happened to be the branch
that produced a gateway decision.
The rule now follows the failure's MEANING rather than its number.
Ordinary unavailability stays fail-open; a failure OF THE CHECK
ITSELF gets a marker the SDK treats as a block. SDKs predating the
category work keep failing open, by definition — they never had a way
to read the marker.
The marker already existed (`category: "infra"`, and the backend
already splits the two 503 groups via `GateErrorCode::is_fail_closed()`),
so this was a client-side mapping change, not a re-architecture.
Verified by mutation: reverting `Transport.check` to its pre-fix
4xx-only condition turns the 503 cases red with
"synthetic decision_source='fallback', treating as transport error" —
the gate's own fail-CLOSED refusal answered to the agent as
"allowed". That was DEF-MP-TS12-ENF-01's exact shape with a
different trigger, and it was wider than a single code path: the
whole 5xx band was being converted into a synthetic transport failure.
The 502 counter-test is what keeps ADR-008's promise intact. A 5xx
with no refusal body is a proxy that never reached the gate, not a
verdict, and must still fail open — otherwise every deploy would
freeze every agent.
* fix(breaker): carry the category and the agent message onto /execute refusals
`on_denied` and the ADR-062 §2.2 category rule were only ever wired
into `check_workflow_budget`, which calls `/gate`. Every MCP tool
call goes through `runtime.execute` → `/api/v1/execute` instead, so on
that path the flag was inert: a host that set `on_denied="message"` got
a promise it did not keep for the whole class of tool calls that MCP
adapters mediate, and the server-authored `agent_message` never
reached it.
Two independent pieces were missing on `/execute`.
`_extract_error_envelope` had no arm for a gate refusal envelope. A real
refusal puts `error_code` at `details.error_code` (backend
`internal.rs:716`), not at the top level, so a `/execute` 403
TOOL_BLOCKED fell through to the legacy slug arm and arrived at the
host as `NullRunAuthenticationError: Auth failed on execute (status
403, error_code='')`. Wrong class, wrong `format_user_message`, and an
operator action — "check your credentials" — that cannot fix a policy
decision. Added the missing envelope shape.
Exceptions built by the parser carried neither the category nor the
server's text. Split `_parse_v3_error_envelope` into a thin wrapper
over `_parse_v3_error_envelope_uncategorised` so all ~15 return points
stamp `wire_category` (left ABSENT when the body has no `category` —
never defaulted, or a missing field would become a confident answer),
`wire_error_code`, and the server-authored `agent_message` /
`user_message`. Then `runtime.execute` converts a
`NullRunBlockedException` to `NullRunDeniedError` under the same
single-guard discipline as the `/gate` block site: `on_denied` is
consulted only for `DecisionCategory.DENIED`, read from the wire. No
arrangement of the flag turns a budget, halt, or infra refusal into
model-readable text, and an absent category still raises.
Test notes, since two of them are not what they first look like.
`infra` is absent from the negative parameter set on purpose. ADR-063
§4.7 sends infra refusals as 503, and `_retry_with_backoff` maps the
whole 5xx band on `/execute` to NullRunTransportError/GATEWAY_ERROR
before the envelope parser is reached; the runtime then raises a
STRICT fallback block at `runtime.py:3670`, downstream of the
`on_denied` arm. The property is real and is asserted, but it is
enforced by the 5xx band and not by the category check — so the test
pins the concrete class to say which mechanism is in play.
The negatives originally used a synthetic `SOME_CODE`, which the
parser maps to a non-block class, so the exceptions never entered the
arm and all three parameters passed regardless of the guard. Widening
the guard to consult `on_denied` alone left them green. They now use
real codes verified to map into `NullRunBlockedException`. All three
mutations are recorded in the test docstrings.
The `DEF-NR-TOOLBLOCKED-PARSER` source pins are repointed in the same
commit because they are not separable from the parser split — the
split alone leaves the suite red, and splitting them across commits
would violate the one-green-commit-per-step rule. They resolved the
parser by name, which after the split points at the wrapper and
asserts nothing. They now resolve the dispatching function by content
(the branch text, not the exception name: the name also appears in the
import block, so resolving on it left `test_branch_in_parser` green
after the branch was deleted — verified). Deleting the branch now
turns 13 of 14 red.
1543 passed, 1 skipped.
* docs(readme): state the three enforcement limitations, verified
The README carried no fail-open claim and no limitations section, so
the two most consequential behaviours of the gate were only
discoverable by reading the source. Adds both.
Each limitation was checked against the code before being written
down, and two of them are not what the shorthand description of them
suggests.
The 503 case fails OPEN on old SDKs, not closed. A failed security
check is served as a 503 carrying a `category`; the current SDK reads
it and refuses. Pre-category there was nothing to read, the 503 became
a synthetic FALLBACK decision, and `check_workflow_budget` fails open
on a FALLBACK source — so the call proceeds. Verified by probing both
cases through a real runtime and by reading the pre-category source
(`git show 83960b4~1:src/nullrun/transport.py`): the 5xx band had no
`is_refusal` condition and fell straight to the FALLBACK return.
The breaker setting is server-side, not an SDK default. `LogOnly` is
`NULLRUN_GATE_CB_TRIP_ENFORCEMENT_MODE` in the backend, not a fallback
mode in this SDK (there is no `LogOnly` anywhere in `src/`). The
production boot check refuses to start unless it is explicitly
`Enforce` or `LogOnly` (backend `main.rs:290-293`); `detect_mode()`
still defaults to `LogOnly` on unset outside production
(`main.rs:294-295`). In `LogOnly` the trip is recorded and alerted but
tripped workflows still pass `/check` — the backend's own words at
`main.rs:286-288`.
Pause and kill really are 403-only, and that is now stated as the
limitation it is: `WORKFLOW_PAUSED` and `WORKFLOW_INACTIVE` are served
as the same 403 from the same key, distinguishable only by
operator-facing text, which the backend deliberately keeps distinct
(`error_codes.rs:3400-3412`). Support tooling should key off the error
code, not the status.
Also adds the fail-open summary, matching the ADR-008 table in
`runtime.py` — which the module docstring requires to be kept in
lockstep with any README claim. There was no README claim to keep in
step before this; there is one now.
* fix(docs): name ADR-064's discriminator, not a marker that is not on the wire
The transport comment and two test docstrings repeated ADR-063 §4.7's
claim that the backend "already distinguishes the two 503 groups
internally via GateErrorCode::is_fail_closed()". That is wrong in the
way ADR-063 §4.7 was wrong: is_fail_closed is an ordinary Rust method
that is never serialised, so no client could have read it. The SDK
never did — the code has always keyed on decision == "block"
(categories.py:168), which is ADR-064's discriminator.
The comment also claimed BUDGET_DATA_UNAVAILABLE "arrives as 503 with
decision=block". It does not, and cannot: grepping the backend, that
string appears only in error_codes.rs and budget.rs, never in
internal.rs or orchestrator.rs, so no gate producer emits it. The one
response that does carry it (budget.rs:246,
BudgetUnavailableResponse) has no decision field at all — which is
correctly treated as a non-refusal and fails open per ADR-008.
Logic unchanged; this is a claim correction, and every ADR-063 §4.7
citation for the 503 rule retargeted to ADR-064, which now owns it.
44 passed.
* fix(gate): a /gate body must state WHO decided before it is a verdict
The refusal classifier splits a 5xx into "the gate ruled" and
"nothing ruled" by reading the body. Every other category change on
this branch rests on the same unexamined premise: that a body on the
SDK's HTTPS connection was written by NULLRUN.
A test for that premise found the premise is false in one direction.
`{"decision": "allow"}` from an arbitrary responder passes every check
`_require_gate_decision` applied -- it is a dict, and `decision` is a
known string -- and was then honoured as a gateway decision, because
the runtime's rule reads a MISSING decision_source as "not fallback"
and therefore as authoritative. So an on-path responder that gets
JSON in front of the SDK: a captive portal on a hotel WLAN, a
corporate TLS-interception proxy. It needs neither the API key nor a
HMAC bypass, only to answer before the real backend. The call is
authorised, /track books its cost against a policy nobody consulted,
the audit trail records an allow, and there is no later point at
which it is caught.
Fix: require `decision_source` to be a known provenance value BEFORE
looking at `decision`. It is the right field because it cannot be
absent from a real answer -- `GateResponse.decision_source` is a
non-Option String with no skip_serializing_if (`gate/internal.rs:638`).
`GateResponseBody` in schemas.rs declares it optional, but that struct
is referenced only from openapi.rs: it is the documentation schema,
not the wire.
Mutation-verified: with the guard absent, the new test reports DID
NOT RAISE on `{"decision": "allow"}`. 18 passed with it.
The same guard caught a drifted fixture, not a second defect:
test_v3_wire_contract.py's require_approval body omitted all three
non-Option fields on GateResponse. Corrected to the real shape.
Full suite 1561 passed / 1 skipped.
* docs(readme): state the provenance boundary of server-authoritative enforcement
Limitation 4. The three existing entries describe what the SDK does
not enforce; this one describes what it cannot verify. Every
enforcement claim in the README rests on a premise it never said out
loud: that the JSON on the HTTPS connection was written by NullRun.
The SDK now checks for it -- a /gate body with no usable
decision_source is rejected rather than acted on -- but that is a
field in the body, not a signature. An on-path responder that can
produce a well-formed body including decision_source defeats it, and
needs neither the API key nor an HMAC bypass. NullRun's answers are
not signed, so there is no after-the-fact detection either: the audit
trail would record the allow.
Also extends the fail-open summary with the not-a-verdict case, which
is newly user-visible as NullRunMalformedGateResponseError.
This is the honest boundary of "server-authoritative": authoritative
against a client, not against an attacker who owns the path.
* style: clear the three lint errors this branch's own test files carry
ruff runs over src/ only in CI, so none of these would have failed a
build -- but all three files are new on this branch (neither exists on
origin/master), so they are this branch's to leave clean.
Two F841: the `rt = make_runtime(...)` binding was never used, because
@Protect resolves the runtime from the context, not the local. The call
is load-bearing; the assignment was not.
One I001: nullrun.runtime sorted after nullrun.toolbox.
* docs(readme): name the trust boundary precisely, and the second version skew
Limitation 4 said "use https://", which understates the problem. It
implied https was sufficient, when the residual is specifically an
on-path responder holding a certificate the OS already trusts for
api.nullrun.io -- i.e. corporate TLS interception, which is ordinary
on a managed network and which https alone does nothing about.
Corrected to state what each layer does and does not give:
- certificate verification is on and CANNOT be disabled by
configuration. `verify_cert = True` (transport.py:572) and the only
override is NULLRUN_TLS_CA_CERT (:568), which swaps in a trust
anchor you chose and is still verification. Grepped exhaustively:
the only TLS env vars in the SDK are NULLRUN_TLS_CA_CERT /
_CLIENT_CERT / _CLIENT_KEY, none of which turns verification off.
- plain http is refused (InsecureTransportError, :526).
- responses are not signed, so no after-the-fact detection: the audit
trail would faithfully record an allow the gate never gave.
The consequence is stated as the operational control it is -- exclude
api.nullrun.io from interception and verify it stays excluded -- and
not as something the SDK does for you.
Also adds the old-SDK note beside limitation 1, since it is the same
version skew pointing the same way and a reader who pins for the 503
guarantee should know they also need it for the forged-allow case:
both a real refusal and a fabricated permission are read permissively
on an older SDK, and neither shows up in the SDK's output.
* test(protect): prove refusals and provenance at the decorator level
Every other file in this cluster calls runtime.check_workflow_budget()
directly. That is the wrong level for the question a user has, because
@Protect is what they call -- and @Protect is a separate code path with
its own try, its own except BaseException, and its own two wrappers
(sync and async).
Reading decorators.py says the refusals propagate. That is exactly the
reasoning that let DEF-MP-TS12-ENF-01 ship: check_workflow_budget's
except Exception looked harmless in isolation too, until it was reached
by a real 401.
14 tests, four properties:
- a denied / budget / unclassifiable refusal, and a body with no
decision_source, do NOT run the decorated body;
- the counter-tests: a real allow, a connection failure under
on_denied="message", and a 502 with no refusal body all DO run it
(a suite of only "did not run" assertions passes against a
decorator that refuses everything);
- on_denied reaches @Protect for category=denied and, proven by the
negative parametrisation, does not reach budget/halt;
- the async wrapper is not a separate property, so the refusal, the
forged allow and the real allow are each proven on it too.
Mutation-verified: deleting the decision_source guard in
_require_gate_decision turns exactly the sync and async forged-allow
tests red (DID NOT RAISE), and the other 12 stay green.
Verified in a clean python:3.11 container on the unmutated tree:
ruff check src/ tests/ -- All checks passed
mypy src/ -- no issues in 37 source files
pytest -q -- 1575 passed, 1 skipped
* docs(changelog): record the second half of DEF-MP-TS12-ENF-01
0.19.0 is already on origin/master and its entry closed the paths
found by reading the code. These commits closed the paths that only
appeared once the properties were asserted end-to-end, and none of
them were in the release notes -- so the two most serious items
(a forged allow authorised by decision_source absence, and the
production-usable sensitive fail-open opt-out) were documented only in
README.
Filed as [Unreleased] rather than inventing a version number. That is
deliberate: the behaviour change is real and needs a migration note,
not a patch number chosen by accident. An unclassifiable refusal now
raises where 0.19.0 let the call proceed, and a hand-rolled /gate
test double must now carry decision_source. Whoever cuts the release
picks 0.19.1 or 0.20.0 with that in hand.
Verified citation: GateResponse.decision_source is a non-Option String
at backend/src/proxy/http/gate/internal.rs:638.
* docs(changelog): migration note for the refusal-category SDK
Six differences from 0.19.0, written as things you can act on rather
than as a list of changed internals. The third is the one worth
reading twice: an unclassifiable refusal now raises where 0.19.0 let
the call proceed, and NullRunUnclassifiedRefusalError is deliberately
a SIBLING of NullRunTransportError, not a subclass -- so an existing
`except NullRunTransportError:` fail-open arm will not catch it.
That is the fix working, and it is also the thing most likely to turn
a working loop into a throwing one.
The note gives the exact import path (nullrun.breaker.categories --
it is not re-exported at the top level, which was verified, not
assumed), the error_code, and the three options as three, with the
consequence of each stated. Widening the arm to NullRunError is
called out as worse than the hole being closed, because it also
swallows real policy refusals.
Every code shape in the note was executed against the installed
package rather than reasoned about:
recommended two-arm -> propagates (recommended)
one-arm permissive -> swallowed == 0.19.0 behaviour, bypass restored
widened to NullRunError -> a real policy refusal is swallowed too
issubclass(NRError, NullRunError) = True
* fix(protect): make @tool/@Protect ordering irrelevant
Priority 1 for DEF-MP-TS12-ENF-01 was "one @Protect, rest under the
hood", and the part nobody had tested is whether that survives the
framework layer. It did not, in one direction.
`@protect` above `@tool` collapsed the tool to a plain function.
`functools.wraps`-based wrapping cannot preserve `.invoke`, `.name`,
`.args_schema` or `.description`, so the agent loop could not bind it --
and a tool the loop cannot see produces no refusal. Measured on
langchain-core 0.3.86:
@Protect # outer
@tool # inner -> StructuredTool
def f(query: str) -> str: ...
# f is now <function>, not StructuredTool
# convert_to_openai_tool(f) -> NameError: name 'Annotated' is not defined
That last line is the failure mode worth naming: the error a user sees
is a langchain type-resolution error, not "your tool is not a tool any
more". Enforcement was gone and the symptom pointed somewhere else.
`protect` now duck-types the argument (`.invoke` + `.name`, then
`.coroutine` before `.func` -- an async tool sets both, with `func` as a
sync fallback that raises if called), wraps whichever callable is live,
and returns the SAME object. Every other attribute is untouched, so the
loop sees the tool it saw before. Duck-typed rather than isinstance
because `langchain_core` is an optional dependency; the three attributes
matched are the ones `BaseTool.run()`/`arun()` dispatch through.
Three properties were established by probe on the real library before
being written down, and all three are now pinned:
- Ordering is irrelevant in BOTH directions. `@tool` outside already
worked; the reverse now does too, and each has an allow counter-test
so "the body did not run" cannot pass against a wrapper that refuses
everything.
- `handle_tool_error=True` cannot turn a refusal into model-visible
text. It only catches `ToolException`; a `NullRunBudgetError` reaches
the bare `except (Exception, KeyboardInterrupt)` arm and is re-raised.
This is the DEF-MP-TS12-ENF-01 shape in a different costume, so it is
pinned rather than assumed.
- `@protect` with no `init()` fails LOUD. The lazy resolver propagates
`NullRunAuthenticationError` (NR-A001) and the body does not run.
FIX-4 removed the fallback that used to swallow this; the AST guard
keeps it removed, and an AST check rather than a substring search
because the docstring describes the very `except` being asserted gone.
`on_denied="message"` is now exercised through a real `@tool` too, in
both orders and both sync/async, with the negative set (budget, halt)
kept at the framework layer: an agent told "that tool is not allowed"
when the truth is "you are out of money" will look for another way to
spend.
Mutation-verified; each mutant ran the suite and the restore was
byte-identical:
sync arm converts refusal to text -> 2 red
async arm converts refusal to text -> 1 red
FIX-4 fallback in lazy resolver -> 2 red
BaseTool in-place wrap removed -> 5 red
on_denied category guard dropped -> 8 red across 3 files
Three earlier mutation attempts were invalid and are not counted: an
inner `try` cannot observe a `_protect_body.__enter__` raise, and one
anchor matched `get_protected_runtime` instead of
`_get_or_create_runtime`.
Suite: 1597 passed, 1 skipped. ruff and mypy clean.
* release(sdk): 0.20.0 — migration note first, ordering documented
Minor, not patch: an unclassifiable refusal now RAISES where 0.19.0 let
the call proceed. That turns a working loop into a throwing one, which
is a behavioural break, so the release note leads with the migration
rather than burying it under the security items.
`### Migration` is the first section under 0.20.0 and enumerates all six
behavioural differences from 0.19.0, then repeats the three inherited
from 0.19.0 that are the same class of break and the same shape of fix
(`mode="inline"` removed, `decision_source` now required in a `/gate`
test double, `handle` -> `guard`). A reader upgrading from 0.18.x sees
the whole break list in one place instead of reconstructing it across
two releases.
The load-bearing item is #3: `NullRunUnclassifiedRefusalError` is a
sibling of `NullRunTransportError`, not a subclass, so an existing
`except NullRunTransportError:` arm will not catch it. Catching it
alongside restores 0.19.0's behaviour exactly -- which restores the
bypass. The note says so and says not to widen to `NullRunError`, which
is a larger hole than the one this closes.
Also records the `@tool`/`@protect` fix under `### Fixed`, since it is a
silent enforcement loss users may have shipped against: enforcement was
absent whenever `@protect` was the outer decorator on a LangChain tool.
README gains the two things a user would otherwise get wrong:
- decorator order, both directions, with the reason (an unbound tool
cannot refuse) and the `NameError` that was the real symptom;
- `on_denied="message"` as the operator-facing "explain, don't crash"
mode -- `category="denied"` only, budget and halt deliberately
excluded -- and why LangChain's `handle_tool_error=True` is not an
equivalent (it catches `ToolException` and does not know which
exceptions are refusals).
Every claim in the new README section was executed before being written
down, including the pre-fix `NameError`, which is reproduced rather than
asserted from memory.
Not published to PyPI. Suite: 1597 passed, 1 skipped; ruff and mypy
clean.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Why
How
Test plan
cd backend && cargo test,cd frontend && npm test)cd frontend && npm run lint)cd frontend && npm run type-check)Risk
Checklist
CONTRIBUTING.md(if present)