chore(vendor): bump tinytools to 82c0d97 and tinyinference to c29d511, with harness acceptance tests - #226
Conversation
tinytools 52e9ab1..82c0d97 (tinyhumansai#25, tinyhumansai#26, tinyhumansai#27, tinyhumansai#29, tinyhumansai#30, tinyhumansai#31): text tool-call parsing fixes, including named <invoke> wrapped in <tool_call>, a call tag on a fence line, and element-form calls for offered tools. tinyinference 6aff95c..c29d511 (tinyhumansai#28, tinyhumansai#37 in tinyinference-llm; the rest is the new, unused tinyinference-decisions crate): a streamed numeric error code is read as the HTTP status, and a 429's Retry-After reaches the provider error. Refs tinyhumansai/openhuman#6722 Refs tinyhumansai/openhuman#6724
…arness
- A named <invoke> wrapped in <tool_call> (the sanitized production shape)
and a todo element call are dispatched under the Python code dialect.
- A streamed {"error":{"code":400}} fails the run on the first attempt,
while a streamed 503 is still retried (control).
Each fails at the previous submodule pins.
Refs tinyhumansai/openhuman#6722
Refs tinyhumansai/openhuman#6724
Tiny Sweeper reviewAdds acceptance tests for stream error codes and Python dialect tool dispatch, but includes test examples that embed tool calls inside bare code fences, reversing an earlier security guard (fenced-code-safety). The reviewer advises against merging until this is addressed. State: Changes requested Review snapshot
Completeness: Complete What changedBumps vendored submodules (tinytools to 82c0d97, tinyinference to c29d511) and adds two new integration test files for harness acceptance tests: e2e_stream_error_code.rs and e2e_tool_dialects.rs. FeaturesNone identified with supported citations. Tests
Findings
Resolved this pass
Before merge
How this fits togetherflowchart LR
n0["new"]:::impacted
n1["push_middleware"]:::impacted
n2["harness_with"]:::impacted
n3["set_default_model"]:::impacted
n2 -->|calls| n0
n2 -->|calls| n1
n2 -->|tests| n1
n2 -->|calls| n3
n2 -->|tests| n3
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
|
Warning Review limit reached
This review includes 4 billable files and costs up to $1.00. Or wait 14 minutes for your next included review. View limit detailsLimit details: You’ve used all 2 included reviews currently available. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (4)
Comment |
There was a problem hiding this comment.
Requesting changes: 3 lane(s) blocking, worst finding is high.
Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.
$0.0354 · 320,918 in / 16,084 out · 22,243 cached (7%) · ladder/vectors, gpt-5.6-luna, deepseek-v4-flash · 780 embedded
critique: $0.0220 · 178,374 in / 9,118 out · 14,053 cached (8%) · gpt-5.6-luna, deepseek-v4-flash
security: $0.0130 · 126,679 in / 2,503 out · 7,166 cached (6%) · gpt-5.6-luna
description: $0.0002 · 10,589 in / 1,449 out · 1,024 cached (10%) · deepseek-v4-flash
|
Independent audit at Static:
Revert-check. I ran it myself (isolated target dir), not taking the report on trust:
Each acceptance test fails at the old pin on its own assertion, and the 503 control stays green at both pins. So the tests pin exactly the behaviour the bump delivers. |
There was a problem hiding this comment.
Requesting changes: 2 lane(s) blocking, worst finding is high.
Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.
$0.0263 · 262,460 in / 13,975 out · 31,433 cached (12%) · ladder/vectors, gpt-5.6-luna, deepseek-v4-flash · 780 embedded
critique: $0.0130 · 121,410 in / 6,558 out · 18,124 cached (15%) · gpt-5.6-luna, deepseek-v4-flash
security: $0.0129 · 120,195 in / 3,687 out · 7,165 cached (6%) · gpt-5.6-luna
description: $0.0002 · 10,032 in / 1,687 out · 1,536 cached (15%) · deepseek-v4-flash
Refs tinyhumansai/openhuman#6722
Refs tinyhumansai/openhuman#6724
Bumps the two vendored submodules so the parser fixes and the stream-error classification reach tinyagents' consumers. It also adds harness-level acceptance tests that fail at the previous pins.
No other submodule moves, and
Cargo.lockis unchanged.Behaviour changes
tinytools
52e9ab1..82c0d97(all intinytools-agent, text tool-call parsing and stream scrubbing)< | DSML | invoke) is tolerated, so such calls now parse.</</followed only by whitespace, so a fragment boundary there no longer releases the bracket as text and drops the call.</tool_call>with no opener) is swept from visible text instead of leaking into the reply.<invoke name=…>wrapped in<tool_call>…</tool_call>is decoded. Previously the block was claimed and dropped as malformed (openhuman#6722). The anchor keeps an invoke quoted inside other body text from executing.```<tool_call>, a named or bare<invoke>) is a call fence, not a protected example.<function …>and XML-namespaced tags are excluded, and language-tagged fences stay protected.<NAME><param>…</param></NAME>decodes to a call whenNAMEis an offered tool with a registry entry (code/P-Format dialects) and the body is only parameter children. It reportsMalformedBlock { source: Element }when claimed but undecodable, which the #224 nudge counts. Anything else stays in the text. The stream no longer stalls on an unclaimable element, and reserved names (tool_call,invoke) are never parameter children. AddsCallSource::Element(the enum is#[non_exhaustive]).tinyinference
6aff95c..c29d511Two changes to
tinyinference-llm, which tinyagents uses:codein 400–599 ({"error":{"code":400,…}}inside an HTTP 200 stream) is read as the HTTP status and classified by it. A 4xx is now non-retryable (it was retried as the default before), and a 5xx stays retryable. A non-status number keeps the message heuristics (openhuman#6724).Retry-Afterheader is read before the body and carried asProviderError::retry_after_ms. The harness retry layer already honours it, capped byRetryPolicy::max_retry_after_ms.Everything else in the range adds the new
tinyinference-decisionscrate and its docs. tinyagents does not depend on it (noCargo.tomlreferences it), so it has no effect here.Acceptance tests (through the harness, not the parser alone)
a_named_invoke_wrapped_in_tool_call_is_dispatched_under_the_python_dialect(e2e_tool_dialects.rs)todoelement closed by a stray</tool_call>, then```<tool_call>and a named<invoke>withstring=attributes. Thesearch_repositoriescall is dispatched withq,sortandper_page: 20decoded.52e9ab1: 0 dispatched (left 0, right 1)a_todo_element_call_is_dispatched_under_the_python_dialect<todo><todos>[…]</todos></todo>is dispatched with both items.52e9ab1: 0 dispatcheda_streamed_numeric_400_error_fails_on_the_first_attempt(e2e_stream_error_code.rs, a realOpenAiModelagainst a loopback SSE server)6aff95c: 4 requests (left 4, right 1)a_streamed_numeric_503_error_is_still_retried(control)The tests use the Python code dialect (
ToolDispatcher::Python) and a recording tool with a real schema, so they assert what the harness actually dispatched.Lanes run locally at this head:
cargo fmt --all -- --checkcargo clippy --workspace --all-targets -- -D warningstinyagents-integration-testspackagetinyagents-harness,tinyagents-graphandtinyagents-orchestrationtestsAll pass, and
--listconfirms the new tests are in the binary.Why a fenced call is asserted as dispatched
The record-27 fixture opens a fence with
```<tool_call>. That is how the model wrote its real call in production (openhuman#6722): the call tag was the fence's info string, and the fence was never closed. tinytools #30 treats a fence whose info string opens with a call tag as a call fence. The quoted-example protection is unchanged: a language-tagged fence (```xml,```xml<tool_call>,```text <tool_call>) still protects its contents, anda_language_tagged_fenced_call_is_not_dispatched_unary_or_streamed(tinyagents #225) pins that on both the unary and streamed paths. The harness-local fence guard that ran only on the unary path was removed in #225 with maintainer sign-off (openhuman#6732). The residual risk of a quoted call executing on the streamed path is openhuman#6733, filed for a product decision.CI note: tinysweeper critique, security and description findings on
4064174call dispute this fence policy; each is answered on its thread. These lanes are advisory; the required check isCI.