Skip to content

chore(vendor): bump tinytools to 82c0d97 and tinyinference to c29d511, with harness acceptance tests - #226

Merged
M3gA-Mind merged 2 commits into
tinyhumansai:mainfrom
M3gA-Mind:chore/bump-tinytools-tinyinference
Sep 28, 2026
Merged

M3gA-Mind merged 2 commits into
tinyhumansai:mainfrom
M3gA-Mind:chore/bump-tinytools-tinyinference

Conversation

@M3gA-Mind

@M3gA-Mind M3gA-Mind commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Refs tinyhumansai/openhuman#6722
Refs tinyhumansai/openhuman#6724

Bumps the two vendored submodules so the parser fixes and the stream-error classification reach tinyagents' consumers. It also adds harness-level acceptance tests that fail at the previous pins.

$ git diff upstream/main --submodule=log | grep ^Submodule
Submodule vendor/tinyinference 6aff95c5..c29d5118
Submodule vendor/tinytools 52e9ab10..82c0d975

No other submodule moves, and Cargo.lock is unchanged.

Behaviour changes

tinytools 52e9ab1..82c0d97 (all in tinytools-agent, text tool-call parsing and stream scrubbing)

PR Change
#25 Whitespace before the DSML/namespace prefix (< | DSML | invoke) is tolerated, so such calls now parse.
#26 The stream scrubber holds back a bare < / </ followed only by whitespace, so a fragment boundary there no longer releases the bracket as text and drops the call.
#27 An orphaned tag-family closer (a habitual </tool_call> with no opener) is swept from visible text instead of leaking into the reply.
#29 A named <invoke name=…> wrapped in <tool_call>…</tool_call> is decoded. Previously the block was claimed and dropped as malformed (openhuman#6722). The anchor keeps an invoke quoted inside other body text from executing.
#30 A fence whose info string opens with a call tag (```<tool_call>, a named or bare <invoke>) is a call fence, not a protected example. <function …> and XML-namespaced tags are excluded, and language-tagged fences stay protected.
#31 New registry-gated element grammar: <NAME><param>…</param></NAME> decodes to a call when NAME is an offered tool with a registry entry (code/P-Format dialects) and the body is only parameter children. It reports MalformedBlock { source: Element } when claimed but undecodable, which the #224 nudge counts. Anything else stays in the text. The stream no longer stalls on an unclaimable element, and reserved names (tool_call, invoke) are never parameter children. Adds CallSource::Element (the enum is #[non_exhaustive]).

tinyinference 6aff95c..c29d511

Two changes to tinyinference-llm, which tinyagents uses:

PR Change
#37 A streamed error event with a numeric code in 400–599 ({"error":{"code":400,…}} inside an HTTP 200 stream) is read as the HTTP status and classified by it. A 4xx is now non-retryable (it was retried as the default before), and a 5xx stays retryable. A non-status number keeps the message heuristics (openhuman#6724).
#28 For a non-success HTTP response, the Retry-After header is read before the body and carried as ProviderError::retry_after_ms. The harness retry layer already honours it, capped by RetryPolicy::max_retry_after_ms.

Everything else in the range adds the new tinyinference-decisions crate and its docs. tinyagents does not depend on it (no Cargo.toml references it), so it has no effect here.

Acceptance tests (through the harness, not the parser alone)

Test Asserts At the old pins
a_named_invoke_wrapped_in_tool_call_is_dispatched_under_the_python_dialect (e2e_tool_dialects.rs) The sanitized production shape behind openhuman#6722: a todo element closed by a stray </tool_call>, then ```<tool_call> and a named <invoke> with string= attributes. The search_repositories call is dispatched with q, sort and per_page: 20 decoded. tinytools 52e9ab1: 0 dispatched (left 0, right 1)
a_todo_element_call_is_dispatched_under_the_python_dialect <todo><todos>[…]</todos></todo> is dispatched with both items. tinytools 52e9ab1: 0 dispatched
a_streamed_numeric_400_error_fails_on_the_first_attempt (e2e_stream_error_code.rs, a real OpenAiModel against a loopback SSE server) The run fails, and the server saw exactly 1 request. tinyinference 6aff95c: 4 requests (left 4, right 1)
a_streamed_numeric_503_error_is_still_retried (control) The run fails after more than 1 request, so the 400 result is not simply "retries off". passes at both pins

The tests use the Python code dialect (ToolDispatcher::Python) and a recording tool with a real schema, so they assert what the harness actually dispatched.

Lanes run locally at this head:

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • the full tinyagents-integration-tests package
  • tinyagents-harness, tinyagents-graph and tinyagents-orchestration tests

All pass, and --list confirms the new tests are in the binary.

Why a fenced call is asserted as dispatched

The record-27 fixture opens a fence with ```<tool_call>. That is how the model wrote its real call in production (openhuman#6722): the call tag was the fence's info string, and the fence was never closed. tinytools #30 treats a fence whose info string opens with a call tag as a call fence. The quoted-example protection is unchanged: a language-tagged fence (```xml, ```xml<tool_call>, ```text <tool_call>) still protects its contents, and a_language_tagged_fenced_call_is_not_dispatched_unary_or_streamed (tinyagents #225) pins that on both the unary and streamed paths. The harness-local fence guard that ran only on the unary path was removed in #225 with maintainer sign-off (openhuman#6732). The residual risk of a quoted call executing on the streamed path is openhuman#6733, filed for a product decision.

CI note: tinysweeper critique, security and description findings on 4064174c all dispute this fence policy; each is answered on its thread. These lanes are advisory; the required check is CI.

tinytools 52e9ab1..82c0d97 (tinyhumansai#25, tinyhumansai#26, tinyhumansai#27, tinyhumansai#29, tinyhumansai#30, tinyhumansai#31): text tool-call
parsing fixes, including named <invoke> wrapped in <tool_call>, a call tag
on a fence line, and element-form calls for offered tools.

tinyinference 6aff95c..c29d511 (tinyhumansai#28, tinyhumansai#37 in tinyinference-llm; the rest is
the new, unused tinyinference-decisions crate): a streamed numeric error
code is read as the HTTP status, and a 429's Retry-After reaches the
provider error.

Refs tinyhumansai/openhuman#6722
Refs tinyhumansai/openhuman#6724
…arness

- A named <invoke> wrapped in <tool_call> (the sanitized production shape)
  and a todo element call are dispatched under the Python code dialect.
- A streamed {"error":{"code":400}} fails the run on the first attempt,
  while a streamed 503 is still retried (control).

Each fails at the previous submodule pins.

Refs tinyhumansai/openhuman#6722
Refs tinyhumansai/openhuman#6724
@tinysweeper

tinysweeper Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Adds acceptance tests for stream error codes and Python dialect tool dispatch, but includes test examples that embed tool calls inside bare code fences, reversing an earlier security guard (fenced-code-safety). The reviewer advises against merging until this is addressed.

State: Changes requested
Priority: high
Reviewed head: 4064174ca866
Updated: 1790606934 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 0 Active findings 1
Tests 2 Noted findings 0
Documentation 0 Resolved findings 12
Configuration 0 Pending checks/questions 0

Completeness: Complete
Test assessment: Test coverage is assessed from changed tests and lane evidence; execution is not claimed without trusted check data.

What changed

Bumps vendored submodules (tinytools to 82c0d97, tinyinference to c29d511) and adds two new integration test files for harness acceptance tests: e2e_stream_error_code.rs and e2e_tool_dialects.rs.

Features

None identified with supported citations.

Tests

  • integration — Verifies that a provider error with code 400 inside a 200 SSE stream causes the harness run to fail and that exactly one provider request is made (no retry).: No issues identified. (crates/tinyagents-integration-tests/tests/e2e_stream_error_code.rs)
  • integration — Verifies that a provider error with code 503 inside a 200 SSE stream causes the run to fail but is retried (more than one request).: No issues identified. (crates/tinyagents-integration-tests/tests/e2e_stream_error_code.rs)
  • integration — Verifies that a tool call wrapped in a `<invoke>` element with named parameters (including `string="false"` decoding as JSON) is dispatched correctly under the Python tool dialect.: Contains a test example where the tool call is inside a bare code fence, flagged as a security concern (fenced-code-safety). (crates/tinyagents-integration-tests/tests/e2e_tool_dialects.rs)
  • integration — Verifies that a `<todo>` element call (with `<param>`) is dispatched once under the Python tool dialect.: No direct issue identified for this test alone, but the test file overall has a critique regarding fenced tool-call examples. (crates/tinyagents-integration-tests/tests/e2e_tool_dialects.rs)

Findings

  • high · critique · Keep fenced tool calls out of dispatch — This fixture places the complete `search_repositories` call inside a Markdown fence, beginning with ```` ```<tool_call> ```` and ending after `</tool_call>`, but the test requires (crates/tinyagents\-integration\-tests/tests/e2e\_tool\_dialects\.rs:1773)

Resolved this pass

  • Keep fenced tool-call examples out of dispatch
  • Keep fenced tool markup out of dispatch
  • Keep fenced tool calls out of dispatch
  • Keep bare fenced markup from triggering tool execution
  • Keep fenced tool-call examples out of dispatch
  • Keep fenced tool markup out of dispatch
  • Keep fenced tool calls out of dispatch
  • Keep bare fenced markup from triggering tool execution
  • Keep fenced tool-call examples out of dispatch
  • Keep fenced tool markup out of dispatch
  • Keep fenced tool calls out of dispatch
  • Keep bare fenced markup from triggering tool execution

Before merge

  • Address Keep fenced tool calls out of dispatch (crates/tinyagents\-integration\-tests/tests/e2e\_tool\_dialects\.rs).

How this fits together

flowchart LR
  n0["new"]:::impacted
  n1["push_middleware"]:::impacted
  n2["harness_with"]:::impacted
  n3["set_default_model"]:::impacted
  n2 -->|calls| n0
  n2 -->|calls| n1
  n2 -->|tests| n1
  n2 -->|calls| n3
  n2 -->|tests| n3
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 2 files; 1 finding. _The code index is behind this pull request (indexed at `6d3f1d68cffe`), so retrieved context may be out of date._ _5 memory call(s) failed (model: cortex: v1/recall: timed out after 10s), so this review saw part of what the engine holds._
  • Evidence: crates/tinyagents\-integration\-tests/tests/e2e\_tool\_dialects\.rs — Keep fenced tool calls out of dispatch

security

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 2 files; 1 finding. (1 already reported on an earlier push) _The code index is behind this pull request (indexed at `6d3f1d68cffe`), so retrieved context may be out of date._ _5 memory call(s) failed (model: cortex: v1/recall: timed out after 10s), so this review saw part of what the engine holds._

tests

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No behavioural change: nothing outside documentation, configuration and tests.

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This PR bumps tinytools and tinyinference submodules (parser fixes, streamed error classification) and adds harness-level acceptance tests that verify the new behavior. The tests are well-structured and the code looks correct; no new issues are introduced. _The code index is behind this pull request (indexed at `6d3f1d68cffe`), so retrieved context may be out of date._ _5 memory call(s) failed (model: cortex: v1/recall: timed out after 10s), so this review saw part of what the engine holds._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No behavioural change: nothing outside documentation, configuration and tests.
Evidence and run details
  • Models: ladder/vectors, gpt-5.6-luna, deepseek-v4-flash
  • Spend: $0.026302
  • Tokens: 262460 input · 13975 output · 31433 cached · 780 embedding
Head State Pass summary
4064174ca866 changes requested 4 active finding(s), 0 resolved finding(s) (at 1790606688)
4064174ca866 changes requested 1 active finding(s), 12 resolved finding(s) (at 1790606934)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

This review includes 4 billable files and costs up to $1.00.

Or wait 14 minutes for your next included review.

Check out review usage here.

View limit details

Limit details: You’ve used all 2 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 9304c928-ce3c-4093-b2fc-5632f99c239c

📥 Commits

Reviewing files that changed from the base of the PR and between 63d5e9d and 4064174.

📒 Files selected for processing (4)
  • crates/tinyagents-integration-tests/tests/e2e_stream_error_code.rs
  • crates/tinyagents-integration-tests/tests/e2e_tool_dialects.rs
  • vendor/tinyinference
  • vendor/tinytools

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 3 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0354 · 320,918 in / 16,084 out · 22,243 cached (7%) · ladder/vectors, gpt-5.6-luna, deepseek-v4-flash · 780 embedded
critique:    $0.0220 · 178,374 in / 9,118 out  · 14,053 cached (8%) · gpt-5.6-luna, deepseek-v4-flash
security:    $0.0130 · 126,679 in / 2,503 out  · 7,166 cached (6%)  · gpt-5.6-luna
description: $0.0002 · 10,589 in  / 1,449 out  · 1,024 cached (10%) · deepseek-v4-flash

Comment thread crates/tinyagents-integration-tests/tests/e2e_tool_dialects.rs
Comment thread crates/tinyagents-integration-tests/tests/e2e_tool_dialects.rs
@tinysweeper tinysweeper Bot added the priority: p1 Next. Wrong behaviour a user will hit, or a security weakness behind a condition. label Sep 28, 2026
@M3gA-Mind

Copy link
Copy Markdown
Collaborator Author

Independent audit at 4064174c (comment only).

Static:

  • git diff --submodule vs tinyagents main 63d5e9da: only vendor/tinytools 52e9ab10 → 82c0d975 and vendor/tinyinference 6aff95c5 → c29d5118 move.
  • Both targets are exactly their repos' current main.
  • Both moves are forward-only (tinytools +15 / 0 behind, tinyinference +27 / 0 behind).
  • Cargo.lock is unchanged. The only other files are the two new acceptance-test files.

Revert-check. I ran it myself (isolated target dir), not taking the report on trust:

pins e2e_tool_dialects: python-dialect dispatch ×2 e2e_stream_error_code: 400 503 (control)
PR head (82c0d97 / c29d511) ✅ ✅ ✅ ✅
tinytools → 52e9ab1 ❌ :1777 "the wrapped invoke is dispatched once" (0 vs 1), ❌ :1794 "the element call is dispatched once" (0 vs 1) n/a n/a
tinyinference → 6aff95c n/a ❌ :110 "a deterministic 400 is not retried: exactly one provider request" (4 vs 1) ✅

Each acceptance test fails at the old pin on its own assertion, and the 503 control stays green at both pins. So the tests pin exactly the behaviour the bump delivers.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0263 · 262,460 in / 13,975 out · 31,433 cached (12%) · ladder/vectors, gpt-5.6-luna, deepseek-v4-flash · 780 embedded
critique:    $0.0130 · 121,410 in / 6,558 out  · 18,124 cached (15%) · gpt-5.6-luna, deepseek-v4-flash
security:    $0.0129 · 120,195 in / 3,687 out  · 7,165 cached (6%)   · gpt-5.6-luna
description: $0.0002 · 10,032 in  / 1,687 out  · 1,536 cached (15%)  · deepseek-v4-flash

Comment thread crates/tinyagents-integration-tests/tests/e2e_tool_dialects.rs
@M3gA-Mind
M3gA-Mind merged commit 8a58b7e into tinyhumansai:main Sep 28, 2026
14 of 16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p1 Next. Wrong behaviour a user will hit, or a security weakness behind a condition.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant