Skip to content

fix(harness): let hosts exclude replayed history from user-input screening - #223

Merged
senamakel merged 4 commits into
tinyhumansai:mainfrom
M3gA-Mind:fix/6710-screen-new-user-input
Sep 28, 2026
Merged

senamakel merged 4 commits into
tinyhumansai:mainfrom
M3gA-Mind:fix/6710-screen-new-user-input

Conversation

@M3gA-Mind

@M3gA-Mind M3gA-Mind commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

prepare_agent_turn screens every user message in an AgentTurnRequest as ContentOrigin::User, on every turn. A host whose text tool dialect replays tool results as user rows therefore re-screens those results under the user-input policy on each turn.

The results were already screened one at a time as Tool output when they ran (agent_loop/tools.rs). Coalesced into a single row, though, they can cross the threshold. On a real OpenHuman thread, four GitHub-issue results scored 0.18, 0.48, 0.42 and 0.18 on their own; combined into one row they scored 0.72, which blocks. Because that row lives in the durable transcript, every later turn fails before its first model call and the thread never recovers. Details are in tinyhumansai/openhuman#6710.

Refs tinyhumansai/openhuman#6710

This PR adds AgentInvocation::with_replayed_prefix(n). The host declares that the first n messages are replayed history it already screened when it admitted them. Only the messages after them are screened as this turn's user input, and only those feed the memory/experience recall query.

Why a boundary set by the host and not an inferred rule such as "last message only": another embedder may legitimately send several new user messages in one request, and an inferred rule would quietly stop screening them.

API Or Behavior Changes

  • New builder: AgentInvocation::with_replayed_prefix(count). It sets a private field; AgentTurnRequest is unchanged from main.
  • Source-compatible. AgentInvocation already has a private field (runtime), so no downstream code can build it with a struct literal. Adding another private field breaks nothing. (An earlier revision put a public field on AgentTurnRequest, which would have broken struct literals.)
  • Default 0 = unchanged behaviour: every user message is screened. A recursive child invocation (from_shared_host) always uses 0.
  • A prefix greater than the message count is rejected with a Validation error; it is always a host bug. A prefix equal to the count screens nothing, which is legitimate for a deferred resume. It is accepted, with a warn! when the last message is a User message.
  • The screened slice's text (input_text) also feeds the memory/experience recall query and TurnSummary::with_text / Experience::new in finish_host_turn. For a host that opts in, those inputs narrow to this turn's new messages.
  • Redaction contract (documented on with_replayed_prefix): opt in only if your replayed history stores user rows exactly as the gate returned them (redacted), e.g. the messages of the AgentRun that admitted them.

Security

  • New input is still screened as ContentOrigin::User.
  • Tool output is still screened one result at a time as Tool when the tool runs. That screen is untouched.
  • Replayed messages are skipped only when the host opts in. By opting in, the host asserts it screened them when it first admitted them.
  • The default is unchanged.

Tests

New tests in runtime/test.rs:

  • hosted_turn_screens_only_the_new_user_input_not_replayed_history: blocked content sits in replayed [Tool results]-style rows, including an interrupted-turn shape where a results row is directly followed by the new input. The turn must succeed.
  • hosted_turn_rejects_a_replayed_prefix_past_the_messages: with the old clamp restored, it fails at its expect_err.
  • hosted_turn_accepts_a_replayed_prefix_covering_every_message: the deferred-resume shape is not rejected.
  • hosted_turn_replays_the_redacted_form_of_an_admitted_user_row: replays the first run's transcript. If the host instead replays the raw input, it fails at !second.contains("secret").
  • tinyagents-runtime: driver_history_is_the_committed_history_plus_exactly_the_turn_input. A host that sets with_replayed_prefix(len - 1) relies on each turn appending exactly one input after the committed history. This test pins that behaviour. Fault check: when the runtime pushes an extra row, the test fails.
  • hosted_turn_still_blocks_new_user_input_and_screens_all_by_default: three cases.
    • First turn: the blocked input is rejected.
    • New input after a replayed prefix: the blocked input is rejected.
    • No prefix set: blocked content in an earlier row is still rejected, which pins the default.

Revert-check: with only the slice in prepare_agent_turn reverted, the first test fails at its .expect("replayed history must not be re-screened as new user input") and the second still passes.

Commands run locally:

  • cargo fmt --all --check
  • cargo clippy -p tinyagents-harness --tests -- -D warnings
  • cargo test -p tinyagents-harness --lib -- runtime::test::hosted_ (16 passed), plus cargo test -p tinyagents-runtime --lib -- test::driver_history_is test::trailing_input (2 passed) and cargo clippy -p tinyagents-runtime --tests -- -D warnings, at b8b0818
  • cargo clippy --all-targets -- -D warnings: not run, left to CI
  • cargo clippy --all-targets --all-features -- -D warnings: not run, left to CI
  • cargo build --all-targets: not run, left to CI
  • cargo build --all-targets --all-features: not run, left to CI
  • cargo test: not run in full, left to CI
  • cargo test --all-features: not run, left to CI

Documentation

Rustdoc on the new field and builder, plus an updated rustdoc for screen_user_messages. No other docs describe screening scope.

Summary by CodeRabbit

  • Bug Fixes
    • Prevented previously replayed conversation history from being screened again, while continuing to screen new user messages.
    • Corrected turn history so each request includes committed prior history and the current user input, without prematurely including the assistant’s response.
    • Added validation to reject replay boundaries that exceed the number of messages.

…ening

`prepare_agent_turn` screened every user message in the request as
`ContentOrigin::User` on every turn. A host whose text tool dialect
replays tool results as user rows re-screens those results under the
user policy each turn. The results were already screened one by one as
`Tool` output when they ran, but coalesced into one row they can cross
the threshold. That turns into a permanent block: the row lives in the
durable transcript, so every later turn fails before the first model
call (openhuman#6710).

Add `AgentTurnRequest::with_replayed_prefix(n)`. The first `n` messages
are history the host already screened when it admitted them, so only
the messages after them are screened as this turn's input and used as
the recall query. The default of 0 screens every user message, exactly
as before, so hosts that do not opt in lose no screening.
@tinysweeper

tinysweeper Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lanes and found 3 active actionable findings. The change adds `AgentInvocation::with_replayed_prefix` to let hosts exclude replayed history from user-input screening. Two earlier findings about API stability and authorization bypass remain unresolved.

State: Changes requested
Priority: high
Reviewed head: b8b08187c225
Updated: 1790601176 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 1 Active findings 3
Tests 2 Noted findings 0
Documentation 0 Resolved findings 9
Configuration 0 Pending checks/questions 0

Completeness: Complete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The PR moves the `replayed_prefix` field from `AgentTurnRequest` to `AgentInvocation`, adds a builder method `with_replayed_prefix`, and modifies the `prepare_agent_turn_bounded` and `screen_user_messages` functions to skip screening of the replayed prefix. Validation rejects a prefix larger than the message count and logs a warning when the prefix covers all messages ending with a user message.

Features

  • Added — AgentInvocation::with_replayed_prefix: Allows hosts to exclude replayed history from user-input screening, preventing re-screening of already-admitted messages which could block the thread (openhuman#6710). (crates/tinyagents-harness/src/runtime/agent.rs#impl<State: Send + Sync, Ctx: Send + Sync> AgentInvocation<State, Ctx> {, crates/tinyagents-harness/src/runtime/agent.rs#pub struct AgentInvocation<State: Send + Sync, Ctx: Send + Sync = ()> {)

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

  • high · security · Do not let callers bypass screening with an arbitrary replay prefix — This public builder lets any caller mark an arbitrary prefix of the request as already screened, so `prepare_agent_turn` skips the security gate for those user messages and sends t (crates/tinyagents\-harness/src/runtime/agent\.rs:300)
  • medium · security · Preserve compatibility for public struct literals — Adding this field changes the public `AgentInvocation` struct shape. Existing downstream code that constructs the public struct with a literal will fail to compile because the new (crates/tinyagents\-harness/src/runtime/agent\.rs:256)
  • medium · tests · Add `#[non_exhaustive]` to `AgentInvocation` to preserve API stability — `AgentInvocation` is a public struct. Adding a new field without `#[non_exhaustive]` breaks external struct literals and prevents the crate from adding future fields without anothe (crates/tinyagents\-harness/src/runtime/agent\.rs:256)

Resolved this pass

  • medium — Preserve compatibility for public struct literals
  • medium — Add #[non_exhaustive] to preserve API stability
  • medium — Preserve compatibility for public struct literals
  • high — Do not let callers bypass screening with an arbitrary replay prefix
  • medium — Add #[non_exhaustive] to preserve API stability
  • Do not let callers bypass screening with an arbitrary replay prefix
  • Preserve compatibility for public struct literals
  • Do not let callers bypass screening with an arbitrary replay prefix
  • Add `#[non_exhaustive]` to preserve API stability

Before merge

  • Address Do not let callers bypass screening with an arbitrary replay prefix (crates/tinyagents\-harness/src/runtime/agent\.rs).

How this fits together

flowchart LR
  n0["start_progress_dispatcher<br/>changed<br/>3 findings"]:::blocking
  n1["prepared_tools_are_dynamic_and_never_reused<br/>changed"]:::changed
  n2["new"]:::impacted
  n3["new"]:::impacted
  n4["...corded_tools_into_a_compaction_generation"]:::impacted
  n5["...xact_tool_turn_neither_merges_nor_records"]:::impacted
  n6["...restores_tools_from_its_write_destination"]:::impacted
  n7["clone"]:::impacted
  n0 -->|calls| n3
  n0 -->|calls| n7
  n1 -->|calls| n2
  n1 -->|tests| n2
  n4 -->|calls| n2
  n4 -->|tests| n2
  n5 -->|calls| n2
  n5 -->|tests| n2
  n6 -->|calls| n2
  n6 -->|tests| n2
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 3 files; 1 finding. (1 already reported on an earlier push) (3 earlier finding(s) still open) _The code index is behind this pull request (indexed at `bc91f448fd49`), so retrieved context may be out of date._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

security

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 3 files; 2 findings. (1 earlier finding(s) still open) (1 observation(s) grouped into shared inline comments) _The code index is behind this pull request (indexed at `bc91f448fd49`), so retrieved context may be out of date._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._
  • Evidence: crates/tinyagents\-harness/src/runtime/agent\.rs — Do not let callers bypass screening with an arbitrary replay prefix
  • Evidence: crates/tinyagents\-harness/src/runtime/agent\.rs — Preserve compatibility for public struct literals

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Adds a `replayed_prefix` field to `AgentInvocation` so hosts can mark prior messages as already screened, avoiding re-screening on every turn. The tests cover the new feature, bounds, and the redacted-form contract. Two earlier API-stability findings remain unfixed: the public struct gains a field without `#[non_exhaustive]`, breaking struct literals and preventing future field additions without another break. (3 earlier finding(s) still open) _The code index is behind this pull request (indexed at `bc91f448fd49`), so retrieved context may be out of date._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._
  • Evidence: crates/tinyagents\-harness/src/runtime/agent\.rs — Add `#[non_exhaustive]` to `AgentInvocation` to preserve API stability

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Adds `AgentInvocation::with_replayed_prefix` to let hosts exclude replayed history from user-input screening. The change is source-compatible, adds validation, and includes comprehensive tests. No defects found in the diff; all earlier concerns are resolved. Safe to merge. No new findings. All earlier findings are resolved. _The code index is behind this pull request (indexed at `bc91f448fd49`), so retrieved context may be out of date._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash
  • Spend: $0.029037
  • Tokens: 216620 input · 24423 output · 24878 cached · 811 embedding
Head State Pass summary
b8fe87051855 changes requested 3 active finding(s), 0 resolved finding(s) (at 1790594717)
1f8cafb2d866 changes requested 7 active finding(s), 14 resolved finding(s) (at 1790597264)
b8b08187c225 changes requested 3 active finding(s), 9 resolved finding(s) (at 1790601176)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: f04912ae-84a0-4b9a-89a9-66c8f6e21613

📥 Commits

Reviewing files that changed from the base of the PR and between b8fe870 and b8b0818.

📒 Files selected for processing (3)
  • crates/tinyagents-harness/src/runtime/agent.rs
  • crates/tinyagents-harness/src/runtime/test.rs
  • crates/tinyagents-runtime/src/test.rs

Included review availability: This review used your included allowance. Your plan provides up to 2 included reviews per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

AgentInvocation now records the number of leading request messages that belong to replayed history. Preparation validates this prefix and screens only user messages after it. Tests cover hosted screening behavior and the history passed to the driver across turns.

Changes

Replayed-prefix screening

Layer / File(s) Summary
Request boundary and screening
crates/tinyagents-harness/src/runtime/agent.rs, crates/tinyagents-harness/src/runtime/test.rs
AgentInvocation defaults the replayed-prefix count to zero and provides with_replayed_prefix to set it. Hosted and streaming preparation reject prefixes longer than the message list and screen only messages after the prefix. Tests cover replayed rows, new input, default screening, boundary counts, and retained redacted input.
Driver history regression
crates/tinyagents-runtime/src/test.rs
A two-turn regression test checks that the driver receives committed history followed by the new turn input, without the current turn’s assistant response.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Bug fix

Suggested reviewers: senamakel

Merge Risk: ⚪ Minimal · up to b8b08

No new merge-blocking issue is established by the supplied evidence.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to b8b08

Opt-in replay can prevent legitimate conversations from being blocked by repeated screening, but its safety depends on hosts replaying only content that was previously admitted in its screened form. The runtime does not verify that condition. No production-host exploit is established.

Retained concerns

  • Medium · security · inferred: An opt-in host can mark arbitrary in-range request rows as previously admitted. The runtime checks the count but not whether those rows were screened and preserved in their effective form before sending them to the model. This is a conditional screening bypass if a host misclassifies or reconstructs history; production attacker reachability remains unknown.
Security review details

Security Blast Radius

  • inferred — For an opted-in root invocation, the independently bypassable portion is the declared prefix, potentially the whole request. Exposure beyond callers able to construct a host-authorized invocation, including tenant or production-service reachability, is not established.

Security Findings and Attack Paths

  • inferred — If a host places new or raw attacker-controlled user text inside the declared replay prefix, that text avoids this turn’s user-input gate and reaches model context. The available evidence does not show a production host making that mistake or an external attacker controlling its declaration.

Trust Boundaries and Controls

  • observed — The host supplies the replay boundary; the runtime verifies its length, not prior admission. Default-zero invocations screen all user rows, the suffix remains screened after opt-in, and child invocations start with no replay exemption.

Resilience and Maintainability Implications

  • inferred — Safe repetition and recovery depend on the host preserving each admitted row in its screened form and computing the boundary from that history. The tests demonstrate a successful redacted replay and a separate committed-history shape, but do not establish that invariant for interrupted, reconstructed, or concurrent production histories.

Hardening Proposals

  • proposed — Bind the exemption to a host-owned admitted transcript or equivalent provenance, and verify the replay boundary against it before omitting screening. Cover reconstructed and deferred-resume histories as well as ordinary second turns.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed Docstring coverage is 90.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 20 functions across 3 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: allowing hosts to exclude replayed history from user-input screening.

A rabbit checks the history line,
Replayed rows stay screened and fine.
New words pass through the gate,
While old turns keep their state.
The driver gets just what is due,
Then hops along to something new.

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 1 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0241 · 186,990 in / 15,441 out · 6,098 cached (3%) · ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash · 466 embedded
critique:    $0.0146 · 110,729 in / 4,411 out  · 4,228 cached (4%) · gpt-5.6-luna, deepseek/deepseek-v4-flash
security:    $0.0056 · 41,777 in  / 1,177 out  · 1,870 cached (4%) · gpt-5.6-luna
tests:       $0.0015 · 17,275 in  / 79 out     · 0 cached (0%)     · deepseek/deepseek-v4-flash
description: $0.0012 · 8,927 in   / 2,421 out  · 0 cached (0%)     · deepseek/deepseek-v4-flash

Comment thread crates/tinyagents-harness/src/runtime/agent.rs Outdated
Comment thread crates/tinyagents-harness/src/runtime/agent.rs Outdated
@tinysweeper tinysweeper Bot added the priority: p1 Next. Wrong behaviour a user will hit, or a security weakness behind a condition. label Sep 28, 2026
@M3gA-Mind

Copy link
Copy Markdown
Collaborator Author

Review at b8fe870 (comment only), on the four points I was asked to check, plus the failing tinysweeper/security check.

1. Can the slice ever skip the turn's NEW input? Not with the default. new() sets replayed_prefix: 0 (agent.rs:221), and the only non-test constructor in the workspace, hosted delegation at agent.rs:566, uses new(). So sub-agent turns keep full screening. .min(len) prevents a panic. But nothing in the harness can tell whether the host's count is right. A host that computes the prefix after appending the new user row passes len, the gate sees an empty slice, and the new input is unscreened with no error. No test covers prefix >= len. Suggestions:

  • Make replayed_prefix > messages.len() a Validation error instead of clamping. It's always a host bug, and failing closed is cheap.
  • == len is legitimate (a deferred-results resume has no new user row), so it can't be rejected. At least log at debug/warn when == len and the last message is Message::User, because that is exactly the off-by-one above.

2. Recall query from the same slice: yes. input_text is the return value of screen_user_messages(&mut messages[replayed..]) and feeds RecallRequest::new(&input_text) / recall_for. Note it ALSO feeds TurnSummary::with_text (agent.rs:1171) and Experience::new(.., &prepared.input_text, ..) (agent.rs:1196). For an opted-in host, the summary and experience input change from "all user rows in history" to "this turn's input". That's an improvement, but the PR body only mentions recall, and it's behaviour a learning or experience host will see.

3. Other callers of prepare_agent_turn: there is only prepare_agent_turn_bounded, which is reached from both invoke_agent (:697) and the streaming entry (:764). Both go through the new slice, and neither loses screening at the default. Replayed tool rows are safe to skip: agent_loop/tools.rs:1121-1145 screens and redacts tool output before its transcript row exists, so the replayed row is already the screened form.

4. Struct-literal break note: accurate. There are no AgentTurnRequest { literals in this workspace. In openhuman, subagent_host/ops/runner.rs:1645 is agent::harness::agent_graph::AgentTurnRequest and triage/evaluator/arm.rs:105 is agent::bus::AgentTurnRequest, both unrelated types. The vendored tinybus mentions are doc comments. The type derives only Clone, Debug, with no serde, so the field can't arrive from a wire message.

On tinysweeper/security (high, "bypass with an arbitrary prefix"): partly right. The "attacker" framing doesn't hold. The field is set only by in-process host code, which is the same party that installs (or omits) the SecurityGate, and it can't be deserialized from outside. But the underlying risk, that a host counting bug silently becomes a gate bypass, is real. Point 1's > len rejection plus the == len && last is User log addresses the part that can be addressed. (That review also reports a stale index and failed memory calls, so treat its wording as degraded; the point stands on the code.)

For the openhuman consumer (not this PR), to verify there: skipping the replay screen is only safe if the persisted user row is the redacted form. The harness redacts user rows in place in request.messages, but openhuman's session driver builds the committed history from request.history and then appends the run conversation (session_host/driver.rs:285,303). If that stores the raw user input, a secret redacted on turn 1 would go to the model unredacted from turn 2 onward, once the host opts in. The consumer PR should have a test: a redacted-then-replayed user row still reaches the model redacted.

…on contract

- A `replayed_prefix` greater than the message count is always a host
  bug. It is now a validation error instead of being clamped, which
  would have silently skipped screening of the whole request.
- A prefix equal to the count screens nothing. That is legitimate for a
  deferred resume, so it is accepted, with a warning when the last
  message is a user message.
- Document the opt-in contract: replayed user rows must be the form the
  gate returned (redacted), or a secret redacted on its first turn
  reaches the model raw later. A test replays a run's own transcript
  and asserts the redacted form is what the model sees.
@M3gA-Mind

Copy link
Copy Markdown
Collaborator Author

Re-checked at c05d69b4: my points are addressed. > len now returns Validation, and hosted_turn_rejects_a_replayed_prefix_past_the_messages asserts Policy with no model request. == len is accepted, with a warn when the trailing row is a user message (that path is covered by hosted_turn_accepts_a_replayed_prefix_covering_every_message, though the log line itself is not asserted). The redaction contract is documented on with_replayed_prefix, and hosted_turn_replays_the_redacted_form_of_an_admitted_user_row pins it. The cargo lanes are green at this head (Lint, Coverage, and Test: all-features, default, multimodal, sqlite, tools). tinysweeper/security has not run at this head yet (only tinysweeper/review exists, in progress), so the earlier red has not been re-evaluated.

…river history

Hosts that set `with_replayed_prefix(len - 1)` screen only the last message
(openhuman#6710). That is sound only while the runtime appends exactly the
turn's single input after the committed history; a second appended row
would be replayed unscreened. Fault-checked by making the runtime push an
extra row.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0475 · 381,567 in / 20,368 out · 52,915 cached (14%) · ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash, deepseek-v4-flash · 756 embedded
critique:    $0.0260 · 181,227 in / 9,464 out  · 6,857 cached (4%)   · gpt-5.6-luna, deepseek/deepseek-v4-flash
security:    $0.0192 · 140,807 in / 4,383 out  · 7,146 cached (5%)   · gpt-5.6-luna
tests:       $0.0013 · 39,751 in  / 2,715 out  · 38,912 cached (98%) · deepseek/deepseek-v4-flash
description: $0.0003 · 11,959 in  / 3,249 out  · 0 cached (0%)       · deepseek-v4-flash

Comment thread crates/tinyagents-harness/src/runtime/test.rs Outdated
Comment thread crates/tinyagents-harness/src/runtime/agent.rs Outdated
Comment thread crates/tinyagents-harness/src/runtime/test.rs Outdated
Comment thread crates/tinyagents-harness/src/runtime/agent.rs Outdated
@M3gA-Mind

Copy link
Copy Markdown
Collaborator Author

Re-checked 1f8cafb2: driver_history_is_the_committed_history_plus_exactly_the_turn_input pins the runtime half of the one-input invariant, and SessionTurnRequest takes a single Message, so it holds structurally. One note for the consumer (openhuman#6726): it derives the prefix as len - 1 of history_to_messages(history) after the core has shaped the driver history. This test can't see anything the core appends after the user row before invoking the harness. A matching assertion on #6726's side (the last message handed to the hosted request is the turn's user input) would close the pair.

`AgentTurnRequest` is a public struct with public fields, and adding
`replayed_prefix` to it broke downstream struct literals. The boundary
now lives on `AgentInvocation` as a private field with a
`with_replayed_prefix` builder. `AgentInvocation` already has a private
field, so no downstream code can build it with a literal, and this
addition is source-compatible. `AgentTurnRequest` is unchanged from
main. A recursive child's invocation (`from_shared_host`) always
screens everything.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0290 · 216,620 in / 24,423 out · 24,878 cached (11%) · ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash · 811 embedded
critique:    $0.0128 · 93,183 in  / 4,628 out  · 4,142 cached (4%)   · gpt-5.6-luna, deepseek/deepseek-v4-flash
security:    $0.0092 · 62,012 in  / 3,317 out  · 0 cached (0%)       · gpt-5.6-luna
tests:       $0.0036 · 37,098 in  / 9,020 out  · 19,456 cached (52%) · deepseek/deepseek-v4-flash
description: $0.0015 · 9,970 in   / 3,957 out  · 1,280 cached (13%)  · deepseek/deepseek-v4-flash

Comment thread crates/tinyagents-harness/src/runtime/agent.rs
Comment thread crates/tinyagents-harness/src/runtime/agent.rs
@senamakel
senamakel merged commit fefbd11 into tinyhumansai:main Sep 28, 2026
14 of 16 checks passed
M3gA-Mind added a commit to M3gA-Mind/openhuman that referenced this pull request Sep 28, 2026
…humansai#6723)

Brings in tinyhumansai/tinyagents#223 (replay boundary, merged as
fefbd11b; tree identical to the previously pinned b8b08187) and
tinyhumansai/tinyagents#224 (text-dialect undecodable-block nudge,
tinyhumansai#6723). No nested submodule moves.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p1 Next. Wrong behaviour a user will hit, or a security weakness behind a condition.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants