Skip to content

feat(harness): tool-result artifact persistence and use_skill tool packs (OpenHuman wave 2) - #239

Merged
senamakel merged 18 commits into
mainfrom
oh-w2-artifacts
Sep 30, 2026
Merged

senamakel merged 18 commits into
mainfrom
oh-w2-artifacts

Conversation

@senamakel

Copy link
Copy Markdown
Member

Two host-independent mechanisms lifted out of OpenHuman's core. Both reach the host through a small seam; the product data and policy stay in the host.

1. artifacts::tool_results (per-tool-result persistence)

ToolResultArtifactStore, apply_per_result_persistence, spill_aggregate_tool_results, artifact_read_target, page_artifact_read. Extends the existing harness::artifacts (reuses ArtifactRedactor/Redacted, no second policy trait).
Seams: the host passes an Arc<dyn ArtifactRedactor>, its file-read tool name, the wrapper tool name (use_skill) and the largest body its reader opens. The envelope and page trailer text is byte-identical (literal fixture test envelope_text_is_byte_stable). The random fallback file name no longer needs uuid.

2. tool::packs (on-demand tool disclosure, use_skill)

PackCatalog (host pack table + not-found marker), ToolPack, PackRegistryHandle, UseSkillTool, render_pack_filtered, scope_use_skill_spec, route_sentence, NoSuchPackTool, named_tool. Tool name, description and schema are pinned by a literal test. Distinct from tool::discover (model-driven BM25 search over deferred schemas).
The host keeps the pack table, group posture (ToolGroups) and registry binding.

3. Gitlink

Bumps nested vendor/tinyinference to tinyhumansai/tinyinference (oh-w2-business-limit), which exposes contains_business_limit. Merge that PR first.

Stacking

Based on main. The OpenHuman pin (ed9cc77, from #230) is already an ancestor of main; 15 commits on top.

Tests

cargo test -p tinyagents-harness --lib artifacts::tool_results: 21 passed; tool::packs: 16 passed. clippy -D warnings and fmt clean.

senamakel and others added 15 commits September 30, 2026 10:24
Introduce a new `tool_results` module within the artifacts crate to store and manage the outcomes of tool executions, enabling better tracking and retrieval of individual tool results during agent runs.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Fix the serialization of tool results to properly handle cases where tool outputs are empty, ensuring that the JSON representation includes the expected structure rather than omitting fields. This resolves a deserialization mismatch that occurred when processing tool results with no output content.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test assertion for tool results was updated to reflect a change in how the harness formats tool output, ensuring the test continues to validate the correct behavior.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Move the test module declaration from the artifacts mod.rs into tool_results.rs, placing it alongside the code it tests. This keeps the test module co-located with its subject and removes a separate test file inclusion from the parent module.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ied redactor and tool names

Moves OpenHuman's tool_result_artifacts mechanics into harness::artifacts::tool_results:
ToolResultArtifactStore, apply_per_result_persistence, spill_aggregate_tool_results,
artifact_read_target and page_artifact_read. The host supplies an ArtifactRedactor,
the read/wrapper tool names and the largest body its reader opens.

Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce a catalog module and associated types to enable structured discovery and registration of tool packs. This change provides a centralized way to list and access available packs, improving extensibility and maintainability of the tool system.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a tool pack is not found, the harness now returns a clear error message instead of panicking. This improves robustness by allowing the system to report the issue and continue operating.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Make the `packs` module publicly accessible from the tool crate by adding a `pub mod packs` declaration, enabling external consumers to use pack-related functionality. Also reformat a method chain in `handle.rs` and collapse a format string in `render.rs` for improved readability, removing trailing newlines in the process.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Relaxed assertions that accepted partial matches are replaced with exact checks, ensuring the rendered schema and dispatched arguments are verified precisely rather than allowing any substring match.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
PackCatalog (host pack table + not-found marker), ToolPack, PackRegistryHandle,
UseSkillTool and the listing/scoping renderers, lifted from OpenHuman's toolpacks.
The pack table, group posture and registry binding stay with the host.

Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…d path

Remove the `use tinytools_agent::dialect::ToolOutcome` import and instead reference the type directly as `tinytools_agent::dialect::ToolOutcome` in both the production code and tests. This eliminates an unnecessary import while keeping the same behavior.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…uplicate root accessor

Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

This review includes 13 billable files and costs up to $3.25.

Or wait 6 minutes for your next included review.

Check out review usage here.

View limit details

Limit details: You’ve used all 2 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 99e44aa9-9762-46d9-8941-78965bbf721f

📥 Commits

Reviewing files that changed from the base of the PR and between 9201327 and f82b6c7.

📒 Files selected for processing (13)
  • crates/tinyagents-harness/src/artifacts/README.md
  • crates/tinyagents-harness/src/artifacts/mod.rs
  • crates/tinyagents-harness/src/artifacts/tool_results.rs
  • crates/tinyagents-harness/src/artifacts/tool_results_test.rs
  • crates/tinyagents-harness/src/tool/mod.rs
  • crates/tinyagents-harness/src/tool/packs/catalog.rs
  • crates/tinyagents-harness/src/tool/packs/handle.rs
  • crates/tinyagents-harness/src/tool/packs/mod.rs
  • crates/tinyagents-harness/src/tool/packs/render.rs
  • crates/tinyagents-harness/src/tool/packs/test.rs
  • crates/tinyagents-harness/src/tool/packs/tool.rs
  • crates/tinyagents-harness/src/tool/packs/types.rs
  • vendor/tinyinference

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T08:42:23.095516Z f82b6c7 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 36dd3a14c1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/tinyagents-harness/src/artifacts/tool_results.rs
Comment thread crates/tinyagents-harness/src/artifacts/tool_results.rs Outdated
Comment thread crates/tinyagents-harness/src/artifacts/tool_results.rs Outdated
@senamakel senamakel self-assigned this Sep 30, 2026
senamakel and others added 3 commits September 30, 2026 11:23
Replace the stale-session check with a recursive `newest_modified` function that walks the session tree without following symlinks, so nested tool writes keep their containing session alive. Add a canonical-path check in `persist` to reject artifact paths that escape the action directory via symlinks. Floor the per-result persistence budget at `MIN_ENVELOPE_ALLOWANCE_BYTES` to guarantee the recovery pointer is always written, even when the caller supplies a tiny budget.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The doc comment for the scope_use_skill_spec function used a `super::` prefix to reference `UseSkillTool::execute_with_context`, which is unnecessary since the type is already in scope. Removing the prefix keeps the documentation cleaner and avoids potential confusion about the path resolution.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@senamakel
senamakel merged commit 3ab9327 into main Sep 30, 2026
9 checks passed
@senamakel
senamakel deleted the oh-w2-artifacts branch September 30, 2026 08:34
@tinysweeper

tinysweeper Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 4 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Incomplete
Priority: critical
Reviewed head: f82b6c798d3e
Updated: 1790757961 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 9 Active findings 4
Tests 2 Noted findings 0
Documentation 1 Resolved findings 0
Configuration 0 Pending checks/questions 23

Completeness: Incomplete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

  • critical · commits · Remove the committed credential (a GitHub personal access token) — Rotate the credential first, then remove it from the working tree and purge it from history. Treat it as compromised: it is in the push, whatever happens next. (crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs:15)
  • critical · commits · Remove the committed credential (a GitHub personal access token) — Rotate the credential first, then remove it from the working tree and purge it from history. Treat it as compromised: it is in the push, whatever happens next. (crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs:46)
  • critical · commits · Remove the committed credential (a GitHub personal access token) — Rotate the credential first, then remove it from the working tree and purge it from history. Treat it as compromised: it is in the push, whatever happens next. (crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs:60)
  • critical · commits · Remove the committed credential (a GitHub personal access token) — Rotate the credential first, then remove it from the working tree and purge it from history. Treat it as compromised: it is in the push, whatever happens next. (crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs:68)

Could not review: crates/tinyagents-harness/src/artifacts/README.md, crates/tinyagents-harness/src/artifacts/mod.rs, crates/tinyagents-harness/src/artifacts/tool_results.rs, crates/tinyagents-harness/src/artifacts/tool_results_test.rs, crates/tinyagents-harness/src/tool/mod.rs, crates/tinyagents-harness/src/tool/packs/catalog.rs, crates/tinyagents-harness/src/tool/packs/handle.rs, crates/tinyagents-harness/src/tool/packs/mod.rs, crates/tinyagents-harness/src/tool/packs/render.rs, crates/tinyagents-harness/src/tool/packs/test.rs, crates/tinyagents-harness/src/tool/packs/tool.rs, crates/tinyagents-harness/src/tool/packs/types.rs

Before merge

  • Complete the critique review for crates/tinyagents-harness/src/artifacts/README.md, crates/tinyagents-harness/src/artifacts/mod.rs, crates/tinyagents-harness/src/artifacts/tool_results.rs, crates/tinyagents-harness/src/artifacts/tool_results_test.rs, crates/tinyagents-harness/src/tool/mod.rs, crates/tinyagents-harness/src/tool/packs/catalog.rs, crates/tinyagents-harness/src/tool/packs/handle.rs, crates/tinyagents-harness/src/tool/packs/mod.rs, crates/tinyagents-harness/src/tool/packs/render.rs, crates/tinyagents-harness/src/tool/packs/test.rs, crates/tinyagents-harness/src/tool/packs/tool.rs, crates/tinyagents-harness/src/tool/packs/types.rs.
  • Complete the security review for crates/tinyagents-harness/src/tool/packs/tool.rs, crates/tinyagents-harness/src/artifacts/mod.rs, crates/tinyagents-harness/src/artifacts/tool_results.rs, crates/tinyagents-harness/src/artifacts/tool_results_test.rs, crates/tinyagents-harness/src/tool/mod.rs, crates/tinyagents-harness/src/tool/packs/catalog.rs, crates/tinyagents-harness/src/tool/packs/handle.rs, crates/tinyagents-harness/src/tool/packs/mod.rs, crates/tinyagents-harness/src/tool/packs/render.rs, crates/tinyagents-harness/src/tool/packs/test.rs, crates/tinyagents-harness/src/tool/packs/types.rs.
  • Address Remove the committed credential (a GitHub personal access token) (crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs).
  • Address Remove the committed credential (a GitHub personal access token) (crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs).
  • Address Remove the committed credential (a GitHub personal access token) (crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs).
  • Address Remove the committed credential (a GitHub personal access token) (crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs).

How this fits together

flowchart LR
  n0["tool"]:::impacted
  n1["ToolDispatch"]:::impacted
  n2["exposure"]:::impacted
  n3["insert_dispatch"]:::impacted
  n4["CanonicalDispatch"]:::impacted
  n5["deferred_schemas_with_families"]:::impacted
  n2 -->|calls| n0
  n3 -->|uses| n1
  n4 -->|implements| n1
  n5 -->|calls| n0
  n5 -->|uses| n0
  n5 -->|calls| n2
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/tinyagents-harness/src/artifacts/README.md, crates/tinyagents-harness/src/artifacts/mod.rs, crates/tinyagents-harness/src/artifacts/tool_results.rs, crates/tinyagents-harness/src/artifacts/tool_results_test.rs, crates/tinyagents-harness/src/tool/mod.rs, crates/tinyagents-harness/src/tool/packs/catalog.rs, crates/tinyagents-harness/src/tool/packs/handle.rs, crates/tinyagents-harness/src/tool/packs/mod.rs, crates/tinyagents-harness/src/tool/packs/render.rs, crates/tinyagents-harness/src/tool/packs/test.rs, crates/tinyagents-harness/src/tool/packs/tool.rs, crates/tinyagents-harness/src/tool/packs/types.rs
  • Lane summary: Reviewed 0 files; 0 findings. 12 files could not be reviewed: crates/tinyagents-harness/src/artifacts/README.md, crates/tinyagents-harness/src/artifacts/mod.rs, crates/tinyagents-harness/src/artifacts/tool_results.rs, crates/tinyagents-harness/src/artifacts/tool_results_test.rs, crates/tinyagents-harness/src/tool/mod.rs, crates/tinyagents-harness/src/tool/packs/catalog.rs, crates/tinyagents-harness/src/tool/packs/handle.rs, crates/tinyagents-harness/src/tool/packs/mod.rs, crates/tinyagents-harness/src/tool/packs/render.rs, crates/tinyagents-harness/src/tool/packs/test.rs, crates/tinyagents-harness/src/tool/packs/tool.rs, crates/tinyagents-harness/src/tool/packs/types.rs.

security

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/tinyagents-harness/src/tool/packs/tool.rs, crates/tinyagents-harness/src/artifacts/mod.rs, crates/tinyagents-harness/src/artifacts/tool_results.rs, crates/tinyagents-harness/src/artifacts/tool_results_test.rs, crates/tinyagents-harness/src/tool/mod.rs, crates/tinyagents-harness/src/tool/packs/catalog.rs, crates/tinyagents-harness/src/tool/packs/handle.rs, crates/tinyagents-harness/src/tool/packs/mod.rs, crates/tinyagents-harness/src/tool/packs/render.rs, crates/tinyagents-harness/src/tool/packs/test.rs, crates/tinyagents-harness/src/tool/packs/types.rs
  • Lane summary: Reviewed 0 files; 0 findings. 11 files could not be reviewed: crates/tinyagents-harness/src/tool/packs/tool.rs, crates/tinyagents-harness/src/artifacts/mod.rs, crates/tinyagents-harness/src/artifacts/tool_results.rs, crates/tinyagents-harness/src/artifacts/tool_results_test.rs, crates/tinyagents-harness/src/tool/mod.rs, crates/tinyagents-harness/src/tool/packs/catalog.rs, crates/tinyagents-harness/src/tool/packs/handle.rs, crates/tinyagents-harness/src/tool/packs/mod.rs, crates/tinyagents-harness/src/tool/packs/render.rs, crates/tinyagents-harness/src/tool/packs/test.rs, crates/tinyagents-harness/src/tool/packs/types.rs. 1 file was not security-reviewed: crates/tinyagents-harness/src/artifacts/README.md (prose or tabular data).

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Adds per-tool-result artifact persistence and on-demand tool disclosure via packs. The test suites for both features are thorough, covering success paths, error handling, edge cases (near-max offsets, zero budgets, missing stores, cross-skill denial), and byte-stable envelope formatting. No test gaps were found in the new code. _The code index is behind this pull request (indexed at `24f963b83bbb`), so retrieved context may be out of date._ _4 memory call(s) failed (model: cortex: v1/recall: timed out after 10s), so this review saw part of what the engine holds._

commits

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: 4 sensitive items entered this pull request's history. _The code index is behind this pull request (indexed at `24f963b83bbb`), so retrieved context may be out of date._ _4 memory call(s) failed (model: cortex: v1/recall: timed out after 10s), so this review saw part of what the engine holds._
  • Evidence: crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs — Remove the committed credential (a GitHub personal access token)
  • Evidence: crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs — Remove the committed credential (a GitHub personal access token)
  • Evidence: crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs — Remove the committed credential (a GitHub personal access token)
  • Evidence: crates/tinyagents\-harness/src/artifacts/tool\_results\_test\.rs — Remove the committed credential (a GitHub personal access token)

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This pull request adds two new mechanisms: per-tool-result artifact persistence and on-demand tool disclosure via `use_skill` packs. The code is well-structured, thoroughly tested, and addresses important concerns like path validation and session pruning. However, it introduces a few issues: use of `anyhow::Result` instead of the preferred typed error, weak path validation in artifact read detection, and a silent no-op in spec scoping. None are blocking, but they should be addressed for consistency and robustness. _The code index is behind this pull request (indexed at `24f963b83bbb`), so retrieved context may be out of date._ _4 memory call(s) failed (model: cortex: v1/recall: timed out after 10s), so this review saw part of what the engine holds._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: ladder/vectors, deepseek/deepseek-v4-flash
  • Spend: $0.010368
  • Tokens: 127459 input · 14083 output · 50688 cached · 1110 embedding
Head State Pass summary
f82b6c798d3e incomplete 4 active finding(s), 0 resolved finding(s) (at 1790757961)

tinysweeper 0.1.0

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f82b6c798d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +324 to +325
let root = self.action_dir.join(ARTIFACT_ROOT);
let entries = match std::fs::read_dir(&root) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Validate the artifact root before pruning

When the model-writable artifacts/tool-results component is a symlink to a directory outside action_dir, read_dir follows it and each returned entry still reports as a normal directory; the later remove_dir_all(entry.path()) can therefore recursively delete an external directory once it is older than max_age. Canonicalize and confine the pruning root against the canonical action directory before enumerating or deleting anything.

Useful? React with 👍 / 👎.

);
}
}
tokio::fs::write(&absolute_path, sanitized.text.as_bytes()).await?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reject symlinks at the final artifact filename

The new parent-canonicalization check is fresh evidence that the previously reported escape was only partially fixed: it validates parent, but this write still follows a pre-existing symlink at the final call.txt component. If an action workspace already contains such a symlink for a predictable tool/call ID, persisting an oversized result overwrites its target outside the workspace; open the leaf with no-follow semantics or explicitly reject an existing symlink.

Useful? React with 👍 / 👎.

Comment on lines +134 to +138
let inner_args = args.get("args").cloned().unwrap_or_else(|| json!({}));
tracing::debug!(tool = tools[idx].name(), "[toolpacks] use_skill dispatch");
tools[idx]
.execute_with_context(inner_args, options, context)
.await

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Apply canonical argument injection before inner dispatch

When a packed tool declares a host- or call-ID-injected argument, harness admission only prepares the outer use_skill declaration, whose injected-argument list is empty, and these nested model-supplied args are then forwarded directly. Consequently required call IDs are absent and model-provided values for host-authoritative fields are never stripped or replaced, allowing forged authority for any such packed tool; route the inner invocation through canonical argument preparation rather than calling it directly.

Useful? React with 👍 / 👎.

Comment on lines +136 to +138
tools[idx]
.execute_with_context(inner_args, options, context)
.await

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Enforce the inner tool's ToolPolicy before dispatch

When a packed tool uses ToolPolicy to require approval, a sandbox, or side-effect restrictions, the agent loop evaluates only the outer use_skill policy before reaching this direct call. UseSkillTool retains the default policy, so the inner tool can execute without the approval deferral and policy middleware behavior it receives when invoked normally; inner dispatch must preserve the canonical policy-admission path rather than forwarding only PermissionLevel.

Useful? React with 👍 / 👎.

Comment on lines +613 to +620
let persisted_output = if looks_like_preview_envelope(&original) {
Ok(PersistedToolResult {
output: original.clone(),
path: "<existing-preview>".to_string(),
original_bytes: original_len,
stored_bytes: original_len,
redacted: false,
})

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Track persisted envelopes without trusting tool output text

If an ordinary tool result happens to begin with [tool_result_preview]\n, this prefix check treats it as an already-persisted envelope even though no artifact exists. During aggregate spilling the result is then truncated in place instead of being written, permanently discarding its tail and potentially leaving a spoofed pointer; carry persistence metadata alongside the result or validate a real store-generated artifact rather than classifying untrusted output by prefix.

Useful? React with 👍 / 👎.

Comment on lines +425 to +432
original_bytes: {}\n\
stored_bytes: {}\n\
artifact_path: {relative_path}\n\
read_with: {read_tool} {{\"path\":\"{relative_path}\"}} (a long read returns one page and names the \"offset\" to continue from)\n\
notes: Full scrubbed output was persisted under the action workspace.{redaction_note}{truncation_note}\n\n\
[preview]\n{preview}",
content.len(),
sanitized.text.len(),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Report when only the compacted fallback was persisted

When full_output exceeds max_readable_bytes but the earlier rewritten content fits, readable_body selects that fallback, yet this envelope still says the full output was persisted and reports the fallback length as original_bytes. The model and artifact index therefore claim recoverability and fidelity that no longer exist; propagate whether fallback selection occurred and label the stored body and byte counts accordingly.

Useful? React with 👍 / 👎.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 1 lane(s) blocking, worst finding is critical.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0104 · 127,459 in / 14,083 out · 50,688 cached (40%)  · ladder/vectors, deepseek/deepseek-v4-flash · 1,110 embedded
tests:       $0.0016 · 48,550 in  / 4,037 out  · 48,384 cached (100%) · deepseek/deepseek-v4-flash
description: $0.0047 · 39,462 in  / 7,254 out  · 2,304 cached (6%)    · deepseek/deepseek-v4-flash


impl ArtifactRedactor for TestRedactor {
fn redact(&self, content: &str) -> Redacted {
let mut out = content.replace("ghp_abcdefghijklmnopqrstuvwxyz123456", "[REDACTED_SECRET]");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority critical commits confident

Remove the committed credential (a GitHub personal access token)

Rotate the credential first, then remove it from the working tree and purge it from history. Treat it as compromised: it is in the push, whatever happens next.

Matched: ghp_… <redacted, 36 chars>

[RULE] github-personal-access-token ·

let raw = format!(
"{} {}",
"x".repeat(4096),
"ghp_abcdefghijklmnopqrstuvwxyz123456"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority critical commits confident

Remove the committed credential (a GitHub personal access token)

Rotate the credential first, then remove it from the working tree and purge it from history. Treat it as compromised: it is in the push, whatever happens next.

Matched: ghp_… <redacted, 36 chars>

[RULE] github-personal-access-token ·

assert!(out.contains("original_bytes:"));
assert!(out.contains("[preview]"));
assert!(out.contains("Credential/PII redaction was applied"));
assert!(!out.contains("ghp_abcdefghijklmnopqrstuvwxyz123456"));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority critical commits confident

Remove the committed credential (a GitHub personal access token)

Rotate the credential first, then remove it from the working tree and purge it from history. Treat it as compromised: it is in the push, whatever happens next.

Matched: ghp_… <redacted, 36 chars>

[RULE] github-personal-access-token ·

)
.unwrap();
assert!(stored.contains("xxxx"));
assert!(!stored.contains("ghp_abcdefghijklmnopqrstuvwxyz123456"));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority critical commits confident

Remove the committed credential (a GitHub personal access token)

Rotate the credential first, then remove it from the working tree and purge it from history. Treat it as compromised: it is in the push, whatever happens next.

Matched: ghp_… <redacted, 36 chars>

[RULE] github-personal-access-token ·

@tinysweeper tinysweeper Bot added the priority: p0 Drop what you are doing. Data loss, a live break, or an exploitable hole. label Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p0 Drop what you are doing. Data loss, a live break, or an exploitable hole.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant