feat(tools,goals): absorb OpenHuman time/goal tool behaviour so the host copies can go - #229
Conversation
Extend CurrentTimeTool and ResolveTimeTool to optionally render their results as compact markdown lists when the caller requests markdown output via ToolCallOptions. This allows agents that prefer human-readable formatted responses to receive structured time information without needing to parse raw JSON. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add test coverage verifying that all time tools are marked as read-only and support markdown output. Also add tests for the `CurrentTimeTool` and `ResolveTimeTool` to confirm that markdown rendering is only produced when explicitly requested via `prefer_markdown()`, and that the rendered output contains the expected fields. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Every goal control now returns a JSON payload with `goal` and `text` fields, replacing the previous split between markdown content and a separate raw value. A new `GoalUpdateHook` callback lets hosts observe writes from `goal_set`, `goal_complete`, `goal_pause`, and `goal_resume`. The tool also reports its permission level as read-only or write based on the control kind, and error messages are simplified for clarity. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…n levels Add comprehensive tests for the goals tool module covering the new JSON payload format that includes both goal and text fields, verify that the update hook fires only on write operations, and confirm that permission levels correctly distinguish read-only Get from write operations. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Export the `GoalUpdateHook` type from the goals module so it is publicly accessible, and simplify the tinytools import in the tool module by removing the multi-line formatting. The time test assertions are reformatted for consistency without changing their logic. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The import of `Result` from `tinyagents_harness::error` was unused in this file, so it has been removed to keep the code clean and avoid compiler warnings. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…yload/hook
Time tools: CurrentTimeTool and ResolveTimeTool now override
execute_with_options, render a markdown form when prefer_markdown is set,
report supports_markdown, return pretty-printed JSON text, log at debug, and
use the longer 'expr is required' hint OpenHuman shipped.
Goal tools: every control answers with {goal, text} (goal null when absent),
attaches text as markdown, reports Write permission for mutating controls,
surfaces store and argument errors as error results, and exposes
GoalTool::with_update_hook so a host can publish an event after a write.
The goal_set description regains the usage guidance and budget wording.
Co-authored-by: Medulla <medulla@tinyhumans.ai>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
Tiny Sweeper reviewThis pull request absorbs OpenHuman time/goal tool behaviour into the host, replacing the separate OpenHuman copies. It modifies goal tools to return structured JSON payloads with goal and text, adds a GoalUpdateHook for write notifications, and updates error messages. Additionally, it introduces a new title module for thread title generation and sanitization, and a summarization module with a ModelSummarizer and FaultTolerantCachingSummarizer for context-window-aware conversation summarization with fault tolerance. Documentation for goals is updated. State: Incomplete Review snapshot
Completeness: Incomplete What changedGoal tools now return structured JSON payloads ({ goal, text }) instead of separate content and raw, added GoalUpdateHook for write notifications, updated error messages to be clearer, added permission_level method, and added tracing. Time tools added markdown support via execute_with_options, supports_markdown method, and markdown formatting functions; error messages improved. New title module (crates/tinyagents-harness/src/title/mod.rs) provides functions for generating and sanitizing thread titles. New summarization module (crates/tinyagents-harness/src/summarization/model_summarizer.rs, resilient.rs) adds LLM-backed summarization with fault-tolerant caching. Documentation updated in docs/modules/graph/goals.md. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred.
Findings
Could not review: crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/title/test.rs, tinysweeper/tests Before merge
How this fits togetherflowchart LR
n0["GoalTool<br/>changed"]:::changed
n1["GoalToolKind<br/>changed"]:::changed
n2["Store"]:::impacted
n3["update_hook_fires_on_writes_only"]:::impacted
n4["run"]:::impacted
n5["new"]:::impacted
n6["run_gate"]:::impacted
n0 -->|uses| n1
n0 -->|uses| n2
n3 -->|calls| n4
n3 -->|tests| n4
n5 -->|uses| n1
n5 -->|uses| n2
n6 -->|uses| n2
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinyagents-graph/src/goals/mod.rs, crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/tools/time.rs, crates/tinyagents-harness/src/tools/time_test.rs, docs/modules/graph/goals.md, tinysweeper/description, tinysweeper/tests.
$0.0000 · 0 in / 0 out · 932 embedded · ladder/vectors
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e48380f877
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Updated the vendored dependencies for tinyinference and tinytools to their latest versions, incorporating upstream fixes and improvements. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Refactor the summarization module to integrate a resilient wrapper around model summarization, improving fault tolerance during inference. The change introduces a new resilient layer that retries on transient failures and adds corresponding test coverage, while updating dependency locks to match. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ilient retry Introduces a new summarization module that provides a model summarizer with resilient retry logic, enabling robust text summarization capabilities within the harness. The module includes a test file to verify the summarizer's behavior. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…agents-harness/src/title/mod.rs Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The model summarizer test was incorrectly using a direct summarizer instead of the resilient wrapper, which caused test failures when the underlying service experienced transient errors. Updated the test to instantiate the resilient summarizer to match the production behavior. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the vendored tinytools dependency to incorporate upstream fixes and improvements. This change ensures the project uses the latest stable version of the library. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b3cbdfa635
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Co-authored-by: Medulla <medulla@tinyhumans.ai>
feat(harness): ModelSummarizer + fault-tolerant summarizer and thread-title helpers from OpenHuman
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 61b61142de
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| loop { | ||
| let remaining = previous_token_estimate + estimate_slice_tokens(&messages[start..]); | ||
| if remaining <= self.fallback_trim_budget || start + 1 >= messages.len() { | ||
| break; |
There was a problem hiding this comment.
Enforce the fallback trim budget for oversized tails
When the newest message alone exceeds fallback_trim_budget—for example, a large tool result—the start + 1 >= messages.len() condition stops trimming while that entire message remains. The same happens when previous_summary alone exceeds the budget. The wrapper then reports successful recovery and returns an oversized checkpoint, so the subsequent model call can still overflow despite the summarizer failure having been handled; truncate or omit indivisible oversized content rather than unconditionally retaining it.
Useful? React with 👍 / 👎.
| let summary = summary.trim(); | ||
| if summary.is_empty() { | ||
| return Err(TinyAgentsError::Model( | ||
| "summarizer returned empty response".into(), | ||
| )); | ||
| } |
There was a problem hiding this comment.
Reject summaries that do not reduce context
When the summarization model returns any nonempty but excessively verbose response, this path accepts it without an output cap or comparison against the source size. A model that echoes the transcript can therefore produce a replacement as large as—or larger than—the compacted head, after which the main model request remains over its context window even though compaction was recorded as successful. Bound the request's output tokens and/or fall back when the returned summary does not meaningfully shrink the input.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs and 3 more.
$0.0019 · 84,030 in / 8,568 out · 55,040 cached (66%) · ladder/vectors, deepseek/deepseek-v4-flash · 1,160 embedded
description: $0.0009 · 26,640 in / 1,952 out · 26,368 cached (99%) · deepseek/deepseek-v4-flash
Ports the behaviour OpenHuman's private copies of the time and goal tools had and this crate lacked, so OpenHuman can delete them and consume these tools directly.
Stacked on #228 (
move-openhuman-tool-helpers); base is that branch so the diff here is only this change. Retarget tomainonce #228 merges.Time tools: markdown rendering under
prefer_markdown,supports_markdown, pretty JSON text output, debug logs, fullerexpr is requiredhint.Goal tools:
{goal, text}payload (UI banner readsgoal), markdown form, Write permission for mutating controls, errors as error results, andGoalTool::with_update_hookfor host events.goal_setdescription restored to OpenHuman's wording.Tests:
cargo test -p tinyagents-harness -p tinyagents-graph(508 + 1398 + doc tests pass), clippy -D warnings and fmt clean.