Skip to content

feat(tools,goals): absorb OpenHuman time/goal tool behaviour so the host copies can go - #229

Merged
senamakel merged 17 commits into
move-openhuman-tool-helpersfrom
oh-dedupe-time-goals-tools
Sep 30, 2026
Merged

senamakel merged 17 commits into
move-openhuman-tool-helpersfrom
oh-dedupe-time-goals-tools

Conversation

@senamakel

Copy link
Copy Markdown
Member

Ports the behaviour OpenHuman's private copies of the time and goal tools had and this crate lacked, so OpenHuman can delete them and consume these tools directly.

Stacked on #228 (move-openhuman-tool-helpers); base is that branch so the diff here is only this change. Retarget to main once #228 merges.

Time tools: markdown rendering under prefer_markdown, supports_markdown, pretty JSON text output, debug logs, fuller expr is required hint.

Goal tools: {goal, text} payload (UI banner reads goal), markdown form, Write permission for mutating controls, errors as error results, and GoalTool::with_update_hook for host events. goal_set description restored to OpenHuman's wording.

Tests: cargo test -p tinyagents-harness -p tinyagents-graph (508 + 1398 + doc tests pass), clippy -D warnings and fmt clean.

senamakel and others added 7 commits September 29, 2026 22:56
Extend CurrentTimeTool and ResolveTimeTool to optionally render their results as compact markdown lists when the caller requests markdown output via ToolCallOptions. This allows agents that prefer human-readable formatted responses to receive structured time information without needing to parse raw JSON.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add test coverage verifying that all time tools are marked as read-only and support markdown output. Also add tests for the `CurrentTimeTool` and `ResolveTimeTool` to confirm that markdown rendering is only produced when explicitly requested via `prefer_markdown()`, and that the rendered output contains the expected fields.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Every goal control now returns a JSON payload with `goal` and `text` fields, replacing the previous split between markdown content and a separate raw value. A new `GoalUpdateHook` callback lets hosts observe writes from `goal_set`, `goal_complete`, `goal_pause`, and `goal_resume`. The tool also reports its permission level as read-only or write based on the control kind, and error messages are simplified for clarity.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…n levels

Add comprehensive tests for the goals tool module covering the new JSON payload format that includes both goal and text fields, verify that the update hook fires only on write operations, and confirm that permission levels correctly distinguish read-only Get from write operations.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Export the `GoalUpdateHook` type from the goals module so it is publicly accessible, and simplify the tinytools import in the tool module by removing the multi-line formatting. The time test assertions are reformatted for consistency without changing their logic.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The import of `Result` from `tinyagents_harness::error` was unused in this file, so it has been removed to keep the code clean and avoid compiler warnings.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…yload/hook

Time tools: CurrentTimeTool and ResolveTimeTool now override
execute_with_options, render a markdown form when prefer_markdown is set,
report supports_markdown, return pretty-printed JSON text, log at debug, and
use the longer 'expr is required' hint OpenHuman shipped.

Goal tools: every control answers with {goal, text} (goal null when absent),
attaches text as markdown, reports Write permission for mutating controls,
surfaces store and argument errors as error results, and exposes
GoalTool::with_update_hook so a host can publish an event after a write.
The goal_set description regains the usage guidance and budget wording.

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: bd86905b-5ebc-4002-bdb1-cbbd78686906

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@tinysweeper

tinysweeper Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

This pull request absorbs OpenHuman time/goal tool behaviour into the host, replacing the separate OpenHuman copies. It modifies goal tools to return structured JSON payloads with goal and text, adds a GoalUpdateHook for write notifications, and updates error messages. Additionally, it introduces a new title module for thread title generation and sanitization, and a summarization module with a ModelSummarizer and FaultTolerantCachingSummarizer for context-window-aware conversation summarization with fault tolerance. Documentation for goals is updated.

State: Incomplete
Priority: medium
Reviewed head: 61b61142dee4
Updated: 1790749598 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 9 Active findings 1
Tests 4 Noted findings 0
Documentation 1 Resolved findings 0
Configuration 0 Pending checks/questions 21

Completeness: Incomplete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

Goal tools now return structured JSON payloads ({ goal, text }) instead of separate content and raw, added GoalUpdateHook for write notifications, updated error messages to be clearer, added permission_level method, and added tracing. Time tools added markdown support via execute_with_options, supports_markdown method, and markdown formatting functions; error messages improved. New title module (crates/tinyagents-harness/src/title/mod.rs) provides functions for generating and sanitizing thread titles. New summarization module (crates/tinyagents-harness/src/summarization/model_summarizer.rs, resilient.rs) adds LLM-backed summarization with fault-tolerant caching. Documentation updated in docs/modules/graph/goals.md.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

  • Unreviewed: tinysweeper/tests

Findings

  • medium · description · Update PR description to cover the new title and summarization modules — The PR description says it only ports OpenHuman's time and goal tool behaviour, but the diff adds entirely new modules: `title/mod.rs`, `title/test.rs`, `summarization/model_summar (\(pull request description\))

Could not review: crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/title/test.rs, tinysweeper/tests

Before merge

  • Complete the critique review for crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/title/test.rs.
  • Complete the security review for crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/types.rs, crates/tinyagents-harness/src/title/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-graph/src/goals/test.rs.
  • Complete the tests review for tinysweeper/tests.

How this fits together

flowchart LR
  n0["GoalTool<br/>changed"]:::changed
  n1["GoalToolKind<br/>changed"]:::changed
  n2["Store"]:::impacted
  n3["update_hook_fires_on_writes_only"]:::impacted
  n4["run"]:::impacted
  n5["new"]:::impacted
  n6["run_gate"]:::impacted
  n0 -->|uses| n1
  n0 -->|uses| n2
  n3 -->|calls| n4
  n3 -->|tests| n4
  n5 -->|uses| n1
  n5 -->|uses| n2
  n6 -->|uses| n2
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/title/test.rs
  • Lane summary: Reviewed 0 files; 0 findings. 10 files could not be reviewed: crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/title/test.rs.

security

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/types.rs, crates/tinyagents-harness/src/title/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-graph/src/goals/test.rs
  • Lane summary: Reviewed 0 files; 0 findings. 10 files could not be reviewed: crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/types.rs, crates/tinyagents-harness/src/title/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-graph/src/goals/test.rs.

tests

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: tinysweeper/tests
  • Lane summary: No reviewer could be consulted.

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The diff ports time/goal tool behaviour as described, but also adds unmentioned `title` and `summarization` modules. The description needs updating to reflect the full scope. _The code index is behind this pull request (indexed at `ab5287d8dbe1`), so retrieved context may be out of date._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._
  • Evidence: \(pull request description\) — Update PR description to cover the new title and summarization modules

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: ladder/vectors, deepseek/deepseek-v4-flash
  • Spend: $0.001851
  • Tokens: 84030 input · 8568 output · 55040 cached · 1160 embedding
Head State Pass summary
e48380f8771e incomplete 0 active finding(s), 0 resolved finding(s) (at 1790712101)
61b61142dee4 incomplete 1 active finding(s), 0 resolved finding(s) (at 1790749598)

tinysweeper 0.1.0

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T06:21:22.000525Z 61b6114 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinyagents-graph/src/goals/mod.rs, crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/tools/time.rs, crates/tinyagents-harness/src/tools/time_test.rs, docs/modules/graph/goals.md, tinysweeper/description, tinysweeper/tests.

$0.0000 · 0 in / 0 out · 932 embedded · ladder/vectors

@tinysweeper tinysweeper Bot added the priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. label Sep 29, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e48380f877

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/tinyagents-graph/src/goals/tool.rs Outdated
senamakel and others added 6 commits September 30, 2026 03:29
Updated the vendored dependencies for tinyinference and tinytools to their latest versions, incorporating upstream fixes and improvements.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Refactor the summarization module to integrate a resilient wrapper around model summarization, improving fault tolerance during inference. The change introduces a new resilient layer that retries on transient failures and adds corresponding test coverage, while updating dependency locks to match.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ilient retry

Introduces a new summarization module that provides a model summarizer with resilient retry logic, enabling robust text summarization capabilities within the harness. The module includes a test file to verify the summarizer's behavior.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…agents-harness/src/title/mod.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The model summarizer test was incorrectly using a direct summarizer instead of the resilient wrapper, which caused test failures when the underlying service experienced transient errors. Updated the test to instantiate the resilient summarizer to match the production behavior.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the vendored tinytools dependency to incorporate upstream fixes and improvements. This change ensures the project uses the latest stable version of the library.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b3cbdfa635

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/tinyagents-graph/src/goals/mod.rs
Comment thread crates/tinyagents-harness/src/tools/time.rs
senamakel and others added 2 commits September 30, 2026 09:08
Co-authored-by: Medulla <medulla@tinyhumans.ai>
feat(harness): ModelSummarizer + fault-tolerant summarizer and thread-title helpers from OpenHuman
@senamakel
senamakel merged commit 938b6af into move-openhuman-tool-helpers Sep 30, 2026
9 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 61b61142de

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +135 to +138
loop {
let remaining = previous_token_estimate + estimate_slice_tokens(&messages[start..]);
if remaining <= self.fallback_trim_budget || start + 1 >= messages.len() {
break;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Enforce the fallback trim budget for oversized tails

When the newest message alone exceeds fallback_trim_budget—for example, a large tool result—the start + 1 >= messages.len() condition stops trimming while that entire message remains. The same happens when previous_summary alone exceeds the budget. The wrapper then reports successful recovery and returns an oversized checkpoint, so the subsequent model call can still overflow despite the summarizer failure having been handled; truncate or omit indivisible oversized content rather than unconditionally retaining it.

Useful? React with 👍 / 👎.

Comment on lines +118 to +123
let summary = summary.trim();
if summary.is_empty() {
return Err(TinyAgentsError::Model(
"summarizer returned empty response".into(),
));
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject summaries that do not reduce context

When the summarization model returns any nonempty but excessively verbose response, this path accepts it without an output cap or comparison against the source size. A model that echoes the transcript can therefore produce a replacement as large as—or larger than—the compacted head, after which the main model request remains over its context window even though compaction was recorded as successful. Bound the request's output tokens and/or fall back when the returned summary does not meaningfully shrink the input.

Useful? React with 👍 / 👎.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinyagents-graph/src/goals/test.rs, crates/tinyagents-graph/src/goals/tool.rs, crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs and 3 more.

             $0.0019 · 84,030 in / 8,568 out · 55,040 cached (66%) · ladder/vectors, deepseek/deepseek-v4-flash · 1,160 embedded
description: $0.0009 · 26,640 in / 1,952 out · 26,368 cached (99%) · deepseek/deepseek-v4-flash

@tinysweeper tinysweeper Bot added the priority: p2 Soon. Real but survivable — a rough edge, a gap, a thing that will bite later. label Sep 30, 2026
@tinysweeper tinysweeper Bot removed the priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. label Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p2 Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant