Skip to content

feat(harness): ModelSummarizer + fault-tolerant summarizer and thread-title helpers from OpenHuman - #230

Merged
senamakel merged 8 commits into
oh-dedupe-time-goals-toolsfrom
oh-extract-leaves
Sep 30, 2026
Merged

senamakel merged 8 commits into
oh-dedupe-time-goals-toolsfrom
oh-extract-leaves

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • tinyagents-harness::summarization: ModelSummarizer (LLM-backed Summarizer), FaultTolerantCachingSummarizer (per-turn circuit breaker, deterministic-trim fallback, single-slot content-hash cache) and summarization_policy / summarization_policy_with. The 0.90 threshold and keep-last-8 are now parameters (DEFAULT_SUMMARIZE_THRESHOLD_FRACTION, DEFAULT_SUMMARIZE_KEEP_LAST are the defaults), with tests using ScriptedModel.
  • tinyagents-harness::title: thread-title shaping (shorten_title, sanitize_generated_title, title_from_user_message, placeholder detection) and build_title_request, moved with their 30+ tests from OpenHuman's threads/title.rs. The [threads:title] log prefix stays in the host.
  • Bumps vendor/tinytools (new tinytools-std: feat(tinytools-std): file_state, url_guard, detect_tools from OpenHuman; parser regression tests tinytools#33) and vendor/tinyinference (scrub_credentials: feat(core): scrub_credentials in sanitize (from OpenHuman) tinyinference#40).

config::required_output already existed here; OpenHuman's duplicate copy is being deleted host-side.

Stacking

Base is oh-dedupe-time-goals-tools (#229, itself stacked on #228). The two nested submodule PRs above are likewise stacked on their pinned branches; merge them first, innermost first.

Verification

cargo fmt --check, cargo clippy -p tinyagents-harness --all-targets -- -D warnings, cargo test -p tinyagents-harness (1439 lib tests passed).

Co-authored-by: Medulla medulla@tinyhumans.ai

senamakel and others added 6 commits September 30, 2026 03:29
Updated the vendored dependencies for tinyinference and tinytools to their latest versions, incorporating upstream fixes and improvements.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Refactor the summarization module to integrate a resilient wrapper around model summarization, improving fault tolerance during inference. The change introduces a new resilient layer that retries on transient failures and adds corresponding test coverage, while updating dependency locks to match.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ilient retry

Introduces a new summarization module that provides a model summarizer with resilient retry logic, enabling robust text summarization capabilities within the harness. The module includes a test file to verify the summarizer's behavior.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…agents-harness/src/title/mod.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The model summarizer test was incorrectly using a direct summarizer instead of the resilient wrapper, which caused test failures when the underlying service experienced transient errors. Updated the test to instantiate the resilient summarizer to match the production behavior.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the vendored tinytools dependency to incorporate upstream fixes and improvements. This change ensures the project uses the latest stable version of the library.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Incomplete
Priority: low
Reviewed head: ab5287d8dbe1
Updated: 1790749661 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 6 Active findings 0
Tests 2 Noted findings 0
Documentation 0 Resolved findings 0
Configuration 0 Pending checks/questions 10

Completeness: Incomplete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

No active actionable findings.

Could not review: crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs

Before merge

  • Complete the critique review for crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs.
  • Complete the security review for crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/types.rs.
Agent review details

critique

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs
  • Lane summary: Reviewed 0 files; 0 findings. 5 files could not be reviewed: crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs.

security

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/types.rs
  • Lane summary: Reviewed 0 files; 0 findings. 5 files could not be reviewed: crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/types.rs.

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The pull request adds LLM-backed conversation summarization via `ModelSummarizer`, a fault-tolerant caching wrapper `FaultTolerantCachingSummarizer`, context-window-aware policy builders, and a thread-title generation module. The tests cover success, failure, caching, fallback trim, and title parsing comprehensively. No critical or high-severity defects found; the changes are safe to merge. Minor gaps exist (e.g., no test for a single word longer than the character ceiling being truncated), but these are low severity and do not block merging. The findings are about test coverage gaps, not code defects. _The code index is behind this pull request (indexed at `b3cbdfa6355f`), so retrieved context may be out of date._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Introduces `ModelSummarizer`, `FaultTolerantCachingSummarizer`, and thread-title helpers with comprehensive tests. The code is functionally sound, but the test file `model_summarizer_test.rs` is placed outside the `model_summarizer` module, violating the repository convention of module-local tests in `test.rs`. _The code index is behind this pull request (indexed at `b3cbdfa6355f`), so retrieved context may be out of date._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: ladder/vectors, deepseek/deepseek-v4-flash
  • Spend: $0.008875
  • Tokens: 118586 input · 16707 output · 48128 cached · 1056 embedding
Head State Pass summary
ed9cc77997d6 incomplete 0 active finding(s), 0 resolved finding(s) (at 1790729177)
ab5287d8dbe1 incomplete 0 active finding(s), 0 resolved finding(s) (at 1790749661)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: e2d9edd4-3621-4d37-a0b4-83c11e4f9233

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T06:16:54.111711Z ab5287d New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ed9cc77997

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/tinyagents-harness/src/summarization/model_summarizer.rs
Comment thread crates/tinyagents-harness/src/summarization/resilient.rs Outdated
Comment thread crates/tinyagents-harness/src/summarization/resilient.rs
Comment thread crates/tinyagents-harness/src/summarization/model_summarizer.rs Outdated

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinyagents-harness/src/lib.rs, crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/title/mod.rs, crates/tinyagents-harness/src/title/test.rs, tinysweeper/tests.

             $0.0025 · 39,316 in / 4,013 out · 0 cached (0%) · ladder/vectors, deepseek/deepseek-v4-flash · 1,023 embedded
description: $0.0003 · 21,110 in / 78 out    · 0 cached (0%) · deepseek/deepseek-v4-flash

@tinysweeper tinysweeper Bot added the priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. label Sep 30, 2026
@senamakel
senamakel merged commit 61b6114 into oh-dedupe-time-goals-tools Sep 30, 2026
9 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ab5287d8db

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +136 to +138
let remaining = previous_token_estimate + estimate_slice_tokens(&messages[start..]);
if remaining <= self.fallback_trim_budget || start + 1 >= messages.len() {
break;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Enforce the deterministic fallback budget

When the prior checkpoint already exceeds fallback_trim_budget, or the final remaining message is itself oversized, start + 1 >= messages.len() exits this loop while remaining is still over budget. The fallback then emits the entire previous summary and final message, so a summarizer outage can leave the retried model request above its context window and fail the turn this adapter is intended to preserve; truncate or discard over-budget content, including the previous checkpoint, before returning.

Useful? React with 👍 / 👎.

Comment on lines +36 to +39
pub use model_summarizer::{
DEFAULT_SUMMARIZE_KEEP_LAST, DEFAULT_SUMMARIZE_THRESHOLD_FRACTION, summarization_policy,
summarization_policy_with,
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Document the new summarization surface

These new public policy builders, along with ModelSummarizer and FaultTolerantCachingSummarizer, are absent from summarization/README.md: its public-surface section still describes only ConcatSummarizer, and its file table omits both new implementation files. Update the module documentation so consumers can discover the intended construction, caching, and failure behavior, as required for public API changes.

AGENTS.md reference: AGENTS.md:L78-L82

Useful? React with 👍 / 👎.

Comment on lines +285 to +286
#[cfg(test)]
mod model_summarizer_test;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep module-local tests in test.rs

This registers unit tests from the new sibling model_summarizer_test.rs, although the repository requires module-local unit tests to live in a dedicated test.rs. Move these cases into the existing summarization test.rs, or make model_summarizer a module directory containing its own test.rs, so the new feature follows the required module layout.

AGENTS.md reference: AGENTS.md:L18-L20

Useful? React with 👍 / 👎.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinyagents-harness/src/summarization/mod.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_test.rs, crates/tinyagents-harness/src/summarization/resilient.rs, crates/tinyagents-harness/src/summarization/types.rs.

             $0.0089 · 118,586 in / 16,707 out · 48,128 cached (41%) · ladder/vectors, deepseek/deepseek-v4-flash · 1,056 embedded
tests:       $0.0054 · 58,915 in  / 5,158 out  · 27,648 cached (47%) · deepseek/deepseek-v4-flash
description: $0.0026 · 19,551 in  / 4,807 out  · 0 cached (0%)       · deepseek/deepseek-v4-flash

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant