Skip to content

feat(extension): dogfood the decision features and add Budget (#241 PR 2) - #247

Merged
yuanhao merged 11 commits into
mainfrom
feat/extension-dogfood
Oct 6, 2026
Merged

yuanhao merged 11 commits into
mainfrom
feat/extension-dogfood

Conversation

@yuanhao

@yuanhao yuanhao commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

PR 2 of #241: the test of the design. If yoagent's own features can't be written as extensions, the API is wrong. They can.

Dogfooding

ToolGate, InputGuard and the decision advisor implement Extension:

Feature Hook Installed by
ToolGate before_tool (same decide) with_tool_gate
InputGuard on_input (same screen) with_input_guard
Advisor before_model note (same advise, memo shared across the agent's runs) with_decision_model
  • Wiring: decision::wire now returns extensions, appended after the agent's own, in Agent::build_config and SubAgentTool.
  • The old trait impls stay (ToolMiddleware, AsyncInputFilter) for installing by hand.
  • Ordering effects (CHANGELOG Changed):
    • the gate now also judges arguments an extension rewrote (it runs after every middleware and extension);
    • the guard screens after all input filters;
    • the advisor's note comes before turn-hook notes, once per turn (memoized per request, so it sends the same request).
  • All existing decision tests pass unchanged.
  • New test: gate_judges_the_arguments_an_extension_rewrote.

extension::Budget

  • What it does: a dollar limit that stops a run in before_model, before the request that would exceed it, with [Agent stopped: budget of $… spent …].
  • How spend is measured: each assistant MessageEnd's usage, through on_event, priced with one CostConfig. feat: Extension — one plug-in contract for the agent lifecycle (#241 PR 1) #246's flush barrier makes this deterministic.
  • Scope: per run, or .across_runs() (one total for a session, or for a whole delegation tree with with_tree_extension).
  • The issue's open question is settled: Budget::for_model(max, &model) returns None for an unpriced model, because a budget without a price is no limit.
  • Tests: tests/extension_budget_test.rs covers per run, across runs, a tree including a sub-agent's spend, and an unpriced model.

Not in this PR

yoagent-rutis collapsing into one impl Extension. It touches the same bridge files as #240 (rutis 0.6), which is on hold. It follows once #240 lands.

Checks

  • cargo test --all-features (58 binaries);
  • clippy with --all-features and --no-default-features;
  • wasm32 clippy with decision;
  • cargo doc -Dwarnings, fmt.

Docs: extensions.md (Budget, "Built on extensions"), decision-models.md, CHANGELOG, CLAUDE.md.

🤖 Generated with Claude Code

https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

yuanhao and others added 11 commits October 6, 2026 22:42
AgentLoopConfig::new(provider, model) sets every other field to the
defaults the struct-literal examples used. The struct is non_exhaustive,
so new fields stop being breaking changes for agent_loop callers; fields
stay public. Every in-tree literal (Agent, SubAgentTool, 30 test sites)
and the docs now use new(). transform_context is kept (#241 needs it until
on_context exists).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
…non_exhaustive] in listings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
Extension (a factory: name, mode, filters_tool_output,
rechecks_modified_calls, start_run, on_event) and RunHooks (fresh per run:
tools, on_input, before_model, before_tool, after_tool, on_stop, finish).
Advisory or required failure handling (a required failure ends the run
with an Error message prefixed EXTENSION_FAILED_PREFIX), partial tool
output withheld while an extension filters it, a rechecking policy judges
rewritten arguments again, on_stop continues capped by
max_stop_continues. Agent / SubAgentTool::with_extension, and
with_tree_extension to apply an extension to every delegated run at any
depth (host policy reaches sub-agents; the tree's tools stay with the
parent). The existing hooks are unchanged and the hook-order pin test
passes unmodified.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

# Conflicts:
#	CHANGELOG.md
#	src/agent.rs
#	src/sub_agent.rs
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
- Required failures are never lost: settle() (flush the event observer,
  then take the failure) at the top of each turn before the limit and
  on_before_turn checks, before on_stop, and after run_loop on any path.
- on_event is deterministic: the observer acknowledges a flush only after
  every event sent before it was observed; finish flushes first.
- A required verifier outvoted by an advisory Continue still fails at the
  cap; an extension that filters tool output and cannot start fails the
  run; started extensions still get finish when a later start fails.
- after_tool gets ToolOutput { result, is_error }; ExtensionMode is
  non_exhaustive; run_label passes to delegated runs (ToolContext::run_label);
  a failure outside a turn gets its own TurnStart/TurnEnd; required
  failures log at error!; ModelHalt replaces the unreachable arm.
- 11 more tests (34): lost-failure, on_event determinism, outvoted
  verifier, broken redactor, finish per ending, transcript after a stop,
  child denial honoured and tree tools not offered, child required failure,
  custom delegation tool, real concurrency (barrier), after_tool is_error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
…eard

- fail_run latches the run as failed: a failure recorded while the first
  is reported (an on_event that fails on those very events) is only logged,
  so there is one Error message and one on_error.
- on_stop Continue sends one line per extension that asked, so a required
  verifier's message reaches the model even when an advisory one asked
  first.
- Tests: a sink that panics on every MessageEnd fails the run exactly once
  (mutation-checked); both continue messages are sent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
…R 2)

- ToolGate, InputGuard and the decision advisor implement Extension;
  with_tool_gate / with_input_guard / with_decision_model install them as
  extensions after the agent's own (the gate runs last and judges
  arguments an extension rewrote). Their old trait impls stay for
  installing by hand.
- extension::Budget: a dollar limit, priced from each assistant
  MessageEnd's usage, checked in before_model; per run or across_runs
  (a session, or a tree with with_tree_extension); for_model is None for an
  unpriced model.
- Tests: the gate sees an extension's rewrite; budget per run, across runs,
  across a delegation tree, unpriced.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

# Conflicts:
#	CHANGELOG.md
#	CLAUDE.md
#	docs/concepts/extensions.md
#	src/agent.rs
#	src/extension.rs
#	src/sub_agent.rs
- before_tool / after_tool take &self and run under a read lock, so the
  calls of one response are judged concurrently again (a slow gate no
  longer serializes parallel calls); the other hooks keep &mut self.
- ToolGate rechecks modified calls, so a gate installed as a tree
  extension also judges an extension's rewrite.
- Budget::spent_usd() is the across-runs total (None per run); the
  CHANGELOG and test name say when Budget really stops; docs note that a
  failed mid-stream attempt reports no usage.
- Tests: parallel calls judged concurrently (a barrier; serializing them
  times out); the tree budget test now fails unless the spend is shared
  (child request count asserted).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
…efore tools run

- The guide's first example used &mut self for before_tool.
- Per-call hooks need interior mutability for state: said in the guide.
- A gate installed as a tree extension can cost a second decision
  request when a later extension rewrites the call: said.
- Extensions observe a turn's response before its tools run, so a
  sub-agent sees the tree's spend so far.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
@yuanhao
yuanhao merged commit 67175f3 into main Oct 6, 2026
14 checks passed
@yuanhao
yuanhao deleted the feat/extension-dogfood branch October 6, 2026 22:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant