Repository navigation
feat: Extension — one plug-in contract for the agent lifecycle (#241 PR 1) - #246
Merged
Merged
Conversation
AgentLoopConfig::new(provider, model) sets every other field to the defaults the struct-literal examples used. The struct is non_exhaustive, so new fields stop being breaking changes for agent_loop callers; fields stay public. Every in-tree literal (Agent, SubAgentTool, 30 test sites) and the docs now use new(). transform_context is kept (#241 needs it until on_context exists). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
…non_exhaustive] in listings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
Extension (a factory: name, mode, filters_tool_output, rechecks_modified_calls, start_run, on_event) and RunHooks (fresh per run: tools, on_input, before_model, before_tool, after_tool, on_stop, finish). Advisory or required failure handling (a required failure ends the run with an Error message prefixed EXTENSION_FAILED_PREFIX), partial tool output withheld while an extension filters it, a rechecking policy judges rewritten arguments again, on_stop continues capped by max_stop_continues. Agent / SubAgentTool::with_extension, and with_tree_extension to apply an extension to every delegated run at any depth (host policy reaches sub-agents; the tree's tools stay with the parent). The existing hooks are unchanged and the hook-order pin test passes unmodified. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG # Conflicts: # CHANGELOG.md # src/agent.rs # src/sub_agent.rs
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
- Required failures are never lost: settle() (flush the event observer,
then take the failure) at the top of each turn before the limit and
on_before_turn checks, before on_stop, and after run_loop on any path.
- on_event is deterministic: the observer acknowledges a flush only after
every event sent before it was observed; finish flushes first.
- A required verifier outvoted by an advisory Continue still fails at the
cap; an extension that filters tool output and cannot start fails the
run; started extensions still get finish when a later start fails.
- after_tool gets ToolOutput { result, is_error }; ExtensionMode is
non_exhaustive; run_label passes to delegated runs (ToolContext::run_label);
a failure outside a turn gets its own TurnStart/TurnEnd; required
failures log at error!; ModelHalt replaces the unreachable arm.
- 11 more tests (34): lost-failure, on_event determinism, outvoted
verifier, broken redactor, finish per ending, transcript after a stop,
child denial honoured and tree tools not offered, child required failure,
custom delegation tool, real concurrency (barrier), after_tool is_error.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
…eard - fail_run latches the run as failed: a failure recorded while the first is reported (an on_event that fails on those very events) is only logged, so there is one Error message and one on_error. - on_stop Continue sends one line per extension that asked, so a required verifier's message reaches the model even when an advisory one asked first. - Tests: a sink that panics on every MessageEnd fails the run exactly once (mutation-checked); both continue messages are sent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
yuanhao
added a commit
that referenced
this pull request
Oct 6, 2026
…R 2) (#247) * feat!: AgentLoopConfig::new + #[non_exhaustive] (#228, for #241) AgentLoopConfig::new(provider, model) sets every other field to the defaults the struct-literal examples used. The struct is non_exhaustive, so new fields stop being breaking changes for agent_loop callers; fields stay public. Every in-tree literal (Agent, SubAgentTool, 30 test sites) and the docs now use new(). transform_context is kept (#241 needs it until on_context exists). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG * docs(AgentLoopConfig): review fixes — comments for default fields, #[non_exhaustive] in listings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG * feat: Extension — one plug-in contract for the agent lifecycle (#241) Extension (a factory: name, mode, filters_tool_output, rechecks_modified_calls, start_run, on_event) and RunHooks (fresh per run: tools, on_input, before_model, before_tool, after_tool, on_stop, finish). Advisory or required failure handling (a required failure ends the run with an Error message prefixed EXTENSION_FAILED_PREFIX), partial tool output withheld while an extension filters it, a rechecking policy judges rewritten arguments again, on_stop continues capped by max_stop_continues. Agent / SubAgentTool::with_extension, and with_tree_extension to apply an extension to every delegated run at any depth (host policy reaches sub-agents; the tree's tools stay with the parent). The existing hooks are unchanged and the hook-order pin test passes unmodified. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG * fix(merge): drop the duplicate AgentLoopConfig::new left by merging main Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG * fix(extension): review fixes for #246 - Required failures are never lost: settle() (flush the event observer, then take the failure) at the top of each turn before the limit and on_before_turn checks, before on_stop, and after run_loop on any path. - on_event is deterministic: the observer acknowledges a flush only after every event sent before it was observed; finish flushes first. - A required verifier outvoted by an advisory Continue still fails at the cap; an extension that filters tool output and cannot start fails the run; started extensions still get finish when a later start fails. - after_tool gets ToolOutput { result, is_error }; ExtensionMode is non_exhaustive; run_label passes to delegated runs (ToolContext::run_label); a failure outside a turn gets its own TurnStart/TurnEnd; required failures log at error!; ModelHalt replaces the unreachable arm. - 11 more tests (34): lost-failure, on_event determinism, outvoted verifier, broken redactor, finish per ending, transcript after a stop, child denial honoured and tree tools not offered, child required failure, custom delegation tool, real concurrency (barrier), after_tool is_error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG * fix(extension): a run fails once; every extension that continues is heard - fail_run latches the run as failed: a failure recorded while the first is reported (an on_event that fails on those very events) is only logged, so there is one Error message and one on_error. - on_stop Continue sends one line per extension that asked, so a required verifier's message reaches the model even when an advisory one asked first. - Tests: a sink that panics on every MessageEnd fails the run exactly once (mutation-checked); both continue messages are sent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG * feat(extension): dogfood the decision features and add Budget (#241 PR 2) - ToolGate, InputGuard and the decision advisor implement Extension; with_tool_gate / with_input_guard / with_decision_model install them as extensions after the agent's own (the gate runs last and judges arguments an extension rewrote). Their old trait impls stay for installing by hand. - extension::Budget: a dollar limit, priced from each assistant MessageEnd's usage, checked in before_model; per run or across_runs (a session, or a tree with with_tree_extension); for_model is None for an unpriced model. - Tests: the gate sees an extension's rewrite; budget per run, across runs, across a delegation tree, unpriced. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG * fix(extension): review fixes for #247 - before_tool / after_tool take &self and run under a read lock, so the calls of one response are judged concurrently again (a slow gate no longer serializes parallel calls); the other hooks keep &mut self. - ToolGate rechecks modified calls, so a gate installed as a tree extension also judges an extension's rewrite. - Budget::spent_usd() is the across-runs total (None per run); the CHANGELOG and test name say when Budget really stops; docs note that a failed mid-stream attempt reports no usage. - Tests: parallel calls judged concurrently (a barrier; serializing them times out); the tree budget test now fails unless the spend is shared (child request count asserted). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG * docs(extension): before_tool takes &self in the guide; flush events before tools run - The guide's first example used &mut self for before_tool. - Per-call hooks need interior mutability for state: said in the guide. - A gate installed as a tree extension can cost a second decision request when a later extension rewrites the call: said. - Extensions observe a turn's response before its tools run, so a sub-agent sees the tree's spend so far. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
10 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR 1 of #241: the
Extensioncontract and its loop wiring. It builds on #242 (hook-order pin test), #244 (loop fixes) and #245 (AgentLoopConfig::new), all merged.The contract (
src/extension.rs,yoagent::extension)Extension, the factory, one per installation. It hasname,mode(Advisory/Required),filters_tool_output,rechecks_modified_calls,start_run(&RunContext) -> Box<dyn RunHooks>, and a syncon_event(&self, run_id, event).RunHooks, fresh per run,&mut self, no-op defaults. It hastools,on_input,before_model,before_tool,after_tool,on_stopandfinish.Types:
RunContext:run_id, the host'slabel,prompts,depth,cancel.InputContext,StopContext; decisionsInputDecision,TurnDecision,StopDecision;RunOutcome,ExtensionError. All the context, decision and outcome types are#[non_exhaustive].Stateless::new(name, hooks)clones the hooks per run.EXTENSION_FAILED_PREFIX,EXTENSION_MESSAGE_PREFIX(inis_loop_injected) andDEFAULT_MAX_STOP_CONTINUES.API
AgentandSubAgentTool:with_extension,with_tree_extension,with_max_stop_continues.Agent::with_run_label.AgentLoopConfiggainsextensions,tree_extensions,max_stop_continuesandrun_label(non-breaking since feat!: AgentLoopConfig::new + #[non_exhaustive] (#228, for #241) #245).ToolContextgainstree_extensions()anddelegation_depth()for custom delegation tools.Semantics, as revised in #241
Failures follow the mode table.
on_inputrejects,before_tooldenies,after_toolwithholds the result.Errormessage prefixed[Extension failed: <name>], callson_error, and keepsTurnStart/TurnEndpaired.after_toolandon_eventare checked at the next boundary.before_modelruns once per turn, before the retry loop.Stopbehaves like a limit (a stop marker).StoporFail.before_toolruns after theToolMiddlewarechain. Arechecks_modified_callspolicy judges the final arguments again when a later extension rewrote them (the gap raised on Design: Extension — one plug-in contract for the agent lifecycle #241).after_toolsees the result before truncation, the output sink andToolExecutionEnd. While an extensionfilters_tool_output,ToolExecutionUpdate/ProgressMessageare withheld.on_stopapplies when the last message is an assistantStopand no follow-ups are queued.Continueappends[Extension <name>] …as a loop-injected user turn, capped. Exceeding the cap fails a required extension and only warns for an advisory one.Tree extensions:
ToolContext, whichSubAgentToolreads;Extension.This closes the existing bypass where a parent's middleware did not gate a sub-agent's own calls.
Deviations from the issue text (reasoned)
on_eventis onExtension(&self), notRunHooks(&mut self). If it locked the run's hooks, an event would wait behind an in-flight async hook. For example, an approvalbefore_toolwaiting on a UI that waits on the event would stall. Events are delivered in order through a forwarder task that exists only when extensions do, and is drained before the run returns.finish(&mut self)instead ofself: Box<Self>.tests/hook_order_test.rspasses unmodified. That is the compatibility guarantee.before_modelnotes land beforeTurnHooknotes, because turn hooks run inside the provider call, once per attempt.Tests
tests/extension_test.rs(23 cases) covers every hook and the mode table:Stop, the rechecking policy;on_stopcap, advisory and required;on_eventseeing every event in consumer order;finishonce per run;continue_looprunning noon_input;RunContextcontents;tests/hook_order_test.rsis unchanged.Checks
cargo test --all-features: 57 test binaries.--all-featuresand--no-default-features.CLIPPY_CONF_DIR=.github/clippy-wasm32,--features decision).cargo doc -Dwarnings, fmt.Docs:
docs/concepts/extensions.md, in SUMMARY;Unreleased → Added;PR 2 (dogfooding ToolGate / InputGuard / Advisor / Budget / yoagent-rutis) follows.
🤖 Generated with Claude Code
https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG