Skip to content

feat: Extension — one plug-in contract for the agent lifecycle (#241 PR 1) - #246

Merged
yuanhao merged 7 commits into
mainfrom
feat/extension-contract
Oct 6, 2026
Merged

yuanhao merged 7 commits into
mainfrom
feat/extension-contract

Conversation

@yuanhao

@yuanhao yuanhao commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

PR 1 of #241: the Extension contract and its loop wiring. It builds on #242 (hook-order pin test), #244 (loop fixes) and #245 (AgentLoopConfig::new), all merged.

The contract (src/extension.rs, yoagent::extension)

Extension, the factory, one per installation. It has name, mode (Advisory/Required), filters_tool_output, rechecks_modified_calls, start_run(&RunContext) -> Box<dyn RunHooks>, and a sync on_event(&self, run_id, event).

RunHooks, fresh per run, &mut self, no-op defaults. It has tools, on_input, before_model, before_tool, after_tool, on_stop and finish.

Types:

  • RunContext: run_id, the host's label, prompts, depth, cancel.
  • Contexts InputContext, StopContext; decisions InputDecision, TurnDecision, StopDecision; RunOutcome, ExtensionError. All the context, decision and outcome types are #[non_exhaustive].
  • Stateless::new(name, hooks) clones the hooks per run.
  • Constants EXTENSION_FAILED_PREFIX, EXTENSION_MESSAGE_PREFIX (in is_loop_injected) and DEFAULT_MAX_STOP_CONTINUES.

API

  • Agent and SubAgentTool: with_extension, with_tree_extension, with_max_stop_continues.
  • Agent::with_run_label.
  • AgentLoopConfig gains extensions, tree_extensions, max_stop_continues and run_label (non-breaking since feat!: AgentLoopConfig::new + #[non_exhaustive] (#228, for #241) #245).
  • ToolContext gains tree_extensions() and delegation_depth() for custom delegation tools.

Semantics, as revised in #241

  • Failures follow the mode table.

    • The single-call guards fail closed in both modes: on_input rejects, before_tool denies, after_tool withholds the result.
    • A required failure ends the run with an assistant Error message prefixed [Extension failed: <name>], calls on_error, and keeps TurnStart/TurnEnd paired.
    • Failures inside after_tool and on_event are checked at the next boundary.
  • before_model runs once per turn, before the retry loop.

    • Notes go on the latest user turn and are never stored.
    • Stop behaves like a limit (a stop marker).
    • The model is never called on Stop or Fail.
  • before_tool runs after the ToolMiddleware chain. A rechecks_modified_calls policy judges the final arguments again when a later extension rewrote them (the gap raised on Design: Extension — one plug-in contract for the agent lifecycle #241).

  • after_tool sees the result before truncation, the output sink and ToolExecutionEnd. While an extension filters_tool_output, ToolExecutionUpdate/ProgressMessage are withheld.

  • on_stop applies when the last message is an assistant Stop and no follow-ups are queued. Continue appends [Extension <name>] … as a loop-injected user turn, capped. Exceeding the cap fails a required extension and only warns for an advisory one.

  • Tree extensions:

    • they are prepended in child runs at any depth, passed through ToolContext, which SubAgentTool reads;
    • their tools are not offered to children;
    • state for one run is fresh per run, and aggregate state lives on the Extension.

    This closes the existing bypass where a parent's middleware did not gate a sub-agent's own calls.

Deviations from the issue text (reasoned)

  • on_event is on Extension (&self), not RunHooks (&mut self). If it locked the run's hooks, an event would wait behind an in-flight async hook. For example, an approval before_tool waiting on a UI that waits on the event would stall. Events are delivered in order through a forwarder task that exists only when extensions do, and is drained before the run returns.
  • finish(&mut self) instead of self: Box<Self>.
  • The older hooks are not re-routed through adapters. They keep their exact code paths and run first at each point, so tests/hook_order_test.rs passes unmodified. That is the compatibility guarantee. before_model notes land before TurnHook notes, because turn hooks run inside the provider call, once per attempt.

Tests

tests/extension_test.rs (23 cases) covers every hook and the mode table:

  • panics contained per mode, start failures, notes, Stop, the rechecking policy;
  • redaction with partial output withheld (and sent when no extension filters it);
  • the on_stop cap, advisory and required;
  • on_event seeing every event in consumer order;
  • per-run isolation across two concurrent agents sharing an extension, with finish once per run;
  • continue_loop running no on_input;
  • RunContext contents;
  • tree extensions: a deny at depth 0, 1 and 2 while an ordinary extension stays with the parent, the tree's tools not offered to children, and a tree-wide budget stopping a child.

tests/hook_order_test.rs is unchanged.

Checks

  • cargo test --all-features: 57 test binaries.
  • clippy with --all-features and --no-default-features.
  • wasm32 clippy (CLIPPY_CONF_DIR=.github/clippy-wasm32, --features decision).
  • cargo doc -Dwarnings, fmt.

Docs:

  • a new docs/concepts/extensions.md, in SUMMARY;
  • CHANGELOG Unreleased → Added;
  • a CLAUDE.md architecture note.

PR 2 (dogfooding ToolGate / InputGuard / Advisor / Budget / yoagent-rutis) follows.

🤖 Generated with Claude Code

https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

yuanhao and others added 7 commits October 6, 2026 22:42
AgentLoopConfig::new(provider, model) sets every other field to the
defaults the struct-literal examples used. The struct is non_exhaustive,
so new fields stop being breaking changes for agent_loop callers; fields
stay public. Every in-tree literal (Agent, SubAgentTool, 30 test sites)
and the docs now use new(). transform_context is kept (#241 needs it until
on_context exists).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
…non_exhaustive] in listings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
Extension (a factory: name, mode, filters_tool_output,
rechecks_modified_calls, start_run, on_event) and RunHooks (fresh per run:
tools, on_input, before_model, before_tool, after_tool, on_stop, finish).
Advisory or required failure handling (a required failure ends the run
with an Error message prefixed EXTENSION_FAILED_PREFIX), partial tool
output withheld while an extension filters it, a rechecking policy judges
rewritten arguments again, on_stop continues capped by
max_stop_continues. Agent / SubAgentTool::with_extension, and
with_tree_extension to apply an extension to every delegated run at any
depth (host policy reaches sub-agents; the tree's tools stay with the
parent). The existing hooks are unchanged and the hook-order pin test
passes unmodified.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

# Conflicts:
#	CHANGELOG.md
#	src/agent.rs
#	src/sub_agent.rs
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
- Required failures are never lost: settle() (flush the event observer,
  then take the failure) at the top of each turn before the limit and
  on_before_turn checks, before on_stop, and after run_loop on any path.
- on_event is deterministic: the observer acknowledges a flush only after
  every event sent before it was observed; finish flushes first.
- A required verifier outvoted by an advisory Continue still fails at the
  cap; an extension that filters tool output and cannot start fails the
  run; started extensions still get finish when a later start fails.
- after_tool gets ToolOutput { result, is_error }; ExtensionMode is
  non_exhaustive; run_label passes to delegated runs (ToolContext::run_label);
  a failure outside a turn gets its own TurnStart/TurnEnd; required
  failures log at error!; ModelHalt replaces the unreachable arm.
- 11 more tests (34): lost-failure, on_event determinism, outvoted
  verifier, broken redactor, finish per ending, transcript after a stop,
  child denial honoured and tree tools not offered, child required failure,
  custom delegation tool, real concurrency (barrier), after_tool is_error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
…eard

- fail_run latches the run as failed: a failure recorded while the first
  is reported (an on_event that fails on those very events) is only logged,
  so there is one Error message and one on_error.
- on_stop Continue sends one line per extension that asked, so a required
  verifier's message reaches the model even when an advisory one asked
  first.
- Tests: a sink that panics on every MessageEnd fails the run exactly once
  (mutation-checked); both continue messages are sent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG
@yuanhao
yuanhao merged commit f60c71b into main Oct 6, 2026
14 checks passed
@yuanhao
yuanhao deleted the feat/extension-contract branch October 6, 2026 21:50
yuanhao added a commit that referenced this pull request Oct 6, 2026
…R 2) (#247)

* feat!: AgentLoopConfig::new + #[non_exhaustive] (#228, for #241)

AgentLoopConfig::new(provider, model) sets every other field to the
defaults the struct-literal examples used. The struct is non_exhaustive,
so new fields stop being breaking changes for agent_loop callers; fields
stay public. Every in-tree literal (Agent, SubAgentTool, 30 test sites)
and the docs now use new(). transform_context is kept (#241 needs it until
on_context exists).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

* docs(AgentLoopConfig): review fixes — comments for default fields, #[non_exhaustive] in listings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

* feat: Extension — one plug-in contract for the agent lifecycle (#241)

Extension (a factory: name, mode, filters_tool_output,
rechecks_modified_calls, start_run, on_event) and RunHooks (fresh per run:
tools, on_input, before_model, before_tool, after_tool, on_stop, finish).
Advisory or required failure handling (a required failure ends the run
with an Error message prefixed EXTENSION_FAILED_PREFIX), partial tool
output withheld while an extension filters it, a rechecking policy judges
rewritten arguments again, on_stop continues capped by
max_stop_continues. Agent / SubAgentTool::with_extension, and
with_tree_extension to apply an extension to every delegated run at any
depth (host policy reaches sub-agents; the tree's tools stay with the
parent). The existing hooks are unchanged and the hook-order pin test
passes unmodified.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

* fix(merge): drop the duplicate AgentLoopConfig::new left by merging main

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

* fix(extension): review fixes for #246

- Required failures are never lost: settle() (flush the event observer,
  then take the failure) at the top of each turn before the limit and
  on_before_turn checks, before on_stop, and after run_loop on any path.
- on_event is deterministic: the observer acknowledges a flush only after
  every event sent before it was observed; finish flushes first.
- A required verifier outvoted by an advisory Continue still fails at the
  cap; an extension that filters tool output and cannot start fails the
  run; started extensions still get finish when a later start fails.
- after_tool gets ToolOutput { result, is_error }; ExtensionMode is
  non_exhaustive; run_label passes to delegated runs (ToolContext::run_label);
  a failure outside a turn gets its own TurnStart/TurnEnd; required
  failures log at error!; ModelHalt replaces the unreachable arm.
- 11 more tests (34): lost-failure, on_event determinism, outvoted
  verifier, broken redactor, finish per ending, transcript after a stop,
  child denial honoured and tree tools not offered, child required failure,
  custom delegation tool, real concurrency (barrier), after_tool is_error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

* fix(extension): a run fails once; every extension that continues is heard

- fail_run latches the run as failed: a failure recorded while the first
  is reported (an on_event that fails on those very events) is only logged,
  so there is one Error message and one on_error.
- on_stop Continue sends one line per extension that asked, so a required
  verifier's message reaches the model even when an advisory one asked
  first.
- Tests: a sink that panics on every MessageEnd fails the run exactly once
  (mutation-checked); both continue messages are sent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

* feat(extension): dogfood the decision features and add Budget (#241 PR 2)

- ToolGate, InputGuard and the decision advisor implement Extension;
  with_tool_gate / with_input_guard / with_decision_model install them as
  extensions after the agent's own (the gate runs last and judges
  arguments an extension rewrote). Their old trait impls stay for
  installing by hand.
- extension::Budget: a dollar limit, priced from each assistant
  MessageEnd's usage, checked in before_model; per run or across_runs
  (a session, or a tree with with_tree_extension); for_model is None for an
  unpriced model.
- Tests: the gate sees an extension's rewrite; budget per run, across runs,
  across a delegation tree, unpriced.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

* fix(extension): review fixes for #247

- before_tool / after_tool take &self and run under a read lock, so the
  calls of one response are judged concurrently again (a slow gate no
  longer serializes parallel calls); the other hooks keep &mut self.
- ToolGate rechecks modified calls, so a gate installed as a tree
  extension also judges an extension's rewrite.
- Budget::spent_usd() is the across-runs total (None per run); the
  CHANGELOG and test name say when Budget really stops; docs note that a
  failed mid-stream attempt reports no usage.
- Tests: parallel calls judged concurrently (a barrier; serializing them
  times out); the tree budget test now fails unless the spend is shared
  (child request count asserted).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

* docs(extension): before_tool takes &self in the guide; flush events before tools run

- The guide's first example used &mut self for before_tool.
- Per-call hooks need interior mutability for state: said in the guide.
- A gate installed as a tree extension can cost a second decision
  request when a later extension rewrites the call: said.
- Extensions observe a turn's response before its tools run, so a
  sub-agent sees the tree's spend so far.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T7iq5hpndSiHQcnAsywKuG

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant