crates.io · Docs · API · GitHub · DeepWiki · Changelog
The agent loop for Rust. Stream from any of 7 LLM protocols, run tools, loop until done.
git clone https://github.com/yologdev/yoagent && cd yoagent
ollama serve &
ollama pull llama3.1:8b # or pass --model <any pulled model>
cargo run --example cli -- --provider ollamaThat's a working coding agent in your terminal — file read/write/edit, shell, ripgrep search, streaming output, skills. No signup, no key, nothing to configure.
yoagent cli — mini coding agent
Type /quit to exit, /clear to reset
model: llama3.1:8b
cwd: /home/user/my-project
> find all TODO comments in src/
▶ search 'TODO' ✓
Found 3 TODOs:
src/main.rs:42: // TODO: handle edge case
src/lib.rs:15: // TODO: add tests
src/utils.rs:8: // TODO: optimize this
tokens: 1250 in / 89 out
Point it at a hosted model instead by swapping the flag:
ANTHROPIC_API_KEY=sk-... cargo run --example cli
GROQ_API_KEY=... cargo run --example cli -- --provider groq --model openai/gpt-oss-120b
cargo run --example cli -- --api-url http://localhost:1234/v1 --model my-model # LM Studio, llama.cpp, vLLM[dependencies]
yoagent = "0.24"
tokio = { version = "1", features = ["full"] }Building for wasm32? Use default-features = false (see
WebAssembly & Cloudflare Workers).
On a native target keep the default native feature: without it reqwest has no TLS, and every
HTTPS provider call fails at runtime.
An agent that actually uses a tool — the thing the crate exists for:
use yoagent::provider::ModelConfig;
use yoagent::{tools, Agent, AgentEvent, StreamDelta};
#[tokio::main]
async fn main() {
// The provider is selected from the config's protocol and the key is read
// from ANTHROPIC_API_KEY. Call `.with_api_key(k)` to pass one explicitly.
let mut agent = Agent::from_config(ModelConfig::claude_sonnet_5())
.with_system_prompt("You are a coding assistant.")
.with_tools(tools::default_tools());
let mut events = agent.prompt("Find every TODO in src/ and summarise them").await;
while let Some(event) = events.recv().await {
match event {
AgentEvent::MessageUpdate { delta: StreamDelta::Text { delta }, .. } => print!("{delta}"),
AgentEvent::ToolExecutionStart { tool_name, .. } => println!("\n▶ {tool_name}"),
AgentEvent::AgentEnd { .. } => break,
_ => {}
}
}
agent.finish().await;
}Swap the model by swapping the config — the provider follows, and the key is read from that provider's conventional env var:
Agent::from_config(ModelConfig::groq("openai/gpt-oss-120b", "GPT-OSS 120B")); // GROQ_API_KEY
Agent::from_config(ModelConfig::google("gemini-3.8-flash", "Gemini 3.8 Flash")); // GEMINI_API_KEY
Agent::from_config(ModelConfig::ollama("http://localhost:11434/v1", "llama3.1:8b")); // no keyyoagent is deliberately narrow. It is the loop, tool execution, and the machinery you need to run that loop in production. It ships no vector stores, embedding pipelines, or task-graph layer — if your problem is retrieval or orchestration, one of these is the better fit:
| If you need | Look at |
|---|---|
| RAG pipelines, vector stores, embeddings, transcription and image generation | rig — "Build modular and scalable LLM Applications in Rust" |
| Typed task graphs and streaming RAG indexing alongside agents | swiftide — "Composable LLM agents and harness, typed task graphs, and streaming RAG pipelines in Rust" |
| A tool-calling loop you host, gate, steer, branch, and record | yoagent |
What that focus bought:
- The loop is a free function.
agent_loop()is stateless and takes everything it needs as arguments.Agentis an optional wrapper that adds history and queues. You can drive the loop yourself without adopting our state model. - 7 native wire protocols, not one OpenAI-compat shim with adapters bolted on. Anthropic Messages, OpenAI Completions, OpenAI Responses, Azure, Gemini, Vertex, and Bedrock each have a real implementation, so provider-specific features (thinking budgets, prompt-cache breakpoints, reasoning deltas) survive instead of being flattened away.
- Every tool call passes one gate.
ToolMiddlewarecan allow, modify, or deny each call at a single choke point shared by all execution strategies — the mechanism behind approval prompts and policy engines. - Steer a run that's already going. Inject guidance mid-flight; it's picked up between tool batches without restarting the turn.
- History is a tree, not a list.
Sessionforks, checkpoints, and seeks. Edit an earlier turn and re-run it without destroying the original branch. - Runs are recordable. With
features = ["gasp"], a run becomes an append-only semantic event log in a git repo — restore is clone + replay. Conformance-checked in CI. - The whole loop is testable offline.
MockProviderscripts multi-turn tool-calling conversations and honours cancellation, so abort and steering paths are testable with no network. 1,197 of our tests run with no network and no key.
yoyo-evolve — a coding
agent that evolves its own source in public. It began as 200 lines of Rust; every commit since has
been agent-written and gated on tests. It runs on this loop with the
openapi feature enabled.
Also built on yoagent:
| Project | What it is |
|---|---|
rab |
A lightweight, extensible Rust coding agent |
greatsage |
"Rimuru's Unique Skill, you know the one" |
yoclaw |
OpenClaw reborn in Rust — a single-binary agent that remembers you |
Built something on yoagent? Open a PR and add it here — we'd like to see it.
The loop & control
- Full event stream:
AgentStart→TurnStart→MessageUpdate(deltas) →ToolExecution*→TurnEnd→AgentEnd; a retried provider attempt is closed byMessageEnd(StopReason::Error) followed byProviderRetry - Parallel tool execution by default;
SequentialandBatched { size }strategies available - Steering — interrupt mid-run; follow-ups — queue work after completion; both queues are inspectable and editable
ToolMiddleware— asyncAllow/Modify(args)/Deny(reason)hooks gating every call. A denial becomes an error tool result the model sees, so the loop keeps goingInputFilter— rewrite or reject user input before it reaches the model (PII redaction, prompt-injection guards);AsyncInputFilterfor filters that awaitTurnHook— an async hook before every LLM request that may add one transient note to the latest user turn (the cached prefix is untouched); middleware can read the conversation (ToolCallRequest::messages,user_request())- Execution limits (max turns, max tokens, wall-clock timeout),
abort(), and lifecycle callbacks (before_turn,after_turn,on_error) - Automatic retry with exponential backoff and ±20% jitter, for rate-limit and network errors only;
retry::retry_safe_eventsholds each attempt's output back until it succeeds, for output that cannot take text back (a pipe, a log)
Providers — 7 protocols, 20+ providers
| Protocol | Providers |
|---|---|
| Anthropic Messages | Anthropic (Claude) |
| OpenAI Completions | OpenAI, xAI, Groq, Cerebras, OpenRouter, Mistral, DeepSeek, MiniMax, Z.ai, Qwen, Meta (Muse Spark), Ollama, local servers, custom compatible APIs |
| OpenAI Responses | OpenAI (Responses API) |
| Azure OpenAI | Azure OpenAI |
| Google Generative AI | Google Gemini |
| Google Vertex | Google Vertex AI |
| Bedrock ConverseStream | Amazon Bedrock |
ModelConfig presets cover the common providers; ModelConfig::openai_compat(..) handles anything
else with a base_url. Per-provider quirks (auth style, reasoning format, max_tokens field name)
live in OpenAiCompat / AnthropicCompat flags — 12 compat profiles ship in the box.
The opencode_zen(..) / opencode_go(..) gateways pick the wire protocol from the model id
automatically, so one config reaches models across several vendors.
Thinking/reasoning controls are wired for all 7 protocols. Client-side cache hints are sent on two:
Anthropic gets cache_control breakpoints, native OpenAI gets a prompt_cache_key. Most others cache
server-side with nothing to configure; Bedrock does not cache automatically, and its explicit
cachePoint blocks are not yet wired. Context-overflow detection is centralised across 15+
provider-specific error strings.
Tools — built-in, custom, MCP, OpenAPI
Built in: bash (timeout, deny patterns), read_file / write_file (line numbers, path
restrictions), edit_file (fuzzy-match hints on failure), list_files, search (ripgrep) —
native hosts only (the default native feature; not on wasm32). Tools return stdout and stderr even on failure, so the model can self-correct.
Custom tools implement one trait:
#[async_trait::async_trait]
impl AgentTool for GreetTool {
fn name(&self) -> &str { "greet" }
fn label(&self) -> &str { "Greet" }
fn description(&self) -> &str { "Greets someone" }
fn parameters_schema(&self) -> serde_json::Value {
serde_json::json!({ "type": "object", "properties": { "name": { "type": "string" } } })
}
async fn execute(&self, params: serde_json::Value, _ctx: ToolContext)
-> Result<ToolResult, ToolError>
{
let name = params["name"].as_str().unwrap_or("stranger");
Ok(ToolResult {
content: vec![Content::Text { text: format!("Hello, {name}!") }],
details: serde_json::Value::Null,
})
}
}A tool that must also build for wasm32 swaps the bare #[async_trait] for the cfg_attr pair in
the wasm guide.
MCP — with_mcp_server_stdio() / with_mcp_server_http() connect to Model Context Protocol
servers over stdio or Streamable HTTP (session ids, SSE framing, incremental parsing) and register
their tools transparently. Stdio needs the default native feature; HTTP works on wasm32 too.
ToolSource — tools resolved at the start of every run (with_tool_source), for tool sets
that change while the agent lives: plugin systems, reconnecting MCP servers. The
yoagent-rutis bridge (not yet on crates.io) builds on it so
rutis plugins can add tools, gate calls, add turn notes and
filter input.
OpenAPI (features = ["openapi"]) — point with_openapi_url() at a spec and every operation
becomes a tool, filtered by OperationFilter.
Sub-agents & shared state
SubAgentTool delegates to a child loop with its own model, system prompt, tools, skills,
middleware, retry policy, and turn limits — a fully independent configuration, not a thin shim.
Run a cheap model for triage and an expensive one for the hard step in the same session.
SharedState is a pluggable key-value store (MemoryBackend, FileBackend (native), or your own via the
SharedStateBackend trait). A parent stores a large artifact once and sub-agents read it by key,
so it never gets re-pasted into every context window. Opt in with .with_shared_state(state) —
it injects the shared_state tool and a state summary into the sub-agent's system prompt.
Context, sessions & skills
-
ContextTracker— hybrid real-usage + estimation, calibrated against actual provider usage -
Tiered compaction — truncate tool outputs → summarise old turns → drop middle turns
-
LlmCompaction— opt-in alternative that summarises the dropped span with a background LLM request instead of discarding it. The request runs off the hot path and is spliced in on a later turn, so the loop never stalls and can never wedge on it — an unfinished or failed summary falls back to the deterministic tiers. Buys retention quality; costs tokens the default never spends, and does not reduce prefix-cache breaks. Both paths report their own cost onAgentEvent::ContextCompactedlet agent = Agent::from_config(ModelConfig::claude_sonnet_5()) .with_compaction_strategy(LlmCompaction::from_config( ModelConfig::claude_haiku_4_5(), // cheap model for summaries ));
-
Session— history as an id/parent tree withappend,seek,checkpoint,branch_tips, and JSONL persistence. Appending after a seek forks a branch; it never overwrites -
Skills — load AgentSkills-standard
SKILL.mddirectories. The agent sees a compact index and reads the full skill on demand, so skills stay cross-compatible with Claude Code, Codex CLI, Cursor, and others -
Structured outputs —
prompt_structured::<T>()returns typed, schema-validated replies, enforced natively where supported (Anthropicoutput_config.formatfor Claude ids from 4.5 — presets,ModelConfig::anthropic, OpenCode — tool-forcing otherwise; OpenAI Chat Completionsjson_schema; GeminiresponseSchema)
Decision models — typed judgments in a few hundred ms (feature decision)
A decision model answers typed questions — yes/no (Noul), one-of-N (Choice), a rating (Score) — with calibrated probabilities instead of prose. The first backend is TypeSafe's Jev; integrations depend on a DecisionBackend trait, not on a vendor. Off by default, no extra dependencies, and nothing is sent until you pick a model:
let jev = DecisionModel::jev(); // key from TYPESAFE_API_KEY
let urgent = jev.noul(message, "Does this convey urgency?").await?; // urgent.p_true()
let jev = jev.or(DecisionModel::logprobs("http://localhost:8080", "llama-3.1-8b-instruct")); // fallback: any logprob LLM, thinking off
let agent = Agent::from_config(ModelConfig::claude_sonnet_5())
.with_skills(skills)
.with_decision_model(jev.clone()) // advisory only: skill / tool hints; needs skills or 40+ tools
.with_tool_gate(ToolGate::new(jev.clone())) // opt-in: deny destructive, unrequested calls; fails closed
.with_input_guard(InputGuard::new(jev)); // opt-in: reject injection / harmful prompts; fails closedDecisionModel::logprobs(url, model) turns any OpenAI-compatible server that returns logprobs (llama.cpp, vLLM, SGLang, LM Studio) into a decision model — approximately calibrated; measure it on your own examples with decision::calibrate. Hosted models see what you send; DecisionModel::local(url) or a loopback logprobs server keeps it on your machine. The tool gate and input guard are defence in depth, not a security boundary — injected content can steer a decision model. See Decision Models.
Production concerns
- Cost tracking —
CostConfigcarries separate input/output/cache-read/cache-write rates plus optional context tiers;session_cost_usd()gives a running total,AgentEvent::AgentEndcarries aSessionStatsrollup, andModelConfig::costis anOption—Nonemeans pricing unknown, never $0. Rates live in a data file (src/provider/prices.json) that first-party constructors look up by provider and model id; override them at runtime withprices::global::install_overrideorYOAGENT_PRICES, per config withwith_prices/reprice, or opt into a live source (PriceTable::fetchfrom models.dev or the checked file onmain) - Loop detection — a model calling one tool with identical arguments forever trips none of the turn/token/duration limits until the whole budget is spent. On by default: steers on the third consecutive repeat, stops on the next, and emits
AgentEvent::LoopDetectedeither way - Retrievable tool output — head-tail truncation discards the middle irrecoverably. Attach a
SharedStateand the full text is stashed, with the marker naming a key the model can fetch - Telemetry —
tracingspans per loop / LLM stream / tool, recording tokens and cost. OpenTelemetry is bridged app-side viatracing-opentelemetry; the library carries no OTel dependency by design - GASP (
features = ["gasp"]) — record runs into a GASP agent repo; yoagent is a tested-conformant runtime, with the 7-check suite running in CI - Serde throughout — every core type is
Serialize/Deserialize/PartialEq, so sessions persist and replay set_model()— hot-swap the model mid-session without rebuilding the agent- WebAssembly —
--no-default-featuresbuilds forwasm32-unknown-unknown(e.g. Cloudflare Workers): the loop, every provider over the host'sfetch, HTTP MCP, in-memorySharedState, sub-agents anddecision. See WebAssembly & Cloudflare Workers
Eleven of the runnable examples in examples/ are below; five need no API key at all. The rest are live-provider harnesses and offline evaluation sweeps.
| Example | What it shows | Key needed |
|---|---|---|
cli |
A 385-line coding agent — all tools, skills, streaming, colored output. Like a baby Claude Code | optional¹ |
rlm |
An LLM that explores a codebase on its own by spawning sub-agents | yes |
code_review |
Three sub-agents reviewing a diff in parallel, results merged | yes |
shared_state |
Passing a large artifact between sub-agents by reference | yes |
sub_agent |
Delegation basics with a per-sub-agent model | yes |
basic |
The smallest possible agent | yes |
callbacks |
Lifecycle hooks and a custom tool | no |
persistence |
Save and restore a session | no |
telemetry |
tracing spans with token and cost fields |
no |
gasp_emit |
Recording a run into a GASP repo | no |
decision |
Decision-model questions in one line, and attaching a model to an agent (feature decision) |
yes |
¹ --provider ollama or --api-url needs no key; hosted providers read their conventional env var.
MockProvider scripts a whole multi-turn tool-calling conversation with no network (guide: Testing Your Agent):
use yoagent::provider::mock::{MockProvider, MockResponse, MockToolCall};
let provider = MockProvider::new(vec![
MockResponse::ToolCalls(vec![MockToolCall {
name: "search".into(),
arguments: serde_json::json!({ "pattern": "TODO" }),
provider_metadata: None,
}]),
MockResponse::Text("Found 3 TODOs.".into()),
]);
let agent = Agent::from_provider(provider, ModelConfig::mock());It emits real StreamEvents and honours the CancellationToken, so abort and steering paths are
testable too.
- 1,197 tests run with no network and no API keys —
cargo test --all-features; 16 more are opt-in live checks and benchmarks - HTTP-level tests with
wiremockacross 19 suites: provider SSE streams, Bedrock auth and eventstream, MCP over HTTP, OpenAPI, decision backends and price fetching clippy --all-targets --all-featureswith-Dwarnings,cargo fmt --check- Linux + macOS test matrix, a Windows compile check, a pinned MSRV 1.86 job, per-feature builds (default,
openapi,gasp,decision,--no-default-features), awasm32-unknown-unknownclippy job plus a wasm32 test suite under Node, theyoagent-rutisbridge jobs, and a GASP conformance job
| Module | What lives there |
|---|---|
agent_loop |
The loop itself — agent_loop, agent_loop_continue, AgentLoopConfig, execution strategies |
agent |
Optional stateful wrapper — history, tool registry, steering/follow-up queues |
rt |
spawn, sleep, timeout, Instant, MaybeSend / MaybeSync — Tokio on native targets, the host executor on wasm32 |
types |
Message, Content, AgentEvent, AgentTool, ToolMiddleware, InputFilter |
provider/ |
StreamProvider trait, ModelConfig, registry, and the 7 protocol implementations + MockProvider |
tools/ |
bash, file, edit, list, search, shared_state_tool |
tool_source |
ToolSource — tools resolved at the start of every run |
sub_agent |
SubAgentTool — delegation to child loops |
shared_state |
SharedState + pluggable backends |
session |
Branching conversation trees with JSONL persistence |
context |
Token tracking, tiered compaction, execution limits |
llm_compaction |
LlmCompaction — summarise the dropped span with a background LLM request |
skills |
AgentSkills SKILL.md loading |
retry |
Backoff with jitter; retry_safe_events filter for append-only consumers |
mcp/ |
MCP client, stdio (native) + HTTP transports, tool adapter |
openapi/ |
OpenAPI 3.0 → tools (feature openapi) |
gasp |
Run recording into a GASP repo (feature gasp) |
decision/ |
Decision models (SystemOne and logprob backends, fallbacks, calibration), advisory hints, the tool gate and the input guard (feature decision) |
- The book — concepts, guides, and a page per provider (source)
- API reference — built with all features enabled
- CHANGELOG — every release
- CONTRIBUTING — how to build, test, and send a PR
MSRV is 1.86, enforced in CI. Raising it is a minor-version change.
MIT — see LICENSE.