Static scanners tell you an AI agent's tool configuration looks dangerous. This one proves whether it actually is.
An AI agent wired to MCP servers can usually do three things at once: read content from the outside world, read something private, and send data somewhere. Individually none of those servers is "vulnerable." Together they are a working exfiltration primitive — the lethal trifecta.
A handful of tools already scan MCP configurations for this. Every one of them infers the risk by reading JSON manifests. None of them run the agent. As the leading scanner states in its own documentation:
Static only. Runtime arguments and LLM jailbreaks still need sandboxing and policy enforcement.
So you get a warning that a configuration is theoretically dangerous, with no way to know whether your agent will actually fall for it — and no way to tell a genuine exposure from a false positive.
mcp-exploit-validator closes that loop. It finds candidate attack paths statically,
then tests each one empirically:
- Builds a sealed range where the only reachable host is a sinkhole that logs
anything sent to it (a loopback sinkhole today; Docker
internal: truenetwork isolation is the next backend). - Plants canary tokens shaped like real secrets where the agent's private-data tool will find them.
- Hides a prompt injection in content the agent's ingress tool genuinely reads — a page, a file, a database row.
- Gives the agent a completely benign task ("summarise this page"). Every malicious instruction arrives through ingested data, never from the user.
- Runs N trials and reports measured attack success rate, with the transcript of any trial where the canary escaped.
The output is not a severity score. It is 3/5 trials exfiltrated plus the receipts.
Why the benign task matters. If you prompt an agent to misbehave and it does, you have measured obedience. Only when the instruction arrives through data the agent was asked to read have you measured a vulnerability. That distinction is the entire experiment.
| Static analysis | Cross-server paths | Proves exploitability | Success-rate metric | |
|---|---|---|---|---|
mcp-scan |
✅ | ✅ | ❌ | ❌ |
MCPhound |
✅ | ✅ | ❌ | ❌ |
Proximity |
✅ | ❌ | ❌ | ❌ |
mcp-audit |
✅ | ❌ | ❌ | ❌ |
mcp-exploit-validator |
✅ | ✅ | ✅ | ✅ |
This tool composes with those scanners rather than replacing them — see
docs/COMPARISON.md for an honest breakdown, including what
they do better.
No install required - open the committed sample outputs in examples/:
sample-scan-report.html- a static scan (inference)sample-range-result.json- a validation-range run:5/5 trials exfiltratedwith the transcript (proof)
Or run the full narrated demo locally (free, offline):
bash scripts/demo.shpip install mcp-exploit-validatorFind every MCP server configured on your machine:
mcpev discoverAnalyse a tool surface — fully offline, no API key, no cost, nothing leaves your machine:
mcpev scan --manifest tests/fixtures/trifecta-manifest.json --html report.htmlProve whether a path is actually exploitable:
mcpev range --scenario exfil-via-webhook --trials 5Static analysis is free and offline. Only the validation range calls an LLM, and you bring your own API key.
| Mode | Cost | What it covers |
|---|---|---|
mcpev scan |
$0 | Discovery, capability classification, poisoning checks, attack paths, full report — offline |
mcpev range (stub agent, CI default) |
$0 | Full range orchestration with a deterministic agent — no key required |
mcpev range --live |
~$0.025/trial (Haiku) | Real model susceptibility measurement |
A --live run prints an estimated cost and requires explicit confirmation before it
spends anything.
This tool executes attack scenarios, so containment is part of its design:
- The instrumented fixture posts only to a loopback sinkhole, regardless of the destination an agent is tricked into using — so even a successful exfiltration provably never leaves the host.
- A
RangeBackendabstraction is in place for Docker network isolation (internal: true); the current default is a process backend with a loopback-only sinkhole (see ADR-0005). - Canaries are synthetic and minted per run. The tool refuses to run if a canary would collide with a real credential in your environment.
- The payload corpus contains generic technique classes, not weaponised exploits against named products.
Only test infrastructure you own or are authorised to test. See
SECURITY.md.
| Document | What it covers |
|---|---|
| PRD | Problem, goals, non-goals, personas, requirements, success metrics |
| Threat model | What is being tested and why, with trust boundaries |
| Architecture | Components, data flow, extension points |
| Methodology | How success rate is measured, and the validity limits |
| Results | Demonstrated controls + the multi-model sweep harness |
| Comparison | Honest assessment against existing tools |
| ADRs | Why the significant design decisions were made |
Active development. See ROADMAP.md for what is built and what is next.
Built by ayertime.