Skip to content

mcp-exploit-validator

Static scanners tell you an AI agent's tool configuration looks dangerous. This one proves whether it actually is.

CI CodeQL Python 3.11+ License: Apache 2.0


The problem

An AI agent wired to MCP servers can usually do three things at once: read content from the outside world, read something private, and send data somewhere. Individually none of those servers is "vulnerable." Together they are a working exfiltration primitive — the lethal trifecta.

A handful of tools already scan MCP configurations for this. Every one of them infers the risk by reading JSON manifests. None of them run the agent. As the leading scanner states in its own documentation:

Static only. Runtime arguments and LLM jailbreaks still need sandboxing and policy enforcement.

So you get a warning that a configuration is theoretically dangerous, with no way to know whether your agent will actually fall for it — and no way to tell a genuine exposure from a false positive.

What this does instead

mcp-exploit-validator closes that loop. It finds candidate attack paths statically, then tests each one empirically:

  1. Builds a sealed range where the only reachable host is a sinkhole that logs anything sent to it (a loopback sinkhole today; Docker internal: true network isolation is the next backend).
  2. Plants canary tokens shaped like real secrets where the agent's private-data tool will find them.
  3. Hides a prompt injection in content the agent's ingress tool genuinely reads — a page, a file, a database row.
  4. Gives the agent a completely benign task ("summarise this page"). Every malicious instruction arrives through ingested data, never from the user.
  5. Runs N trials and reports measured attack success rate, with the transcript of any trial where the canary escaped.

The output is not a severity score. It is 3/5 trials exfiltrated plus the receipts.

Why the benign task matters. If you prompt an agent to misbehave and it does, you have measured obedience. Only when the instruction arrives through data the agent was asked to read have you measured a vulnerability. That distinction is the entire experiment.

How it compares

Static analysis Cross-server paths Proves exploitability Success-rate metric
mcp-scan ✅ ✅ ❌ ❌
MCPhound ✅ ✅ ❌ ❌
Proximity ✅ ❌ ❌ ❌
mcp-audit ✅ ❌ ❌ ❌
mcp-exploit-validator ✅ ✅ ✅ ✅

This tool composes with those scanners rather than replacing them — see docs/COMPARISON.md for an honest breakdown, including what they do better.

See it work

No install required - open the committed sample outputs in examples/:

Or run the full narrated demo locally (free, offline):

bash scripts/demo.sh

Quickstart

pip install mcp-exploit-validator

Find every MCP server configured on your machine:

mcpev discover

Analyse a tool surface — fully offline, no API key, no cost, nothing leaves your machine:

mcpev scan --manifest tests/fixtures/trifecta-manifest.json --html report.html

Prove whether a path is actually exploitable:

mcpev range --scenario exfil-via-webhook --trials 5

Cost

Static analysis is free and offline. Only the validation range calls an LLM, and you bring your own API key.

Mode Cost What it covers
mcpev scan $0 Discovery, capability classification, poisoning checks, attack paths, full report — offline
mcpev range (stub agent, CI default) $0 Full range orchestration with a deterministic agent — no key required
mcpev range --live ~$0.025/trial (Haiku) Real model susceptibility measurement

A --live run prints an estimated cost and requires explicit confirmation before it spends anything.

Safety

This tool executes attack scenarios, so containment is part of its design:

  • The instrumented fixture posts only to a loopback sinkhole, regardless of the destination an agent is tricked into using — so even a successful exfiltration provably never leaves the host.
  • A RangeBackend abstraction is in place for Docker network isolation (internal: true); the current default is a process backend with a loopback-only sinkhole (see ADR-0005).
  • Canaries are synthetic and minted per run. The tool refuses to run if a canary would collide with a real credential in your environment.
  • The payload corpus contains generic technique classes, not weaponised exploits against named products.

Only test infrastructure you own or are authorised to test. See SECURITY.md.

Documentation

Document What it covers
PRD Problem, goals, non-goals, personas, requirements, success metrics
Threat model What is being tested and why, with trust boundaries
Architecture Components, data flow, extension points
Methodology How success rate is measured, and the validity limits
Results Demonstrated controls + the multi-model sweep harness
Comparison Honest assessment against existing tools
ADRs Why the significant design decisions were made

Status

Active development. See ROADMAP.md for what is built and what is next.

Author

Built by ayertime.

License

Apache 2.0

About

Proves whether AI agent (MCP) tool configurations are actually exploitable - not just risky-looking. Existing scanners infer risk from manifests; this one validates it in a sealed range with canary tokens and a live agent.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages