Skip to content

Python: model MCP server handler parameters (mcp, fastmcp) as remote sources #22702

Description

@Alex-Hofer

Description of the issue

MCP (Model Context Protocol) servers expose tools to LLM agents. Every parameter of a tool,
resource or prompt handler is chosen by the model, which prompt injection can steer, or by any
client that can reach the server over HTTP. That makes these parameters remote input. CodeQL
has no models for the Python MCP SDKs, so in an MCP server the security queries have no source to
start from.

Evidence. mcp-vulnbench pins real, publicly
disclosed vulnerabilities in open-source Python MCP servers as vulnerable/fixed commit pairs with
function-level ground truth. It covers command injection, path traversal, SSRF, SQL injection and
code injection.

  • With the security-extended suite, CodeQL 2.27.1 detects 1 of 26 cases in v0.1.0, and that
    single hit comes from py/shell-command-constructed-from-input through its library-input source.
  • Example misses:
  • With the prototype models below (v0.2.0), the same CodeQL finds 13 of 28 cases. The models were
    written on a development half of 12 cases and frozen before the other 16 were measured; on that
    held-out half CodeQL goes from 0 to 9 detected cases (56 %), at 0.27 to 1.09 alarms per KLOC.
    A detection means a finding of the right class inside a function the fix changed or the sink
    function; nothing CodeQL found before is lost. Details:
    docs/results-v0.2.0.md.

Proposal. Source models of kind remote as Models-as-Data in
python/ql/lib/semmle/python/frameworks/, next to the existing Stdlib.model.yml sources and the
openai/anthropic sink models, plus typeModel rows for the import aliases. With them the
existing queries work unchanged and interprocedurally:
py/command-line-injection, py/path-injection, py/full-ssrf, py/code-injection and
py/sql-injection. Coverage:

  • mcp 1.x FastMCP: tool(), resource(), prompt(), add_tool(fn).
  • mcp 2.x MCPServer (the renamed FastMCP): the same entry points.
  • The low-level Server, two generations:
    • 1.x decorators call_tool(), read_resource(), get_prompt();
    • 2.x constructor handlers on_call_tool and friends, where params.arguments / params.uri
      is the source.
  • fastmcp 2.x–4.x:
    • tool and prompt (bare, called or as a plain call), resource();
    • add_tool, add_prompt, Tool.from_function, fastmcp.tools.tool;
    • the transport headers via get_http_headers() and get_http_request().headers.
  • Both SDKs: the bearer token of the Authorization header, where token verifiers receive it
    (verify_token in subclasses of TokenVerifier, and of AuthProvider in fastmcp) and through
    get_access_token().token. Real servers build file paths and queries from it.

A prototype pack with a test fixture (one handler per row) is in
models/codeql/mcp:

  • without the models the fixture has 0 alerts, with them every expected alert appears;
  • documented limits:
    • Pydantic-typed parameters (attribute reads are not taint steps; a QL model like the one for
      FastAPI's Pydantic parameters would cover them);
    • headers read through the Context object;
    • handlers behind a project decorator (auth or error handling), which is common in real
      servers. With functools.wraps the API graph loses the handler. Without it, the source
      lands on the wrapper, and the call func(*args, **kwargs) does not lead back to the
      handler. The Flask and FastAPI modeling avoids this by reading the decorator list
      (result.getADecorator()).

The library test would follow library-tests/frameworks/asyncpg/MaDTest.ql: a MaDTest.ql
importing experimental.meta.MaDTest, and a fixture where every row has a handler with a
# $ mad-source__remote=... expectation.

One observation while preparing this. Data-extension sources are ThreatModelSources but not
RemoteFlowSources. Three stable queries, py/nosql-injection, py/xml-bomb and py/xxe, take
RemoteFlowSource directly, while the other security customizations take
ActiveThreatModelSource. Remote sources from data extensions therefore never reach those three
queries. I checked this with CodeQL 2.27.1 (security-extended) on a small file:

  • the same three sinks, once fed from a Flask request and once from an MCP tool parameter
    (a data-extension source of kind remote);
  • a control where the MCP parameter reaches os.system.

The Flask variants alert (py/xxe, py/xml-bomb, py/nosql-injection) and so does the control
(py/command-line-injection). The MCP variants of the three queries stay silent.

Reproducer
import os
import xml.etree.ElementTree as ET

import pymongo
from flask import Flask, request
from lxml import etree
from mcp.server.fastmcp import FastMCP

mcp = FastMCP("probe")
app = Flask(__name__)
users = pymongo.MongoClient().db.users


# MCP tool parameters: MaD sources of kind remote
@mcp.tool()
def mcp_xxe(xml: str) -> str:
    parser = etree.XMLParser(resolve_entities=True)
    return str(etree.fromstring(xml, parser=parser))  # MCP-XXE


@mcp.tool()
def mcp_bomb(xml: str) -> str:
    return ET.fromstring(xml).tag  # MCP-BOMB


@mcp.tool()
def mcp_nosql(name: str) -> str:
    return str(list(users.find({"$where": "this.name == '" + name + "'"})))  # MCP-NOSQL


@mcp.tool()
def mcp_shell(command: str) -> int:
    return os.system(command)  # MCP-SHELL (control: the models are active)


# Flask request data: RemoteFlowSource, the control for the three queries
@app.route("/xxe")
def flask_xxe():
    parser = etree.XMLParser(resolve_entities=True)
    return str(etree.fromstring(request.args["xml"], parser=parser))  # FLASK-XXE


@app.route("/bomb")
def flask_bomb():
    return ET.fromstring(request.args["xml"]).tag  # FLASK-BOMB


@app.route("/nosql")
def flask_nosql():
    name = request.args["name"]
    return str(list(users.find({"$where": "this.name == '" + name + "'"})))  # FLASK-NOSQL

Questions

  1. Is Models-as-Data the form you prefer here? A QL module (semmle/python/frameworks/Mcp.qll)
    could also:
    • find handlers behind project decorators through the decorator list, as Flask.qll does;
    • cover Pydantic-typed parameters;
    • reach the three queries above.
      In the benchmark's development half, 2 of 8 misses are servers with such decorators.
  2. Is remote the right threat-model kind? Over stdio the input comes from the local client
    process, but its content is chosen by an LLM that reads remote content.
  3. Should the three queries above move to ActiveThreatModelSource? I can do that in a separate PR.

I'm happy to open a PR with the models, the library test and a change note.

Activity

  1. jketema commented on Sep 30, 2026

    @jketema
    Contributor

    Hi @Alex-Hofer,

    Thanks for opening this issue and the PR. Someone should eventually get around to reviewing your PR, but it might take a while, because we have seen quite a influx in external contributions over the last few weeks. Feel free to ping here if you believe it's taking too long.

  2. Alex-Hofer commented on Oct 3, 2026

    @Alex-Hofer
    Author

    I opened #22749 with the same kind of models for JavaScript and TypeScript: @modelcontextprotocol/sdk 1.x, @modelcontextprotocol/server 2.x and fastmcp, plus flow summaries for zod, which handlers of the low-level server validate their arguments with.

    Two things that relate to this issue:

    • The gap described above for py/xxe, py/xml-bomb and py/nosql-injection has no counterpart in JavaScript: there, sources from data extensions are ordinary RemoteFlowSources.
    • In JavaScript I found a different limit: js/path-injection does not follow a summary of kind taint from a data extension, it does follow one of kind value. The PR explains where that matters.

    The measurements behind both PRs are now in mcp-vulnbench v0.5.0: 66 cases, with the Python results unchanged from what I reported here.

  3. Alex-Hofer commented on Oct 4, 2026

    @Alex-Hofer
    Author

    #22750 proposes a second way to do this for Python: a QL class that extends RemoteFlowSource::Range, where #22703 uses data extensions. To make the two easier to compare I ran both on the 28 Python cases of mcp-vulnbench (CodeQL 2.27.1, python-security-extended; #22750 at commit 21a5a92, patched into the python-all of that bundle).

    cases detected
    CodeQL as shipped 1 / 28
    with the class of #22750 9 / 28
    with the rows of #22703 (as a model pack) 13 / 28
    with both 13 / 28

    The benchmark is mine, and the rows were written against 12 of these cases. On those 12 both detect 4; the difference is in the other 16.

    Each covers things the other does not:

    • Only the rows: handlers of the low-level Server (@server.call_tool(), @server.read_resource(); three of the four cases that only the rows detect), a tool decorator applied as a call, mcp.tool()(fn) (the fourth), and HTTP headers and bearer tokens.
    • Only the class: handlers with a project decorator between @mcp.tool and the function, where data extensions do not reach the handler. That changes no case in the table, but in four of the repositories it reaches handlers that the rows miss, and in one case it reports the vulnerable flow itself, which the rows do not. Its sources are RemoteFlowSources, so py/xxe, py/xml-bomb and py/nosql-injection see them (the gap described in the issue text above), and it leaves out Context parameters, which rows cannot.
    • The two PRs have no file in common, and with both active the alerts are exactly the union.

    To me they fit together, but that is the maintainers' call. I can leave #22703 as it is next to the class, or port what the class lacks to QL in a follow-up, whichever is easier to review.

    @Tito0015, thanks for the QL version; it addresses two of the limits I had listed in the issue. Method and the per-case table are in docs/comparison-codeql-22750.md, and I can rerun it when #22750 changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions