Skip to main content

What Does Your MCP Attack Surface Look Like From the Outside?

· 8 min read
Bandana Kaur
APISec Research Labs

What is the first step of attacking an MCP server? Nearly every single one answers a question for free, to anyone who asks: what can you do here? The response, tools/list, is usually treated as a harmless discovery mechanism. We wanted to challenge that and see how far schema-based recon can really get us. A tool list is a capability disclosure, and reading it carefully, without ever calling a single tool, can tell you more about a server's risk than you'd expect. This article develops a threat model for exactly that, then we test it against 234 public MCP servers to see how much of it holds up.

What even is tools/list?​

Most security discussion of MCP servers begins and ends with one question: does the server require authentication? That question is necessary but incomplete. It treats tools/list as a locked or unlocked door, when in practice an open tool list is closer to a floor plan.

An MCP tools/list response is a machine-readable description of a server's entire interface including tool names, parameter names and types, free-text descriptions, and in some cases explicit behavioral annotations (readOnlyHint, destructiveHint). Some may say it looks eerily similar to an OpenAPI spec, except it is also written in natural language to be consumed directly by a language model.

This sort of schema is what an attacker (or a good hacker like me) needs before deciding where to spend effort, and what an agent needs before deciding what to trust. In this article from APIsec Labs, we'll show you a threat model that organizes what is learnable from an MCP schema into distinct classes. What does the attack surface look like for MCP servers IRL? We also conducted a pilot measurement showing that several of these classes are computable in bulk, passively, across real public infrastructure.

Threat Model​

Observers​

Who reads a tool listing? We can define three parties who read a tool list, each with a different objective:

ObserverObjectiveCapability required
Passive observerCatalog what a server exposesRead access to tools/list only
Targeting attackerDecide which server or tool to pursueSame, plus intent to act on the schema
Consuming agentDecide what to trust and invokeIngests schema and descriptions as context

Tool descriptions and server instructions read into the model's context and can shape its behavior. The same field that functions as reconnaissance for an attacker functions as an injection surface for an agent. So tool descriptions shouldn't just be read as some inert metadata.

Top 5 things the schema gives away​

We organize what a schema reveals into five classes, further mapped as observed (read directly from the schema, no interpretation required) or inferred (requires a judgment call with attached uncertainty).

ClassDescriptionStatus
Object-addressing surfaceWhich tools take parameters that reference an object by ID, path, or name, mostly the structural precondition for Broken Object Level AuthorizationObserved
Identity modelWhether identity is derived server-side or supplied by the caller (e.g., owner_key, client_id as a request parameter)Inferred
Effect classWhether a tool reads, writes, or destroys data, from annotations, naming, and description textInferred (observed where annotations are present)
Backend leakageInfrastructure and storage hints surfaced in parameter names (e.g., r2path, database_id)Observed
Agent-facing trustUnsigned instructions, resource, and prompt text ingested by downstream agentsObserved

In practice, this could look like:

{
"name": "get_listing_details",
"description": "Fetch full details for a marketplace listing, including owner contact info.",
"inputSchema": {
"type": "object",
"properties": {
"listing_id": { "type": "string" }, // object-addressing (explicit_id)
"owner_key": { "type": "string" }, // identity model: client-asserted
"r2path": { "type": "string" } // backend leakage (storage hint)
},
"required": ["listing_id", "owner_key"]
},
"annotations": {
"readOnlyHint": false, // effect class: not declared read-only
"destructiveHint": false
}
}

The observed classes can be reported as counts, while the inferred classes require a measured classifier and a stated precision before they can be reported as anything stronger than a heuristic signal. We treat them accordingly.

Methodology​

We used REAP, an open-source reconnaissance tool for agentic surfaces, to discover and probe public MCP endpoints.

  • Discovery frame: 6,951 endpoints identified via public registries and public code/documentation.
  • Probed: 247 of the discovery frame.
  • Confirmed MCP endpoints: 234, defined as targets where confirm_state resolved to confirmed or confirmed_auth_gated.

All our measurement was passive. We captured tools/list responses and their input schemas where readable. No tool was invoked at any point, and no authorization check was tested.

We then classified parameters into identifier categories by token matching against a fixed vocabulary (explicit_id, named_object, path_like), with session and credential tokens explicitly excluded. explicit_id matches (e.g., *_id, *_key) are treated as high-confidence; named_object matches (e.g., channel, handle, task) are treated as a heuristic upper bound, since several of these tokens are ambiguous without inspecting the surrounding schema.

What do real, public MCP servers tell us?​

What could an attacker point at?​

Across 234 confirmed endpoints, we inspected 3,253 tools. Of these, 1,191 (36.6%) contained at least one object-addressing parameter.

MetricValue
Tools inspected3,253
Object-addressing tools1,191 (36.6%)
Endpoints with ≥1 object-addressing tool174 / 234 (74.4%)
Object-addressing tools per exposed endpointmedian 5, max 127
Identifier categoryTools
named_object941
explicit_id758
path_like28

The most frequent identifier parameters were agent_id (252 tools), proof_escrow_id (218), slug (197), and run_id (161). Several others like agent_name, contact_uri, and owner_key appeared on an identical 112 tools each, and listing_id and listing_slug each appeared on 109.

That clustering is informative in itself: identical counts across unrelated parameter names are more consistent with a small number of large platforms or templated deployments contributing many tools apiece than with hundreds of independently designed servers.

Agent-facing trust​

ExposureEndpointsRate
Server instructions returned161 / 23468.8%
Resources or prompts exposed19 / 2348.1%

All of this text is unauthenticated, unsigned, and by design, ingested directly into agent context. Nothing in the protocol distinguishes it from any other untrusted input an agent might retrieve.

Classical transport hygiene, the boring checks​

For comparison, standard web-transport checks on the same 234 endpoints were uniformly clean:

CheckFlagged / ApplicableRate
Plaintext transport0 / 2340.0%
TLS certificate health0 / 2340.0%
Session-ID entropy0 / 1390.0%
Redirect-URI laxity0 / 2340.0%

This is a useful contrast, where the exposures in this study are protocol-specific, not symptoms of generally poor operational hygiene. Registry-listed, HTTPS-reachable servers appear to be a filtered population with respect to transport security.

What do you do with this?​

If you build MCPs:

  • Derive identity server-side from authenticated credentials; do not accept scoping identifiers as caller-supplied tool parameters.
  • Populate behavioral annotations (readOnlyHint, destructiveHint) accurately, they cost little and benefit both human reviewers and agents.
  • Treat the tool list as documentation an adversary reads first; avoid leaking backend structure through parameter naming.

If you build agents:

  • Treat tool descriptions, instructions, and prompt/resource text as untrusted input, subject to the same scrutiny as any other retrieved document.
  • Pin or hash reviewed schemas and alert on drift between sessions.

If you're a security nerd like me:

  • Triage before interacting: sort tools into public-catalog, tenant-scoped, and client-asserted-identity categories directly from the schema, and prioritize the last category.

Limitations​

Our pilot study shows that object-addressing surface is computable from schemas alone, at scale, with zero interaction, and that the parameter vocabulary is rich enough to separate public catalog identifiers (slug, listing_slug) from tenant-scoped ones (owner_key, client_id, database_id), the latter being the sharper BOLA-precondition signal. It also shows that agent-facing text like instructions, resources, and prompts is exposed on most servers with no provenance mechanism attached.

What it doesn't show is that a tool taking file_id is a confirmed BOLA, since no tool was ever invoked and no downstream authorization check was tested. The near-identical parameter counts suggest the true count of independently designed exposures is lower than the raw tool-level figures imply, so we don't treat these as a clean prevalence estimate, and the results shouldn't be generalized to the broader MCP ecosystem.

Also, the identifier classifier's precision hasn't yet been measured (named_object matches are a heuristic upper bound pending a hand-labeled sample), and no clustering adjustment has been applied to correct for the shared-platform effect noted above. The inference classes, particularly identity model and effect class, still need validation against servers with known, controlled configurations to establish real precision and recall. And the sample itself is small relative to the discovery frame: 247 of 6,951 endpoints probed (3.6%), drawn only from public registries and public code or documentation per the study's pre-registration.

Zero tool calls later…​

Whether an MCP server requires authentication is not the only question worth asking. Before a single tool is ever called, its schema already discloses its object model, hints at where authorization has been delegated to the caller, and hands an agent unsigned text it will treat as context. This pilot shows that reconnaissance against that disclosure is not theoretical, it is computable, passively, across real public infrastructure, at a scale that makes manual review impractical and automated triage necessary. If a schema can tell a stranger this much after you ship, imagine what it tells ai-surface before you do… ;)