What Does Your MCP Attack Surface Look Like From the Outside?

What is the first step of attacking an MCP server? Nearly every single one answers a question for free, to anyone who asks: what can you do here? The response, tools/list, is usually treated as a harmless discovery mechanism. We wanted to challenge that and see how far schema-based recon can really get us. A tool list is a capability disclosure, and reading it carefully, without ever calling a single tool, can tell you more about a server's risk than you'd expect. This article develops a threat model for exactly that, then we test it against 234 public MCP servers to see how much of it holds up.
What even is tools/list?
Most security discussion of MCP servers begins and ends with one question: does the server require authentication? That question is necessary but incomplete. It treats tools/list as a locked or unlocked door, when in practice an open tool list is closer to a floor plan.
An MCP tools/list response is a machine-readable description of a server's entire interface including tool names, parameter names and types, free-text descriptions, and in some cases explicit behavioral annotations (readOnlyHint, destructiveHint). Some may say it looks eerily similar to an OpenAPI spec, except it is also written in natural language to be consumed directly by a language model.
This sort of schema is what an attacker (or a good hacker like me) needs before deciding where to spend effort, and what an agent needs before deciding what to trust. In this article from APIsec Labs, we'll show you a threat model that organizes what is learnable from an MCP schema into distinct classes. What does the attack surface look like for MCP servers IRL? We also conducted a pilot measurement showing that several of these classes are computable in bulk, passively, across real public infrastructure.
Threat Model
Observers
Who reads a tool listing? We can define three parties who read a tool list, each with a different objective:
| Observer | Objective | Capability required |
|---|---|---|
| Passive observer | Catalog what a server exposes | Read access to tools/list only |
| Targeting attacker | Decide which server or tool to pursue | Same, plus intent to act on the schema |
| Consuming agent | Decide what to trust and invoke | Ingests schema and descriptions as context |
Tool descriptions and server instructions read into the model's context and can shape its behavior. The same field that functions as reconnaissance for an attacker functions as an injection surface for an agent. So tool descriptions shouldn't just be read as some inert metadata.
Top 5 things the schema gives away
We organize what a schema reveals into five classes, further mapped as observed (read directly from the schema, no interpretation required) or inferred (requires a judgment call with attached uncertainty).
| Class | Description | Status |
|---|---|---|
| Object-addressing surface | Which tools take parameters that reference an object by ID, path, or name, mostly the structural precondition for Broken Object Level Authorization | Observed |
| Identity model | Whether identity is derived server-side or supplied by the caller (e.g., owner_key, client_id as a request parameter) | Inferred |
| Effect class | Whether a tool reads, writes, or destroys data, from annotations, naming, and description text | Inferred (observed where annotations are present) |
| Backend leakage | Infrastructure and storage hints surfaced in parameter names (e.g., r2path, database_id) | Observed |
| Agent-facing trust | Unsigned instructions, resource, and prompt text ingested by downstream agents | Observed |
In practice, this could look like:
{
"name": "get_listing_details",
"description": "Fetch full details for a marketplace listing, including owner contact info.",
"inputSchema": {
"type": "object",
"properties": {
"listing_id": { "type": "string" }, // object-addressing (explicit_id)
"owner_key": { "type": "string" }, // identity model: client-asserted
"r2path": { "type": "string" } // backend leakage (storage hint)
},
"required": ["listing_id", "owner_key"]
},
"annotations": {
"readOnlyHint": false, // effect class: not declared read-only
"destructiveHint": false
}
}
The observed classes can be reported as counts, while the inferred classes require a measured classifier and a stated precision before they can be reported as anything stronger than a heuristic signal. We treat them accordingly.
Methodology
We used REAP, an open-source reconnaissance tool for agentic surfaces, to discover and probe public MCP endpoints.
- Discovery frame: 6,951 endpoints identified via public registries and public code/documentation.
- Probed: 247 of the discovery frame.
- Confirmed MCP endpoints: 234, defined as targets where
confirm_stateresolved toconfirmedorconfirmed_auth_gated.
All our measurement was passive. We captured tools/list responses and their input schemas where readable. No tool was invoked at any point, and no authorization check was tested.
We then classified parameters into identifier categories by token matching against a fixed vocabulary (explicit_id, named_object, path_like), with session and credential tokens explicitly excluded. explicit_id matches (e.g., *_id, *_key) are treated as high-confidence; named_object matches (e.g., channel, handle, task) are treated as a heuristic upper bound, since several of these tokens are ambiguous without inspecting the surrounding schema.
What do real, public MCP servers tell us?
What could an attacker point at?
Across 234 confirmed endpoints, we inspected 3,253 tools. Of these, 1,191 (36.6%) contained at least one object-addressing parameter.
| Metric | Value |
|---|---|
| Tools inspected | 3,253 |
| Object-addressing tools | 1,191 (36.6%) |
| Endpoints with ≥1 object-addressing tool | 174 / 234 (74.4%) |
| Object-addressing tools per exposed endpoint | median 5, max 127 |
| Identifier category | Tools |
|---|---|
named_object | 941 |
explicit_id | 758 |
path_like | 28 |
The most frequent identifier parameters were agent_id (252 tools), proof_escrow_id (218), slug (197), and run_id (161). Several others like agent_name, contact_uri, and owner_key appeared on an identical 112 tools each, and listing_id and listing_slug each appeared on 109.
That clustering is informative in itself: identical counts across unrelated parameter names are more consistent with a small number of large platforms or templated deployments contributing many tools apiece than with hundreds of independently designed servers.
Agent-facing trust
| Exposure | Endpoints | Rate |
|---|---|---|
Server instructions returned | 161 / 234 | 68.8% |
| Resources or prompts exposed | 19 / 234 | 8.1% |
All of this text is unauthenticated, unsigned, and by design, ingested directly into agent context. Nothing in the protocol distinguishes it from any other untrusted input an agent might retrieve.
Classical transport hygiene, the boring checks
For comparison, standard web-transport checks on the same 234 endpoints were uniformly clean:
| Check | Flagged / Applicable | Rate |
|---|---|---|
| Plaintext transport | 0 / 234 | 0.0% |
| TLS certificate health | 0 / 234 | 0.0% |
| Session-ID entropy | 0 / 139 | 0.0% |
| Redirect-URI laxity | 0 / 234 | 0.0% |
This is a useful contrast, where the exposures in this study are protocol-specific, not symptoms of generally poor operational hygiene. Registry-listed, HTTPS-reachable servers appear to be a filtered population with respect to transport security.
What do you do with this?
If you build MCPs:
- Derive identity server-side from authenticated credentials; do not accept scoping identifiers as caller-supplied tool parameters.
- Populate behavioral annotations (
readOnlyHint,destructiveHint) accurately, they cost little and benefit both human reviewers and agents. - Treat the tool list as documentation an adversary reads first; avoid leaking backend structure through parameter naming.
If you build agents:
- Treat tool descriptions,
instructions, and prompt/resource text as untrusted input, subject to the same scrutiny as any other retrieved document. - Pin or hash reviewed schemas and alert on drift between sessions.
If you're a security nerd like me:
- Triage before interacting: sort tools into public-catalog, tenant-scoped, and client-asserted-identity categories directly from the schema, and prioritize the last category.
Limitations
Our pilot study shows that object-addressing surface is computable from schemas alone, at scale, with zero interaction, and that the parameter vocabulary is rich enough to separate public catalog identifiers (slug, listing_slug) from tenant-scoped ones (owner_key, client_id, database_id), the latter being the sharper BOLA-precondition signal. It also shows that agent-facing text like instructions, resources, and prompts is exposed on most servers with no provenance mechanism attached.
What it doesn't show is that a tool taking file_id is a confirmed BOLA, since no tool was ever invoked and no downstream authorization check was tested. The near-identical parameter counts suggest the true count of independently designed exposures is lower than the raw tool-level figures imply, so we don't treat these as a clean prevalence estimate, and the results shouldn't be generalized to the broader MCP ecosystem.
Also, the identifier classifier's precision hasn't yet been measured (named_object matches are a heuristic upper bound pending a hand-labeled sample), and no clustering adjustment has been applied to correct for the shared-platform effect noted above. The inference classes, particularly identity model and effect class, still need validation against servers with known, controlled configurations to establish real precision and recall. And the sample itself is small relative to the discovery frame: 247 of 6,951 endpoints probed (3.6%), drawn only from public registries and public code or documentation per the study's pre-registration.
Zero tool calls later…
Whether an MCP server requires authentication is not the only question worth asking. Before a single tool is ever called, its schema already discloses its object model, hints at where authorization has been delegated to the caller, and hands an agent unsigned text it will treat as context. This pilot shows that reconnaissance against that disclosure is not theoretical, it is computable, passively, across real public infrastructure, at a scale that makes manual review impractical and automated triage necessary. If a schema can tell a stranger this much after you ship, imagine what it tells ai-surface before you do… ;)