Skip to main content

It's not a vulnerability until the money moves

· 7 min read
José Haro Peralta
APIsec Research Labs

Proof of exploitability, or it didn't happen​

One of the most pervasive problems in cybersecurity is false positives. Organisations receive dozens of bogus vulnerability reports about their websites every day, and security teams burn hours trudging through the noise that automated scanners generate. Worse, CISOs can't prioritise what actually matters, because a false positive is so often hard to tell apart from a real finding.

How did we go so wrong? The ingredient missing from a false positive is proof of exploitability, and a rigorous one at that. Proof of exploitability, also called chain-of-attack proof, is one of the most important elements of any vulnerability report. Without it, a penetration testing report isn't worth your money. Reliable bug bounty platforms like HackerOne place heavy emphasis on proving a vulnerability is exploitable; without that ingredient, a report simply doesn't have a case.

What counts as sufficient proof?​

The question is, what is proof of exploitability, and more specifically, what is sufficient proof of exploitability? Proof of exploitability is something that proves that the vulnerability observed in an application has a security implication. For example, if a website has a CORS misconfiguration with overtly permissive rules, it could be vulnerable to anything from data breaches to account impersonation. The key word here is could. The question is whether this risk is real and we can prove it with an attack chain that results in the predicted outcome.

Consider the case of a website that fails to parametrise a URL parameter that goes straight into a database query. Say, an e-commerce platform where you can filter and sort products through an endpoint like GET /products?filter=toaster&sort_by=price. If the sort_by parameter reaches the database without parametrisation, the application is clearly vulnerable to database injection. But how exploitable is it? If the parameter lands in a GROUP BY or ORDER BY clause and the database is properly configured, there is little room for full-scale exploitation.

That example highlights something important: proof of exploitability shows that a vulnerability has a security impact, and it also tells us how impactful the exploit is. Not all database injections are the same. An injection proven to let a threat actor mutate and leak data from our system is very different from one where the payload reaches the query but can't be used to cause any harm.

Where it gets hard: the business layer​

The examples so far are fairly straightforward. If CORS is misconfigured, you must prove it can be exploited. If a database isn't parametrising a value, you must prove how serious it is. Things get more subtle when we reach the business layer. For example, when the vulnerability we're testing mutates a resource in a way that produces undesirable side-effects. To illustrate why this raises the bar for proving exploitability, let's walk through a specific example.

Imagine an e-commerce application where you can place an order online and update it during the twenty-four hours period before it ships. In that window you can change the number of items, add or remove products, and change the delivery date. Suppose the request looks something like this:

PUT /orders/878a17fd

{
"items": [
{
"product_id": "1ad5ecf4",
"quantity": 1
}
],
"delivery_date": "2027-09-01"
}

The order ID is 878a17fd. On this request, we tell the API we want to update the quantity of product 1ad5ecf4 to one and set the delivery date to the 1st of September 2027. If the request succeeds, the API responds with a 200 status code and a full representation of the order, reflecting the new price:

{
"id": "878a17fd",
"items": [
{
"product_id": "1ad5ecf4",
"quantity": 1
}
],
"delivery_date": "2027-09-01",
"total_price": 207,
"status": "pending"
}

The application updates the order's total price based on the new selection of items and sets the status to pending. After this, it would take us to the payment step to confirm the change, at which point the status changes to processing. Now, say we're pen-testing this endpoint and want to see whether it's vulnerable to mass assignment. We attempt to override the order's status:

{
"items": [
{
"product_id": "1ad5ecf4",
"quantity": 1
}
],
"delivery_date": "2027-09-01",
"status": "returned"
}

The API processes the request successfully and responds with a 200 status code and the following payload:

{
"id": "878a17fd",
"items": [
{
"product_id": "1ad5ecf4",
"quantity": 1
}
],
"delivery_date": "2027-09-01",
"total_price": 207,
"status": "returned"
}

That's a successful response with the injected value reflected in the payload. At this point we could take it as evidence that the API is vulnerable to mass assignment, send the report and the invoice to our customer, and call it a day. But before you rush to send that invoice, slow down and answer one question: is this all we can do to prove the endpoint is vulnerable to mass assignment? Does this response prove we can manipulate the state of an order? Or are there other checks that would make our case stronger?

The answer is: we can, and we must. The response alone doesn't prove the order's status actually switched to returned. The real proof comes from the side-effect of manipulating the order state. Here, that means verifying we can claim a refund on the returned order. In other words, it's not a vulnerability until the money moves.

Reflection is not proof. But then, what is?​

Value reflection is weak proof of mass assignment and a common source of false-positive reports. Earlier research from APISec Labs demonstrates this (see The Ghost in the Shopping Cart by Bandana Kaur), and our experience running mass assignment tests on the APISec platform bears it out.

So what does it take to prove a mass assignment attack was effective? We have several options. Sometimes the response itself is the proof, not because the injected field is reflected back, but because the injection makes the API leak data that shouldn't be there. For example, in September 2024, MTN Group disclosed a report where a threat actor could bypass authentication and take over other user accounts by setting gateway to true (HackerOne #1709881).

More often, though, exploit verification means stepping out of the response context and confirming the side-effects through a flow. For example, in August 2024, Rocket.Chat disclosed a report where a threat actor could escalate privileges by overriding their role (HackerOne #501081); the proof of exploitability came from verifying that the user could then reach admin-only functions. At APISec, we call this flow-based exploit validation.

Proving exploitability sometimes requires exercising complex flows, like in the Rocket.Chat example. Often, though, it's simpler than it looks. In June 2022, Reddit disclosed a report where a threat actor could bypass ad-campaign verification and approval by setting admin_approval to APPROVED (HackerOne #1543159); here, verification was done by reading back the state of the campaign and watching it switch from PENDING to ACTIVE.

Why this breaks automated scanners, and how APISec solves it​

Why is this such a big deal? Because offering proof of exploitability via user-based flows requires a deep understanding of how the application works. For a human, this is fairly straightforward. You simply interact with the application a few times to understand how it works.

The challenge is for automated security scanners. Why? Because automated scanners struggle to build an understanding of the application under test, and without that understanding, it's impossible to infer the flows needed to prove an exploit. At APISec, we are solving this problem by building an application model and an intent model of the application under test. This allows us to understand how the application is meant to be used, and which flows can be used to achieve certain user goals, or to prove an exploit. If you haven't tried it yet, give it a go!

How about you? How do you prove that the vulnerabilities you find are exploitable?