PXL Security LTD, Sofia, Bulgaria Offensive security since 2014[email protected]
APIAccess ControlCloud Security

Your tenant boundary is an HTTP header

By PXL Security3 October 202621 min read

In most composed SaaS platforms, the thing keeping one customer's data away from another's is not a boundary at all. It is a string in an HTTP header, written by a caller nobody on the receiving side has authenticated. The partition key began as a property of a verified session and, somewhere between the browser and the fourth internal service, became an argument.

How the partition key escapes

Start at the only place where tenancy is genuinely known. A user signs in. The identity provider issues a token bound to a subject, and the subject belongs to exactly one tenant — or to a small, enumerable set with explicit grants. At that instant the tenant identifier is derived: a fact the platform established, not a value anyone supplied.

Then the request begins its journey. The browser talks to a backend-for-frontend, because single-page applications need an aggregating layer that speaks their shape rather than the shape of a dozen downstream domains. The BFF unpacks the session and fans out to internal services — ordering, pricing, inventory, documents. Those services are themselves composed; they call each other.

Each hop needs to know which tenant it is working for. The honest answer is that the tenant should be re-derived at every hop from a credential that hop can verify — a forwarded assertion, a service-to-service token with the tenant claim inside it, something signed. The practical answer, shipped under deadline, is that the first hop reads the tenant from the session and passes it along as a header, because headers are free, because the downstream team asked for it in that form, and because at the moment the code was written the only caller was the BFF.

Hop 1  browser  ->  BFF
       Authorization: Bearer <session token>
       (no tenant header; tenant is derived from the token)

Hop 2  BFF  ->  order-sourcing service
       X-Tenant-Id:     1001        <- read out of the session, now a string
       X-User-Id:       4471
       X-Tenant-Schema: tenant_1001

Hop 3  order-sourcing  ->  inventory / pricing
       X-Tenant-Id:     1001        <- copied verbatim from hop 2
       X-Tenant-Schema: tenant_1001

Notice what changed between hop 1 and hop 2. At hop 1 the tenant is a conclusion; at hop 2 it is a parameter. The BFF knows it derived the value properly; the order-sourcing service knows only that a value arrived. By hop 3 the provenance is gone entirely — the value has been copied twice and the signed artefact it came from was never forwarded.

The failure is what happens next, and it is almost always operational rather than a code change. The internal service becomes reachable: the private ingress gets an external route for a partner integration, a mobile client is pointed at it to shave a hop, or the gateway's header-scrubbing rules were written for the two headers that existed at the time. Now the client supplies X-Tenant-Id itself, and the service does what it has always done: it believes it.

Nobody decided to trust the client. The BFF team decided to forward a value it had validated. The service team decided to accept a value from a trusted caller. The platform team decided which paths were internal. Each decision was defensible in isolation, and the trust accumulated across them until it reached the browser.

Note. A tenant-schema header is a particularly sharp version of this: it usually reaches a connection factory or a SET search_path rather than a WHERE clause, so it stops selecting rows and starts selecting the database namespace the query runs in.

Nobody owns the check

Ask a platform team who enforces tenant isolation and you will get a confident answer from everyone, and two different answers in total.

The gateway team's model: the gateway authenticates the caller, terminates the session, and hands downstream a clean, normalised request. Authorisation of business objects belongs to the service that owns them, because only it knows what the object is. The gateway is not going to pretend to know whether order 88120 is yours.

The service team's model: anything arriving on the internal network has already passed the gateway. The gateway is the policy enforcement point — that is the point of having one. The service's job is to serve the request it was given, applying the per-object rules it owns.

Both models are coherent. Together they leave the tenant check in the gap, and the gap is well camouflaged, because the header looks authoritative. It arrived over mTLS from a known internal CIDR, with a service identity, carrying a well-formed value that matches the logs of the hop above. Everything about its presentation says "this was computed by someone who had the right to compute it". Nothing about it is evidence of that.

There is a subtler asymmetry underneath. Per-object authorisation is now a solved habit on most mature teams. Engineers have internalised that GET /orders/88120 must verify that this subject may read that order; the lesson of a decade of IDOR findings landed. Per-tenant authorisation did not land the same way, because it does not feel like authorisation at all. It feels like routing — pointing the query at the right partition, the right schema, the right shard. Routing is infrastructure, infrastructure is somebody else's layer, and so the one check that governs the largest blast radius in the system is the only one no team has written down as theirs.

That is also why the per-object check often survives a cross-tenant attack intact. If the service scopes its query by the supplied tenant and then checks that the object belongs to that tenant, the two checks agree perfectly and the request is authorised. The object really does belong to tenant 1002. You simply had no right to be acting as tenant 1002 — the question nobody asked.

Reading another tenant's data

An abstracted walkthrough from an order-sourcing service, with illustrative values. The service accepted the tenant identifier, the acting user identifier and the tenant schema name as request headers, with no binding whatsoever to the authenticated session that accompanied them.

The baseline, as the legitimate client sends it:

POST /sourcing/v2/orders/search HTTP/1.1
Host: orders.internal.example.test
Authorization: Bearer <valid session for a user in tenant 1001>
X-Tenant-Id: 1001
X-User-Id: 4471
X-Tenant-Schema: tenant_1001
Content-Type: application/json

{"status":"OPEN","limit":25}

HTTP/1.1 200 OK    -> 25 orders, all belonging to tenant 1001

Now change one header and nothing else. Same session token, same user, same everything:

X-Tenant-Id: 1002
X-User-Id: 4471
X-Tenant-Schema: tenant_1002

HTTP/1.1 200 OK    -> orders belonging to tenant 1002

Two hundred, with another customer's records in the body. The session was never consulted on the question of tenancy. The headers were treated as the answer.

The diagnostically important part came next, and it is the part most testers skip because the finding already looks finished. We sent tenant identifiers that could not possibly resolve:

X-Tenant-Id: 999999        (no such tenant)   -> 200, empty result set
X-Tenant-Id: not-a-number  (non-numeric)      -> 200, empty result set
X-Tenant-Id: 1002'         (quote appended)   -> 200, empty result set
X-Tenant-Id: (header absent)                  -> 200, empty result set

Read that carefully, because it settles the root cause rather than demonstrating the symptom. A service that validates a tenant identifier must look it up — to confirm it exists and that the session's subject is entitled to it — and a failed lookup produces a 403, 404 or 400. What we got instead was a clean 200 with an empty body: the signature of a value inserted straight into a query or schema selector and never checked against anything. The absent-header 200 is the same evidence from the other direction — no required-field assertion either.

So the distinction worth writing into the report is this. An empty result for a nonsense tenant is not a sign that validation held; it is proof that no validation ran. Had the control existed and merely failed to stop you, nonsense input would have hit it and been rejected. Nonsense input sailing through means the code path has no tenant authority check in it at all — so every tenant identifier in the estate is reachable, not just the one you happened to guess.

How to harden this

  • Derive tenancy at every hop from something that hop can verify cryptographically — a forwarded signed assertion or a service-to-service token carrying the tenant claim. Never from a bare header.
  • If a tenant header must exist for compatibility, treat it as an assertion to be reconciled, not an input: compare it with the tenant derived from the credential and reject on mismatch, loudly. Strip and re-stamp these headers at the trust boundary from a strict allow-list of names; deny-lists age badly.
  • Never let a client-influenced string reach a schema name, search path, connection selector or shard key. Map an identifier you derived to a partition through a server-side lookup.
  • Make an unresolvable tenant identifier a hard error, not an empty set. Empty results hide the absence of the check from your own engineers as effectively as from your auditors.
  • Enforce the partition below the application — row-level security, per-tenant credentials, or a data-access layer that cannot emit an unscoped query — so a missed check in one handler is not the only thing between tenants.

The write that was only stopped by a typo check

This is the most important lesson in the article, so here it is first, plainly: on the same service, a cross-tenant write passed authorisation completely. It failed on request-body validation. It was initially read as "blocked".

The attempt carried the forged tenant identifier, exactly as the read did, and a payload that was deliberately imperfect — a field in the wrong format, say a malformed date or an out-of-range enumeration:

POST /sourcing/v2/orders HTTP/1.1
Authorization: Bearer <valid session for a user in tenant 1001>
X-Tenant-Id: 1002
X-User-Id: 4471
X-Tenant-Schema: tenant_1002
Content-Type: application/json

{"sku":"AX-9","qty":5,"requestedDate":"31/02/2026"}

HTTP/1.1 422 Unprocessable Entity
{"errors":[{"field":"requestedDate","message":"invalid date format"}]}

A triager glancing at that transcript sees a non-2xx status and a rejection, writes "cross-tenant write blocked", and downgrades the finding. That reading is wrong, and seeing why is the single most useful thing a tester can carry into a tenancy engagement.

A request rejected by body validation is a request whose authorisation check already succeeded. Body validation is downstream of authorisation in every sane pipeline: a service decides whether you may act before it decides whether your payload is well-formed. For the server to have an opinion about the shape of requestedDate, it had already accepted that this caller, acting as tenant 1002, was permitted to create an order there. The 422 is not the boundary stopping you; it is the service telling you how to fix the one line between you and a persisted cross-tenant write. Correct the date and the record lands in another customer's tenant.

So how do you tell a genuine authorisation denial from a validation rejection that only looks like one? Read the response for what it is reasoning about, not merely its status code:

  • What is the error about? A body-field error — a format, a length, a missing property, an enum value — is proof the authority gate was already passed. A real denial talks about the actor, the tenant or the resource, not about your JSON. In status terms, 400/422 sit after the authority gate; 401/403 sit on it.
  • Does a well-formed body change the verdict? Fix the field and resend. If the status moves from 422 to 201, authorisation was never the thing stopping you. If a corrected body is still refused with the same actor-or-tenant reason, that is a real control.
  • Is the rejection order-invariant? Send a body that is both cross-tenant and malformed. A service that checks authority first returns the authority error; one that returns the format error is telling you authority was not checked first, or at all.
  • Did anything persist? A 4xx with a side effect — a row written, a counter moved, an event emitted — is an authorisation failure wearing a validation costume. Check the read path afterwards from the victim tenant's side.

Framed that way, the write finding is not a weaker cousin of the read finding; it is the more severe one. A cross-tenant read exposes data; a cross-tenant write corrupts it, and it did so here with the authorisation layer fully bypassed and only a date parser between the attacker and persistence.

How to harden this

  • Order the pipeline so authority is decided before the body is ever parsed, so a format error can never be the only thing that stopped a cross-tenant mutation.
  • Write test cases that assert the order of failure, not just that a request failed: a cross-tenant, malformed request must return the authorisation error, and the test must fail if it returns the validation error instead.
  • Give triage an explicit rule: a 4xx on a cross-tenant attempt is "blocked" only when the error concerns the actor, tenant or resource. A body-shape error is an authorisation bypass with a validation backstop — not blocked.
  • Re-test every cross-tenant write with a corrected, well-formed payload before closing the finding. A clean 422 in the first transcript proves nothing about what a clean body would have done.

The identity the token never claimed

The tenant header is the loud version of this problem. The quiet version is a user identifier in a request body.

A checkout flow: a front-end service, an order service, a payment component behind that. The call that created the order took the buyer as a parameter. Not derived — supplied:

POST /api/checkout/order HTTP/1.1
Authorization: Bearer <token for user 7781, tenant 1001>
Content-Type: application/json

{
  "userId": "7782",
  "cartId": "c-9f12",
  "paymentMethodId": "pm-4411"
}

Nothing compared userId to the token subject. The token was validated — signature, expiry, issuer — then set aside. The authenticated principal decided whether the call was allowed; the body decided whose order it was. Two notions of identity in one handler, never reconciled.

The request failed. Three hops later the payment component looked up pm-4411, found it belonged to 7781, found the order belonged to 7782, and rejected the mismatch. That lands as a medium, and the team read it as reassurance. But the flow has no authorisation control; it has a referential integrity check in a component solving a different problem. Payment refused because a payment method and an order disagreed about their owner, not because a caller asserted an identity it had no right to. Identical from outside; completely different in blast radius.

Now move that downstream check. Payment methods become shareable across a corporate account, and ownership stops being one-to-one. A new order type ships carrying no payment method, so the rescuing comparison has nothing to compare. Payment is replaced by an integration that trusts the order service's assertion of the buyer, because that service is internal. Each is a routine roadmap item, and each turns a medium into account takeover by proxy.

A control that lives in one component and protects a vulnerability in another is not a control, it is a coincidence with good timing. It has no owner, appears in no threat model, and the commit that removes it passes review — in its own file, it is dead weight.

How to harden this

  • Treat any field naming a user, account or owner as untrusted until the receiving service has compared it with the token subject. Where it is redundant, delete it from the contract.
  • Where acting-on-behalf-of is real, model it as an explicit scope and log every use.
  • Put the regression test at the layer that owns the check, not the layer that currently happens to fail.
  • Of every rejection, ask which component refused and whether it was refusing its job. If not, you have found an unowned control.

Existence is information

The same programme turned up a management API that was, functionally, behaving correctly. A resource belonging to another tenant returned 403; one that existed nowhere returned 404; malformed input returned 400.

GET /api/admin/accounts/1001-A4F2   -> 403  (exists, not yours)
GET /api/admin/accounts/1001-A4F3   -> 404  (does not exist)
GET /api/admin/accounts/not-an-id   -> 400  (malformed)

Three distinct answers to a question the caller has no right to ask. The 403 is the leak: it confirms the identifier is real while declining to serve it. An identifier space meant to be unguessable now has a membership test, and a membership test is all you need to turn it into an enumerated list.

The economics are underrated. These identifiers are rarely uniform random; they are structured — a tenant prefix, a type code, a sequence, a short random tail — and the structure is visible from two or three legitimate examples pulled from your own tenant. That collapses the space from astronomical to tractable, and the oracle answers at the speed of the endpoint: a few hundred requests a second against a few million candidates is an afternoon. Nor is there usually a rate limit, because a denial does not look like abuse.

What comes out is not a vulnerability. It is a target list: real identifiers, grouped by prefix into real tenants, ready to feed every other endpoint in the estate. This is the first move of the attack described earlier, not a separate low-severity curiosity — a cross-tenant read needs an identifier belonging to another tenant, and the oracle is where it comes from. Leave both and you have shipped the reconnaissance phase as a feature.

How to harden this

  • Collapse "exists but denied" and "does not exist" into one response: identical status, body and headers. Which status matters far less than picking one.
  • Scope the lookup by tenant in the query, so an out-of-tenant identifier is genuinely not found — identical by construction rather than by discipline.
  • Watch timing: a scoped miss in two milliseconds against a denial in forty has moved the oracle into the clock.
  • Alert on denial volume per principal. Building a target list produces a near-perfect stream of rejections.

Authorisation has to happen where the data is

All of this is one architectural mistake in different clothes: a service accepted a caller's claim about scope, then fetched data with it. The rule that removes the class is short. Every hop that can reach tenant-scoped data must derive the tenant from the authenticated principal, and nothing else. Not a header, not a body field, not a path segment, not a value a trusted neighbour passed along.

The partition key becomes a derived value, never an input: the validated token carries the tenant claim and the data layer takes the key from that claim alone. The strongest form makes an unscoped query structurally impossible — a tenant-scoped repository, row-level security driven by a session variable, a builder that refuses an unscoped predicate. Controls relying on each developer remembering a WHERE clause fail at the rate developers forget things.

Between hops, propagate context that cannot be forged: the original token forwarded intact, or a short-lived, audience-restricted internal token minted at the edge and validated downstream. A forwarded X-Tenant-Id is a claim with no provenance — the receiver cannot tell whether a gateway derived it from a token, whether middleware defaulted it, or whether the client sent it and nobody stripped it. A signed assertion answers that locally, without trusting the topology.

Which brings us to the argument that ends most of these discussions: that endpoint is internal. Internal describes network reachability; authorisation describes what a principal may do. One has never implied the other. The internal network carries your other tenants' traffic, your service accounts, every library you import, and whatever arrives through the first SSRF, leaked credential or misconfigured ingress. A boundary that collapses when any of those happens is not a boundary, it is a schedule. In a backend-for-frontend architecture it is often not even true: the "internal" header proves reachable from the browser, because the layer meant to strip it was also the layer meant to set it.

How to harden this

  • Derive the partition key from the validated token at every hop. Strip client-supplied tenant and user identifiers at the edge — strip, not override, so a stray value fails loudly rather than losing a precedence contest.
  • Make unscoped access unavailable rather than discouraged: row-level security, a scoped repository layer, a build-failing lint rule.
  • Authenticate and authorise internal service endpoints. Mutual TLS is transport identity, not permission.
  • Record which component authorises which object class, and make sure no row says "the one before it".
  • Log the tenant the request asserted beside the one the service derived. Divergence is an attack or a bug.

How to test multi-tenancy properly

The result here depends almost entirely on the setup. One tenant with one role finds none of this, however good the tester, because there is no second tenant whose data could leak.

Provision two tenants and at least two roles in each — an administrator and an ordinary user at minimum, plus any role with unusual reach. Populate both with realistic data of every object class and record the identifiers. Four principals is the floor, because you must separate three failures that look alike: crossing a tenant boundary, climbing within a tenant, and a permission simply wrong for the role.

Build the matrix before touching the application: object classes down one axis, principals across the other, what should happen in each cell — reads separately from writes, because they diverge constantly. The usual pattern is a list endpoint correctly filtered by tenant while the update endpoint on the same object is not, because filtering happened in the query and write authorisation was assumed to follow.

object: order 1001-A4F2 (tenant 1001)

principal              GET    PATCH   DELETE
t1001 admin            allow  allow   allow
t1001 user (owner)     allow  allow   deny
t1001 user (other)     deny   deny    deny
t1002 admin            deny   deny    deny
t1002 user             deny   deny    deny

Attempt every cell, not a sample. Skipped cells are the ones teams consider obviously safe, and obviously safe is where the bugs live. The tenant 1002 administrator is the most productive principal here: their own tenant grants them so much that a missing filter is invisible from inside.

Then replay: take a request that succeeds for tenant 1001 and reissue it with 1001's credentials and 1002's identifiers, at every hop you can reach — gateway, backend-for-frontend, the service endpoint behind it, the webhook, the export job, the GraphQL resolver. Each is a separate decision, often written by a different team. Vary one thing at a time — identifier in the path, identifier in the body, tenant header — then together, because precedence between them is frequently undefined and occasionally favours the attacker. Do all of it against the API, not the interface: the UI builds the tenant header for you, hides the object you want to name, and enforces client-side rules that exist nowhere on the server.

Finally, the discipline that separates a useful result from a misleading one: distinguish a rejection on authorisation from a rejection on validation. A 400 from a schema check, a 409 from records disagreeing about ownership, a 500 from something downstream choking — none is an access control decision, and logging them as "blocked" is how a critical finding becomes a clean report. When a cross-tenant attempt fails, establish which component refused and on what grounds, then satisfy the validation and try again. The cross-tenant write described earlier was found that way.

Note. Keep the matrix after the test. It is the regression suite, the brief for the next assessment, and the only honest answer to "has this been tested?" — a question that otherwise gets answered with a feeling.

What to insist on from a vendor

If you buy composed SaaS rather than build it, you cannot read the code. Three questions still give diagnostic answers.

How is the tenant key derived, and where is it enforced? The answer you want names the authenticated token as the source and names the enforcement point, ideally at the data layer and by construction. An answer describing a gateway that populates a header, or using the word "context" without saying where that context is validated, describes the architecture in this article. Follow up: which services can reach tenant data without validating a token themselves?

What evidence is there of cross-tenant testing between real tenants? Not a line saying authorisation was reviewed. Ask whether two tenants were provisioned, which object classes were covered, whether reads and writes were tested separately, and whether the tests run in CI. A vendor doing this has the matrix and will discuss it; one who does not will reach for certifications, which attest that a process exists — not that tenant 1001 cannot read tenant 1002's orders.

Treat "it's internal" as the start of a conversation. It is a claim with a testable shape: internal to what, enforced by what, and what happens when something inside that perimeter is compromised? Ask whether internal service calls are authenticated, whether client-supplied tenant headers are stripped at the edge, and whether anyone has tested those endpoints directly rather than through the front door. If the honest answer is that the network is the control, you know the shape of the breach in advance — and can price it, scope it into your own testing, or require it fixed before you sign.

Can one tenant reach another's data?

An API-focused penetration test answers that concretely — every boundary, both directions, reads and writes tested separately.

Scope an API test