SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Checklist 23 mins

Cross-Tenant Agent Security: A Negative-Test Matrix

A practical negative-test matrix for checking whether an AI agent can expose one customer’s data, memory, tools or cached responses to another.

The PADISO Team ·

A cross-tenant security test should try to make an agent cross a customer boundary, then show exactly where the attempt was stopped. A normal successful conversation does not test this. The important cases are mismatched identities, stale context, indirect tool access, cached outputs and concurrent requests that interleave.

This checklist is for teams operating a multi-customer agent or adding tenant separation to an existing one. It focuses on negative tests across retrieval, memory, tools and caches—not on cloud infrastructure isolation. Each item asks for observable evidence: the test input, the expected denial or scoped result, and the trace or record that proves what happened.

Treat a test as failed if the agent produces another tenant’s data, performs an action under the wrong tenant, or returns an output whose tenant scope cannot be established. A refusal is not enough evidence by itself: the test should also establish that the protected data or action was not reached through another path.

1. Define the boundary and the test identity

Start by defining what “tenant” means in the application. It may be an organization, account, workspace or customer environment. Record which identifier is authoritative, where it comes from, and which components are expected to use it. If two services use different identifiers, document the mapping and test it; a boundary that exists only in a diagram is not a usable test oracle.

  • Write down the tenant identity used for each request.

    Specify the source of identity, such as an authenticated session or a trusted service assertion, and identify the component that resolves it. Do not treat a tenant name supplied in free text as proof of identity. Create test accounts for at least two distinct tenants, and record the expected tenant ID for each. The test runner should be able to assert that the identity resolved by the application matches the account used to authenticate.

  • Separate tenant identity from user identity and role.

    Include the tenant ID, user ID and relevant role as distinct test fields. Verify that changing a user or role within one tenant does not silently change the tenant, and that changing the tenant requires an independently valid identity. This catches code paths where a user ID, email domain or conversational statement is incorrectly used as a substitute for an organization boundary.

  • Choose data that makes leakage unmistakable.

    Seed each test tenant with a unique marker that is not present in shared documentation, such as a distinctive project code and a fictional account name. Keep the markers synthetic. A pass requires that Tenant A cannot retrieve, summarize, quote or act on Tenant B’s marker through the tested route. Avoid relying only on a phrase such as “do not reveal this”; the test must inspect the actual result and relevant access trace.

  • Define the expected result for denial and for valid access.

    Decide whether an out-of-scope request should return a generic denial, an empty result or a controlled clarification. The exact user-facing wording may vary, but the protected resource must not be fetched or used to construct a response. For an in-scope request, assert that the correct tenant’s marker can be found. Testing both sides helps catch an implementation that appears secure only because it blocks all retrieval.

Record the expected tenant, user, role, resource and outcome for every case. This makes failures reproducible and prevents an ambiguous result—such as a generic answer that contains no marker—from being mistaken for proof that the tenant filter worked.

2. Check the request path before testing agent behavior

An agent may receive identity through several application components before it reaches retrieval or a tool. Tests need to verify the handoff, not just the final answer. A model may respond plausibly even when the request context was incomplete, while a correct denial may conceal a broken identity mapping that would fail on another route.

  • Test the tenant-binding step with missing and conflicting values.

    Submit a request with no resolved tenant, an unknown tenant ID, and a tenant ID that conflicts with the authenticated account. The expected result should be a controlled failure before protected work begins. Include a case where the conversation text asks the agent to “switch” to another organization; free-text instructions must not replace the application’s resolved identity.

  • Check that every agent hop receives the intended scope.

    For a multi-step request, record the tenant context at the initial request and at each server-side handoff to retrieval, memory or a tool. Test a path that invokes more than one component. The values should stay consistent, or an explicit, validated transition should be visible. If only the first hop has a tenant ID, later calls may operate with a broader default scope.

  • Test retries, continuation and resumed conversations.

    Start a request under Tenant A, force a controlled retry or continuation, and verify that the resumed work remains bound to Tenant A. Then attempt to continue the same conversation using Tenant B’s authenticated session. The system should either reject the mismatched continuation or create a distinct, correctly scoped conversation. A conversation identifier alone should not be treated as proof that its current caller is authorized for its contents.

  • Record the policy decision, not just the final answer.

    Capture a correlation ID, resolved tenant, component, resource scope and allow-or-deny result in the test trace. Exclude secrets and unnecessary customer content. The record should let an operator distinguish “retrieval returned no matches” from “retrieval was never permitted.” For a fuller discussion of trace design and agent behavior signals, see AI Agents in Production: Agent Observability.

A simple decision path helps make the test boundary explicit:

flowchart TD
    A["Incoming request"] --> B["Resolve trusted tenant"]
    B --> C["Check requested scope"]
    C --> D{"Scope matches?"}
    D -->|"Yes"| E["Run scoped operation"]
    D -->|"No"| F["Deny and record"]
    E --> G["Verify tenant-bound result"]

    accTitle: Tenant boundary test path
    accDescr: Resolve tenant identity before checking scope. Matching scope permits a scoped operation whose result is verified; mismatched scope is denied and recorded.

The yes branch is not a blanket authorization: it means the requested scope matches the identity and policy resolved for this operation. The operation still needs to constrain its own resource access. The no branch should stop the protected operation rather than merely instruct the model not to disclose the result. The final verification step checks that a technically successful operation did not return data outside the resolved tenant.

Retrieval tests should cover both the search request and the material that can influence the answer. A filter in one retrieval route does not prove that every search route applies the same constraint. Search indexes, generated queries, fallback retrieval and secondary tools can all create alternate paths to content.

  • Attempt direct retrieval of another tenant’s unique marker.

    Authenticate as Tenant A and ask for Tenant B’s marker by exact name, by a partial phrase and by a descriptive paraphrase. Repeat through each supported retrieval entry point used by the agent. The expected result is no access to Tenant B’s content, not merely a response that omits the marker after it has already been retrieved.

  • Test the same search with a deliberately mismatched tenant parameter.

    Where a search service accepts both an authenticated context and a tenant filter, provide conflicting values in a controlled test. The application should derive or validate the effective scope rather than allowing an untrusted parameter to widen access. Record which value governed the query and make a mismatch an explicit test failure.

  • Probe alternate retrieval paths and fallback behavior.

    Test keyword search, semantic search, reranking inputs, fallback indexes and any retrieval route invoked after an empty result. If the primary route correctly returns no Tenant B documents but a fallback searches a shared corpus without the tenant constraint, the overall request has crossed the boundary. Preserve the route name and resource identifiers in the trace so a failure can be localized.

  • Check generated queries and administrative endpoints separately.

    If a component generates or exposes queries, test the normal application path and any query console, API or administrative route that can access the same data. A filter attached to one dataset or interface does not establish that every route is filtered. For example, Apache Superset describes row-level security as applying clauses to generated queries for configured datasets and subjects. For this portal design, validate SQL Lab and API access separately; embedding does not itself establish authorization. Review the Superset security guidance.

  • Verify the content actually supplied to generation.

    Inspect the selected document IDs, tenant scope and relevant snippets passed into the generation step, subject to your logging and data-handling rules. The answer alone is an incomplete signal: an agent might omit a retrieved marker this time while still receiving another tenant’s material. A pass requires that the inputs were scoped correctly as well as that the response did not disclose out-of-scope content.

Use paired cases: an identical query that returns a Tenant A marker under Tenant A identity and does not return it under Tenant B identity. If both requests return nothing, the test has not proved isolation; it may only show that the fixture or search query is broken. If both return the same record, stop the release investigation and identify the earliest component that admitted it.

4. Negative-test conversation and durable memory

Memory can be stored as conversation history, summaries, user preferences, extracted facts or another application-managed record. Its risk is not limited to an explicit “remember this” feature. Any persistent context that influences later requests needs a defined owner and retrieval scope.

  • Test memory reads under a different tenant identity.

    Create a memory entry under Tenant A using a distinctive synthetic marker. Start a separate request under Tenant B and ask for the same fact directly, indirectly and through a request to summarize previous work. Tenant B should not receive the entry or a derived summary of it. Verify the memory lookup scope, not only the final wording.

  • Test writes that attempt to attach to another tenant.

    Under Tenant A, ask the agent to save information for Tenant B, or submit a write request with a Tenant B identifier that conflicts with the authenticated context. The write should be denied or safely confined to the caller’s authorized scope. Check the stored record after the attempt; an agent message claiming that it did not save anything is not proof that the write was rejected.

  • Test conversation reuse across accounts.

    Open a conversation under one tenant, then attempt to read, append to or resume it under another tenant. Confirm that ownership is checked at the point of access and that the conversation’s prior messages are not included in the second tenant’s prompt. Include both a direct conversation ID and a copied link or reference if your application supports them.

  • Test summaries and derived memory after source access changes.

    A summary can preserve a sensitive fact even after the original conversation is no longer used. Create a test summary from Tenant A’s marker and then exercise the ordinary read path as Tenant B. Also verify the expected behavior when a source record is removed, expires or is no longer within the current tenant scope. The goal is to find a derived copy that outlives the boundary enforced on its source.

  • Check that memory identifiers cannot act as authorization.

    Try a known synthetic memory key from the wrong tenant and an unguessable or nonexistent key. Both should remain inaccessible outside the correct scope. Keys can be useful for locating records, but a key supplied by a caller should not by itself grant access to the record it names.

Include a test for empty or partially available memory. If the store is unavailable, the application should not silently substitute a broader store or a shared default context. Define whether the request fails, proceeds without memory, or asks the user to retry, then assert that the behavior stays within the same tenant boundary.

5. Negative-test tools and delegated actions

A tool call can expose data through a read operation or cause a consequential write. Test both. The agent’s natural-language response is not the authorization boundary; the application component that executes the operation must validate the identity and scope attached to that specific call.

  • Attempt cross-tenant reads through every data tool.

    Use Tenant A credentials to request a record, report or object identified as belonging to Tenant B. Include IDs supplied in user text, IDs found in an agent-generated plan and IDs returned by another tool. The operation must reject or scope the request before returning protected content. Record the tool, resource identifier and effective tenant without recording unnecessary payload data.

  • Attempt a cross-tenant write with a harmless test payload.

    In a non-production environment, try to create or modify a synthetic record under another tenant. If the tool supports a preview or validation-only route, test that route and the execution route separately. A safe response from the model does not pass the test if the underlying action was performed. Where no harmless test operation exists, do not execute a destructive probe; use a controlled simulation or a representative test double.

  • Test arguments that try to alter the authorization context.

    Provide a valid resource request alongside a conflicting tenant, account or owner field. Check that the server-side executor uses the authenticated scope and rejects incompatible arguments instead of trusting a model-generated parameter. Also test missing scope. A broad implicit default should be treated as a failure condition unless the operation is explicitly designed and authorized to be cross-tenant.

  • Verify that tool output cannot authorize a follow-on action.

    Make a test tool return content that requests another tenant’s data or tells the agent to use a broader identity. The agent should treat tool output as data, not as proof that a new scope is authorized. MCP security guidance identifies token handling and confused-deputy threats as reasons to establish explicit authorization boundaries; tool output should not be trusted automatically. Read the MCP security best practices.

  • Separate permission to request an action from permission to execute it.

    For actions with external effects, test a request that is denied, one that is valid, and one whose payload changes between review and execution. Any required approval should be bound to the exact proposed payload and expire; a general approval for an earlier or broader request must not authorize a changed action. Then verify the business result independently of the model’s claim that it succeeded.

A useful test trace follows the call from intent to outcome: resolved tenant, authorization decision, exact resource scope, tool invocation, tool result status and any external effect. This is not an invitation to log secrets or entire customer records. Capture the minimum fields needed to explain whether the correct operation ran under the correct scope.

6. Test cache boundaries and stale responses

Caches can return a result without repeating the retrieval or tool checks that originally produced it. For testing, treat every cache layer as a distinct data path: response caches, retrieval caches, prompt or context caches managed by the application, and cached tool results may have different keys and lifetimes.

  • Repeat an identical request under two tenant identities.

    Warm the cache with a request from Tenant A, then issue the same request text as Tenant B. The second response must be generated or retrieved only from Tenant B’s authorized scope. Use a query likely to produce a tenant-specific answer, and check for the synthetic marker as well as the cache-hit trace. A shared cache key based solely on normalized question text is an important failure to investigate.

  • Reverse the order of the tenant requests.

    Run the same test with Tenant B first and Tenant A second. This catches one-way leakage where a cache entry created under one tenant is available to another, but a different tenant’s entry behaves differently because of key composition or cache population order. Repeat in a fresh cache namespace where feasible so stale test results do not mask the outcome.

  • Exercise near-duplicate prompts and changed identities.

    Test exact matches, small wording changes and a request that reuses a conversation or request reference under a different account. Confirm that the cache’s lookup scope is appropriate for the data returned. A cache may safely share a truly tenant-neutral response in some designs, but the test must establish that the result contains no tenant-specific data and cannot inherit tenant-specific context.

  • Test invalidation and tenant membership changes.

    Warm a cache while the test user has access to a synthetic record, then change the test fixture’s permitted scope or remove the record according to the application’s normal process. Check that a later request does not return a stale result merely because the cache has not been refreshed. Define the expected maximum staleness for this application rather than assuming every cache invalidates immediately.

  • Observe cache hits without exposing cached content.

    The test record should identify whether a cache was hit, which scope was used to address it and which tenant context was resolved. Do not emit full cached responses into general-purpose logs just to prove a test. A useful audit signal establishes that the lookup was scoped and whether the returned object passed the same tenant checks as a fresh result.

When a cache miss is expected but a hit occurs, do not stop at the cache key. Trace which content populated the entry, whether the entry was tenant-neutral or tenant-specific, and whether any downstream scope check ran. A correctly partitioned key reduces accidental reuse, but it does not replace validation of the cached object and its authorization context.

7. Exercise concurrency, retries and partial failures

Single-request tests can pass while concurrent work still mixes state. Race conditions are especially relevant when a service reuses mutable request context, conversation state or a shared cache-fill operation. Design tests to overlap requests, not merely to send them one after another.

  • Overlap requests from two tenants.

    Start a Tenant A request that retrieves its marker and a Tenant B request that retrieves a different marker at nearly the same time. Repeat with different response delays so one request completes while the other is still waiting on retrieval or a tool. Each final response and each downstream call must remain associated with its own tenant. If timing is uncontrolled, use test barriers or a stubbed dependency to make the interleaving reproducible.

  • Force a retry after a dependency timeout.

    Simulate a timeout at retrieval or a tool boundary, then allow a retry. Verify that the retried call uses the original resolved tenant and does not inherit context from another in-flight request. Check both the retry’s arguments and the final result. A request that happens to return a safe answer after a timeout is not a pass if the retry ran without an explicit scope.

  • Test partial success across a multi-tool sequence.

    Have an early operation succeed and a later operation fail. Confirm that the failure does not cause the agent to repeat the earlier action under a broader scope, switch to an unfiltered fallback or disclose partial data from another tenant. If an external write may already have happened, record and verify the resulting state rather than assuming a retry can safely repeat the operation.

  • Verify context cleanup after completion and cancellation.

    Cancel or terminate a request while a dependency is still pending, then immediately submit a request under another tenant. Ensure that abandoned work cannot append its result to the new response or populate an incorrectly scoped cache entry. Test cleanup after normal completion too; leaked state does not require an error to become a security defect.

Run concurrency tests repeatedly with controlled scheduling where possible, and retain the seed or event sequence that reproduces a failure. A single successful run is weak evidence against an intermittent race. If the test cannot reliably expose the relevant interleaving, document that limitation instead of treating the absence of a failure as proof.

8. Make failures diagnosable and recovery testable

A negative test has operational value when the team can tell what failed and contain the affected path. This section is about evidence for the cross-tenant test—not a replacement for a full incident process. Keep the response proportionate: investigate the component that admitted the wrong scope, prevent repeated exposure and verify the correction with the same test case.

  • Preserve a privacy-conscious failure trace.

    For each case, retain a test ID, request correlation ID, authenticated tenant, resolved scope, component, decision and outcome. Store synthetic markers rather than real customer secrets. The trace should answer whether the leak began in identity resolution, retrieval, memory, tool authorization, cache lookup or response assembly. If your current traces cannot distinguish these paths, add that observability before relying on a large test suite.

  • Give each failure a reproducible fixture and expected result.

    Save the tenant setup, synthetic records, request shape and expected denial or scoped result. Record the application version and configuration relevant to reproducing it. Keep the fixture small enough for engineers to understand what crossed the boundary; a test with a large dataset and no clear expected marker is hard to diagnose and easy to misinterpret.

  • Verify containment before declaring the test repaired.

    After correcting a failure, rerun the exact failing case, the paired valid-access case and at least one adjacent route using the same data source. Check that the unauthorized content was not also copied into memory or cached output during the failed attempt. If a test exposed a path that could have been reachable in production, use your organization’s established incident process to decide what investigation is needed. A separate agent incident runbook covers incident handling in greater depth.

  • Set a release gate that reflects the tested boundary.

    Identify which cases block release: cross-tenant retrieval, memory access, tool effects, tenant-specific cache reuse and reproducible concurrency leakage should not be silently downgraded to warnings. Assign an owner and a status to each failed case. A reliability objective should measure completed, correctly scoped work rather than just successful responses; see AI Agent SLOs: Define Reliability in Terms of Completed Work for that distinct operational question.

Use your existing response process if a test indicates real exposure. This article does not prescribe legal reporting decisions or establish regulatory requirements. The immediate engineering questions are narrower: which tenant boundary failed, what path returned or acted on the data, what must be disabled or corrected, and what test demonstrates that the path is contained.

9. Worked hypothetical: a cached retrieval crosses accounts

Consider a hypothetical support agent used by two fictional customers, Northstar and Bluehaven. Each has a synthetic project marker in its internal knowledge base. A support engineer asks the agent to summarize Northstar’s marker while authenticated to Northstar. The answer is correct. A second tester, authenticated to Bluehaven, submits the same normalized question. The agent returns Northstar’s marker because the response cache uses the normalized question as its lookup key and does not include a tenant scope.

The initial negative test fails even if the retrieval service itself applies a correct tenant filter. The second request receives a cached response and may never call retrieval. This distinction matters: changing the retrieval filter alone would leave the demonstrated path open. The test record should show the two identities, the shared request shape, the cache-hit status, the source tenant associated with the entry and the returned marker. Use synthetic data and a controlled test environment so no real customer content is involved.

A recovery trace for this hypothetical case can be represented as a sequence of acceptance checks. First, reproduce the failure with a fresh cache and capture the failing cache lookup. Next, change the proposed design so tenant-specific entries are partitioned by the resolved tenant and the returned object is checked against the current scope. Then rerun both request orders, an exact prompt and a near-duplicate prompt, and confirm that Bluehaven receives neither the Northstar marker nor a derived summary. Finally, verify Northstar still receives its own valid result after the cache change.

The pass condition is not simply “the second answer says access denied.” It is that no cross-tenant cache entry was served, the trace shows the correct tenant scope at lookup, and no alternate retrieval or memory route supplied the marker. If the agent returns a generic answer, inspect the cache and downstream inputs to establish that the protected content was not present in the request context. If the cache is bypassed, the case has not tested the suspected failure path and should be rerun with the cache behavior confirmed.

A tempting counterexample is to disable all response caching and declare the boundary fixed. That could remove this one route, but it does not prove that retrieval, memory or tool calls are scoped, and it may discard a useful application behavior unnecessarily. Another weak fix is to add the tenant name to the user’s prompt. A prompt value is not a trusted identity source and does not establish which scope the cache or data store enforced. The test should validate the application’s resolved tenant, not ask the model to police its own context.

10. Turn results into a repeatable release decision

A test matrix is useful only if it stays connected to the actual routes and changes that can affect tenant boundaries. Keep a compact case inventory in version control or in the test system your team already uses. Tie each case to the component it covers and the release gate it informs. Do not let a single “agent security passed” label hide that only retrieval was tested while memory, tools or caching were skipped.

  • Map each test to its entry point and protected component.

    For every test, record whether it exercises the public request path, a direct service interface, a background continuation, a tool executor or a cache. This helps reviewers identify uncovered alternate routes. Mark tests that depend on a particular feature or deployment configuration so teams do not mistake an unrun case for a passing case.

  • Rerun relevant cases when scope-affecting behavior changes.

    Changes to identity mapping, retrieval queries, memory handling, tool arguments, cache keys, retries or conversation continuation can alter the boundary. Include the relevant negative tests in the change’s verification plan. Token limits and multi-hop execution can also change which components are called or retried; when that affects this path, coordinate these tests with Token Budget Management Across Agent Hops, rather than treating token management as a substitute for authorization.

  • Review failures and skipped cases explicitly.

    A skipped test should have a named reason, owner and next action. A fixture failure, unavailable dependency or missing trace does not count as a pass. For a broader production-readiness decision beyond cross-tenant boundaries, use the separate CTO decision checklist as a related planning resource, not as evidence that this matrix has passed.

  • Assign engineering ownership for the paths that lack reliable tests.

    If the matrix exposes a gap across identity propagation, retrieval, cache partitioning or tool enforcement, assign an owner to close it and a reviewer to confirm the negative and valid-access cases. Teams that need support designing testable boundaries across application components can explore production platform engineering.

A proposed release rule is straightforward: block release on any confirmed cross-tenant read, write, memory exposure or cached response; block on a missing or contradictory tenant identity where protected work can still proceed; and require a documented owner and retest plan for any untested route that can access tenant data. Teams may add their own severity thresholds, but should not reclassify an unexplained cross-tenant result as acceptable noise.

Printable worksheet: cross-tenant negative-test matrix

Use this compact worksheet to plan a run and retain its outcome. It is formatted for printing or copying into a test ticket; it is not a separate downloadable file. Add rows for every route your agent can use, including background and fallback paths.

Test areaCaller and attempted scopeExpected resultEvidence to retainStatus
IdentityTenant A; missing, unknown or conflicting tenantProtected operation stopsResolved identity and decisionPass / fail / not run
RetrievalTenant A requests Tenant B markerNo foreign document or snippet is suppliedRoute, scope, document IDsPass / fail / not run
MemoryTenant B reads or resumes Tenant A contextNo foreign memory or derived summaryMemory scope and conversation ownerPass / fail / not run
Tool read/writeTenant A requests a Tenant B resource or effectDenied before access or executionTool, resource, authorization resultPass / fail / not run
CacheSame tenant-specific prompt under two identitiesOnly the current tenant’s result is servedCache hit, scope and source tenantPass / fail / not run
ConcurrencyOverlapping requests from Tenant A and BResults and calls remain separately scopedCorrelation IDs and event sequencePass / fail / not run
RecoveryRe-run a confirmed failure after correctionOriginal failure no longer reproduces; valid access still worksBefore-and-after test recordsPass / fail / not run

For each row, fill in the application version, test date, fixture identifiers, test owner and reviewer. Record whether the test passed, failed or was not run; never convert “not run” into an implicit pass. For a failure, attach a concise reproduction sequence, identify the earliest component with an incorrect scope, and name the retest that will close it.

Prioritize execution by exposure path, not by how easy a test is to run. First verify identity binding and retrieval, then test memory and tool execution, and include cache and concurrency cases wherever those features can carry tenant-specific state. Before release, rerun the failed case and the corresponding valid-access case. That pairing demonstrates both that the boundary blocks cross-customer access and that the intended customer workflow remains available.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call