Prerequisites
Before designing tenant isolation, write down what “tenant” means in your product. It may be a customer organization, a business unit, a regulated subsidiary or a workspace belonging to a customer. The boundary must be stable enough to identify every request, stored object and consequential action. If a user can belong to several tenants, the active tenant must be selected and validated for each operation; a user identity alone is not a tenant boundary.
This guide assumes an agent application hosted on Azure that serves more than one customer or organizational tenant. It addresses separation of data, conversational state and execution. It does not prescribe a particular Foundry deployment, MCP server design, observability stack, quota policy or compliance framework. For platform selection and broader architecture prerequisites, see Microsoft Foundry, Bedrock AgentCore, or Gemini Enterprise: Choosing an Enterprise Agent Platform.
Gather a list of data stores, state stores, tools and external systems that the agent can reach. Include indirect paths: a tool that calls another service, a retrieval index populated by a batch process, and a queue consumer are all part of the tenant boundary. Identify who owns each resource and who can change its access rules. If the inventory is incomplete, begin with the paths already used by a real workflow rather than designing an abstract platform in isolation.
Agree on a small set of test tenants and synthetic records. Each record should carry an unmistakable tenant marker, such as tenant-red or tenant-blue, and each tenant should have at least one resource or action that the other must not see or perform. Test accounts should have realistic role differences, but should not contain live customer information.
Pro tip: Treat tenant identity as validated request context, not as a value the model may infer from conversation text. A person asking the agent to “switch to the Acme account” does not establish that the person may act for that account.
1. Draw the boundary around the whole request
Start with the path from a signed-in person to the business effect. A typical path includes an application entry point, an agent runtime, retrieval or application data, conversation state, a tool or connector, and the target business system. Add any asynchronous work that continues after the initial response. The isolation design is incomplete if one branch of that path is missing.
For each request, establish a server-validated context containing at least the authenticated principal, the tenant identifier, the requested operation, and a correlation identifier. Where a user can act in several tenants, include the selected tenant and the evidence that authorizes the selection. Validate that combination before invoking the agent. Do not ask the model to choose a tenant, turn a natural-language organization name into an authority decision, or repair a missing tenant value by guessing.
Pass the validated context through an application-controlled boundary. The model can help select an allowed operation or prepare arguments, but it should not create or alter the trusted tenant context. At every downstream boundary, either pass a verifiable identity or pass a context whose integrity the receiving service can validate. A plain tenant string supplied by an untrusted caller is not proof of authorization.
Separate identification from authorization. A tenant key tells a service which partition or namespace to address; it does not prove the caller can access it. Authorization checks should bind the authenticated principal and permitted operation to that tenant at the service that owns the resource. That way, a mistaken routing value cannot independently grant access.
2. Choose isolation boundaries for data
Inventory data by how it is created, queried, cached and deleted. Include source records, search indexes, embeddings, uploaded files, generated summaries, tool results, logs and exports. The practical test is not whether the primary database has a tenant column. It is whether a request, background task or operator action can retrieve another tenant’s information through any maintained path.
For shared storage, make tenant scope explicit in the data-access interface. A query should receive validated tenant context as a required argument, and the storage layer should apply the constraint rather than relying on each caller to remember it. Review alternate access paths such as administrative search, bulk export and repair jobs. A repository method that returns all matching records and filters them later in application code creates a larger failure surface than a method that cannot issue an unscoped tenant query.
Partitioning choices involve tradeoffs. A shared store with tenant-scoped records can be operationally efficient and makes common maintenance easier, but raises the importance of correct query enforcement and isolation tests. Separate stores or indexes can reduce cross-tenant exposure from some classes of query error, while increasing the work needed to provision, migrate, monitor and restore each tenant. A hybrid design may reserve stronger physical separation for tenants with a justified operational or contractual need. No single layout removes the need to authenticate and authorize each request.
Keep derived data in the same tenant boundary as the source unless a deliberate transformation creates a new, separately governed dataset. An agent-generated summary can contain the same confidential information as the records it summarizes. A shared cache key should therefore include tenant scope when its contents are tenant-dependent. Search filters, cache keys and result assembly all need tests that prove one tenant’s marker never appears in another tenant’s output.
Deletion and correction deserve particular attention. A customer record may appear in a source store, an index, a cached response and a persisted conversation. Define how a change reaches each location, how operators detect a missed update, and how the application behaves while propagation is incomplete. Isolation is not only about preventing a read at request time; stale derived material can expose information after access should have ended.
3. Separate conversational and workflow state
State includes more than visible chat transcripts. It can include thread identifiers, tool-call arguments, uploaded-file references, workflow checkpoints, cached credentials, retry payloads and resumable jobs. For each state type, decide whether it is tenant-specific, user-specific, shared configuration or ephemeral. Record its owner, retention rule and authorized readers.
Make tenant context part of the server-side association between a conversation and its state. On every read or update, verify that the caller can use that conversation in the selected tenant. Do not treat possession of an opaque conversation identifier as authorization. If an identifier is accidentally exposed or copied from another session, the lookup still needs to reject the wrong tenant or principal.
When a person changes tenants, do not silently carry over tenant-specific history, retrieval results or pending tool arguments. Start a correctly scoped conversation or explicitly reauthorize and reconstruct only the state that is safe to move. A user may be legitimate in both organizations while the content and permissions of each remain distinct. The interface should make the active tenant visible enough to catch an accidental switch before a consequential operation.
Persist only the state needed to resume a workflow. Long-lived state expands the period in which a missed authorization check or stale permission can matter. For a resumable action, revalidate permission when the work resumes, not only when it was first queued. A workflow can remain pending after a user’s role changes or a tenant connection is disabled.
4. Separate identity for user actions and application work
Choose identity according to whose authority the operation represents. If an action must be limited by the signed-in user’s own access, carry delegated user authority to the resource and have that resource make its authorization decision. If a background operation is legitimately performed by the application, use an application identity and constrain its permissions to the necessary workload and tenant context. These are different trust models, not interchangeable ways to make a request succeed.
Microsoft Entra’s OAuth on-behalf-of flow distinguishes delegated user authority from application identity; the resource boundary remains responsible for authorization. Design around that distinction rather than treating a successful token exchange as proof that a particular record or action is permitted (Microsoft Entra OAuth on-behalf-of flow). A service should check the operation and target against the identity and tenant context it receives.
If an application identity can act across tenants, the receiving service must still enforce the intended tenant boundary. A broad application credential combined with a caller-supplied tenant identifier is a particularly important path to test: a bug could keep the credential valid while routing a request to the wrong customer. Reduce the scope of each workload identity where the architecture permits, and avoid passing credentials through model-visible content or conversation state.
For every tool call, define which identity reaches the target system, what that identity can do, and where tenant authorization is evaluated. The agent’s decision to call a tool is not the authorization decision. This distinction is central to a multi-tenant tool layer; for its separate treatment, see Multi-Tenant MCP Servers: Tenant Isolation That Holds.
5. Isolate execution and network paths
An agent can be prevented from reading another tenant’s records and still cause a cross-tenant effect if execution context is shared incorrectly. A tool request may contain a target account, a resource name or a destination supplied by the user or model. Validate these values against the authorized tenant’s allowed targets before execution. Never use a model-generated identifier as the sole basis for selecting a customer’s account or endpoint.
Treat inbound reachability and outbound reachability as separate design questions. A private inbound endpoint does not, by itself, establish outbound isolation; outbound paths and compatibility of required tools need their own review (Foundry networking options). Document which components can call which destinations, how those destinations are selected, and what happens when the intended path is unavailable. This is an architecture constraint to verify for the chosen tool path, not a blanket claim about every Azure network design.
A tenant-aware execution wrapper can enforce a sequence: confirm the request context, authorize the action, resolve a destination from trusted configuration, validate the tool arguments, and only then dispatch. The tool response should be associated with the same tenant and workflow before it is stored or returned. If any step cannot establish the expected tenant, fail closed for that operation and return a useful error without revealing another tenant’s existence or data.
Approval must be bound to the precise action, not to a vague intention such as “the agent may update the account.” A sound approval record identifies the tenant, target, operation and exact payload that will be executed, and it expires after a defined interval. If the model or a later workflow step changes any material argument, require a new approval. Recheck authorization immediately before dispatch because a long pause may make the original decision stale.
Do not promise exactly-once effects from retries or model behavior. A timeout can happen after the target system has applied an action but before the agent receives confirmation. Use a target-supported idempotency mechanism where available, or reconcile the outcome before retrying. When neither is possible, route ambiguous outcomes to a controlled recovery path rather than repeating a potentially harmful operation automatically.
6. Proposed architecture and responsibility matrix
The diagram is a proposed logical design, not a deployment prescription. The application validates identity and tenant context before invoking the agent. Any data or state access passes through a tenant-aware boundary; any consequential operation passes through authorization and execution controls. The diagram deliberately keeps the target system’s authorization separate from the model’s tool choice.
flowchart TD
A["User request"] --> B["Validate identity and tenant"]
B --> C["Agent with scoped context"]
C --> D["Tenant data and state boundary"]
C --> E["Authorize exact action"]
E --> F["Execute against target"]
D --> G["Return scoped result"]
F --> G
accTitle: Proposed tenant isolation path for an Azure agent
accDescr: A request is validated for identity and tenant before reaching the agent. Reads and state use a tenant boundary, while actions require exact authorization before execution. Scoped results return to the request path.
If validation fails at the second node, stop before agent invocation. If the data boundary cannot establish the tenant, do not fall back to an unfiltered read. If action authorization fails, do not invoke the target. These failure branches are not separate services in the drawing; they are mandatory refusal points at the corresponding boundary. A successful model response cannot bypass them.
| Boundary | Primary responsibility | Evidence to retain for design review |
|---|---|---|
| Application entry | Authenticate caller, select and validate tenant context | Test cases for allowed and mismatched tenant selections |
| Agent invocation | Receive only validated context; avoid treating model text as authority | Input contract and a case where the model requests a different tenant |
| Data and state services | Scope reads, writes, caches and conversation lookups | Tenant-scoped query examples and cross-tenant negative tests |
| Execution wrapper | Authorize the exact operation and resolve an allowed target | Action policy, target mapping and approval binding |
| Business target | Enforce its own resource-level authorization | Identity used, permission boundary and denied-action evidence |
| Operations | Detect failures and recover without widening access | Runbook for ambiguous outcomes, tenant mismatch and stale state |
This matrix is an artifact for design discussion. Assign an accountable owner to each row, even when one team owns several boundaries. The owner should be able to show how the boundary is tested and how a failure is handled, rather than merely point to a diagram or a configuration setting.
7. Implement the boundary in numbered steps
-
Define the tenant identifier and selection rule. Choose a stable internal identifier that is not derived from a display name. Specify how a user selects a tenant, how membership is verified, and what the system does when the selection is missing or ambiguous. Test a user with access to two tenants and a user with access to only one.
-
Map every data and state path. List source data, retrieval material, caches, conversations, files, queues, logs and exports. For each, record how tenant context enters the operation, where access is checked, and whether a background process can bypass the normal request path. Make an explicit decision for each item instead of assuming that a database partition covers adjacent state.
-
Create a typed request context at the trusted edge. Have the application construct the context only after authenticating the caller and verifying tenant membership. Keep it separate from user text and model output. Reject absent, malformed or contradictory context. If downstream components cannot verify the context’s origin, redesign the handoff rather than relying on a convention that the tenant field will remain honest.
-
Make scoped access the normal storage interface. Require tenant context for reads and writes to tenant-owned resources. Review administrative and maintenance paths for equivalent checks. Add negative tests that attempt to query a
tenant-bluerecord usingtenant-redcontext, including through search, cache and state retrieval. A test should assert that the data is absent, not merely that a user interface hides it. -
Bind execution to an authorized target. Map each approved operation to destinations allowed for the active tenant. Validate the target and arguments independently of the model’s rationale. Where human approval is required, bind it to the tenant, exact target, operation and payload; expire it and repeat approval if material inputs change. Recheck permission immediately before the effect.
-
Choose and document the identity used at each hop. Record whether the operation represents the signed-in user or the application. Confirm what the destination checks and which tenant resources the identity can reach. Test an otherwise valid identity against a disallowed tenant target. A successful authentication response is not a substitute for this negative authorization test.
-
Exercise interruption and recovery paths. Simulate a timeout after dispatch, a worker resuming with stale state, a user losing membership, a missing tenant marker, and an unavailable destination. Define whether each case stops, retries safely, reconciles or requires operator action. Do not silently broaden access or replay an uncertain business operation to make the workflow appear complete.
-
Review network paths and tool compatibility. Trace inbound and outbound connections separately for every component in the request path. Check that required tools can use the chosen network route and that a failure cannot cause an undocumented route or credential fallback. Record the intended destination and the behavior when it cannot be reached.
-
Run cross-tenant tests before expanding access. Seed unique marker records and actions for at least two tenants. Test direct reads, retrieval, conversation lookup, cached responses, background work and tool execution. Include an adversarial request that asks to switch tenants, a valid user choosing an unauthorized tenant, and a wrong-tenant identifier that is otherwise well formed.
-
Make the test repeatable after change. Run the same boundary cases when changing prompts, tools, indexes, identity configuration, storage queries or execution wrappers. A harmless-looking change to result caching or workflow resumption can reopen a path that earlier tests covered. Keep expected outcomes attached to the component that owns the boundary.
Warning: Do not use a successful demonstration with one tenant as evidence of isolation. The critical evidence is that requests using the wrong tenant, principal, state identifier or target are rejected at the owning boundary.
8. Work through a hypothetical cross-tenant case
Consider a hypothetical distributor platform with two customers, Northwind and Contoso. A staff member belongs to both, but the active workspace is Northwind. The agent can retrieve order status and prepare a shipment change. Northwind and Contoso each have a record with a similar order number, and both use the same application entry point. This example is illustrative; the names and workflow do not describe a real deployment.
The application authenticates the staff member and validates Northwind as the selected tenant. It constructs a request context containing the person, Northwind’s internal tenant identifier, the requested operation and a correlation identifier. The agent receives that context but cannot replace it with a tenant named in the conversation. When the person asks for Contoso’s order, the application must treat that as a new authorization request, not as a harmless change in search terms.
A read for Northwind’s order passes the validated tenant context to the data boundary. The query retrieves a Northwind record, and the response is associated with the Northwind conversation. If the order number resolves only in Contoso, the Northwind-scoped lookup returns no matching record. It must not widen its search to all tenants and then ask the model which result looks right.
Now suppose the person requests a shipment-date change. The model can prepare a proposed payload, but a server-side policy verifies that the selected tenant permits this operation and that the target order belongs to Northwind. If an approval step applies, the displayed approval includes the exact order, new date and tenant. Editing the date after approval invalidates that approval. Just before dispatch, the execution boundary checks that the action is still authorized and that the target resolves to Northwind.
Assume the external system times out after receiving the request. The agent must not tell the person that the change definitely succeeded because it generated a plausible confirmation. Nor should it automatically submit the same operation again without checking whether the first request took effect. The workflow records an ambiguous result and follows a reconciliation path appropriate to the target system. If reconciliation cannot establish the outcome, it stops for controlled follow-up.
The worked case yields concrete acceptance criteria: the Northwind request cannot return Contoso’s order; a conversation identifier from Contoso is rejected in Northwind context; a model-generated tenant change does not alter server context; the action target must belong to Northwind; approval cannot survive a payload change; and an ambiguous timeout does not produce an unverified success message or an unsafe replay.
9. Test failure timelines, not just static access
Isolation defects often appear across time. A user can be authorized when a workflow is created and unauthorized when it resumes. A record can be correctly scoped in a primary store but remain in a cache after tenant membership changes. An action can be approved, modified and then dispatched by a delayed worker. Write tests that represent these transitions, not only clean requests that begin and finish within one session.
A useful failure timeline starts with a valid Northwind request, persists a pending action, removes the user’s Northwind membership, and then resumes the work. Expected behavior: revalidation fails before the target is called. Another timeline approves a specific payload, changes one target field, and retries. Expected behavior: the earlier approval no longer authorizes dispatch. A third sends an action and loses the response. Expected behavior: the workflow reports uncertainty and uses a safe reconciliation procedure.
Test confused-deputy conditions explicitly. An application identity may be technically capable of reaching several tenants, while a user is authorized for only one. Confirm that the application does not use its broader permission to satisfy an operation the user was not allowed to request. Test both a valid tenant with an invalid operation and a valid operation aimed at an invalid tenant; the two cases expose different policy mistakes.
Treat refusal quality as part of correctness. A denied request should not reveal whether another tenant has a matching customer, order or document. It should identify the active tenant clearly enough for the user to correct a mistake, while avoiding details that disclose another tenant’s data. For operators, preserve sufficient diagnostic context to identify the failing boundary without copying sensitive content into a general-purpose error message.
10. Decide what evidence is enough to proceed
Do not try to prove isolation with one all-purpose security test. Build evidence by boundary: tenant selection, data access, state lookup, identity handoff, target authorization and recovery. Each test should name the expected tenant, principal, operation and result. A reviewer should be able to tell which boundary rejected a bad request and whether downstream components were called.
A minimum release gate for a new tenant-aware workflow can include: no cross-tenant marker in read results; no cross-tenant state lookup; no execution for an unauthorized target; stale permissions checked again at resume; changed payload invalidates approval; and uncertain external outcomes do not become claimed business success. These are proposed acceptance criteria, not a certification or a guarantee that every possible defect has been eliminated.
Keep the evidence close to the design. A concise record might link a boundary to its owner, its negative test, the result of the latest run, and the conditions that require retesting. Retest when a change affects identity, tenant selection, data access, tool arguments, destination resolution or persisted state. This keeps review focused on changes that can move a boundary rather than on broad claims that the system is “multi-tenant.”
If the architecture spans several Azure services or teams, enterprise platform engineering can help make these contracts repeatable across application and infrastructure boundaries. The immediate goal is not a larger diagram; it is a testable interface that makes missing tenant context, stale permission and unauthorized targets fail in a predictable place.
Printable isolation worksheet
Use this worksheet in a design review. Fill one row for every data store, state type, tool or target that participates in the workflow. “Not applicable” is acceptable only when the reason is stated.
| Worksheet field | Record for this workflow |
|---|---|
| Workflow and owner | Name the user journey and accountable engineering owner |
| Tenant definition | State what one tenant represents and its stable internal identifier |
| Tenant selection | Explain how the active tenant is selected and membership verified |
| Data and derived material | List source records, indexes, caches, files and exports |
| Conversation and job state | List identifiers, checkpoints, pending actions and resumption paths |
| Caller identity | Identify the principal and whether an operation represents the user or application |
| Read boundary | Name the component that enforces tenant scope and its negative test |
| Execution boundary | Name the component that authorizes operation, target and exact payload |
| Network path | Record inbound and outbound paths and tool compatibility checks |
| Failure handling | State behavior for missing context, stale access, timeout and ambiguous outcome |
| Release evidence | Record the cross-tenant test cases, owner and change triggers for rerunning them |
A worksheet is complete only when each listed resource has a boundary and a failure behavior. If the team cannot say where a tenant check occurs, mark that path unresolved rather than assuming a neighboring component covers it. If the action can outlive the request, record how identity and authorization are checked when it resumes.
Summary: make separation observable and testable
Tenant isolation for an enterprise agent is the combined behavior of its data paths, persisted state, identities and execution routes. A tenant field in a request or a private network path alone cannot establish the whole boundary. Validate tenant selection before agent invocation, enforce scope where data and state are accessed, and authorize the exact action and target before a business effect.
The practical implementation sequence is to map the request end to end, define trusted tenant context, make scoped access unavoidable, separate user and application authority, bind execution and approval to a precise payload, then test denial and recovery paths across time. Use the matrix and worksheet to assign owners and keep evidence specific to each boundary. For deeper treatments of capacity planning, tracing and shared tool ownership, see planning quotas, queues and backpressure, tracing a request to its business outcome and ownership and access for a shared MCP tool layer; those are separate implementation concerns from the isolation boundaries covered here.
When a design has an unclear identity handoff, an unscoped state path or an external action that cannot be safely retried, resolve that specific boundary before expanding the workflow to more tenants. That is a more defensible next step than relying on a successful demo or a broad statement that the agent is isolated.