Contents
- 1. Define recovery for an agent, not just its hosting
- 2. Separate the state that must survive
- 3. Set recovery objectives around business work
- 4. Map Azure dependencies and regional capability
- 5. Design recovery for agent state
- 6. Rebuild identity and network paths deliberately
- 7. Work through a hypothetical regional failure
- 8. Write the runbook around decision points
- 9. Test recovery and learn from failure
- 10. Printable decision worksheet and next steps
1. Define recovery for an agent, not just its hosting
A regional recovery plan for an AI agent is not complete when a second copy of its application can start. The agent is a business process assembled from code or configuration, conversation context, tools, data, identity, network paths, and external systems that may be changed by its actions. After a region failure, those pieces must come together in a way that lets the right work resume without silently losing, duplicating, or misrepresenting a business outcome.
That distinction matters because an agent can appear healthy while being unable to perform useful work. Its interface may load, but its retrieval index could be stale. Its model call may succeed, but the downstream system may reject the agent’s identity. A conversation may resume from a saved transcript while losing an approval decision or the record of a payment already submitted. Recovery therefore needs an application-level definition of “working,” not a server-level one.
Begin by describing the service in terms a business owner can verify. For example: “A user can submit a service request, receive a grounded draft, and—after an authorized person approves the exact action—create one case in the system of record.” That statement identifies a start, a result, a human decision, and an externally visible effect. Each can be tested after failover.
This guide concentrates on state recovery, regional capability checks, and operational runbooks. It does not compare agent platforms or provide a general production-readiness review. For the platform choice and its control implications, see Microsoft Foundry, Bedrock AgentCore, or Gemini Enterprise: Choosing an Enterprise Agent Platform. For tracing an individual request through to a business result, see Tracing a Foundry Agent from User Request to Business Outcome, and for the broader release review, A Production Readiness Review for Microsoft Foundry Agents.
2. Separate the state that must survive
“Agent state” is not one database. It is a collection of information with different lifetimes, authorities, and recovery requirements. Treating all of it as a single backup problem produces either an expensive replica of data that can be regenerated or a fragile recovery plan that omits the one record needed to prevent an unsafe repeat action.
Classify state by its role
Definition state describes what the agent is supposed to do: prompts, tool definitions, application code, policies, configuration, schemas, and deployment manifests. It should be versioned and releasable as a known package. Recovery should identify the exact approved version, not merely restore “the latest” configuration from an unrelated branch.
Reference state is the information used to answer or act: documents, indexes, product catalogs, policy data, and other source material. Some of it may be reproducible from an authoritative source; some may have manual corrections or a refresh lag that must be preserved. Record its source, freshness timestamp, transformation version, and any non-reconstructible changes.
Interaction state includes conversation turns, workflow progress, tool results, and the user’s current task. Decide which interactions need to resume, which can restart safely, and how much lost progress is acceptable. A transcript is not necessarily a complete workflow record: it may not say whether a remote system accepted a request.
Control and execution state includes approvals, action intents, idempotency keys, execution receipts, retries, and reconciliation status. This is often the highest-risk category. It determines whether an external operation is pending, completed, rejected, or uncertain. Keep it tied to the business transaction, rather than relying only on conversational text.
Observability state includes traces, logs, and operational events. It may not be required to answer the user, but it is essential to investigate what happened before failover and to reconcile uncertain actions afterward. Set retention and recovery expectations according to the investigation window the business needs.
Name the authoritative copy
For each state class, identify the authoritative system and the recovery copy. “Primary region” is not an authority model. A restored database can be older than a source system; a replica can contain writes whose upstream event has not yet been acknowledged; a rebuilt index can reflect a different source snapshot. The runbook needs a rule for which version wins and who can declare it authoritative.
A useful state register records the owner, source of truth, region or location, replication or rebuild method, acceptable age, validation query, and recovery action. If a field is not known, write “unknown” and assign an investigation rather than assuming it is covered by a platform default. This creates an actionable gap list before an outage does.
The most important distinction is between replayable work and committed effects. A prompt generation can often be rerun, subject to data freshness and cost controls. A request that may have created a shipment, changed an account, or sent a customer message cannot be assumed safe to replay. Its status must be checked against the system that owns that effect.
3. Set recovery objectives around business work
Recovery time objective (RTO) is the target time to restore a defined service. Recovery point objective (RPO) is the acceptable amount of data loss measured in time or business events. These are targets to validate, not properties implied by deploying components in more than one region. Microsoft’s reliability guidance emphasizes defining recovery objectives and validating recovery rather than treating redundancy as proof of recoverability. Azure Well-Architected disaster recovery guidance
Set objectives for meaningful units. “The app is reachable in 30 minutes” may be a useful infrastructure measure, but it does not say whether users can complete a case, whether old approvals are intact, or whether an agent may safely resume a partially completed action. A more useful service objective might include: accept new work, retrieve sufficiently current material, continue eligible in-progress work, and reconcile actions whose completion is uncertain.
Choose different targets for different state. A search index that can be rebuilt from a verified source may tolerate more downtime than an action ledger used to prevent duplicate transactions. Conversely, a rapidly replicated index can still be unfit to serve if its source snapshot is incomplete. RPO should describe the data the business needs, not simply the replication interval of a storage component.
Convert targets into observable checks
For every objective, define a clock start, a clock stop, and evidence. RTO starts when the incident is declared or when a specified service condition is first detected; choose one definition and use it consistently. It stops only when an agreed end-to-end transaction passes, not when the deployment reports success. Evidence could include a successful retrieval against a known test record, a controlled workflow continuation, or a reconciled action receipt.
RPO also needs a measurement method. Compare the last confirmed business event in the failed region with the latest event present in recovery, using stable event identifiers and timestamps from an agreed authority. For a dataset rebuilt from source, measure both source freshness and the time required to regenerate the usable version. If no one can calculate the loss after an exercise, the RPO is not operationally measurable yet.
Do not promise precision that dependencies cannot support. If a downstream business system cannot confirm whether a timed-out action committed, mark the result as uncertain and route it for reconciliation. A model’s assertion that it completed the task is not evidence that the business system did so. The plan should preserve uncertainty rather than convert it into a false “success” or “failure.”
4. Map Azure dependencies and regional capability
An Azure deployment diagram should show more than application boxes. For each dependency, record where it is configured, what region or scope it uses, how the recovery environment reaches it, and how its behavior is verified. Avoid treating a service name as a guarantee of regional availability, feature parity, quota, or access. Those details can vary by service, subscription, configuration, and time; confirm them for the actual deployment before adopting a recovery design.
Build a dependency inventory
Include the agent runtime and its release artifact; model or inference endpoint; prompt and tool configuration; conversation and workflow storage; retrieval source and index; secrets and identity configuration; network paths; queues or schedulers, if used; and every external tool that can create a business effect. Include monitoring and operator access, because a recovery environment that cannot be observed or safely administered is not ready for production traffic.
For each dependency, answer five questions: Is it needed to serve new requests? Is it needed to resume existing work? Is it regional, global, or otherwise scoped? Is its state replicated, restored, or rebuilt? What concrete check proves it is usable in the recovery environment? If the answer to the last question is only “the resource exists,” define a functional test.
A second region is not automatically a second working system. Configuration may point both deployments at a primary-region endpoint. A recovery network may resolve a name but be blocked from the target. An identity may exist in one environment but lack permission to the restored resource. A secret may be present but expired, or refer to a resource that was never recreated. These are dependency-chain failures, so test the complete path from agent to business effect.
Check capabilities before choosing an architecture
Create a regional capability record for each essential dependency. Confirm the required service or model is available to your subscription and intended region; confirm the relevant feature and deployment configuration; verify quota or capacity through the appropriate operational channel; test authentication and network reachability; and record any limits on data movement or recovery timing that affect your design. Recheck before a planned failover exercise and after material architecture changes.
Where a dependency cannot be used in the recovery region, choose explicitly among a different supported design, a degraded service mode, a manual process, or an accepted recovery gap. “We will switch regions” is not a decision. It leaves responders to discover constraints during an incident, when the business may already be waiting.
An Azure-specific reference design should label boundaries and ownership rather than imply that every Azure service behaves alike. The following flow shows the recovery decision sequence; it does not claim that any dependency is automatically replicated or available.
flowchart TD
A["Declare regional incident"] --> B["Check dependency capability"]
B --> C{"Recovery path usable?"}
C -->|"Yes"| D["Restore and validate state"]
C -->|"No"| E["Choose degraded or manual mode"]
D --> F["Reconcile pending effects"]
E --> G["Resume approved work"]
F --> G
accTitle: Regional recovery decision flow
accDescr: After incident declaration, responders check whether the recovery dependencies are usable. If they are, they restore and validate state, then reconcile pending effects. If not, they choose a degraded or manual mode. Work resumes only after the applicable path is approved.
The key gate is not “does the recovery deployment respond?” It is whether the required dependencies are usable for the intended mode. If capability is incomplete, the degraded path must be a designed operating mode with clear limits, not an improvised attempt to call unavailable systems.
5. Design recovery for agent state
The recovery method should follow each state class. Some state should be redeployed from version control, some restored from a recovery copy, some rebuilt from an authoritative source, and some reconciled with the system that owns the business effect. A single “restore everything” procedure obscures these differences and may overwrite newer or more authoritative information.
Version and redeploy definition state
Package code, prompts, tool schemas, policy configuration, and infrastructure definitions into identifiable releases. Record the release identifier and the source revisions needed to reconstruct it. Keep environment-specific values separate, but make the selected values auditable. During recovery, deploy a known approved version and compare its configuration with the intended recovery manifest before admitting requests.
Microsoft distinguishes Foundry hosted agents, which package custom code as containers, from prompt agents, which use declarative prompts and tools. Hosted agents overview For recovery planning, the operational question is what must be captured to recreate the selected agent form and its dependencies. Do not infer traffic-splitting behavior from the deployment model; if a canary or controlled cutover is required, design and validate the external routing mechanism that would provide it.
Restore, rebuild, or reconcile data
For interaction and workflow data, decide what constitutes a resumable checkpoint. A useful checkpoint may include the conversation or task identifier, last confirmed step, input references, relevant tool results, approval state, and the status of any external action. Store structured workflow status independently of natural-language summaries where safety depends on a precise state transition.
For reference data, establish the source snapshot and a freshness rule. If rebuilding takes hours, the plan should say whether the agent waits, serves a marked limited mode, or uses an older verified snapshot. Set a cutoff for stale material and define which requests must be declined or escalated when that cutoff is exceeded. “The index is available” is insufficient if its contents are not fit for the task.
For external effects, use a durable action record before execution. One illustrative record could include work_id, action_id, payload_hash, approval_id, approval_expiry, submission_status, external_reference, and last_reconciled_at. The approval should bind to the exact payload and expire; if the payload changes, require a new approval. Before retrying, query the system that owns the effect using a stable reference where available. If the outcome cannot be established, hold the action for reconciliation rather than submit a blind duplicate.
No recovery design should promise exactly-once effects across independent systems. A timeout can occur after a remote system commits but before the agent receives confirmation. The practical objective is to make the state transition visible, make retries safe where the receiving system supports that pattern, and make uncertain cases detectable and reviewable.
Define a recovery mode, not just a destination
A recovery region can support several operating modes: full service, read-only or draft-only service, limited workflows, or no user traffic while operators reconcile state. Choose these modes based on dependency capability and business risk. For example, an agent might continue answering from a verified snapshot but disable tools that change records until identity, action history, and downstream reachability are checked.
State the user-facing behavior for each mode. Explain whether existing conversations can continue, whether new work is queued, which actions need human handling, and how users will learn that a response is incomplete. Hidden degradation is worse than a clear temporary limitation because it encourages users to treat a draft as a completed business outcome.
6. Rebuild identity and network paths deliberately
Identity and networking are recovery dependencies, not finishing touches. Map the principal used by each runtime, the resources it must access, the scope of each permission, and the network route used to reach them. Identify which elements are centrally managed and which are tied to a regional resource. Confirm that the recovery runtime uses the intended identity rather than assuming the primary environment’s settings carry over.
The accompanying responsibility matrix is an architecture artifact to complete for the actual environment. “Platform team” and “application team” are example role labels, not assertions about organizational ownership. Assign named teams in the operational version and record the evidence each team must provide before traffic moves.
| Area | Primary responsibility | Recovery-region check | Evidence to retain |
|---|---|---|---|
| Agent release and configuration | Application engineering | Intended release and environment values are present | Release identifier and configuration comparison |
| Runtime identity | Identity/platform owner | Correct principal can authenticate from recovery runtime | Non-destructive access check and identity reference |
| Data and action permissions | Data or application owner | Required reads and writes are scoped as intended | Resource-by-resource permission test |
| Network path | Network/platform owner | Name resolution, routing, and required access work end to end | Timestamped connectivity results and route notes |
| Secrets and certificates | Assigned secrets owner | Required values are present, current, and mapped to recovery resources | Rotation/expiry status and reference identifiers |
| External tools | Owning system team | Tool endpoint and business authorization work | Controlled test or documented manual fallback |
| Operations and monitoring | Service operations | Responders can observe and administer the recovery service | Dashboard access and alert-path confirmation |
Apply least privilege in both environments. A recovery role should not receive broad permissions simply because an outage makes access inconvenient. At the same time, a narrowly scoped role that cannot perform the required recovery action is not useful. Test permissions with a non-destructive check first, then use a controlled business test only when the effect is understood and contained.
Network verification must cover the agent’s actual call path, not only a responder’s workstation. Check name resolution, outbound or private routes as applicable, required destination access, and return behavior. Avoid recording secrets or sensitive payloads in the runbook; record resource identifiers and secure retrieval instructions instead. Keep a clear escalation owner for a failure at each boundary so responders do not spend the first hour guessing which team owns it.
For teams coordinating identity, networking, infrastructure definitions, and release practices across several services, enterprise platform engineering can be a relevant implementation discussion. The recovery design still needs to reflect your own service boundaries, dependency owners, and business acceptance criteria.
7. Work through a hypothetical regional failure
Consider a hypothetical internal service agent that drafts a supplier change request from approved policy material. A person reviews the exact proposed change, and an approved request is submitted to a separate business system. This example is illustrative: it does not describe a PADISO implementation, a tested Azure topology, or a promise about any service’s regional availability.
Assume the business wants new drafts available within two hours, can tolerate losing up to fifteen minutes of resumable conversation progress, and will not accept an unverified duplicate supplier change. These figures are illustrative targets chosen to show the method, not recommended service levels. The business system remains authoritative for whether a change was committed; the agent’s conversation history does not override it.
Before the incident, the team has captured a release identifier, a versioned prompt and tool configuration, a reference-data refresh timestamp, and a structured action ledger. Each approved action is bound to a payload hash and an expiration time. The team has also documented which recovery-region resources must be checked and who can authorize a switch to draft-only mode.
At 09:00, operators declare a regional incident. They do not immediately redirect all users. First they check the capability record for the recovery region and verify the deployed runtime, required model path, storage, identity, network access, and external tool reachability. If a required capability is unavailable or cannot be proven, the service remains in the defined limited mode rather than presenting a partially functioning agent as fully recovered.
At 09:35, assume the checks show that the recovery environment can serve the draft workflow, but the team cannot yet confirm all pending submissions from the failed region. Responders deploy the known release, restore or rebuild the approved reference snapshot, and compare its refresh time with the service’s freshness rule. They admit draft-only traffic after a known-record retrieval and a workflow test pass. They do not enable supplier changes yet.
The action ledger now becomes the critical source for the transition. Suppose one submission was sent shortly before the failure and its response was lost. The agent’s saved conversation says “submitting,” which is not enough to determine completion. The operator searches the business system using the action’s stable reference or other agreed reconciliation method. If a matching change exists, record the external receipt and mark the action complete. If the system confirms no change, revalidate the still-current approval and submit according to the normal control. If the result remains unclear, keep it held for the business owner.
At 10:20, after the external tool path and action reconciliation have passed, the incident lead can consider enabling the approved write workflow. The decision is based on evidence: dependency checks, state freshness, identity and network tests, and a bounded business transaction whose result can be confirmed. The team also records which users were served in draft-only mode and whether any requests need follow-up.
This sequence may take longer than simply directing traffic to a second deployment. That is not necessarily a failure: the service has preserved a crucial invariant by declining to create uncertain duplicate effects. If the business’s two-hour draft target is met but write actions remain paused for longer, report those as separate recovery outcomes rather than declaring the entire agent “up.”
Counterexample: a healthy endpoint with unrecoverable work
Now consider an alternate design that replicates the application deployment but stores task progress only in a regional session store. Its recovery instance starts quickly and returns a successful health check. Users can open the agent, but in-progress tasks are missing. A few submitted actions have uncertain status, and no durable action ledger links them to the business system. Operators either ask users to resubmit everything or risk repeating effects.
This design has recovered compute, not the workflow. The failure was not that the secondary deployment failed to start; the failure was that recovery objectives and state ownership were never defined at the business-process level. A better plan might accept a restart for low-risk drafts while preserving a separate durable record for approvals and external effects. That tradeoff should be made before an outage, with the business owner explicitly accepting the lost progress that cannot be reconstructed.
8. Write the runbook around decision points
A useful runbook is a sequence of observable gates with named owners, not an essay about the architecture. It should let an incident lead decide what is safe to do next and let a responder capture evidence without inventing missing steps. Keep detailed resource identifiers in an access-controlled operational location; the published runbook should contain enough references for authorized staff to find them.
Suggested incident sequence
-
Declare and scope the incident. Record the detection time, affected region or dependency, incident lead, and the business workflows affected. Separate confirmed facts from assumptions. Do not route work to a recovery environment simply because the primary is unavailable.
-
Select the permitted recovery mode. Use the service’s predefined decision rules: full service, limited service, manual handling, or pause. Confirm who can authorize each mode and what user message applies. If a critical capability is unknown, treat it as unavailable until checked.
-
Verify the recovery path. Confirm the intended release, configuration, identity, network routes, data sources, model or inference dependency, monitoring, and operator access. Test dependencies from the recovery runtime. Record pass, fail, or not checked; do not compress these states into a single green status.
-
Restore and validate state. Restore workflow records, rebuild eligible reference data, and verify freshness against the service rule. Compare counts or stable identifiers where appropriate. Preserve the source and recovery timestamps so the team can calculate actual data loss later.
-
Reconcile actions in flight. Extract pending and uncertain action identifiers. Check each against its authoritative business system. Mark confirmed completion, confirmed non-completion, or unresolved status. Do not retry an unresolved action until the defined reconciliation decision is made.
-
Run a bounded business transaction. Use an approved test record or controlled transaction with a clear cleanup or reversal procedure where applicable. Validate the complete journey, including the business result, not only a model response or HTTP success. If no safe test exists, define a non-mutating verification and state that limitation.
-
Admit traffic in stages. Start with the permitted user group or workflow scope selected by the incident lead. Watch the agreed health and business signals. Where controlled routing is required, use the external routing design that has been independently validated; do not assume the agent hosting model supplies it.
-
Communicate, reconcile, and close. Tell users which work resumed, which was lost or delayed, and what needs manual attention. Continue reconciliation until uncertain effects have an owner and disposition. Close the incident only after the business acceptance criteria and evidence record are complete.
Record facts that make the runbook executable
For each step, include the role responsible, the source of truth, the command or portal path maintained by the service team, expected result, failure branch, and escalation contact. This guide intentionally does not invent Azure command names, API calls, quotas, or service-specific recovery behavior. Add those only after validating them for the deployed services and subscription.
The runbook should also say how to stop. If a test transaction produces an unexpected effect, if state freshness fails its threshold, or if action status cannot be reconciled, responders need an explicit hold condition. A stop condition protects the business from turning an infrastructure incident into a series of uncontrolled downstream changes.
9. Test recovery and learn from failure
A recovery plan is a hypothesis until responders have exercised it. Validation should include the state, dependency, and business checks that make the service useful. A successful deployment drill proves only that the steps exercised by that drill worked; it does not prove that every dependency, data loss scenario, or uncertain external action is covered.
Start with a tabletop that follows the decision points in the runbook. Give participants a timeline: a region becomes unavailable, a dependency check returns unknown, an action response is lost, and a reference source is older than expected. Ask who selects the operating mode, what evidence they need, and which work must stop. Tabletop exercises expose missing ownership and ambiguous language without making production changes.
Progress to technical exercises within a controlled scope. Validate recovery access, state restoration, configuration comparison, and non-destructive connectivity. Then test a bounded end-to-end workflow in an environment where the effect is controlled and understood. Record the exact release, configuration, data snapshot, identities, and timestamps used, so a later exercise can distinguish a real improvement from a changed test setup.
Measure actual RTO from the agreed start to the accepted business outcome. Measure RPO using business events or records, not only infrastructure replication status. Track the time required to reconcile uncertain actions, because a service that resumes new requests while unresolved effects accumulate may still leave operators with an unacceptable backlog.
After each exercise, convert observations into owned work: a missing permission check becomes a named identity task; a stale recovery index becomes a freshness rule or rebuild improvement; an ambiguous retry becomes a revised action-state design; and a slow decision becomes a clearer mode-selection gate. Retest the changed step. A closed ticket is not evidence that the recovery path now works.
10. Printable decision worksheet and next steps
Use this worksheet in a design review or recovery exercise. Complete it for one business workflow at a time; a single worksheet for an entire agent often hides different objectives and state needs. The prompts are intentionally concrete so teams can turn answers into runbook steps and measurable acceptance criteria.
Recovery definition
- Business outcome: Name the action or answer the user must successfully receive. Define what counts as completion in the system of record.
- Permitted modes: Write down full, limited, manual, and paused modes that apply. For each, identify what users can do and which effects are disabled.
- RTO and RPO: Set targets for the workflow and each critical state class. Specify the clock boundaries and the evidence used to measure them.
- Acceptance owner: Name the business role that can confirm recovery is fit for use, alongside the incident lead who coordinates restoration.
State and dependency record
- Definition state: Record release, prompt/tool configuration, policy, and infrastructure references needed to reconstruct the intended agent.
- Reference state: Name its authority, snapshot or refresh marker, acceptable age, rebuild method, and stale-data behavior.
- Interaction state: Identify which sessions resume, which restart, and what progress loss the business accepts.
- Action state: Record the durable identifiers and statuses needed to distinguish pending, complete, rejected, and uncertain effects.
- Regional capability: For every essential dependency, record the confirmed recovery-region capability, access path, known limitation, and the date last checked.
- Identity and network: Identify the runtime principal, required resource access, network path, owner, and test that proves the complete path works.
- Operational access: Confirm that responders can observe the service, reach the required recovery instructions, and escalate to the dependency owners.
Runbook evidence and closure
- Decision gates: Record the condition for full service, limited service, manual processing, and stopping recovery.
- State validation: Define the comparison, freshness check, or business-event reconciliation used after restore.
- Uncertain effects: Specify the authoritative check, retry rule, and escalation path. Keep unresolved actions held.
- Business test: Choose a bounded end-to-end test and state what result proves success without relying on a model’s claim.
- Communications: Prepare concise messages for resumed, delayed, lost, and manually handled work.
- Exercise record: Capture actual recovery times, observed data loss, exceptions, decision owners, and corrective actions.
- Next review trigger: Set a review point after material changes to the agent, identity, network, data source, external tool, or regional design.
Printable summary: Before approving a regional recovery design, confirm that (1) every critical state has an authority and recovery method, (2) regional capability is verified for the actual dependency chain, (3) uncertain external effects cannot be blindly replayed, and (4) a tested runbook reaches a business-accepted result. If any statement is false, document the gap, its owner, and the operating limitation until it is corrected.
Summary and next steps
Start with the workflow’s business outcome, then map the state that makes the outcome safe to resume. Assign separate recovery methods to definition, reference, interaction, action, and observability state. Set measurable objectives, confirm the recovery region’s actual capabilities, and include identity and network checks in the same dependency map as the agent runtime.
Most importantly, make pending business effects visible. Preserve enough structured information to determine whether an external action completed, require approval to apply to the exact payload, and hold ambiguous cases for reconciliation. Restore only into a mode whose dependencies and state have passed their checks. A service can offer useful limited operation before it is safe to resume every action, but that limitation must be explicit.
The practical next step is to choose one high-value workflow, complete the worksheet with its owners, and run a tabletop against a regional failure timeline. Turn every unknown into either a verified capability, an accepted limitation, or a specific engineering task. Then exercise the runbook and measure recovery by the business result it restores—not merely by whether another deployment starts.