SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Checklist 21 mins

A Production Readiness Review for Microsoft Foundry Agents

A practical release checklist for Microsoft Foundry agents, covering security, quality, operations and cost gates before production.

The PADISO Team ·

How to use this release checklist

This checklist is for teams deciding whether a Microsoft Foundry agent is ready to handle real work, not just produce plausible answers in a development environment. It covers four release gates: security, quality, operations and cost. Each item asks for observable evidence, an accountable owner and a defined response when the evidence is missing or fails.

Treat a checked box as a claim that can be reviewed, not an intention or a feature request. A team should be able to point to a configuration, test result, operating procedure, owner or decision record that supports it. If an item is not relevant, record why and who accepted that boundary; do not silently skip it.

The review applies to the complete path from the user request through the agent and its tools to the business outcome. A model response that sounds correct is not proof that an external action was safe, completed, or correct. Keep those judgments separate throughout the release decision.

A release owner can run the checklist in a working session with engineering, security, the business process owner and operations. The final decision is not a score. A serious unresolved security or correctness issue should block release even if every other category looks strong. For a broader platform-selection decision, see Microsoft Foundry, Bedrock AgentCore, or Gemini Enterprise: Choosing an Enterprise Agent Platform; this review begins after a platform direction has been chosen.

1. Define the release boundary

Before reviewing controls, agree on exactly what the team is releasing. An agent that drafts a response for a person to approve has a different risk boundary from one that can change a customer record or submit a payment. The review should name the user groups, business process, data classes, tools, action types and environments covered by this release. If any of those are still vague, the release boundary is not stable enough to assess.

  • Name the business task and intended outcome. Describe the task in operational language, such as “prepare a return request for an agent to review,” rather than “answer return questions.” Identify the business record or decision that should change, and specify what counts as completion. This definition anchors quality checks and prevents a fluent response from being mistaken for a successful process.

  • List permitted and prohibited actions. Record each external action the agent may propose or execute, the data it may read or change, and the actions that remain outside its authority. Make exclusions concrete: for example, changing a bank account may be prohibited even if looking up an invoice is allowed. A broad instruction such as “use tools responsibly” is not a usable boundary.

  • Identify the release population and environment. State whether the release is limited to internal users, a specific department, or external customers, and identify the production environment and the systems it can reach. Include any meaningful differences between test and production. A restricted audience is a control only if the restriction is actually enforced and has an owner.

  • Assign a business owner and a technical release owner. The business owner accepts the process outcome and defines when human judgment is required. The technical owner is responsible for the deployed configuration, operational response and rollback decision. Record who can make the release decision when either person is unavailable; a shared mailbox or team name alone may not be enough to establish accountability.

  • Write down the assumptions that could invalidate approval. Examples include a particular source system being authoritative, a tool returning current data, a user having a certain role, or an action being reversible. Give each assumption a way to verify it. If an assumption changes, the release owner should know whether it requires a new review rather than treating the original approval as permanent.

A practical release boundary is deliberately narrower than the product roadmap. If the team has not evaluated a capability, leave it out of this release rather than implying that the whole agent has been approved. That makes the eventual decision explainable and gives the next review a clear change list.

2. Security gate: identity, data and network paths

Security review should follow the actual request and action paths. Start with who can invoke the agent, which identity is used to access dependencies, what data crosses each boundary, and what happens when access is denied. Do not infer safety from a product label or from a successful test using an administrator account. The relevant question is whether the deployed identities and routes are constrained to the work this agent is approved to do.

  • Document user authentication and authorization at the application boundary. Show how the application establishes the caller’s identity and how it decides whether that caller may use this workflow. Test both permitted and denied cases using representative roles. If the agent receives requests through another service, record where caller identity is checked and how the original user context is preserved or intentionally not preserved.

  • Inventory runtime identities and their permissions. For each identity used by the application, agent orchestration or connected tool, list the resources and operations it can access. Separate read permissions from write permissions, and remove permissions that are not needed for the approved task. Test denied operations as deliberately as allowed ones; a successful happy-path call does not show that access is appropriately limited.

  • Trace sensitive data through prompts, tool inputs, outputs and logs. Identify what personal, confidential or business-sensitive fields may appear in each stage. Decide which fields are necessary for the task and which should be withheld, transformed or excluded. Confirm that the team understands what is retained by its own application and observability components; do not assume a field is harmless merely because it is absent from the final user-facing answer.

  • Review inbound and outbound network routes separately. Record who can reach the agent-facing application and where the runtime can connect when calling tools or other dependencies. Private inbound access does not by itself isolate outbound traffic; verify the intended egress path and that each required tool is compatible with it (networking options). Test the expected route from the deployment environment rather than relying only on a network diagram.

  • Define how secrets are provisioned, rotated and revoked. Identify the owner and operational procedure for every credential used by the application or tools. Confirm that credentials are not embedded in prompts, source files or routine diagnostic output. Include the response to suspected exposure: who can revoke access, what dependent workflow may stop, and how the team will establish that access has been restored safely.

  • Set a boundary for untrusted content. Identify user-supplied files, retrieved content, tool responses and other inputs that may contain misleading instructions or malformed values. Specify which inputs can inform an answer and which are allowed to influence an action. Test whether unexpected content can move the workflow outside its permitted task; do not treat a prompt instruction as a substitute for access control on the tool itself.

  • Decide what evidence is retained for investigation. Select the minimum request, decision and action details needed to reconstruct a consequential event, while avoiding unnecessary sensitive content. Define who can inspect that evidence, how it is protected, and the applicable retention decision through the organization’s normal process. The goal is useful, controlled investigation—not collecting every prompt indefinitely.

The network and identity review should result in an ownership map, not merely a diagram of boxes. An Azure-specific responsibility matrix can make the handoffs explicit. Adapt the rows to the actual deployment; the named responsibilities below are a proposed design, not a claim that any particular configuration is present by default.

Boundary or controlAccountable ownerEvidence to review before releaseFailure response
User access to the applicationApplication ownerRole rules and allowed/denied access casesReject request; investigate unexpected access
Runtime identity and permissionsPlatform or application ownerIdentity inventory and permission reviewDisable or narrow access; assess affected actions
Egress to tools and dependenciesNetwork owner with tool ownerApproved route, destination list and connectivity checkStop dependent action; restore only after route review
Data carried through the workflowData or process ownerField inventory and retention decisionLimit processing; follow the organization’s incident process
External business actionBusiness process ownerExact action record and verification methodHold or compensate under the approved procedure

Use the matrix to expose shared responsibility. For example, the network team may own a route while the application team owns which tool calls it, and neither team alone can confirm the end-to-end boundary. If a row has no named owner or no testable evidence, keep it open. A design that assumes another team is responsible is not a completed control.

3. Quality gate: correct work, not convincing language

Quality should be judged against the business task and the consequences of mistakes. Build a representative set of requests that includes ordinary work, incomplete information, conflicting records, denied actions and unusual but plausible inputs. For each case, specify the expected answer or behavior before reviewing outputs. That makes the assessment more useful than asking reviewers whether a response “looks good.”

  • Define acceptance criteria for task completion. State what information must be correct, what must be present, what must not happen, and which cases should be handed to a person. For a record update, the criteria might include matching the correct customer, selecting an allowed status and preserving specified fields. Tie each criterion to a business consequence so that reviewers know which failures are release-blocking.

  • Create a test set from real process variation without exposing unnecessary sensitive data. Include ordinary requests, ambiguous wording, missing fields, conflicting source values, unsupported requests and tool errors. Use approved test data or appropriately prepared examples. A small, carefully chosen set that covers risk boundaries is more useful than a large collection of repetitive easy prompts.

  • Evaluate tool selection and action arguments separately from the final response. Check whether the agent chose an allowed tool, supplied the correct target and fields, and refrained from making an action when required evidence was missing. A polished summary can conceal a wrong identifier or an unauthorized argument, so inspect the consequential intermediate values as well as what the user sees.

  • Test refusal, clarification and escalation behavior. Include cases where the correct result is to ask for a missing value, explain a limitation, or route the task to a person. Decide which uncertainties permit a safe clarification and which should stop the workflow. Do not reward completion for its own sake when the input does not support a reliable action.

  • Separate agent claims from verified business results. Define how the application or an operator confirms that a requested change actually occurred in the system of record. A message saying “updated” is not sufficient evidence by itself. Compare the intended result with the authoritative business record or another agreed confirmation signal before labeling the workflow complete.

  • Set explicit acceptance thresholds and a review owner. The process owner should determine which error classes are unacceptable and which quality indicators are monitored. Record the test population, criteria, known gaps and reviewer decision. Avoid a single average score that can hide a critical error category; a high pass rate does not compensate for a failure that can cause an unauthorized or materially incorrect action.

  • Recheck quality after material changes. Identify changes that trigger focused retesting, such as a different prompt, tool definition, model or data source, permission boundary, or business rule. Keep the criteria stable enough to compare results, and document what changed. A previously approved test result applies to the configuration it assessed, not automatically to every later revision.

A quality release decision should include at least one explicit “do not proceed” case. If the agent receives a request to update a record but cannot establish which record is intended, the expected outcome may be a clarification or handoff—not a best guess. This negative case is especially valuable because successful routine examples can create unwarranted confidence in uncertain situations.

4. Operations gate: detect, contain and recover

Operational readiness means the team can see when the workflow is unhealthy, identify the stage where it failed, and respond without creating a second business problem. It does not require every incident to be prevented. It requires clear signals, an owner, an actionable response and a safe way to pause or resume the affected work.

  • Define the operational unit being monitored. Choose a unit that maps to a business request or task and can be counted consistently. Distinguish started, completed, failed, handed off and outcome-verified work. This prevents a system from appearing healthy because it generated responses while the underlying business tasks remained incomplete.

  • Set useful service and quality signals. Select measures that reveal user impact, such as completion rate, handoff rate, tool failures, time to resolve, or mismatches found during outcome verification. Define how each measure is calculated and who reviews it. Pair volume or latency signals with correctness and safety signals so that faster processing is not mistaken for better service.

  • Assign an on-call or incident response owner for the release. Specify where an operator receives an alert, how to identify the affected workflow and who has authority to pause new actions. Include a backup owner and a way to contact the business process owner when a technical incident may have changed customer or business records.

  • Write the pause and resume procedure. Identify the control that stops new consequential actions, the conditions for keeping the system paused, and the checks required before resuming. Test the procedure in a non-production setting where practical. Restarting a service is not the same as confirming that queued work, partial actions or repeated requests are safe to process.

  • Define handling for timeouts, retries and partial completion. For each external action, decide how the application distinguishes a confirmed failure from an unknown outcome. If a request times out after reaching a downstream system, retrying blindly may duplicate an effect. Record the reconciliation step, the person or system that performs it, and how the user is told whether the action is pending, completed or unresolved. Never promise exactly-once external effects.

  • Retain enough context to reconstruct a failure. Ensure an operator can connect a user request, the agent’s proposed action, any approval and the resulting business state using appropriate identifiers. Minimize sensitive payload retention, and make sure access to investigation material is controlled. If the team cannot determine which action was attempted, it cannot reliably decide whether to retry or compensate.

  • Establish a change and rollback path. Identify what can be reverted directly, what requires a compensating business action, and what cannot be undone. Keep the previous approved configuration or an equivalent recovery option available under the organization’s deployment process. Name the decision-maker for rollback and the conditions that trigger it, including quality regressions as well as infrastructure faults.

A detailed regional recovery design is a separate decision from this release gate. Keep this review focused on the everyday operational controls and link the deeper failure-boundary discussion to Disaster Recovery for Azure Agents: What Must Survive a Region Failure?. Do not mark the release ready merely because a regional plan exists; the operational owner still needs a workable incident and reconciliation procedure for the actual action path.

5. Cost gate: bound expected use and unexpected use

Cost readiness is a workload-planning exercise, not a promise that a particular monthly amount will hold. Identify the drivers of spend, the limits that matter to the business, and the response when traffic or processing differs from the forecast. The estimate should include the agent’s surrounding application and tool path where those costs are relevant, rather than treating model consumption as the whole service.

  • Estimate demand by task type and user population. Record expected task counts, peak periods, and the proportion of work that may require clarification, repeated processing or human review. Use business volume assumptions that an owner can challenge. Separate a normal planning case from a plausible high-demand case so that the release does not rely on an average month to stay within its operating envelope.

  • Measure the work required per completed task. In an appropriately controlled evaluation, record the number of model interactions and the approximate input and output volume needed for representative tasks. Include retries and unsuccessful requests in the estimate. A task that fails after substantial processing still consumes resources, and an agent that repeatedly seeks extra context can cost more than its successful path suggests.

  • Use current limits for the actual account and deployment. Quotas vary by resource, model and deployment, so check the current limits that apply to the account and planned configuration rather than relying on a universal number (quotas and limits). Record who will recheck those constraints when demand or deployment assumptions change.

  • Set an internal operating envelope and escalation point. The business owner should define acceptable consumption for the planned workload and what signal prompts investigation, throttling, a narrower release or a pause. Use current internal and vendor-side cost information for the actual estimate; this checklist does not prescribe a price or universal threshold. Make the decision owner and the response to an approaching limit explicit.

  • Test how the workflow behaves under constrained capacity. Decide what the user sees when the service cannot complete work in the expected time, what work is safe to defer, and whether an operator can identify requests that remain pending. Avoid designs that conceal resource pressure by silently dropping work or repeatedly submitting the same action.

  • Review cost after changes in use or behavior. Revisit the estimate when task volume, prompt context, tool usage, retry behavior or the release population changes materially. Review both total consumption and consumption per completed, verified task. A fall in cost per request is not necessarily an improvement if more requests are failing or being handed off.

For an illustrative calculation, assume—not as a forecast—that a pilot expects 80,000 tasks in a month, averages 1.7 model interactions per task, and plans a 10% allowance for retries and rework. The planning volume is 80,000 × 1.7 × 1.10, or 149,600 interactions. The team would then apply its current account-specific rates and the relevant supporting-service inputs to that workload. If the actual task mix doubles the average interactions, the estimate should be recalculated rather than defended as a fixed budget. The useful result is a visible assumption chain, not a fabricated price.

6. Worked release review: a hypothetical service-desk workflow

Consider a hypothetical organization preparing an agent to help employees resolve access requests. It can look up a request, gather missing information and prepare a permitted status change for an operator. The release team has not approved automatic changes to account privileges. This boundary makes the example useful: it distinguishes information gathering from a higher-impact action without claiming any particular Foundry configuration or outcome.

The business owner defines a successful task as identifying the correct request, presenting the relevant evidence and leaving a clear next step. Engineering lists the permitted read operation and the one status change the agent may propose. Security reviews which user roles can initiate the workflow and which runtime identity is used for the lookup. The team also confirms that the route to the lookup dependency works from the intended environment, and that an unavailable dependency leads to a visible handoff rather than a guessed answer.

For quality review, the team prepares cases for a valid request, an ambiguous employee identity, a missing approval detail, an already-closed request, a lookup timeout and a request asking for an unapproved privilege change. Before reviewing outputs, the process owner records what the agent may say and do in each case. The ambiguous identity should stop the workflow until clarified. The unsupported privilege change should be refused or routed through the established human process, not transformed into a status update that appears close enough.

For action control, suppose the business process permits an operator to approve a specific status change. The proposed design requires the agent to prepare the exact target record and exact change, then obtain a human decision tied to that payload. The approval must expire, and any change to the target or action requires a new decision. The system should not treat an old approval as permission to execute a modified proposal. After execution, a separate check confirms the status in the business system before the task is reported as complete.

Operations owners then define what happens if the external system times out after receiving the change. They do not instruct the agent to repeat the request automatically when the result is unknown. Instead, the process checks the target record, establishes whether the change occurred, and only then decides whether further action is needed. The operator can pause new changes while the team investigates, and the user receives a truthful pending or unresolved status rather than an unsupported success claim.

For the cost plan, the team estimates task volume separately for lookups, clarifications and human handoffs. It uses measured interaction counts from its own controlled evaluation and current account-specific inputs for the estimate. If the pilot handles more requests but a higher fraction require repeated clarification, the team examines both total use and completed verified tasks before expanding access. This example shows how the four gates share a boundary: the permitted action determines what must be tested, monitored, approved and costed.

The release record for this hypothetical case would include the scoped task, allowed and prohibited actions, test cases and decisions, identity/network matrix, operational pause procedure, cost assumptions, and named sign-offs. A checked box without that evidence would not resolve the underlying question. A documented gap can be acceptable only when the owner has narrowed the release or explicitly accepted a bounded risk within their authority.

7. Counterexample and failure analysis

A weak release proposal might say: “The agent is internal, the network is private, all sample questions passed, and the model reports that it completed the update.” Each clause sounds reassuring, but none establishes the relevant release outcome. Internal users can still have different permissions. Private inbound access does not establish outbound isolation. Sample questions may omit ambiguous targets. A model’s statement does not prove that a downstream record changed correctly.

Imagine that a request reaches the wrong employee record because two people have similar names. The agent prepares a status change, a reviewer approves it without seeing the exact target, and the external system accepts the change. A timeout prevents the application from recording a confirmation, so an automatic retry submits the same action again. The user receives a success message because the agent generated one, while the operator has no clear signal that the target and final state need checking.

The corrective review is not simply “add a human.” It asks what the human sees, what decision is bound to the action, whether the approval expires, how a changed payload invalidates approval, and how the business result is independently verified. It also asks how a timeout is reconciled, whether repeated execution is safe, and which owner can pause further changes. If any answer is missing, the team can reduce authority: allow lookup and drafting only, require a person to perform the change, or defer the release until the control path is complete.

A different counterexample is an apparently excellent average quality result that hides one high-consequence failure. If almost every routine request is handled correctly but the agent sometimes changes the wrong record, a simple pass percentage can obscure the decision. Keep high-impact error categories visible and define blocking criteria before testing. The release is judged against its business risk, not by an average that masks a failure the owner has declared unacceptable.

8. Make the release decision and preserve the record

Use this flow to organize the decision. It describes a proposed release process, not a claim about a built-in platform feature. The stop path is intentional: a failed access check, invalid action proposal, missing approval or unverified outcome should not be converted into a successful task by wording alone.

flowchart TD
accTitle: Agent release decision and action flow
accDescr: A request passes access and scope checks, then an exact action is validated and approved before execution. The outcome is verified. Any failed gate goes to stop and repair rather than proceeding.
    A["Request"] --> B["Check identity and scope"]
    B -->|"Pass"| C["Validate exact action"]
    B -->|"Fail"| G["Stop and repair"]
    C -->|"Valid"| D["Approve exact payload; approval expires"]
    C -->|"Invalid"| G
    D -->|"Approved"| E["Execute action"]
    D -->|"Denied or expired"| G
    E -->|"Result available"| F["Verify business outcome"]
    E -->|"Unknown result"| G
    F -->|"Verified"| H["Close task"]
    F -->|"Mismatch"| G

The diagram contains eight nodes, so it stays within the requested seven-node maximum? No: the flow above includes eight named nodes. To keep the diagram within the seven-node limit, combine the request and identity check as one entry node and retain the same stop and verification branches:

flowchart TD
accTitle: Agent release decision and action flow
accDescr: A request passes identity and scope checks, then an exact action is validated and approved before execution. The outcome is verified; failures stop for repair.
    A["Request: check identity and scope"] -->|"Pass"| B["Validate exact action"]
    A -->|"Fail"| G["Stop and repair"]
    B -->|"Valid"| C["Approve exact payload; approval expires"]
    B -->|"Invalid"| G
    C -->|"Approved"| D["Execute action"]
    C -->|"Denied or expired"| G
    D -->|"Result available"| E["Verify business outcome"]
    D -->|"Unknown result"| G
    E -->|"Verified"| F["Close task"]
    E -->|"Mismatch"| G

Before closing the review, assemble a decision record that identifies the exact release boundary, configuration revision, test evidence, accepted limitations, operating owners and open actions. Record which issues block release and which are explicitly deferred by an accountable owner. Set a review trigger for material changes to permissions, tools, data, workload or business authority so that approval does not silently outlive the system it assessed.

  • Security decision recorded: identity and permissions are understood, data and network paths have owners, and unresolved security issues have a release disposition.
  • Quality decision recorded: acceptance criteria and high-impact failure cases have been reviewed, and success claims depend on business-result verification where applicable.
  • Operations decision recorded: monitoring, pause, incident ownership, timeout reconciliation and rollback or compensation responsibilities are usable by the people assigned to them.
  • Cost decision recorded: workload assumptions, current account-specific limits, operating envelope and escalation owner are documented.
  • Release authority recorded: the business and technical owners have made the decision for the stated scope, and any restriction or deferred capability is visible to users and operators.

A release should proceed only when the evidence supports the actual authority being granted. If the evidence supports drafting but not execution, release drafting. If it supports one user group but not another, keep the narrower audience. If an action cannot be verified or reconciled, remove that action from scope until the process can handle it safely.

Printable release worksheet

Copy this section into the team’s normal review record or print it for a release meeting. It is a worksheet inside this article, not a separate downloadable file. Add links to the actual evidence in the organization’s approved record system; avoid attaching sensitive payloads where a controlled reference is sufficient.

Release identification

  • Agent/workflow and configuration revision: ______________________________
  • Business task and intended outcome: _____________________________________
  • Included users, environment and permitted actions: ______________________
  • Explicit exclusions and known assumptions: ______________________________
  • Business owner / technical release owner: _______________________________

Gate decisions

  • Security evidence reviewed; identity, data, network and action boundaries have named owners.
  • Quality criteria and high-impact failure cases reviewed; unresolved cases have a disposition.
  • Operational signals, pause procedure and unknown-outcome reconciliation reviewed.
  • Workload assumptions, current account limits and cost escalation response recorded.
  • Business outcome verification is distinct from the agent’s statement that work is complete.
  • Release scope and any restrictions are visible to the people who will use and operate it.

Decision and follow-up

  • Decision: [ ] Approve stated scope [ ] Approve with recorded restriction [ ] Hold
  • Blocking issues and owner: ______________________________________________
  • Accepted limitations and approving owner: _______________________________
  • Review trigger or next review date: _____________________________________
  • Business owner decision/date: ___________________________________________
  • Technical owner decision/date: _________________________________________

If the team needs help turning these boundaries into an operable Azure design, engage enterprise platform engineering for a focused implementation discussion. Keep the requested work tied to the open release gates rather than treating a broad platform review as a substitute for a decision on this agent’s actual scope.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call