Overview: compare the operating design, not the product names
Microsoft Foundry and Amazon Bedrock AgentCore are often shortlisted for the same reason: a team wants software that can interpret a request, consult business information and initiate a bounded action. Their names alone do not tell you which will fit. The useful comparison is whether each option can support the same workload inside the controls your organization is prepared to operate.
That means defining the work before comparing platform features. Specify who can submit a request, which information is in scope, what the system may propose, what a person must approve, and how the organization will confirm the business action succeeded. Then compare the implementation and ownership required to uphold those boundaries.
This article focuses on Foundry and AgentCore, not on general agent-platform selection or a tour of every hyperscaler. It uses one hypothetical internal case-management workflow to make the comparison concrete. The scenario is illustrative: it does not represent a PADISO deployment, product test, or vendor benchmark.
The two options are not exact equivalents at every architectural layer. Foundry supports distinct approaches for prompt-based agents and custom-code hosted agents: hosted agents package custom code as containers, while prompt agents use declarative prompts and tools. Microsoft’s hosted-agent overview describes those distinctions; it does not establish built-in traffic splitting, so any canary design needs a separate routing mechanism.
AgentCore is described as modular agent infrastructure, and its runtime is a different layer from an agent framework. The AWS Prescriptive Guidance overview is useful for keeping those concepts separate. The practical implication is that the team still needs to decide what framework, application logic, identity boundaries and operational procedures sit around the infrastructure.
These limited product distinctions matter, but they are not a procurement verdict. A team can choose a managed platform and still build a fragile agent, or choose a less opinionated arrangement and make its interfaces easier to change. Compare the whole system that you intend to run, including the parts the vendor does not own.
Define one workload for both candidates
Imagine a US-based distributor whose support staff receive requests to review delayed shipments. A staff member enters an order reference and a short explanation. The agent may retrieve shipment status and the relevant service policy, draft a response, and propose a bounded update to a case record. It must not issue a refund, change a delivery address, or send a customer message without an authorized person confirming the exact action.
For this comparison, assume the distributor already runs its order and case systems in AWS, but its identity team wants to assess both Foundry and AgentCore. This is a deliberately awkward but realistic choice: AWS data gravity favors keeping the workflow near its existing systems, while organizational standards or a broader cloud strategy may still make Foundry a candidate.
The intended result is not an autonomous support representative. The system should return a structured draft containing the order reference, the evidence it used, the proposed case update, and any uncertainty that requires staff attention. A separate application presents the draft for review. The human decision is recorded before the application submits an action to the case system.
The workflow has five boundaries to compare. First, staff identity must determine who can submit a request. Second, retrieval must be restricted to records the staff member is entitled to see. Third, the agent must propose rather than silently execute consequential changes. Fourth, the action service must verify the approved payload and the caller’s authority. Fifth, the business application must confirm whether the case update actually happened.
Those boundaries are more useful than asking whether either platform can produce a convincing answer. A fluent draft can still disclose the wrong record, propose an invalid action, or leave the case system in an uncertain state. The selection exercise should therefore include access denials, tool failures, duplicate submissions and incomplete business outcomes—not only a successful demonstration.
A concrete acceptance test could use a staff member with access to one regional queue and two sample orders. The agent should retrieve only the permitted order, include the policy version used in its rationale, and produce a draft for the correct case. A second account without that regional entitlement should not obtain the same record merely by supplying its order reference.
A separate action test should present a draft containing a case identifier, status change and customer-facing note. The reviewer approves that exact content. If any field changes after approval, the action must return for review rather than treating the earlier approval as reusable. This is an application design requirement for the comparison, not a claim that either platform supplies it automatically.
Feature-by-feature comparison: control and implementation boundaries
1. Agent construction and the framework boundary
For Foundry, first decide whether the proposed agent is a prompt-and-tools configuration or custom code packaged as a hosted agent. That choice affects where application logic lives and which team owns changes. Hosted code can make existing engineering practices relevant, while prompt-oriented configuration can suit a simpler declarative workflow. Neither label removes the need to define tool inputs, validate outputs or test the business rules around an action.
For AgentCore, avoid treating the runtime as if it were the complete agent application. The distinction between modular infrastructure and an agent framework is important: the runtime boundary and the framework boundary answer different design questions. Your team must identify which component manages the agent’s reasoning loop and state, which component exposes tools, and which component applies business authorization.
The comparable criterion is not the number of abstractions in either diagram. It is whether engineers can point to the owner of each decision. If a model proposes an order update, one named application boundary should validate the proposal, and a separate trusted path should verify approval before making the change. Do not let a platform choice blur responsibility between orchestration, business policy and execution.
A useful evaluation exercise is to ask both implementation teams to sketch the same request path, then label every component that can read data or cause an external effect. If a team cannot identify where an input is validated, where a decision is recorded, or how an action is stopped after a timeout, the design is not ready for a platform scorecard.
2. Identity and authorization
Foundry should be assessed against the identity model the organization already operates, including how an authenticated staff request becomes a narrowly authorized data lookup and a separately authorized action. AgentCore should be assessed against the equivalent AWS-side boundary. Do not infer that a platform’s cloud association automatically enforces the application’s user-level permissions; test the end-to-end identity path your design actually proposes.
For the distributor example, the agent process should not have broad, standing permission to read every case or update every shipment. The retrieval service should receive a scoped request and return only records allowed for that staff member and task. The execution service should independently decide whether the approved change is permitted. This prevents a prompt instruction from becoming the authority to access or modify business records.
The strongest comparison question is where authorization is enforced and what evidence is retained. Ask each team to demonstrate a denied access request, an expired staff session and an attempted update to a case outside the user’s region. Record which component blocks each request and how an operator can later reconstruct why.
A control that exists only in a prompt is not an access-control boundary. Prompts can guide behavior, but application services should reject unauthorized identifiers and action fields even if the agent produces them. This is a design recommendation for either platform, not a platform-specific feature claim.
3. State, memory and request isolation
A support workflow needs some state: the original request, retrieval references, draft response, reviewer decision and action result. The key decision is which of those values must persist, for how long, and which should be recomputed. Persisting everything indefinitely increases exposure and makes stale information harder to distinguish from current facts. Persisting too little can make a retry unsafe or an investigation inconclusive.
Treat the business record as the authority for case status. Agent conversation context may help compose a draft, but it should not become the authoritative source for whether an update succeeded. Store a durable action record in the application’s state boundary with a request identifier, a digest or version of the approved payload, the reviewer identity, approval time, execution attempt and confirmed business result.
Isolation matters if the workflow serves multiple companies, business units or customer tenants. A request identifier must not be enough to retrieve another tenant’s conversation or action record. Define the tenant boundary at each data access point, including logs, caches and any shared task store. For additional design considerations on this separate problem, see Multi-Tenant MCP Servers: Tenant Isolation That Holds.
Compare the platforms by mapping the state you intend to keep to the system that owns it, not by assuming a platform’s agent memory should own business truth. A practical design can keep workflow state in the application while treating each agent invocation as a bounded step. That makes retries and audits easier to reason about even if the agent implementation changes.
4. Tool execution and approval
The agent should return a proposed action as data, not execute a consequential side effect simply because it generated a plausible instruction. For the case workflow, the proposal can contain a case identifier, a requested status, a short note and references to the evidence used. The application validates the schema, checks the current case version and presents the exact payload to an authorized reviewer.
Approval must bind to that exact payload and expire. If the note, status, case identifier or relevant underlying case version changes, the earlier approval no longer authorizes execution. The execution service should re-check current authority immediately before the action. These controls reduce the chance that an approved draft is altered or replayed in a different context.
Neither an approval screen nor a successful tool response proves the business result. The case system might accept a request and then fail to commit it, or a network timeout may leave the caller unsure whether it did. The workflow should query the authoritative case record after execution and show a confirmed result, a confirmed failure, or an unresolved state requiring operator review.
Design for at-least-once attempts and uncertain outcomes; do not promise exactly-once external effects. If the case system supports a safe idempotency mechanism, the application can use it according to that system’s contract. If it does not, the execution boundary needs a reconciliation procedure that detects whether the intended change already occurred before retrying.
5. Observability and responsibility
A useful trace is not just a record of model text. It should let an operator connect the staff request to the retrieved evidence, the generated proposal, the approval decision, the execution attempt and the business result. Keep sensitive content out of logs where it is not needed, and make the identifiers sufficient to connect events without copying an entire customer record into every diagnostic system.
Compare how your proposed implementation will answer a specific incident: a staff member says a case was changed without approval. The investigation needs to establish which identity submitted the request, what exact payload was approved, whether it changed before execution, which execution attempt occurred, and what the case system ultimately stored. If those facts are split across teams or unavailable, the platform selection has not solved operational accountability.
Do not confuse agent traces with business outcome verification. A trace can show that a tool call was attempted; the case system’s authoritative record determines whether the update took effect. For a deeper treatment of connecting a request with its business outcome, read Tracing a Foundry Agent from User Request to Business Outcome. This comparison keeps the point narrow: whichever platform is chosen, the application must retain a trustworthy path from proposal to verified result.
For shared tool ownership on Azure, A Shared MCP Tool Layer in Foundry: Ownership, Versions and Access covers versioning and permissions around the tool interface.
For an Azure deployment, Tenant Isolation for Enterprise Agents on Azure follows tenant scope through identities, runtime access and data boundaries.
An AWS-specific reference design for the same workload
The following is a proposed reference design for the hypothetical distributor, not an assertion about a packaged AgentCore architecture or a tested implementation. It separates identity, state and service boundaries so a team can compare the AWS-hosted option with a Foundry proposal using the same control requirements.
flowchart TD
accDescr: Workflow stages and decisions: Staff identity, Request gateway, Agent runtime, Scoped retrieval, Draft and state, Human approval, Action service. The adjacent text explains the conditions and exceptions.
accTitle: Foundry vs AgentCore — Compare the Same Workload and Control Requirements workflow
A["Staff identity"] --> B["Request gateway"]
B --> C["Agent runtime"]
C --> D["Scoped retrieval"]
C --> E["Draft and state"]
E --> F["Human approval"]
F --> G["Action service"]
G --> E
The flow begins with an authenticated staff identity entering through a request gateway. The gateway applies request validation and passes a user and tenant context to the application. The agent runtime receives only the task context it needs. It requests evidence through a scoped retrieval service and returns a draft; it does not receive a direct route to modify the case system.
The draft and review state live in an application-controlled state boundary. That boundary records the proposed payload and its version, then presents it for human approval. The action service receives an approved payload, checks that the approval still matches, and submits the bounded change. The state boundary records the action result after the application checks the authoritative case record.
The arrows represent the intended control flow, not a required vendor topology. A production diagram might add monitoring, a queue or a separate policy service, but those additions should respond to a concrete requirement. Avoid expanding the diagram merely to imply maturity. Each extra service adds an owner, an interface and a failure mode that someone must operate.
In the AWS-specific comparison, place the proposed agent infrastructure inside the account and service boundary selected by the organization, and define the identity trust path before connecting it to business data. Use the organization’s chosen AWS identity and permission controls to scope each service role; do not grant the agent an all-purpose credential that can read and write business systems. The application gateway, state store, retrieval service and action service should have deliberately distinct duties.
A Foundry design should be drawn with the same boxes and acceptance tests. Replacing the agent-runtime box does not change the need for scoped retrieval, approval-bound payloads, a durable action record and verification against the case system. That is how the comparison avoids giving one candidate credit merely because its cloud matches the hypothetical data location.
Worked scenario: what happens on a delayed shipment request
Consider staff member Jordan, who can view orders for the western region. Jordan enters order W-1842 with the request to explain a delay and update the internal case if the policy permits it. The gateway authenticates Jordan and passes the request to the workflow with the relevant tenant and regional context. The design should reject a request that lacks those attributes rather than asking the model to infer authorization from the text.
The retrieval boundary checks Jordan’s access before returning the order status and applicable policy excerpt. It returns only the fields needed to draft an explanation. The agent proposes a concise response and a case status change, and identifies the policy reference. It does not receive permission to alter the record. If the source data conflicts—for example, the order system says delayed while the case record says delivered—the workflow should mark the discrepancy for staff review rather than choosing one source silently.
The application validates that the proposal names the requested order and a permitted status. It displays the proposed note, evidence references and any unresolved discrepancy. Jordan can approve, edit or reject it. If Jordan edits the note, the application records the new payload version and requires approval of that version. A change between review and execution invalidates the approval.
After approval, the action service checks that Jordan still has the required authority, that the approval is current, and that the case has not changed in a way that makes the update unsafe. It submits the change and then reads the case record to confirm the resulting status. The interface tells Jordan whether the change is confirmed, rejected or unresolved. A timeout must not be presented as success.
For an illustrative capacity calculation, assume 40 staff members each submit 25 requests per workday, and each request triggers one retrieval and one proposed case action. That is 1,000 requests and up to 1,000 action attempts per workday before retries or exceptions. This is arithmetic from stated assumptions, not a forecast, price estimate or performance claim. The useful design question is whether review staffing, service limits and reconciliation capacity can handle the assumed volume with an exception queue that does not grow indefinitely.
A platform comparison should make the same scenario concrete in both candidate designs. Ask each implementation team to show what data crosses each boundary, what is persisted, what can act, and how an unresolved result is handled. Then estimate staffing and operating work from your own request patterns; do not treat a model-token estimate as the full cost of the workflow. For a focused cost discussion, see What an AI Agent Actually Costs on Bedrock AgentCore.
Pros and cons by option
Microsoft Foundry
Pros. Foundry gives a team a meaningful choice between prompt-oriented agents and custom-code hosted agents. That distinction can help match the implementation to the team’s preferred division of work: declarative configuration for a bounded prompt-and-tools design, or containerized custom code where the application needs a code-centric implementation. It also gives evaluators a concrete question to resolve early: which parts belong in the agent configuration and which belong in the surrounding application?
A second advantage is that teams can evaluate Foundry without pretending the agent boundary is the whole system. The same request gateway, retrieval policy, approval record and outcome verification can be retained in application services. This makes it possible to compare Foundry’s role in the design while keeping business-control criteria consistent across candidates.
Cons. The choice of construction style can create divided ownership if platform operators, prompt authors and application engineers assume someone else is responsible for validating tool inputs or maintaining business rules. A team can also mistake a hosted agent for a complete deployment strategy. If it wants to route a small portion of production traffic to a new version, that routing must be designed outside the built-in behavior described in the source; do not assume the platform provides traffic splitting.
Foundry may also be a poor fit when the organization’s essential records, execution services and engineering operations are firmly AWS-centered and there is no independent reason to introduce another control plane. That is not a claim that Foundry cannot be used in such a setting. It is a warning that a technically workable option can still add an operating boundary whose value must justify its coordination and data-flow cost.
Amazon Bedrock AgentCore
Pros. AgentCore’s modular framing encourages the team to distinguish agent infrastructure from its framework and business application. That is useful for a workload whose most important controls—authorization, persistent state and action execution—belong in services the business already operates. It can also make a fair comparison possible: evaluate the infrastructure layer without assuming it supplies the full reasoning framework or application policy.
For the hypothetical distributor, keeping the proposed design near AWS-hosted business services may reduce the number of organizational boundaries the team needs to cross. This is an architectural hypothesis based on the stated AWS-centered scenario, not a guarantee about latency, cost, availability or deployment effort. Confirm it by mapping actual data paths and ownership rather than relying on the cloud label.
Cons. Modularity moves decisions onto the engineering team. Someone must select and maintain the framework, define state ownership, construct the request and approval flow, and connect operational evidence across components. The runtime should not be mistaken for the framework, and neither should be mistaken for the business authorization layer. A team with limited platform capacity may find the freedom burdensome rather than useful.
An AWS-centered design can also deepen dependence on existing AWS conventions and make later migration more involved if the application becomes tightly coupled to infrastructure-specific interfaces. Keeping business rules, action schemas and durable workflow state in application-owned contracts can reduce that dependence, but it does not make the underlying platform interchangeable by itself.
Operational failure analysis and a counterexample
The common failure is not necessarily a spectacular model error. It may be a timeout after the action service submits a case update. The agent or gateway retries, the case system receives a duplicate attempt, and the user sees two notes—or the first change succeeds but the interface reports failure. Design the action path to record attempt identifiers and reconcile the authoritative record before retrying. If the outcome cannot be established, put the request into an operator-visible unresolved state.
A second failure begins with a permissions mismatch. The staff member is authorized to open the support application but has no access to the particular order. If retrieval uses a service identity without applying the staff member’s scope, the agent may receive data the user could not otherwise see. The relevant control is not a stronger prompt; it is authorization enforced at the data access boundary, with a denial that the application can explain without exposing the protected record.
A third failure concerns stale approval. A reviewer approves a draft, but the case changes before execution. If the action service only checks that some approval exists, it may apply an outdated decision. Bind approval to the payload version and relevant source-record version, then require a fresh review when either changes. The system should make the invalidation visible rather than quietly regenerating a proposal and treating the old approval as current.
A fourth failure is poor operational evidence. The agent response looks correct, and the tool call is logged, but no one can establish whether the case system committed the update. A trace of the model’s intention is not confirmation of the business effect. The workflow needs a result check against the system of record and a way to distinguish confirmed success from a pending reconciliation.
A useful counterexample tests whether the workload actually needs an agent. Suppose the request is only to show the current shipment status and the applicable policy paragraph. If a deterministic application can retrieve those fields and render a standard response, adding an agent may introduce variability without adding meaningful decision value. In that situation, neither Foundry nor AgentCore is the first decision; simplify the workflow and reserve generative behavior for drafting or exception handling that demonstrably benefits from it.
Another counterexample is a highly variable support request where staff need to explore several records and make judgment calls, but no external action should be automated. A human-facing assistant that prepares evidence and a draft may be appropriate, while an elaborate action-execution pipeline would add controls for capabilities the business has deliberately excluded. Match the platform design to the allowed scope rather than assuming that a more agentic implementation is the objective.
Decision artifact: a side-by-side evaluation worksheet
Use the following worksheet in a design review. Fill it in separately for Foundry and AgentCore using the same workload, users, data and acceptance tests. Record a named owner and evidence for each answer; an untested assumption should remain marked as an assumption rather than being scored as a capability.
| Decision area | What the team must write down | Evidence to request |
|---|---|---|
| Work boundary | Inputs, allowed outputs and prohibited actions | A request and response example, including an out-of-scope request |
| Identity | How the staff identity and tenant scope reach retrieval and execution | A successful request and an unauthorized request through the proposed path |
| State | Which component owns drafts, approvals, attempts and final results | A state diagram and a retention decision for each record type |
| Approval | How approval binds to the exact action and expires | A test where the payload changes after review |
| Execution | How the service validates, submits and reconciles a change | A timeout scenario and the behavior before any retry |
| Evidence | How an operator connects request, proposal, approval and business result | A trace sample using synthetic data and a verified record outcome |
| Operations | Who handles exceptions, updates and unresolved actions | An escalation path with an owner for each unresolved state |
| Portability | Which contracts remain application-owned and what must be rebuilt | A list of platform-specific components and a credible replacement boundary |
| Workload fit | Why an agent is preferable to a deterministic workflow | A comparison against the simplest non-agent implementation |
Do not compress these answers into a single weighted score before the teams have agreed on the hard constraints. A strong average can conceal a fatal gap, such as inability to enforce the organization’s user-level access boundary or no acceptable way to reconcile a timed-out action. First identify pass-or-fail controls, then compare the remaining trade-offs.
For a printable summary, copy this short decision record into the architecture review: Workload: what the agent may do. Data owner: authoritative records and access boundary. Action rule: what requires human approval. Failure rule: how uncertain outcomes are resolved. Platform choice: Foundry or AgentCore, with the reason. Revisit trigger: a concrete change in data location, control requirements, operating ownership or workload volume that would justify reviewing the decision.
A useful acceptance threshold is not a vendor feature count. It is whether the candidate design can pass the same representative tests: authorized retrieval, denied retrieval, stale approval, duplicate attempt, conflicting records and confirmed business outcome. Teams should define expected behavior before running those tests so that a polished demo does not redefine success after the fact.
Verdict: choose by operating fit
Choose Foundry when the team has a clear reason to use its agent construction options and can assign ownership across configuration or hosted code, application policy and business execution. It is a stronger candidate when the proposed design’s engineering and operating model fits the organization, not merely because a feature list appears broad. Require the team to show how any production rollout routing will work externally if a canary is part of the plan.
Choose AgentCore when the organization wants its modular infrastructure to sit within an AWS-oriented design and has the engineering capacity to specify the framework, state and service boundaries around it. This is a reasonable direction for the hypothetical distributor to assess first because its records and action services are assumed to be AWS-centered. That preference remains conditional on the identity path, approval binding and outcome reconciliation passing the same tests as the Foundry design.
Choose neither yet if the organization cannot explain who may retrieve each record, where the approval is stored, or how a timed-out action is reconciled. Those are unresolved system requirements, not minor implementation details to postpone until after procurement. If neither team can demonstrate a safe, auditable boundary for the action, start with a read-only assistant or a deterministic workflow and return to platform selection when the control design is ready.
For teams with mixed cloud estates, do not force portability at every layer. Keep business action schemas, authorization decisions and workflow records explicit and application-owned where that supports your operating model. Accept platform-specific implementation where it has a clear benefit, but document what would need to change if the workload moved. This is more practical than promising cloud neutrality that the design has not earned.
The next step is a short, bounded architecture exercise: have both teams map the same request path, run the same failure scenarios with synthetic data, and document ongoing owners before comparing commercial proposals. If you need help aligning these boundaries with a wider engineering platform, cloud platform engineering can be a relevant next step. Keep deeper production-readiness and regional-recovery work distinct: A Production Readiness Review for Microsoft Foundry Agents and Disaster Recovery for Azure Agents: What Must Survive a Region Failure? address those separate questions.
At higher demand, Scaling Foundry Agents: Plan for Quotas, Queues and Backpressure connects quota constraints to queues and backpressure.