Start with the boundary, not the tool
An agent browser or code interpreter gives an application a way to perform work beyond generating text. A browser can visit pages and interact with web content; a code environment can process inputs and produce files or results. These capabilities are useful precisely because they operate on information and systems outside the model’s own response. That also means their safety depends on the boundaries around them, not on the apparent harmlessness of a prompt.
A sandbox is an execution environment separated from other systems by limits on what it can access or affect. It reduces the consequences of mistakes, malicious content, or unexpected code. It does not, by itself, decide whether a task is permitted, make a user’s credentials safe to hand to a tool, or determine whether an output should be exported. Those are separate controls that the surrounding application must design.
AWS documents a managed Browser Tool for AgentCore. In the architecture proposed here, treat browser execution and business authorization as separate responsibilities, and verify the supported isolation configuration. Treat that distinction as a design constraint: a managed session does not replace the application’s decision about who may do what. AWS Browser Tool documentation
AWS describes Code Interpreter as using a managed isolated environment. Our design recommendation is to authorize data exports separately from that execution boundary. In practical terms, containment and permission are different questions: where code runs is not the same as what data it may receive or where its outputs may go. AWS Code Interpreter documentation
This article focuses on those boundaries: sandboxing, credentials, network access, and artifacts created or retrieved by browser and code work. It does not attempt to cover workload scaling, end-to-end tracing, or agent memory design. Those need their own decisions; see scaling agent workloads on AWS, tracing agents across systems, and designing retention and tenant boundaries for agent memory for those distinct subjects.
A useful mental model is a workshop with a locked tool cabinet, a limited workbench, and a controlled loading bay. The workbench is the sandbox. The cabinet represents credentials and data access. The loading bay represents network routes and artifact export. A safe design asks what each task needs at each boundary, rather than giving every task the same broad access because the tools are convenient.
Four boundaries to make explicit
The first boundary is execution: which work can run in the browser or code environment, and what the environment can reach while it runs. A task that summarizes public pages has different needs from one that manipulates a private business application. If both use a shared, broadly connected environment, the simpler configuration can create a much larger blast radius than either task requires.
The second is identity: which application identity requested the work, and which narrowly scoped permissions the application delegates to it. Identity is not just a username attached to a session. It includes the authority represented by credentials, tokens, roles, or other access material. A browser session logged in as a powerful employee may carry more authority than the agent’s task requires.
The third is network access: where requests can go and which destinations can return data. Network access covers both outbound connections and the consequences of receiving content from a remote system. A site can present misleading instructions, redirect a request, or return a file. A design that checks only whether a destination is reachable misses what the returned content can cause the agent to do next.
The fourth is artifact handling: how inputs, downloads, generated files, and extracted results move between the sandbox and business systems. An artifact is any persistent output or input file associated with the task, including a spreadsheet, report, image, archive, or generated script. Once an artifact leaves the isolated environment, it may be retained, shared, executed, or treated as authoritative by another system.
These boundaries should be expressed separately in a task contract. For example, “summarize these three public pages” can specify approved URLs, no authenticated browsing, text-only output, and a short retention window. “Prepare a draft reconciliation from these two uploaded spreadsheets” can specify the exact input files, allowed transformations, output schema, and a destination that does not post transactions. The second task needs more data access, but it still does not need permission to commit financial changes.
A design should also distinguish the tool’s capability from the task’s authorization. Capability describes what a component can technically do. Authorization describes what it is permitted to do for a particular request. If the tool can browse while the task is only authorized to inspect a fixed set of public pages, enforce the narrower task scope outside the model’s discretion.
A reference flow for an AWS application
The following is a proposed reference design, not a claim about a built-in AgentCore integration or a tested deployment. It keeps the application responsible for authorization and treats browser or code output as untrusted until checked. The design can be implemented with AWS identity and storage services selected to fit the organization’s existing platform, but the exact service configuration and supported tool settings must be verified for the chosen environment.
flowchart TD
accDescr: Workflow stages and decisions: Task request, Authorize scoped work, Run in sandbox, Collect candidate artifact, Validate output and destination, Approve external effect, Store, export, or quarantine. The adjacent text explains the conditions and exceptions.
accTitle: AgentCore Browser and Code Execution — Designing Safe Boundaries workflow
A["Task request"] --> B["Authorize scoped work"]
B --> C["Run in sandbox"]
C --> D["Collect candidate artifact"]
D --> E["Validate output and destination"]
E --> F["Approve external effect"]
F --> G["Store, export, or quarantine"]
accTitle: Proposed boundary flow for agent browser and code work accDescr: A task is authorized before sandbox execution. The resulting artifact is validated, any external effect is approved, and the artifact is then stored, exported, or quarantined.
The request enters through an application service that knows the authenticated caller, tenant, task type, and allowed business operation. The service converts that context into a bounded task description. It should not ask the model to invent its own permission scope from a natural-language request, because natural language may be ambiguous and untrusted content can influence the model’s interpretation.
The authorization step checks the requested task against application policy. It can reject the task, narrow its scope, or produce a scoped execution record. That record should identify the tenant, permitted inputs, allowed destinations, data classification, expiry, and whether an external side effect is permitted. It should be structured enough for software to enforce, not merely appended as a sentence that the agent is expected to obey.
The sandbox performs only the work needed for the authorized task. The application should keep credentials and privileged business operations outside the general-purpose execution environment unless a specific, reviewed design requires otherwise. Browser and code work should not automatically inherit the application service’s broad identity simply because that service initiated the task.
The resulting candidate artifact passes through validation before it is placed in a durable location or sent to another system. Validation can check expected file type, size limits, schema, tenant association, malware scanning where appropriate, and whether the output contains data outside the task’s permitted scope. The right checks depend on what the artifact will be used for; a text summary and a script destined for execution do not deserve the same treatment.
If a result would cause an external effect—such as publishing content, changing a customer record, or submitting a transaction—an approval must precede that effect. The approval should bind to the exact payload and destination, and expire if it is not used promptly. If the payload changes, the previous approval no longer applies. This avoids a vague authorization to “proceed” being reused for materially different actions.
The final disposition is explicit: store the result in a controlled location, export it to an approved recipient, or quarantine it for review. “The model said it succeeded” is not a disposition and is not evidence that a downstream business system accepted the result. Verify the business outcome independently through the receiving system or a suitable application check.
Sandboxing: containment is not a permission model
A managed isolated environment can reduce exposure between execution sessions or between execution and the surrounding application, depending on its supported configuration. The important design question is not whether a sandbox exists, but what isolation properties are available and how they map to the task. Confirm the actual configuration options and boundaries for the environment in use rather than assuming that a generic “sandboxed” label covers storage, network, process, or session separation.
Separate work by trust level and data sensitivity. A public-web research task should not automatically share the same inputs, output location, or credentials as a task processing confidential customer files. Likewise, code that transforms an uploaded spreadsheet should not inherit access to every internal dataset just because another agent workflow needs that access. Where the platform’s isolation controls cannot express the separation you need, add boundaries in the surrounding application or use a different execution design.
Treat task inputs as potentially hostile, even when they originate from an employee. A file may contain formulas, embedded content, misleading instructions, or structures that trigger unexpected behavior in a parser. A webpage can contain text designed to redirect an agent’s behavior. The practical response is to constrain the operation: limit the input set, use a narrow task instruction, avoid passing unrelated secrets into the same context, and validate outputs before they affect other systems.
A counterexample illustrates why “it ran in a sandbox” is insufficient. Suppose a code task receives a confidential workbook, calculates a summary, and then writes a second file containing both the summary and hidden source columns. The execution may have been isolated from the host, yet the exported file still discloses fields that the recipient did not need. The failure occurred at the data-selection and export boundary, not because the sandbox failed to contain the process.
Design for lifecycle as well as isolation. Decide when temporary files are removed, what task state must persist, and which outputs become durable records. Do not preserve every intermediate file simply because it may be useful for debugging. If a workflow genuinely needs a temporary artifact for a retry, give it a defined owner, expiry, and access scope. A retry should not silently revive an expired authorization or expose the previous task’s inputs to a new tenant.
Credentials: give the task authority, not a password drawer
A credential is any material that allows a system to act or obtain protected information: passwords, API tokens, session cookies, signed requests, or role-based access. The safest default is not to place broad credentials into prompts, code inputs, downloadable files, or shared task state. Those surfaces may be logged, copied, summarized, or accidentally included in an output.
Begin with the user’s requested outcome and identify the minimum authority required. A read-only report should not receive write access merely because the underlying account has it. A task that needs one record should not receive an unrestricted export of a whole tenant’s records. If the application can retrieve the permitted data itself and provide only the necessary fields to the sandbox, that often creates a clearer boundary than giving the sandbox direct access to the system of record.
Where delegated access is necessary, keep delegation task-scoped and short-lived. The application should bind the authority to the authenticated caller, tenant, operation, and resource set, then prevent it from being reused for a different task. Do not treat a model-generated statement such as “I will only use this token for the requested page” as enforcement. The enforcement point must be a component that can constrain the actual access.
A practical separation is to distinguish the control identity from the task identity. The control identity is used by application components to authorize and coordinate work. The task identity represents the limited authority available to a specific operation. A control service may be able to request a scoped operation, but the browser or code task should not automatically receive all of that service’s permissions. This separation limits the consequences of a compromised input or mistaken tool action.
For multi-tenant applications, attach tenant identity to every input and output record, and verify the association at each transfer. Do not rely on a filename, folder name, or model summary as proof that an artifact belongs to a particular tenant. A service can check a stable task identifier and tenant field before it accepts an output, then reject mismatches rather than attempting to infer the intended owner.
Credentials also have a failure timeline. A task may start while a credential is valid, pause, and resume after the user’s access has been revoked. If the application treats the original authorization as permanent, a delayed tool call can outlive the intended permission. Recheck expiry and revocation-sensitive conditions at the point of use, particularly before a state-changing operation. A pending approval should expire too, and a changed payload should require fresh authorization.
Network access: control both the destination and the response
Network restrictions are strongest when they reflect the task’s purpose. A research task may require access to a known set of public sites. A document transformation may need no network connectivity at all. A workflow that reaches an internal business service should use an explicit, narrow path rather than general access to internal networks. These are design recommendations; validate what the selected execution environment can enforce and add controls in the application or network architecture where needed.
An allowlist is a set of destinations the task may contact. It is more useful when it identifies the actual resources and protocols required, rather than permitting broad domains for convenience. A wildcard domain can include unexpected subdomains, and a permitted site can redirect to another destination. Think through redirects, embedded resources, downloads, and links discovered in page content. A rule that allows the first URL but ignores the eventual connection is not a complete destination policy.
Do not assume that an agent’s description of a destination is reliable. It may say it is visiting an approved site while following a link supplied by page content. The application can pass a constrained set of URLs, check destinations outside the model, and refuse requests that escape the task’s scope. When an operation requires a human-selected destination, preserve that choice as structured input rather than inferring it from a generated explanation.
Inbound content matters too. A response from an allowed website is still untrusted input. It may contain inaccurate data, malicious instructions, or a download that is irrelevant to the task. Separate the retrieval step from the decision to act on retrieved content. A page can provide evidence to summarize; it cannot grant itself authority to access another system or to export a file.
Be especially careful when a task combines browser access and code execution. A browser can retrieve a file that code then processes, and code can produce a URL or payload that another component uses. Each handoff should carry the task identity and allowed scope. Without that continuity, one component may enforce a narrow policy while a later component interprets the artifact as trusted and performs a broader action.
A useful failure test is to follow the task through an unexpected redirect, a missing destination, a timeout, and a partial download. Define whether the task fails closed, retries within the same scope, or pauses for intervention. Avoid “retry with broader access” as an automatic recovery strategy. If the approved destination cannot be reached, the safer result is often a clear failure with the input and reason retained only as permitted by the task’s retention policy.
Artifacts: treat files as data with a chain of custody
Artifacts cross boundaries more easily than permissions do. A file can be created in an execution environment, downloaded from a browser, copied into object storage, attached to a ticket, or passed to a downstream service. At each transition, ask who can read it, who can alter it, how long it persists, and what the recipient will assume about its trustworthiness.
Use separate locations or logical namespaces for temporary inputs, candidate outputs, and approved records. A candidate output should not become a business record merely because it was successfully written. Keep the transition explicit: validate its association with the task, inspect the expected structure, and decide whether it is eligible for the next system. A file that fails validation should have a defined quarantine or deletion path, not remain in an ambiguously shared folder.
Give artifacts stable metadata: task identifier, tenant, creator or requesting identity, creation time, classification, source references where appropriate, expiry, and validation status. Metadata should be assigned or checked by application code, not accepted solely from fields that a model or uploaded file supplies. This makes it easier to reject a cross-tenant artifact or an output that has been detached from its authorization context.
Limit what is exported. If a report needs three columns, do not export the entire source workbook and rely on a recipient to ignore the rest. If code produces a generated script, treat that script as executable content, not as a harmless document. If a browser downloads a file, record its task association and validate it before downstream processing. The appropriate controls vary by file type and use; a file extension alone is not reliable evidence of content.
Retention should follow the artifact’s purpose. Temporary working files may need a short lifetime; an approved business record may need to follow the organization’s established retention process. Avoid retaining intermediate files indefinitely as a substitute for observability. If operational diagnosis requires a record, prefer metadata and bounded event details that do not unnecessarily preserve sensitive source contents.
Do not let a generated artifact choose its own destination or recipient. The task policy should determine where an eligible output can go. If a user asks to send a report to an address found inside a downloaded document, treat that as a new destination requiring authorization, not as an instruction embedded in the artifact. The same principle applies when a script proposes a network endpoint or a generated spreadsheet contains a link to an external service.
Worked example: preparing a supplier discrepancy report
Consider a hypothetical mid-market distributor whose operations team wants an agent to compare a supplier’s public product listing with a company-provided inventory extract and prepare a discrepancy report. The team wants a draft for a human analyst, not automatic changes to purchasing records. This example is illustrative; it does not describe a tested implementation or a product capability.
The request contains two inputs: a company-approved list of supplier URLs and a tenant-owned inventory extract. The application checks that the caller is permitted to request a report for that business unit. It narrows the inventory data to the product identifiers and fields needed for comparison, omitting unrelated customer or employee information. The task contract records the allowed URLs, permitted fields, report schema, destination, and expiry.
The browser work is limited to the approved supplier pages. The code work receives the retrieved values and the reduced inventory extract, then compares identifiers and selected attributes. The task does not receive purchasing credentials, permission to log into an internal procurement portal, or permission to post a purchase-order change. Those powers are unnecessary for a draft discrepancy report.
The candidate artifact is a structured report with fields such as product identifier, company value, supplier value, source URL, comparison status, and a short explanation. The application checks that each row refers to an expected product and allowed source, that the file matches the agreed schema, and that it does not include omitted fields. It then stores the report in a review location associated with the business unit and task identifier.
If the analyst later asks the system to update purchasing records, that is a separate operation. It requires a fresh authorization decision based on the proposed changes, and any approval must identify the exact records and values to be changed. Approval of the discrepancy report is not approval to alter inventory or submit a purchase order. If a supplier page changes between analysis and action, the business operation should use the verified data presented for approval rather than silently fetching new values and treating the old approval as applicable.
The scenario exposes a subtle tradeoff. A highly restricted workflow that cannot reach a supplier’s site may produce an incomplete report; a broadly connected workflow may expose internal data or follow an unapproved link. The appropriate decision is not to maximize either access or restriction in isolation. It is to give the task the narrowest useful inputs and routes, then make uncertainty visible—for example, marking an unavailable supplier value as unresolved instead of fabricating a match.
Now consider a counterexample: the agent downloads a supplier spreadsheet, processes it, and exports a report to a location supplied by a link inside that spreadsheet. The spreadsheet was an allowed input, but its embedded destination was not part of the approved task. If the output is sent there, the failure is an artifact-routing and authorization error even if browsing and computation were properly isolated. The application should reject the unapproved destination and keep the candidate report in its controlled location.
Failure timelines and operational decisions
Boundary failures often unfold over several steps. A user requests a legitimate report; a retrieved page contains instructions aimed at the agent; the agent follows a link outside the intended scope; a downloaded file is processed; and a generated output is placed in a shared location. No single step needs to look dramatic. The useful response is to make each transition explicit and record enough task metadata to identify where the scope changed.
A second failure sequence begins with a transient error. The browser cannot reach an approved page, so a retry is attempted. The retry uses a different destination or a broader credential because that seems likely to work. The task now has different authority from the one that was originally authorized. Recovery logic should preserve the original scope; if a broader route is genuinely required, stop and request a new authorization rather than expanding access silently.
A third sequence involves stale approvals. An analyst approves a specific report, but a delayed worker later acts on a modified payload. The approval no longer describes the action being taken. Bind approval to a payload representation and destination, set an expiry, and reject a changed payload. This is a design principle, not a promise of exactly-once effects: distributed systems can fail between sending a request and learning whether the receiving system accepted it. Reconcile the business outcome with the receiving system before retrying a potentially consequential operation.
Operational records should help answer concrete questions: which task was authorized, which inputs were accepted, which destination was permitted, what artifact was produced, which checks passed, and whether an external effect was approved. Avoid recording secrets or whole documents merely to answer those questions. The appropriate record is usually a bounded account of decisions and identifiers, with access and retention aligned to the data involved.
When a task fails, define the safe state for each boundary. A rejected destination should not trigger an automatic alternative. An invalid artifact should not enter the approved record location. An expired credential should not be replaced with a broader one without authorization. A timeout after a state-changing request should be treated as an uncertain outcome to reconcile, not proof that nothing happened.
The operational owner should be able to pause a workflow without deleting necessary business records or leaving a tool with continuing authority. Decide how to stop new tasks, revoke or expire delegated access where applicable, and quarantine pending outputs. Test those procedures with the actual deployment configuration before relying on them. A design document that names a kill switch but cannot identify which component enforces it is incomplete.
A practical decision worksheet
Use this worksheet when choosing whether browser or code work belongs in a given task. Fill it in for a task type, not for the platform in the abstract. The same application may make a different decision for public research, confidential document processing, and a state-changing workflow.
| Decision | Record for this task | Reject or redesign when |
|---|---|---|
| Purpose | The business result and who requested it | The desired result cannot be stated without granting open-ended discretion |
| Inputs | Exact files, fields, pages, and tenant | Inputs are broader than the result needs or have no verified owner |
| Execution | Browser, code, or neither; required isolation controls | Required boundaries cannot be confirmed or enforced |
| Identity | Caller, task authority, permitted operations, expiry | A shared or privileged credential is the only available option |
| Network | Allowed destinations, redirects, downloads, and fallback | The task needs unrestricted access without a bounded justification |
| Artifact | Output schema, location, metadata, retention, disposition | The output can choose its recipient or become trusted without validation |
| External effect | Exact payload, destination, approval condition, verification | Approval is vague, reusable after a change, or absent for a consequential action |
| Failure path | Timeout, partial output, rejection, uncertain completion | Recovery silently broadens scope or retries an uncertain effect blindly |
For each row, write the enforcement point beside the decision. “Only approved URLs” is a policy statement; the worksheet should identify which application or network component checks the destination. “No sensitive fields” should identify where fields are removed and how the output is checked. If the team cannot name an enforcement point, record the gap as unresolved rather than treating the intention as a control.
Then walk one task from request to final disposition. Include an ordinary success, a blocked destination, an invalid artifact, an expired authorization, and a timeout after a potential state change. For every branch, state whether the task stops, retries within the same scope, quarantines the output, or asks for a new decision. This exercise usually reveals more than reviewing a list of sandbox settings in isolation.
Finally, decide what evidence is needed to operate the workflow. A support engineer may need task identifiers, policy outcomes, validation status, and the reason an output was quarantined. They may not need the raw confidential workbook or a reusable credential. Define this before deployment so the organization does not solve diagnostic problems by retaining more sensitive content than the workflow requires.
From reference design to implementation
Implementation should begin with one bounded task and its data path. Identify the application component that authenticates the caller, the point that authorizes the task, the environment that performs browser or code work, and the service that accepts or rejects the resulting artifact. Draw the identity and data transitions separately. A single arrow labelled “agent” tends to hide which component holds authority and which component decides where output goes.
Confirm the selected environment’s supported isolation and configuration rather than assuming that a managed sandbox covers every boundary. Map each requirement to an actual control: input restriction, credential handling, destination enforcement, artifact validation, expiry, and failure behavior. Where a required control is not available at the tool boundary, decide whether the application can enforce it before or after execution; if neither can, the task is not ready for that design.
Keep the first release’s business effect modest. A draft report that a person reviews is easier to bound than automatic changes to a system of record. That is not a permanent rule against automation; it is a way to validate the data path and failure behavior before granting authority to cause changes. As the workflow matures, add permissions only when a specific business need and enforceable boundary justify them.
The platform decision should account for the surrounding identity, network, and artifact design, rather than comparing agent runtimes only by their model or tool labels. Teams evaluating the wider choice can use a comparison of Microsoft Foundry, Bedrock AgentCore, and Gemini Enterprise. If the architecture needs a review of AWS account boundaries, identity paths, and deployment controls, cloud platform engineering is a relevant next step.
For a broader implementation plan, estimate the operational work around execution, storage, review, and recovery as well as the tool call itself; the related analysis of what an AI agent actually costs on Bedrock AgentCore can help frame that separate decision. The key boundary decision remains concrete: which task receives which inputs and authority, through which routes, with what artifact checks, and what happens when any one of those assumptions fails.
A safe browser or code workflow is not defined by a claim that its environment is isolated. It is defined by the complete path from authorized request to controlled disposition. Keep task authority narrow, treat retrieved content and generated files as untrusted until checked, and make every external effect depend on a fresh, payload-specific decision. Where the design cannot enforce a boundary, reduce the task’s scope or redesign the path before relying on the sandbox.