Browser agents often need to act inside applications that employees already reach through single sign-on (SSO). That does not mean an agent needs an employee’s password, a permanent browser profile, or access to every application available to that employee. The safer design is to let an authorized person establish access through the organization’s normal sign-in process, then give the agent a narrowly scoped browser session for a defined task.
This guide shows how to design that boundary. You will decide who delegates access, how a session is created and contained, what the agent may do, and how to prove that the session has ended. The emphasis is session design—not how to build a browser agent generally, assess its answers, or approve consequential transactions. For the broader implementation context, see AI Agents in Production: Browser-Use Agents and Designing Agents for Human Handoff.
Prerequisites
Before changing an application or agent, identify the system owner, the employee or service identity that will delegate access, and the exact task the agent is meant to perform. Write the task as an observable outcome, such as “read the status and due date of ticket 4812,” rather than “use the support portal.” The narrower description will help you set boundaries and determine whether the resulting session is appropriate.
You also need a test account or other approved non-production path, a browser environment you can isolate from everyday user activity, and a way to observe the task’s outcome. Do not use a coworker’s real account as a shortcut to a test. If the application cannot provide a safe test environment, restrict initial trials to read-only activity and obtain the relevant application owner’s direction before attempting changes.
Finally, decide who is responsible for starting and ending the delegated session. A human may sign in and launch a task; an orchestrator may create an isolated browser context; an operator may review the outcome. Those responsibilities can belong to different people or components. They should still be explicit, especially when the application’s normal session can reach data beyond the intended task.
Warning: Do not put an employee’s password, recovery code, or one-time sign-in code in a prompt, configuration file, task transcript, or browser-agent instruction. The point of SSO delegation is to keep those credentials with the person and sign-in system that handle them, not to move them into the agent’s reach.
Step 1: Define the delegation boundary
Start by writing down what is being delegated. A useful boundary statement names the human or organizational identity, the application, the purpose, and the allowed operations. For example: “A support operator delegates a session for the agent to read the current status of one named ticket and return its displayed due date.” This is more useful than “the agent may use support,” because it gives the operator and the implementation team something concrete to test.
Separate identity from authority. The employee’s account establishes who is signed in; it does not automatically make every action available through that account appropriate for the agent. If the application exposes many projects, customer records, or administrative actions, decide whether the agent can be restricted through the application, a separate account, a dedicated browser environment, or a combination of controls. If none of those measures can meaningfully narrow reach, do not describe the session as narrowly delegated merely because the task prompt is narrow.
Record the session’s intended lifetime as part of the boundary. A task that should take minutes should not depend on a browser profile that remains available indefinitely. Define when a session begins, what event ends it, and who can stop it early. The end condition should be operational—for example, after the requested record is read and the result is checked—not simply “when the agent is done,” which leaves completion and cleanup ambiguous.
For each allowed action, also write at least one prohibited action. If the task is to read a record, examples might include editing it, opening unrelated records, or following a link into a different workspace. These limits do not make an agent immune to error. They make the intended boundary inspectable, and give you a reason to stop or redesign a task that routinely crosses it.
Step 2: Choose how the person delegates access
The most straightforward model is human-started access. The person signs in through the organization’s ordinary SSO flow in a controlled browser environment, confirms the intended application and task, and then starts the agent. In that arrangement, the person handles sign-in; the agent receives a usable browser session, not the password. The human-started model is often easier to understand and review, but it does not by itself limit what the signed-in account can reach.
A second model uses a dedicated identity for a defined operational purpose. That can separate agent work from an employee’s ordinary account, but it creates account lifecycle and access-management work of its own. A dedicated identity is not automatically least-privileged: if it has broad access, or no clear owner, it can become a persistent route into the application. Assign responsibility for its use and review its actual application reach rather than relying on the label “agent account.”
A third possibility is a delegated access mechanism offered by the application or the organization’s identity setup. Whether that is available, how it works, and what it permits depend on the systems involved; do not assume that a particular SSO configuration supports it. Ask the application and identity owners to confirm the real behavior before designing around it. If the only workable route is for an employee to sign in interactively, build the task around that constraint instead of scripting credential entry.
Choose the least complicated model that gives you a verifiable boundary. Human-started access is a reasonable first design for a short, supervised task. A dedicated identity may fit a recurring workflow when its access can be kept appropriately narrow and its lifecycle owned. If the application provides no way to separate the intended task from broad account access, the right result may be to narrow the workflow further or decline to automate it.
Step 3: Establish a session without handing over credentials
Have the authorized person complete sign-in through the normal SSO experience in the browser environment intended for that task. Keep credentials and authentication factors out of the agent’s instructions and accessible artifacts. The agent should begin work only after the person or surrounding workflow can determine that the intended application is open and the required signed-in state is present.
Do not make a successful-looking page load the sole test of sign-in. A browser can show a sign-in page, a timeout notice, a partially loaded application, or an account-selection screen that resembles a ready state. Define a visible readiness condition tied to the task: the expected workspace is displayed, the intended identity is apparent where the application shows it, and the requested record or starting point is available. If the condition is not met, pause and return control to the person rather than guessing at a sign-in step.
An implementation that reuses saved browser state deserves particular care. Playwright’s authentication guidance notes that stored browser state can include sensitive cookies or headers and advises separating accounts and sessions. Source Treat any stored state as a sensitive credential-bearing artifact: limit who and what can read it, keep it out of source control and ordinary task logs, and avoid sharing it across identities or unrelated work. These are practical handling recommendations; the precise protections available depend on the environment you build.
Where a browser workflow needs a session across several actions, keep the session associated with one task and one intended identity. Avoid copying the state into a second worker simply because that is convenient, or preserving it for a later task without a clear reason. Reuse can blur the boundary between the original delegation and subsequent work, making it difficult to say which human authorized which action.
Step 4: Isolate the browser session
Use a distinct browser context or an equivalently isolated environment for the delegated task. The design goal is that the agent’s task cannot casually inherit tabs, signed-in accounts, downloads, or session state from an employee’s everyday browser. A fresh, task-specific context also makes it easier to attribute what happened to a particular delegation.
Isolation needs to cover the full path the agent can use, not just the first tab. Decide whether opening a new tab, following a link, or navigating to another application is allowed. If the task should stay within one application and one record, make those limits part of the agent’s task boundary and the surrounding controls. A prompt alone is not a substitute for a technical or operational restriction when the consequences of wandering are material.
Keep parallel work from becoming accidental shared access. If two tasks run at once, give each a clearly separate identity and session context where feasible. Otherwise, the second task may act under the first task’s account, or one task may inherit state changed by the other. When separate contexts cannot be provided, serialize the work and make the limitation visible to operators; do not claim that concurrent tasks are independent.
The implementation should also decide where browser outputs go. Screenshots, downloaded files, traces, and logs can reveal records or session details even after the visible browser is closed. Keep only what the task needs for review, restrict access to retained artifacts, and set a practical deletion or retention rule. The key distinction is between the live session and the evidence you choose to keep: closing one does not automatically remove the other.
Step 5: Give the agent a bounded task
Describe the task in terms of a target, permitted actions, and a stopping condition. For a read-only task, that might mean opening a specified ticket, reading two named fields, and stopping after returning those values. Avoid broad instructions such as “look around for anything important.” Broad goals invite unnecessary navigation and make it harder to tell whether the agent stayed within the delegated purpose.
Specify what the agent should do when the page does not match expectations. Examples include a missing record, an unexpected account, a permission error, a changed layout, or a request for additional verification. The safe response is to stop and hand control back, not to invent a workaround or ask the user for their password. A predictable stop rule is especially valuable when a human has already completed sign-in and may not be watching every browser action.
Keep consequential actions outside the scope of a task unless they have been designed and authorized separately. This article does not define an approval process for payments, refunds, or other irreversible actions. For those decisions, use a deliberately designed control path; approval gates for browser-based payments and refunds addresses that separate problem. Session delegation should not silently expand into permission to submit or change a consequential record.
Also distinguish the agent’s report from the application’s state. “The agent says it updated the field” is not evidence that the field now contains the intended value. Even for a read-only task, verify that the returned details came from the intended record and session. For a distinct treatment of checking browser-agent outcomes, see Benchmarking Browser Agents: Check the End State, Not the Transcript.
Step 6: Define the end of the session
Before launch, state the event that ends the delegation. It might be a completed read-and-report task, an explicit operator stop, or a timeout chosen for the workflow. The important property is that the session does not remain available simply because no one remembered to close it. If a task pauses for human input, decide whether the browser remains open, is stopped, or must be re-established after the person returns.
At task end, close the task’s browser context or otherwise remove the agent’s access to it. If the environment uses saved session state, ensure that the task cannot keep using a copy after the visible browser has closed. The right cleanup steps depend on the implementation and identity setup; validate them with the responsible application and platform owners rather than assuming that closing a window revokes all access everywhere.
Plan explicitly for an interrupted task. A worker can stop unexpectedly, an operator can close the browser, or the application can time out in the middle of a workflow. Decide what the system does in each case: stop the agent, mark the task incomplete, notify the responsible operator, and avoid automatically continuing from an uncertain state. If you cannot tell whether the task has ended or whether its session can still be used, treat that uncertainty as a reason to prevent further work until it is resolved.
Where the application offers a way to end or revoke the relevant signed-in access, include that in the cleanup procedure when appropriate. Do not claim that a local browser action revokes an organization-wide session unless the identity and application owners confirm it. In particular, a local context ending and a server-side sign-in ending are different events; decide which one your risk and operating model require.
Step 7: Verify identity, scope, and outcome
Verification should answer three separate questions. First, was the session associated with the intended identity? Second, did the agent stay within the allowed application and task? Third, did the task produce the requested business result? A single transcript or completion message cannot reliably answer all three. Choose evidence that is suitable for the question: for example, an operator check of the displayed account, a record of the task’s starting and ending points, and an independent view of the final requested fields.
Use a small set of explicit acceptance conditions. For a read-only ticket lookup, those might be: the browser displayed the approved workspace; the requested ticket identifier matched; the agent returned the displayed status and due date; no edit or submission action occurred; and the delegated session was closed at the end. These are proposed design criteria, not a universal standard. Adjust them to the application and task, and make sure each condition can actually be observed.
Capture enough information to reconstruct a failure without collecting everything the browser saw. A useful event record could include a task identifier, the authorized identity label, the application name, task start and stop times, completion or stop reason, and whether the expected result was verified. Avoid placing passwords, authentication codes, reusable session material, or unrelated record contents in that record. If a screenshot or trace is needed, decide who can view it and when it should be removed.
Treat an unexpected navigation or identity change as a boundary failure, even if the agent eventually returns the requested answer. The task should stop, the event should be recorded at an appropriate level, and the session should be checked before any retry. If the same condition recurs, revise the session design or task scope; simply adding a more emphatic instruction may leave the underlying access problem unchanged.
Step 8: Exercise the design with a controlled fixture
A controlled browser fixture makes session decisions concrete without relying on a real customer record. Create a non-production page or approved test workflow with a small, known set of states: a visible signed-in identity, one target record, one unrelated record, a session-expired state, and a route that indicates access is outside scope. The fixture should reveal which state the browser is in without containing live credentials or sensitive business data.
The worked interaction trace below is illustrative. Assume a support operator has signed in to a test support portal and asks an agent to read the status and due date for ticket T-1042. The fixture displays the test account identity and a record with status Waiting and due date 2026-10-12. A separate ticket, T-1043, is unrelated. The agent’s task is to read only the named ticket and report those two fields.
| Trace point | Expected observation | Decision |
|---|---|---|
| Session opened | Test portal and intended test identity are visible | Continue only if both match |
| Target located | Ticket identifier is exactly T-1042 | Read only the requested fields |
| Result returned | Status is Waiting; due date is 2026-10-12 | Check against the fixture’s expected values |
| Boundary changes | Identity or ticket differs, or session-expired state appears | Stop and return control; do not guess |
| Task ends | Result has been checked and browser context is closed | Mark complete only after cleanup |
The trace is not proof that an implementation has been tested. It is a design artifact: a compact description of the observations that a future test or operator review should cover. When you adapt it, use test identifiers and expected values that are approved for your environment. Include at least one failure path, because a fixture that shows only a successful sign-in can conceal what the agent or operator does when the boundary is unclear.
Here is the same decision flow at a glance:
flowchart TD
accDescr: Workflow stages and decisions: Human requests bounded task, Confirm identity and scope, Expected signed-in session?, Stop and return to human, Run task in isolated session, Verify result and boundary, Close session and record outcome. The adjacent text explains the conditions and exceptions.
accTitle: Browser Agents and SSO — Designing Sessions Without Sharing Passwords workflow
A["Human requests bounded task"] --> B["Confirm identity and scope"]
B --> C{"Expected signed-in session?"}
C -- "No" --> D["Stop and return to human"]
C -- "Yes" --> E["Run task in isolated session"]
E --> F["Verify result and boundary"]
F --> G["Close session and record outcome"]
The first decision is whether the session matches the identity and application the delegation requires. If it does not, the flow returns control rather than trying to repair sign-in through the agent. If it does, the task runs inside its isolated session; verification checks both the result and the boundary before cleanup. Adjacent prose and operational rules still matter: the diagram does not claim that closing a browser context revokes access beyond that context.
Step 9: Plan for common failure patterns
The wrong account is already signed in. This can happen when a browser environment is reused or a person selects an unexpected account during sign-in. A task prompt cannot make the wrong identity correct. Stop before accessing the target record, return control to the operator, and re-establish the intended session in the proper environment. Record the mismatch without copying authentication material into the incident note.
The application asks for more sign-in steps. The agent may reach a sign-in, verification, or account-selection screen instead of the expected workspace. Do not pass a code or password into the agent to get past it. Hand back to the authorized person, who can decide whether to complete sign-in and re-launch the task. If this happens repeatedly, make the human handoff part of the expected workflow instead of treating it as an exceptional nuisance.
The session appears to cross a boundary. An unexpected tab, workspace, or record may mean the task scope is too broad, the application has redirected, or the environment reused state. Stop the task and inspect what was reached before retrying. For risks arising from untrusted page content, see Prompt Injection in Browser Workflows: Treat Web Pages as Untrusted; this guide’s focus remains the identity and session boundary.
The browser closes but access may remain. A local close is useful cleanup, but it may not answer whether any reusable state or broader signed-in session remains available. Consult the owners of the actual identity and application setup and test the cleanup procedure in an approved environment. If you cannot establish what remains usable, do not carry that session forward to another task.
A task appears to finish after an interruption. A completion message might arrive after the operator has stopped the visible browser, or a worker might resume from an uncertain point. Treat the task as incomplete until the session state and requested result are independently checked. Avoid an automatic retry that might repeat an action; although this guide centers on session boundaries, retries can still create confusing or duplicate effects in systems that change state.
The fixture passes, but a real workflow behaves differently. A test page can confirm that the intended decisions are understandable; it cannot establish that every production sign-in, timeout, or application route behaves the same way. Introduce real workflows only after the relevant owners confirm the scope and the task has an appropriate test or supervised rollout path. Do not infer a guarantee from a successful fixture trace.
Step 10: Turn the design into an operating worksheet
Use this worksheet before an initial rollout and when a task’s identity, application, or intended action changes. It is a printable aid within the article, not a substitute for the application owner’s decisions. Mark a field “unknown” rather than filling it with an assumption; unknowns about who is signed in, what the agent can reach, or how access ends are reasons to pause.
Delegation
- Task outcome: Write one observable result the agent is meant to produce.
- Delegating identity: Name the person or dedicated identity that establishes access.
- Application and starting point: Specify the approved application and the task’s entry point.
- Allowed actions: List the minimum actions required; separate reading from changing data.
- Prohibited actions: Name at least one action or destination that is outside the task.
Session handling
- Sign-in method: Confirm that the person uses the normal approved sign-in route; credentials and one-time codes stay outside the agent’s instructions.
- Isolation: Identify how this task’s browser state is separated from unrelated work.
- Parallel work: State whether concurrent tasks use distinct sessions or must be serialized.
- Start condition: Define a visible, verifiable condition that means the intended session is ready.
- Stop condition: Define completion, early stop, and interruption behavior.
Verification and cleanup
- Identity check: State how an operator or implementation confirms the intended identity.
- Outcome check: Identify the application state or record fields that demonstrate the requested result.
- Boundary check: Specify what evidence would reveal an unexpected account, workspace, or record.
- Artifact handling: Name the owner and handling rule for logs, screenshots, downloads, or saved session state.
- Session close: Record what local cleanup does and whether any separate end-of-access action is required.
- Escalation path: Name who resolves an unexpected sign-in prompt, scope mismatch, or uncertain session state.
A worksheet is ready for operational use only when the answers refer to observable behavior. “The agent should know which account to use” is not a start condition; “the expected test identity is visible in the approved portal” is. “Close the session when done” is not an interruption plan; “stop the task, mark it incomplete, and require an operator to check the session before retrying” is. Converting vague expectations into evidence makes handoffs between engineering, operations, and application owners less dependent on memory.
Summary: delegate a session, not a password
Begin with a specific task and identity, then decide whether human-started sign-in or a dedicated identity gives the right boundary. Establish access through the approved SSO flow without placing credentials in the agent’s reach. Isolate the task’s browser state, set explicit stop conditions, and verify the identity, scope, and result separately. Close the session deliberately and account for the artifacts left behind.
For a small, supervised workflow, the controlled fixture and worksheet provide a practical starting point. For a recurring workflow, review how the chosen identity is managed, how concurrent tasks remain separate, and what happens after interruptions. If you cannot determine which identity is active or whether a session remains usable, stop rather than assume the boundary held. Teams planning an implementation can discuss their requirements through AI workflow automation without treating a particular session design as universally suitable.