SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Tutorial 17 mins

Playwright and AI Browser Control: A Hybrid Workflow Walkthrough

Playwright and AI Browser Control: A Hybrid Workflow Walkthrough. Practical examples, tradeoffs and implementation guidance for technology leaders.

The PADISO Team ·

What this walkthrough builds

This tutorial develops a hybrid browser workflow with a deliberately narrow division of responsibility: Playwright performs known login and navigation actions; an AI-assisted step may help interpret an unexpected page state; and deterministic checks decide whether the workflow may continue. The model is not the browser’s source of truth, and its interpretation is never treated as proof that a business operation succeeded.

The worked example is explicitly hypothetical. Imagine an internal operations team using a web portal to locate a pending request and open its details. The normal route is stable: sign in, open the request queue, search for a supplied reference, and verify the detail page. Occasionally, a notice or a revised layout interrupts that route. The design below pauses at that boundary, collects a bounded description of what is visible, and permits a model-assisted classification only when the page is still within a known set of safe states.

The example intentionally stops before approving, submitting, paying, or otherwise changing a record. Those actions would need their own policy, payload binding, and pre-execution confirmation. A model saying “the request is ready” is not evidence that a consequential action is safe or complete.

The useful outcome is not an autonomous agent that can improvise indefinitely. It is a traceable workflow with an ordinary path, a defined exception path, and a clear stop condition. You can adapt the same design to other low-risk navigation tasks by changing the page contract and the exception vocabulary, not by broadening the model’s authority.

Prerequisites and setup

Use a maintained Playwright project, a test or staging account, and a portal where automation is permitted. Keep the account’s permissions limited to the pages needed for the exercise. Use synthetic or otherwise approved records rather than a live customer transaction. Store credentials outside source control, and do not place passwords, session cookies, or unredacted customer data in prompts or logs.

You also need a model-access layer if you intend to run the optional interpretation step. This walkthrough does not prescribe a vendor, API, model, or hosting arrangement. Implement that layer behind your own application boundary: it should accept a small, redacted observation and return a constrained classification. If there is no approved model service, leave the exception path in manual-review mode; the deterministic path remains useful on its own.

Install Playwright using the project’s standard package manager and browser setup. The exact commands depend on whether the codebase uses JavaScript or TypeScript and how dependencies are managed. Do not copy credentials into a command, fixture, or committed environment file. Before running the browser, confirm that the base URL points to the intended test environment and that the test account cannot reach an unintended production workflow.

Use a separate browser context for each test run. Isolated tests make one run less likely to inherit cookies or page state from another, while checks should target observable behavior through resilient locators rather than fragile implementation details. These are core Playwright testing practices; see the Playwright guidance on resilient, isolated tests.

For the walkthrough, establish these configuration values outside the code: a test-only base URL, a username, a password, and a non-sensitive request reference that exists in the test environment. The examples below use environment-variable names as placeholders, not as a claim about any particular deployment. If a value is missing, fail before opening a page rather than silently falling back to a real account or a production URL.

Step 1: Write down the page contract

Before automating clicks, specify what the workflow must observe. A page contract is a short list of conditions that distinguish a valid state from a merely plausible screen. For login, that might mean a visible signed-in navigation element. For a search result, it might mean a row containing the exact requested reference. For the detail page, it might mean a heading and reference value that agree with the input.

Avoid defining success as “the button was clicked” or “the page stopped loading.” A click only records an attempted interaction. The application may reject the request, preserve the old page, display a validation message, or navigate somewhere unexpected. Web-first assertions can retry until their expected condition is met or a timeout occurs, which is useful for asynchronous interfaces; a click itself still does not prove a business outcome. See Playwright’s assertion guidance.

Write down what is allowed at the exception boundary as well. For example, the workflow could recognize a known maintenance notice, a known session-expired page, or an unfamiliar state. These labels have different consequences: a maintenance notice may be reported and stopped; an expired session may require restarting login; and an unfamiliar state should normally stop for human review. Do not group all three into a broad “recoverable” category.

A useful contract includes both positive and negative evidence. Positive evidence says what must be present, such as the exact request reference. Negative evidence says what must not be present, such as an authentication error or a page that asks for a different account. Negative checks prevent the workflow from interpreting an incomplete or misleading screen as a valid destination.

Step 2: Build the deterministic login path

Start with the predictable route and keep it explicit. The following illustrative TypeScript-shaped snippet shows the control flow, not a complete runnable project. Adapt locator text and page structure to the test portal, and confirm the locators against its actual interface. The code assumes that the environment has already supplied page, baseURL, username, and password through your test harness.

await page.goto(baseURL);

await page.getByLabel("Username").fill(username);
await page.getByLabel("Password").fill(password);
await page.getByRole("button", { name: "Sign in" }).click();

await expect(
  page.getByRole("navigation", { name: "Account navigation" })
).toBeVisible();

The example prefers labels and accessible roles because they describe user-facing controls rather than depending on incidental CSS classes. If the real portal has no usable label or role, first consider whether the interface can be improved. A CSS selector tied to a changing layout can be a practical fallback, but it should be reviewed as a dependency that may break when the page changes.

Treat the post-login assertion as a gate. If it times out, do not continue to the request queue and do not ask a model to infer whether authentication succeeded. Record a failure category, preserve only the diagnostics permitted by your data-handling rules, and stop. A login problem is not an invitation to let an agent explore arbitrary pages or retry credentials without a limit.

For a repeatable test, arrange cleanup so each run starts from a fresh context and does not depend on a previous run’s login state. If the portal uses a second factor or another interactive challenge, use the approved test-environment procedure; do not try to bypass it by extracting someone’s session or automating an unauthorized route. The example is about control boundaries, not circumventing access controls.

Expected result: the account navigation element becomes visible within the assertion timeout, and the workflow records that the authenticated page contract passed. If that result is absent, execution ends at login. The distinction matters operationally: a failure at this point should not produce a search result, a model-generated guess, or a misleading “completed” status.

Step 3: Navigate and verify the requested record

After login has passed its contract, navigate through the known user-facing path. The exact route and controls belong to the portal, so this illustrative code uses role-based labels rather than inventing a URL pattern or application-specific endpoint.

await page.getByRole("link", { name: "Request queue" }).click();
await expect(
  page.getByRole("heading", { name: "Request queue" })
).toBeVisible();

await page.getByLabel("Search requests").fill(requestReference);
await page.getByRole("button", { name: "Search" }).click();

const result = page.getByRole("row", { name: new RegExp(requestReference) });
await expect(result).toBeVisible();
await result.click();

await expect(
  page.getByRole("heading", { name: "Request details" })
).toBeVisible();
await expect(page.getByText(requestReference, { exact: true })).toBeVisible();

The second assertion on the detail page is not redundant. A heading can appear on a generic page, or a click can open a different record than the one requested. Matching the exact reference connects the input to the destination. If the interface has duplicate reference text, scope the locator to the details region and verify a second stable field that the business process considers identifying.

Avoid turning a missing result into a model search problem. If the exact reference is absent, the workflow should report “no verified matching result” and stop or follow a separately approved recovery path. A model may be able to interpret a page, but it cannot make a different record the right record. Fuzzy matching is particularly risky when references share prefixes or when users can enter similar names.

Expected result: the queue heading appears, a result matching the complete reference is visible, and the detail page displays that same reference. Store these as separate trace events, not a single opaque success flag. The sequence makes it possible to tell whether a later issue arose during navigation, search, or destination verification.

At this point the deterministic path has done its job. If the page contract passes, there is no benefit in invoking a model to re-describe a known screen. Keep model use conditional on a concrete exception; unnecessary model calls add latency, cost, and another component whose response must be handled without improving the verified result.

Step 4: Define the exception boundary

The model-assisted path begins only when a deterministic observation fails in a way that may be safely classified. It must not take over the browser by default. For this tutorial, allow the workflow to ask for one of three labels: KNOWN_NOTICE, SESSION_EXPIRED, or UNKNOWN. The labels are illustrative application-level values, not platform features.

First capture a bounded observation. Prefer a short, relevant text excerpt and a small amount of page context over a full-page dump. Remove credentials, personal data, account identifiers, and any content not necessary to distinguish the permitted states. Avoid sending screenshots or page contents to an external service unless the organization has explicitly assessed that data path and approved it.

A useful observation might include the current workflow step, the page title, a short visible message, and whether the signed-in navigation marker remains present. It should not include the password, hidden form values, session tokens, or unrelated page content. Attach a correlation identifier generated by your workflow so the classification can be tied to the run without exposing sensitive record details in routine logs.

The model’s response should be constrained and parsed as data. Reject missing labels, extra claims, malformed output, and any label outside the allowed set. Do not accept a free-form instruction such as “click the first button” as an alternative to a classification. The exception handler may select among prewritten application paths; it should not convert model prose directly into browser actions.

Here is a small flow of the intended control logic:

flowchart TD
  A["Run verified login and navigation"] --> B["Check page contract"]
  B -->|"Pass"| C["Record verified destination"]
  B -->|"Fail"| D["Collect bounded observation"]
  D --> E["Classify allowed exception"]
  E -->|"Known state"| F["Apply safe response or stop"]
  E -->|"Unknown or invalid"| G["Stop for review"]

  accTitle: Hybrid browser exception flow
  accDescr: Deterministic browser actions reach a page contract check. Passing records a verified destination. Failing permits a bounded observation and constrained classification; known exceptions take a predefined safe response or stop, while unknown or invalid classifications stop for review.

The diagram’s pass branch never calls the model. On the fail branch, classification changes only which predefined response is selected; it does not establish that the underlying page is correct. A known notice may lead to a clear stop message, while a session-expired classification may trigger a controlled restart only if the workflow has an explicit retry limit. An unknown state, a malformed response, or a disagreement between the label and the observed page stops for review.

Step 5: Connect classification to bounded actions

Keep the model call behind a function with a narrow contract. The following pseudocode intentionally leaves out vendor-specific request syntax: the approved model service and its interface depend on your environment. The important design feature is that the function returns a validated category, not browser instructions.

type ExceptionLabel = "KNOWN_NOTICE" | "SESSION_EXPIRED" | "UNKNOWN";

async function classifyException(observation: SafeObservation): Promise<ExceptionLabel> {
  const result = await yourApprovedClassifier(observation);
  if (!isAllowedLabel(result)) return "UNKNOWN";
  return result;
}

const label = await classifyException(observation);

switch (label) {
  case "KNOWN_NOTICE":
    await recordStatus("stopped_on_known_notice");
    break;
  case "SESSION_EXPIRED":
    await recordStatus("session_expired_review_or_limited_restart");
    break;
  default:
    await recordStatus("unknown_state_manual_review");
}

yourApprovedClassifier, SafeObservation, and isAllowedLabel are placeholders for application code, not named services or ready-made Playwright APIs. Define the input schema, output parser, timeout behavior, and logging policy in your own repository. In particular, the parser should reject a response that contains an allowed label plus an unapproved action request. Do not silently extract a label from arbitrary prose and discard the rest if that prose could influence downstream control.

A known notice does not necessarily justify an automatic retry. If it says the portal is temporarily unavailable, the safer response may be to stop and let an operator decide when to try again. A session-expired state may justify starting a fresh login attempt, but only after the page is confirmed to be the expected sign-in page and the attempt counter remains below a small configured maximum. A classification alone cannot authorize a retry against an arbitrary page.

Make the state machine explicit in code. Track the current step, the observation identifier, the selected label, the permitted response, and whether the workflow stopped or resumed. If the classifier times out or the model service is unreachable, treat that as an unknown state. The deterministic workflow should fail closed at the exception boundary rather than switching to an unbounded browser agent.

Expected result: each classifier outcome maps to a known, reviewable branch. The ordinary path remains deterministic; the exception path cannot issue a click, choose a record, submit a form, or claim task completion. A human operator can see why the run stopped without having to infer what the model meant.

Step 6: Produce a useful interaction trace

A trace is a compact record of state transitions, not a transcript of everything the browser saw. For each step, record a timestamp, a run identifier, a step name, the expected condition, whether that condition passed, and the resulting status. Include the request reference only if it is appropriate for the log’s access and retention rules; otherwise use a suitably protected or redacted correlation value.

A hypothetical trace might read: login_marker: passed; queue_heading: passed; exact_reference_result: absent; exception_observation: captured_redacted; classification: UNKNOWN; workflow_status: stopped_for_review. This is an illustrative trace, not a claim about a tested portal. Its purpose is to show what an operator needs: where progress stopped, what was verified, and whether a model classification affected the branch.

Keep model output separate from observed evidence. Store the normalized label and, where permitted, a reference to the redacted input; do not present the model’s explanation as if it were a page assertion. If an operator later reviews the run, they should be able to distinguish “the classifier returned a known notice” from “the page displayed the expected notice text.” The latter needs a deterministic observation of its own.

An interaction trace can also support diagnosis without becoming a data exhaust system. Record only the fields needed to reconstruct control flow and investigate a failure. Set retention and access according to the sensitivity of the portal and the organization’s policies. Avoid keeping raw HTML or screenshots by default; if a specific incident requires richer diagnostics, handle it as an explicit, controlled exception.

Step 7: Check failure behavior before expanding scope

Test the exception design with controlled fixtures in the test environment. At minimum, exercise the normal login and navigation route, an expected notice, an expired-session page, an unfamiliar page, a missing exact result, and an unavailable or malformed classifier response. Confirm that each fixture produces the intended branch and that none turns an unverified state into a completed task.

A normal-path check should prove the exact requested reference appears at the destination. A notice check should prove the workflow reports a stop and does not continue to a detail action. An expired-session check should prove that any restart is limited and begins only from a confirmed sign-in page. An unfamiliar-page check should prove that the browser does not follow model-proposed actions. A classifier outage should have the same safe terminal status as other unrecognized exceptions.

One counterexample illustrates why the boundary matters. Suppose the search result is missing, and a model sees a nearby row with a similar reference. An unsafe design might treat the row as a likely match and continue. A safer design records that the exact reference was absent and stops. The model can help explain that the page appears to contain a different result, but it cannot relax the identity condition established by the workflow.

Another failure occurs when a click succeeds technically but the application silently rejects the operation. This tutorial avoids consequential submission, but the lesson still applies to navigation: verify the destination’s observable condition rather than treating input delivery as completion. For any future operation that changes business data, require a separately designed precondition check and a post-action business-state check. Never infer success from a model’s summary or from a button click alone.

Treat repeated timeouts as a signal to improve the page contract or investigate the environment, not as an invitation to increase retries indefinitely. A short retry can be justified for a known transient condition, but retries need a limit and must not duplicate an external effect. This example does not provide exactly-once behavior and should not be used to claim that an action happened only once.

Step 8: Troubleshoot common problems

If the login assertion times out, check that the base URL and test credentials are correct, that the expected account navigation label matches the actual page, and that the test account can access the portal. Inspect the permitted test diagnostics to distinguish a validation error, access denial, loading issue, or unexpected page. Do not send credentials or session data to the classifier to diagnose the failure.

If a locator is unstable, confirm that it identifies the intended user-visible control and is not matching multiple elements. Prefer a specific accessible name or a locator scoped to the relevant region. When the interface changes, update the contract deliberately and rerun the fixture cases; do not simply loosen assertions until the test passes. A weak assertion can hide a wrong destination as effectively as no assertion.

If search returns no verified row, check the input reference for whitespace or data-entry mistakes, confirm that the test fixture exists, and inspect whether the portal requires a search submission or a filter change. The recovery path should still require an exact match before opening a record. Do not substitute a guessed result or let an AI-generated similarity assessment stand in for identity verification.

If the classifier returns an invalid label, times out, or produces extra text, record the classifier failure and stop. Check the input and output schema, the service configuration, and the timeout policy without broadening the allowed response. If the same page consistently appears but is classified as unknown, add a deterministic page check only after a maintainer has confirmed what that page means and what response is safe.

If a run appears successful but the trace lacks a destination assertion, treat it as unverified rather than reconstructing success from a screenshot or a model summary. Repair the trace and assertion sequence, then rerun in the test environment. A clean record of failure is more useful than a false completion that conceals which condition was never checked.

Step 9: Use the decision worksheet before adapting the pattern

Use this compact worksheet when deciding whether a new browser task fits the design. It is a working artifact to complete during implementation review, not a downloadable file or a substitute for your own operational controls.

DecisionRecord before enabling the workflow
Intended taskOne sentence describing the read or navigation task, without vague terms such as “handle the request.”
Starting stateThe exact sign-in or already-authenticated state the workflow expects.
Verified destinationObservable page conditions, including an exact record identifier where relevant.
Allowed exceptionsA short list of recognized states and the predefined safe response for each.
Model inputThe minimum redacted observation needed to distinguish those states.
Model outputA finite set of labels; no free-form browser commands.
Stop conditionsMissing identity evidence, unknown state, invalid response, timeout, or retry limit reached.
Trace fieldsRun identifier, step, assertion result, classification status, and terminal outcome.
Test fixturesNormal route, each allowed exception, unknown page, missing result, and classifier failure.

If the team cannot describe the destination in observable terms, the task is not ready for this hybrid pattern. If the permitted model output cannot be reduced to a short label, the exception boundary is too broad. If there is no fixture for a high-consequence failure mode, keep the workflow manual until that gap is addressed.

For a printable summary, copy this line into the implementation ticket and fill every field: Task: ___ | Start state: ___ | Destination proof: ___ | Allowed exceptions: ___ | Model input/output: ___ | Stop rules: ___ | Trace: ___ | Fixtures: ___. A blank field should prompt a design discussion, not an assumption that the omitted control is unnecessary.

Step 10: Decide what belongs outside this walkthrough

This workflow is intentionally narrower than a general browser-use agent. It does not decide whether a browser is the right integration method for a business process, design a general-purpose agent, compare tool protocols, or provide a complete regression-fixture program. Those are separate design questions. For production concerns around browser-use agents, see AI Agents in Production: Browser-Use Agents. For the separate choice between browser interaction and a system interface, see how to assess a legacy booking workflow.

Likewise, do not assume that exposing a tool interface changes the need for deterministic checks. Teams considering that layer can review how browser automation teams should distinguish WebMCP from MCP and the tradeoffs between MCP servers and REST APIs. This example stays focused on the interaction boundary inside one controlled Playwright flow.

When you adapt the pattern, preserve its core sequence: deterministic action, observable assertion, bounded exception observation, constrained classification, and a predefined response or stop. Expand one boundary at a time. First make the ordinary path reliable; then validate each exception; only afterward consider whether any exception warrants a limited restart or another controlled action.

If your team needs help designing an exception boundary, redacting observations, or connecting the workflow to operational review, explore AI workflow automation. The implementation should still be judged against its own page contracts, failure fixtures, and business controls rather than a promise that an AI-assisted browser can handle every variation.

The completion criterion for this tutorial is modest and concrete: a test run either verifies the intended destination or stops with an interpretable reason. The model may help classify a bounded interruption, but it never substitutes for the page evidence needed to continue. That division keeps the browser interaction understandable when the interface behaves as expected and, more importantly, when it does not.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call