SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Tutorial 19 mins

Browser Agent Regression Tests: A Website-Change Fixture Pack

Build a repeatable Playwright fixture pack for browser agents facing renamed controls, expired sessions and layout shifts—then verify the business state.

The PADISO Team ·

What this fixture pack is for

A browser agent can fail because a website changed, because its session expired, or because the automation itself behaved inconsistently. Those causes need different remedies. A useful regression test makes the website state reproducible, runs the agent against that state, and records what the agent actually changed. It does not treat a fluent explanation or a successful click as proof that the task completed.

This walkthrough builds a small, controlled fixture pack for a hypothetical supplier portal. The task is to save a draft purchase request. Four local page states cover the normal path, a renamed button, an expired session and a shifted layout. The examples are illustrative; they are not a tested integration with any particular agent or supplier system.

The fixture pack is deliberately narrower than an end-to-end benchmark. It does not rank models or estimate general task success. It gives an engineering team a repeatable way to answer a more immediate question: when a known website condition changes, does the agent still reach the intended state, stop safely, or fail in a way that can be diagnosed?

A page fixture controls the inputs the agent sees: visible text, page structure, session state, viewport and test data. The test harness controls the consequences of actions. In this example, a draft submission is intercepted and recorded in a test-only in-memory object rather than sent to a real supplier. That separation lets the test check both interaction behavior and business-state evidence without creating external effects.

The main boundary is important: the fixture should reproduce a condition, not tell the agent what answer to give. Do not expose labels such as renamed-button-case in the page content, or inject a hidden instruction that explains the expected action. Keep the scenario name in the test runner and the fixture manifest; expose only the page the agent would ordinarily encounter.

For implementation patterns that combine deterministic browser control and agent decisions, see a hybrid browser-control walkthrough. This article focuses on the changed-site test data and the evidence each fixture should capture, rather than on choosing an agent-control architecture.

Prerequisites and test design

You need a Node.js project, Playwright Test, a browser installed for that project, and a way to call the browser agent under test. The agent adapter is application-specific: it might accept a page, a browser session, or a remote-control handle. Keep that adapter behind one small test interface so the fixture definitions do not become entangled with a vendor SDK.

The examples below assume the test can provide an agent with a page and a task, then receive a concise run record. That record should include a status such as completed, blocked or failed, plus a trace or event log if your system makes one available. It is an interface contract for the sample, not a claim that any particular agent provides those fields.

Create a project and add the test dependency using your team’s normal package-management process. For a minimal npm project, the commands are:

npm init -y
npm install --save-dev @playwright/test
npx playwright install chromium

Add a test command to package.json, for example "test:fixtures": "playwright test tests/browser-fixtures.spec.ts". Use the browser and operating-system combinations your actual agent supports. For the first fixture pack, keep the matrix small: one browser, one fixed viewport, one test account state, and the four named page conditions. Expanding variables before the basic cases are understandable makes failures harder to interpret.

Playwright recommends testing observable user behavior with resilient locators and isolated tests, rather than depending on incidental implementation details (Playwright testing guidance). Apply that principle to your fixture assertions: verify visible page state and recorded business state, not a particular screenshot pixel or internal CSS class unless that implementation detail is itself what you are investigating.

Web-first assertions retry until their condition is met or the timeout is reached; a click alone does not establish that the business outcome occurred (Playwright assertion guidance). Accordingly, the tests below wait for a recorded draft and separately inspect the agent’s run status. Keep those assertions distinct: one checks what happened in the application, and the other helps explain the agent’s behavior.

Step 1: Define the fixture contract

Before writing HTML, define what each fixture represents and what counts as evidence. A fixture should have a stable identifier, a purpose, its starting state, the condition under test, and an expected safe outcome. The contract also needs an explicit list of effects that must not happen. For an expired session, for example, the correct outcome is to stop at the sign-in boundary; a recorded draft would be a failure, even if the agent reports that it completed the task.

Use a manifest that humans can review alongside the tests. The following example captures the essential distinctions:

Fixture IDPage conditionIntended observationExpected safe result
draft-baselineOriginal form and “Create draft” controlForm is availableOne draft recorded with the supplied reference and amount
draft-renamed-controlSame form; control reads “Save draft”Action label changedOne correct draft recorded; no duplicate
draft-session-expiredSign-in gate replaces the formTask is blocked by session stateNo draft; run records a blocked outcome
draft-shifted-layoutSame form moved into a narrow side panelPosition and surrounding layout changedOne correct draft recorded; no accidental neighboring action

Do not make the expected result vague. “Agent handles change” is not a useful assertion. In the renamed-control fixture, specify the exact request reference and amount that must be stored. In the session fixture, specify that no request is stored. In the layout fixture, specify the one business action that is permitted and the neighboring action that must remain untouched.

Choose inputs that make accidental success unlikely. For example, use a unique reference such as FX-260930-07 and a distinctive amount such as 137.25. These are illustrative values, not business recommendations. A second reference can be used in a later run to detect stale state leaking between tests. Avoid using a value that already exists in a shared test environment.

Keep the task instruction the same across the four cases: “Create a draft purchase request for reference FX-260930-07, amount 137.25, and leave it as a draft.” If the instruction changes along with the page, a passing result cannot tell you whether the page fixture or the prompt caused the difference. The task should name the intended outcome and relevant data, but should not tell the agent which button to click.

Step 2: Build local pages with controlled state

Create one small HTML page per condition, or render variants from a shared template with an explicit fixture key in test code. Separate files are often easier for reviewers to inspect because the changed text and layout are visible in a diff. The pages can submit to a fake endpoint that the test intercepts; they should not contain production URLs, real credentials or live account data.

The baseline page needs a heading, two labeled inputs, a submit button and a visible confirmation region. Give fields meaningful labels such as “Request reference” and “Amount”. The original button reads “Create draft”. On submission, the page sends the form values to a local test endpoint and displays the returned reference. Keep the application behavior deliberately small so the fixture tests the browser interaction rather than a large collection of unrelated business rules.

The renamed-control page should differ in one relevant respect: change the visible action to “Save draft”. Leave the field labels, values, structure and response behavior unchanged. That one-variable change helps attribute a failure. If the action text, form location and field labels all change in a single fixture, a failed run tells you that something broke, but not which condition mattered.

The expired-session page should present a clear sign-in boundary and omit the request form. Do not leave the form present but visually covered if the scenario is meant to represent an expired session; that can create ambiguity about whether an action was technically available. The page should show enough user-facing context for the agent to recognize that it cannot continue without authentication, and the test should verify that no draft reaches the fake endpoint.

The shifted-layout page should keep the same labels and controls as the baseline while moving the form to a narrow side panel or a different part of the viewport. Set a fixed viewport in the test so the layout is repeatable. Avoid artificial obstacles such as overlapping controls or impossible scrolling unless those are the explicit condition under test. The purpose is to detect fragility to a realistic repositioning, not to manufacture a broken page.

One useful variation is to keep semantic labels stable while changing visual position. Another, separate fixture can alter the accessible name while leaving the visual arrangement stable. Those test different dependencies: a layout shift tests how the agent locates an otherwise recognizable control; a changed label tests whether its action selection relies on outdated text. Add them as distinct cases if the distinction matters to your system.

Step 3: Intercept effects and run the agent

The harness should isolate external effects. In the example, a route handler receives a draft request and appends its payload to an in-memory array. The agent sees a local page, and the test can later assert what the page attempted to submit. In your project, adapt the URL and data shape to the application under test; do not point this interception at a production endpoint.

Here is an illustrative TypeScript outline. The agent fixture is intentionally an adapter supplied by your test project. It must connect the agent to the page in the way your system supports. The outline is not a complete executable project and has not been run; use it to define the separation between page state, agent invocation and outcome verification.

import { test, expect } from "@playwright/test";

const task =
  "Create a draft purchase request for reference FX-260930-07, " +
  "amount 137.25, and leave it as a draft.";

test("renamed control still creates the requested draft", async ({ page, agent }) => {
  const drafts: Array<{ reference: string; amount: string }> = [];

  await page.route("**/api/drafts", async route => {
    const payload = route.request().postDataJSON();
    drafts.push({ reference: payload.reference, amount: payload.amount });
    await route.fulfill({
      status: 201,
      contentType: "application/json",
      body: JSON.stringify({ reference: payload.reference, status: "draft" })
    });
  });

  await page.goto("http://fixture.local/draft-renamed-control");
  const run = await agent.run({ page, task });

  await expect(page.getByText("FX-260930-07", { exact: true })).toBeVisible();
  expect(drafts).toEqual([{ reference: "FX-260930-07", amount: "137.25" }]);
  expect(run.status).toBe("completed");
});

The route’s request capture is test instrumentation, not the definition of success by itself. It proves that the form submitted a particular payload to the fake endpoint. The visible confirmation and stored test record provide separate evidence that the fixture’s modeled application accepted it. In a real system, choose the narrowest trustworthy read-only signal available to confirm the intended state; do not infer business completion from the agent’s own summary.

For the baseline and layout fixtures, reuse the task and payload assertions. For the session fixture, assert that the page remains at the sign-in boundary, the draft array stays empty, and the run record indicates a blocked or otherwise non-completed result consistent with your adapter. Do not require a particular internal error string unless your product contract guarantees it; error wording is often less stable than the observable state.

Step 4: Verify the state transition, not the click

A browser interaction trace tells you what the agent attempted. A state assertion tells you whether the application reached the outcome the task required. Keep both because they answer different debugging questions. A trace may show that the agent selected “Save draft”; a recorded request may show that it submitted the wrong amount. Conversely, the correct draft may exist even if the agent’s final status message is poorly phrased.

For every success case, assert the complete payload that matters to the task. At minimum, the reference, amount and draft status should match the fixture’s expectation. If the form contains a currency or supplier identifier that affects meaning, include it too. Do not assert every incidental field simply because it exists; noisy assertions can make harmless presentation changes look like business failures.

Also assert cardinality. A repeated click or retry could create two drafts with identical inputs. An assertion that merely finds at least one matching record would miss that duplicate. In a controlled fixture, the expected number is exactly one submission. This does not promise exactly-once behavior in an external service; it checks that this test’s intercepted path observed one effect.

For the expired-session case, include a negative assertion: zero draft submissions. Pair it with evidence that the sign-in boundary remains visible. Negative checks matter because an agent could encounter an expired session, then act on stale page content or an unintended route. A blocked status without an empty effect record is not enough; the test needs to establish both what was visible and what did not happen.

An example of a compact run record is: fixture ID, task version, agent build identifier, browser and viewport, start time, completion status, captured action trace reference, submitted payloads, final visible confirmation and test outcome. Record only the content needed for diagnosis. Avoid logging session tokens, personal data or unrestricted page dumps; the fixtures should use synthetic values so a failure report can be shared with the people who need to fix it.

Step 5: Add a shifted-layout case without changing the task

A layout fixture is easy to get wrong because teams often change several things at once. Keep the button text and field labels identical to baseline. Move the form into a side panel, place it below a longer introductory block, or adjust the viewport so the form requires scrolling. Pick one realistic change and record it in the manifest. The question is whether the agent can still identify the same form and its intended action under that change.

Fix the viewport explicitly rather than relying on the developer’s current window size. Also fix the page’s initial scroll position and make sure fonts and content are local or otherwise deterministic. A changing banner, remote image or delayed third-party script can shift the target during a run and turn a layout regression into a timing problem. Where possible, remove those unrelated sources of variation from the fixture.

A useful acceptance rule is: the agent submits the requested draft once, with the correct values, and does not activate any neighboring control. Place a harmless neighboring action in the fixture only if it represents a plausible confusion risk, such as “Discard draft”; then assert that it was not activated. Do not add tempting controls merely to make the agent fail. The fixture should model a credible page, not a puzzle.

If the agent uses coordinates, preserve a screenshot or equivalent visual trace around the interaction when your harness can do so. If it uses semantic browser controls, retain the observed accessible name and target description. These records help distinguish a stale coordinate from an outdated label, but they are diagnostic evidence, not a substitute for checking the final state.

Step 6: Make fixture execution repeatable

Run each fixture in a fresh browser context or an equivalent isolated session. Otherwise cookies, local storage, service state or previous form values can silently change the starting conditions. The expired-session case is especially sensitive: a reused authenticated context can make a supposed sign-in gate disappear. Isolation gives each test a known beginning and makes parallel execution safer.

Reset the in-memory effect recorder before every run. If test workers share a server process, give each run a unique namespace or route so simultaneous tests cannot write into the same array. A fixture that occasionally reads another test’s payload is not a useful regression test, even if most runs pass.

Keep the browser viewport, locale-dependent content, time-sensitive page behavior and fixture data stable where they affect the interaction. Do not freeze every environmental detail by default; freeze the conditions that can change the meaning or location of controls. For example, if the production issue concerned a label translated according to locale, locale is part of the fixture input and should be explicitly set.

Separate a fixture failure from an agent failure. If the expected form is missing before the agent starts, classify the run as a fixture or harness failure. If the fixture is correct but the agent does not identify the action, classify it as an agent regression. If the request arrives with incorrect values, classify it as an interaction or task-completion failure. Those categories help a team route the repair to the relevant component rather than treating every red test as a model-quality problem.

flowchart TD
  A["Select named fixture"] --> B["Serve controlled page"]
  B --> C["Check starting state"]
  C --> D{"Preconditions match?"}
  D -->|"No"| F["Record result and evidence"]
  D -->|"Yes"| E["Run agent and capture effects"]
  E --> F
  accTitle: Browser fixture regression flow
  accDescr: Select and serve a named page, verify its starting state, run the agent only when preconditions match, and record evidence for both valid and invalid starts.

The precondition check is a gate, not an agent pass condition. If it fails, stop before invoking the agent and record a harness or fixture problem. If it passes, run the agent and capture effects, then compare the observed state with the case’s expected result. This prevents a malformed fixture from generating a misleading agent regression. The flow records an outcome in either branch; it does not treat a skipped run as a successful task.

Step 7: Read failures as evidence

A changed-label failure can have several causes. The agent may have memorized the old button name, the new label may be ambiguous, or a form submission may have been rejected for a different reason. Inspect the starting-state assertion first, then the action trace, submitted payload and final page state. If the fixture itself is valid and the agent never targets the new control, the likely issue is target selection. If it targets the control but sends the wrong fields, investigate value extraction or form association instead.

An expired-session failure needs a stricter interpretation. If a draft is recorded, the agent crossed a boundary that the fixture expected it to respect. Determine whether stale form content remained available, whether the fixture accidentally retained an authenticated state, or whether the agent attempted an action after seeing the sign-in page. Preserve the event sequence and page state immediately before the attempted effect. Do not “fix” the test by accepting a completed status when the business state says the session gate should have stopped the task.

A shifted-layout failure may be caused by the page geometry, a viewport mismatch, or a target ambiguity. Compare the test’s recorded viewport and initial scroll position with the fixture manifest. Then inspect whether the intended control was visible and uniquely identifiable. If a nearby action has the same label or role, the fixture may be exposing an ambiguity that deserves a separate test; if the target is simply off-screen when the agent begins, confirm that this is an intentional scenario rather than accidental setup drift.

A timeout is not automatically a website regression. It can mean the fixture never served, the agent did not start, the page took too long to reach a state, or the assertion waited for an effect that never occurred. Keep separate timeouts or timing records for page readiness, agent execution and business-state confirmation. One blanket timeout obscures which layer stopped making progress.

A common counterexample is a test that clicks the button by a stable selector and then passes. That may establish that the test script can click a button; it does not establish that the agent selected the right control, filled the correct values or completed the task. Another counterexample is using a screenshot similarity threshold as the sole pass condition. A page can look nearly identical while its action label or form submission behavior has changed. Use screenshots and traces to diagnose, with explicit state assertions as the acceptance evidence.

Step 8: Keep the fixture pack maintainable

Treat each fixture as a versioned test asset. Give it an owner in the code repository, a short reason for existing and a link to the issue or change that introduced the scenario. Review fixture modifications as carefully as test-code changes: changing an expected value can silence a regression just as easily as changing production code can introduce one.

When a real website change is observed, reduce it to the smallest reproducible page condition. Preserve the old fixture if it represents a supported behavior, add the new state as a separate named case, and update the manifest. If the old state is no longer relevant, remove it deliberately with a note explaining why. Avoid silently overwriting the only fixture that captured the previous behavior; a before-and-after pair helps diagnose whether the agent adapted or the test merely moved its target.

Use failure artifacts proportionately. A concise trace, a fixture ID and a small set of state assertions are usually more actionable than a huge bundle of logs. If page content can contain untrusted instructions, keep the fixture focused on the browser-change condition and handle the trust-boundary question separately; guidance on treating web pages as untrusted input covers that distinct concern. Similarly, the choice between browser control and a direct API is an architectural decision, not a fixture design shortcut; see how to assess a browser agent against an API option.

For a team beginning from a production workflow rather than a test harness, first define the agent’s browser boundary and observable outcomes. This overview of browser-use agents in production provides related implementation context. Keep this fixture pack small enough that an engineer can understand a failing case without reconstructing the entire production workflow.

Troubleshooting the first run

If the agent reports completion but the test sees no draft, check whether the fixture form submitted to the intercepted URL, whether the route handler is registered before navigation and whether the page’s response updates the confirmation region. Next inspect the captured trace to determine if the agent never submitted or if the test harness failed to observe the request. Do not weaken the assertion until those possibilities are separated.

If the draft appears twice, check for repeated submission, automatic form retries in the fixture, and shared state between test runs. Use a fresh recorder per test and include a run-specific identifier in diagnostic output. The regression assertion should still fail on two submissions; changing it to accept “one or more” would hide the behavior the fixture is meant to catch.

If the expired-session test unexpectedly succeeds, verify that it loads the intended page variant and that the agent context has no existing authentication state. Then inspect whether any hidden form or stale route is available. A clear visible sign-in gate plus an empty effect record should be the expected blocked result, not a special-case prompt telling the agent to stop.

If the shifted-layout test passes locally and fails in automation, compare viewport, device scale and font availability, then check for content that loads after the initial render and moves the form. Reduce unrelated variability before changing the agent task. If the layout is genuinely responsive, add a second explicit viewport fixture rather than letting the environment pick one unpredictably.

If the test times out on an assertion, report the last known state with the failure. A timeout message that only says “expected visible” is less useful than one that also identifies the fixture, expected target, agent status and whether the fake endpoint received a request. Keep diagnostic output bounded and synthetic so it can be retained safely.

Printable fixture worksheet and acceptance order

Use this worksheet when adding a case. It is intended to be copied into a test issue or code review; it is not a downloadable artifact.

Fixture identity

  • Name the case with a stable ID and record the page condition it represents.
  • State the one website change being tested and identify unchanged controls or content that should remain constant.
  • Record browser, viewport, locale, starting session state and synthetic input values.
  • Confirm that the page does not expose the fixture name or an instruction that gives away the expected action.

Expected interaction and state

  • Write the task once and reuse it across comparable fixtures.
  • Define the permitted business effect, including relevant fields and expected count.
  • Define prohibited effects, especially for expired-session or blocked cases.
  • Identify the page evidence and independent state evidence that will establish the result.
  • Specify what should happen if the fixture’s own starting conditions are invalid: stop before the agent runs.

Repeatability and diagnosis

  • Use a fresh context and reset effect recording for each run.
  • Make external calls impossible or intercept them with a test-only handler.
  • Save a concise trace reference and the fixture ID with the result.
  • Distinguish fixture, harness, agent-targeting, value-entry and business-outcome failures.
  • Review the changed page and expected assertions together whenever the fixture is edited.

Prioritize cases by the cost of an incorrect action and the likelihood of the website condition recurring. For a small initial pack, implement the expired-session case first if an unintended submission would be consequential, then the renamed control and shifted layout. That ordering is an illustrative prioritization, not a universal risk score: a team with frequent responsive redesigns may reasonably begin with layout coverage. Add variants only when they represent a distinct condition that changes the expected observation or safe outcome.

The practical result is a small regression asset that can be rerun when the site, agent or browser harness changes. Each case begins from a named state, changes one meaningful condition, captures the agent’s interaction and checks the resulting business state. If your team needs help designing a fixture harness around a real workflow, AI workflow automation is a relevant next step.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call