SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
How-to Guide 20 mins

How to Measure an AI Operations Pilot Across Multiple Sites

A practical, site-by-site method for measuring AI answer accuracy, booking completion, corrections and staff time before deciding whether to expand.

The PADISO Team ·

Prerequisites

Before collecting pilot results, assemble four things: a written definition of an eligible request, a baseline period, a consistent event log, and named owners for measurement and operational exceptions. You do not need a complex analytics stack to begin. A carefully maintained spreadsheet can work for a small pilot if every site records the same fields and the team can reconcile records against the underlying booking system.

Choose the sites and operating period before reviewing results. Include the locations intended to represent the eventual rollout, not only the easiest site or the team most enthusiastic about automation. Record relevant differences such as opening hours, service mix, booking policies, staffing patterns and seasonal demand. Those differences do not automatically disqualify a site; they determine what needs to be compared separately.

Agree on a human resolution route for any request the workflow cannot safely complete. Staff should know how to identify the unresolved request, what information to preserve and who will follow up. Measurement must not encourage an operator to record a booking as complete merely because the AI said it was complete.

Set a start and end date for both the baseline and the pilot. If the pilot is being compared with a baseline, capture the same weekdays and comparable operating hours where possible. Record changes during either period, including staff training, policy updates, promotions, outages and changes in how requests arrive. These are potential explanations for a result, not inconvenient details to discard.

For broader investment context, use a structured AI ROI framework to connect pilot evidence to a spending decision. This guide focuses more narrowly on four operational measures: answer rate, booking completion, corrections and staff time.

Step 1: Define the unit you are measuring

A request is the denominator for most of the measures in this guide. Define it in language that a site manager can apply consistently. For example: one distinct customer need that requires an answer or a booking action, whether it arrives by phone, chat or another channel included in the pilot.

Decide how to handle a conversation containing several needs. If a customer asks for opening hours and a table booking in the same interaction, count one interaction but two intents only if your workflow and reporting can reliably distinguish them. Otherwise, use one interaction as the unit and tag its primary purpose. The key is to avoid one site counting messages while another counts calls or bookings.

Set rules for duplicates and incomplete contacts. A repeated call about the same unresolved booking might be one service case for correction tracking, but it may be two separate requests for workload measurement. Document the choice. Exclude spam, test traffic and requests outside the pilot’s intended scope consistently, and retain a reason for each exclusion.

Create a short eligibility statement that staff can use without interpretation. For instance: “Count each genuine customer interaction received during staffed pilot hours that asks a question the workflow is intended to answer or requests a booking action.” If a request falls outside that description, log it as out of scope rather than silently dropping it.

Keep the unit stable when channels differ. A phone call may include clarification that a short message does not, so recording only the number of contacts can conceal different work. Where practical, add a request-type field—such as information-only, booking request or mixed—without changing the primary denominator. This makes later comparisons more useful while keeping the core measure understandable.

Step 2: Define the four measures before the pilot

Baseline answer rate is the share of eligible requests that receive a correct, relevant answer without staff having to correct or replace it. Define “correct” using a review rule that fits the request type. An answer that is fluent but wrong does not qualify. Nor should an answer count as successful simply because no customer complaint was recorded.

For a practical manual review, draw a pre-agreed sample of eligible requests and compare the answer with the applicable source of truth or the result staff actually needed to provide. Record whether the answer was correct, incomplete, misleading or not answered. If the pilot handles several request types, report the answer rate by type as well as overall. A high volume of easy opening-hours questions should not conceal weak answers to questions that affect bookings.

Booking completion rate is the share of eligible booking requests that result in a booking confirmed in the authoritative booking record. The denominator is booking requests, not all customer interactions. A request is not complete because the system proposed a time, displayed a success message or generated a conversational confirmation. The booking record must show that the intended reservation exists and matches the agreed details.

Correction rate is the share of completed pilot-handled requests that require a material staff correction within a defined observation window. For bookings, a correction might be a change to the date, time, party size or customer details because the recorded result did not match the customer’s request. State whether cancellations caused by a customer’s later change are excluded; they are not automatically workflow errors. Apply that distinction consistently and record the reason for each correction.

Staff time per eligible request is the staff effort attributable to handling or checking the request, divided by the number of eligible requests. Include the work needed to review an answer, complete an unfinished booking, correct an error and record an exception. If one person handles several requests at once, estimate or sample time using a repeatable method rather than assigning the entire shift to the pilot.

Keep these measures separate. A workflow can answer more questions correctly while completing fewer bookings, or complete bookings while generating costly corrections. Staff time can fall because fewer requests arrive, because staff stop checking results, or because automation genuinely removes work. The accompanying counts and exception reasons help distinguish those explanations.

The unit-economics principle is to relate costs to a defined unit of value and include indirect costs, not only direct spend (FinOps Foundation). For this pilot, the useful unit is usually a resolved eligible request or a verified completed booking. Include review and correction time when interpreting staff effort.

Step 3: Capture a credible baseline

Measure the current operating process before the pilot changes it. Use the same eligibility rules, request categories and correction window planned for the pilot. A baseline assembled from memory or a manager’s estimate can provide context, but it is not a sound numeric comparator for a measured pilot.

For each baseline request, record the site, date and time, request type, whether an answer was provided, whether it was correct under the review rule, whether a booking was requested, whether the booking was verified, whether a correction followed, and staff minutes. Add a short exception reason when a request is not resolved normally. Use a request identifier that allows authorized staff to reconcile the event with operational records without putting unnecessary customer details into the measurement sheet.

If retrospective records do not contain all these fields, run a prospective baseline: have staff record requests using the new form while the existing process remains in place. Explain the logging task in advance and keep it as light as possible. A recording process that materially changes how staff handle requests can itself affect the baseline, so note that limitation.

Compare operating periods rather than just calendar dates. If one site’s baseline includes a weekend promotion and the pilot does not, the resulting difference may reflect demand or request mix rather than the workflow. Record unusual conditions and compare like with like where feasible. When perfect matching is not possible, show the period and caveat alongside the result rather than implying a controlled experiment.

Choose a review sample before inspecting outcomes. Review all requests if volume is low enough; otherwise define how the sample is selected, such as a fixed random sample within each site and request type. Do not let a site select only successful or memorable interactions. Preserve the reviewed request identifiers and reviewer decision so that another person can understand how the rate was produced.

Step 4: Instrument the pilot workflow

Use a compact event record with one row per eligible request. A useful starting set of fields is:

FieldWhat it records
Request IDA non-meaningful identifier for reconciliation
Site and timestampLocation and time received, including local time zone if sites span zones
Request typeInformation, booking or mixed, using agreed categories
Answer statusCorrect, incomplete, incorrect, unanswered or not applicable
Booking statusNot requested, requested, verified complete, incomplete or unresolved
CorrectionWhether a material correction occurred and its reason
Staff minutesHandling, checking, escalation and correction time under the agreed method
Exception codeA short category for failure or human resolution

Record the observed outcome, not just what the workflow reports. For a booking, verify the result against the record staff rely on to manage reservations. For an answer, apply the review rubric to the actual response and request. This makes business result verification independent of any model-generated statement that the task succeeded.

If the workflow has an automated event log, compare it with a sample of real operational records before using it as the measurement source. Dayrun describes an Insights capability; treat any such reporting as a candidate input and verify its definitions, coverage and records against original acceptance tests before relying on it (Dayrun Insights). The important question for this pilot is whether the information supports your agreed measures, not whether a dashboard presents a metric with a familiar label.

Make exceptions visible. Use a small, consistent set of codes such as “answer uncertain,” “availability mismatch,” “required detail missing,” “record not found” and “staff follow-up.” Tailor the list to the workflow, but do not create a different category for every individual incident. A short note can preserve nuance; a stable code makes site-level patterns countable.

Step 5: Run the operational path and exception path

The workflow should make the customer’s request traceable through an answer or booking attempt and into a verified outcome. If a request cannot be resolved, staff should take over and the record should show that it was unresolved by automation. That is a valid operational outcome, not a reason to omit the request from measurement.

flowchart TD
  accTitle: Pilot request measurement flow
  accDescr: An eligible request is recorded, its answer and booking outcome are checked, exceptions go to staff resolution, and the result is included in site-level review.
  A["Record eligible request"] --> B["Assess answer"]
  B --> C["Check booking record"]
  C --> D{"Outcome verified?"}
  D -->|"No"| E["Log exception; staff resolve"]
  D -->|"Yes"| F["Record correction and time"]
  E --> F
  F --> G["Review by site and period"]

The diagram separates an attempted action from a verified outcome. In practice, not every eligible request needs a booking check: mark booking status “not requested” for information-only requests, and calculate booking completion only from booking requests. The shared path represents common measurement discipline, not a claim that all requests have the same operational steps.

When an answer is uncertain or a booking cannot be confirmed, the exception path should preserve the original request, the reason it could not be completed, the staff action and the final result. Staff should be able to finish the task through the normal business process. Measure the time spent on the handoff and follow-up, not only the time before escalation.

A booking confirmation must be checked against the relevant booking record before it is marked complete. Keep measurement of booking results distinct from decisions about payment authority. For a deeper discussion of deposits and refunds, see how to keep payment authority explicit; this pilot’s measure is whether the requested booking was completed and recorded accurately.

For multi-site operation, also distinguish a shared rule from a local exception. If sites have different booking policies, capture which rule applied so reviewers do not classify a correct local outcome as an error—or excuse an error as a local difference. Broader handling of headquarters rules and local overrides is covered in the franchise governance discussion, rather than being folded into this measurement method.

Step 6: Calculate results consistently

Calculate each rate from its own eligible denominator. For answer rate, divide correctly answered requests in the reviewed sample by all reviewed eligible requests for which an answer was expected. Report unanswered and out-of-scope cases separately. If sampling was used, state the sample size and method; a percentage without its denominator is not decision-ready.

For booking completion, divide verified completed bookings by eligible booking requests. Show unresolved requests rather than quietly excluding them. If a request is abandoned by the customer before a booking attempt, classify it under a pre-agreed rule and report the count. Do not change the rule after seeing which treatment makes the result look better.

For correction rate, divide requests with at least one material correction during the observation window by completed requests whose full correction window has elapsed. State the window—for example, whether it ends at the close of the next business day—and apply it to both periods. A recent request whose window remains open should not be treated as correction-free.

For staff time, add the minutes attributed to requests in the period and divide by eligible requests. Show total minutes and request count alongside the average. If a small number of difficult cases dominate the average, a median or range can add context, but neither replaces the total effort needed for planning. Include staff time spent on sampling, routine checking and exception handling if those tasks are expected to continue after the pilot.

Present the site results before the network-wide result. A pooled rate can be calculated by adding the relevant numerators and denominators across sites; do not average site percentages equally unless equal site weighting is the intended question. If one site handles most requests, its outcome will dominate a pooled measure. That may accurately describe total operational volume while concealing poor performance at a smaller location.

Step 7: Compare sites without hiding differences

Use a comparison table with one row per site and a separate network row. For each site, show baseline and pilot counts, the four measures, request mix and significant exceptions. Include the measurement dates. A small site with few booking requests may show a large percentage swing from one case, so counts belong beside rates.

Separate three kinds of variation. First, operational variation: a site may have different hours, demand or staffing. Second, workflow variation: a local configuration or process may produce a different result. Third, measurement variation: staff may classify or log requests differently. Investigate these separately before recommending a change. A difference is a signal to understand, not proof of a cause.

Do not turn a site ranking into a performance league table. The aim is to locate where the workflow works, where it fails, and what operational conditions matter. A site with a lower completion rate may be handling a more difficult request mix; a site with a high answer rate may be serving fewer booking requests. Show category-level results when the volumes support a meaningful comparison.

A practical review sequence is to inspect the weakest result in each measure, then inspect the strongest result for a different pattern. Review the underlying request records for both. If a high-performing site has few corrections because staff intervened before an error reached the customer, that is useful operational evidence—but its staff time must include that intervention. If a low-performing site logs corrections more diligently, its apparent rate may reflect better detection. Compare logging practice, not just the final percentages.

Where a central team changes a shared rule during the pilot, timestamp the change and identify affected sites. Do not combine pre-change and post-change periods as though the process were constant. A split-period view may be more informative than a single before-and-after number, especially if the change was intended to fix a specific failure.

Step 8: Work through an illustrative multi-site example

The following scenario is hypothetical, and all figures are illustrative. Assume a four-site operator runs a two-week baseline followed by a three-week pilot. Each period covers comparable weekdays and staffed hours. The sites log eligible requests prospectively, and a reviewer checks a preselected sample of answers. Booking completion is reconciled against the booking record; corrections are observed through the end of the next business day.

The baseline contains 800 eligible requests across the four sites, including 300 booking requests. In the answer sample, 360 of 500 reviewed requests are correct under the rubric, for a baseline answer rate of 72%. Of the 300 booking requests, 162 result in a verified booking, or 54%. Twenty-four completed requests require a material correction within the observation window, giving a correction rate of 24 divided by 300, or 8% if all 300 are completed; if the denominator is verified completed requests instead, the calculation must use 162. To keep the example internally consistent, define correction rate here over all booking requests that reached a completed or attempted booking outcome, and report the 24 cases as 8% of 300. In a real pilot, choose one denominator in advance and do not switch it.

Suppose staff record 5,200 minutes of handling and follow-up time for those 800 requests. The baseline average is 6.5 staff minutes per eligible request. The pilot records 900 eligible requests, including 360 booking requests. Reviewers find 414 correct answers in a sample of 540, for 76.7%. The booking record confirms 216 of 360 requests, or 60%. Twenty-seven of the 360 booking requests need a material correction, or 7.5%. Staff record 5,400 minutes, which is 6.0 minutes per eligible request.

On the surface, the pilot has a higher answer rate and booking completion rate, a slightly lower correction share and lower staff minutes per request. The total staff time is higher because the pilot period contains more requests; that is not inconsistent with a reduction in average time. The calculation does not establish that the workflow caused the differences. Request mix, staffing, logging completeness and changes in demand still need review.

Now inspect site-level results. Suppose three sites improve on booking completion, while the fourth falls from 58% to 43%. The network result can still rise to 60% if the improving sites have more booking volume. That pooled figure is not a reason to ignore the fourth site. Review its booking failures, operating conditions and exception logs; determine whether the cause is a local process issue, an inconsistent rule or a workflow limitation. The appropriate next action may be to pause that site’s expansion while retaining the pilot elsewhere.

There is also a counterexample to an apparently positive result. Imagine staff at one site begin completing the measurement sheet only for requests they successfully resolve. Unresolved requests disappear from the denominator, raising answer and completion rates while reducing recorded staff time. The dashboard could show improvement even as customers receive worse service. Check request counts against an independent source such as the site’s contact log, investigate missing timestamps and compare the share of requests with no final status. Incomplete records are a measurement failure, not evidence of success.

Step 9: Diagnose operational failures before deciding

A failure timeline often reveals more than a headline rate. Consider a booking request received during a busy shift. The interaction is logged, but the requested time is unavailable. The workflow presents an alternative without confirming whether the customer accepts it. Staff later see a booking record for the alternative and mark the request complete. A subsequent review finds that the customer wanted the original time and the alternative was never agreed.

Classify the failure at the point it occurred: misunderstanding the request, offering an unacceptable alternative, treating an attempted action as a completed booking, or missing the later correction. Each category suggests a different remedy. Better request capture will not fix a verification gap; more verification will not fix a policy that leaves staff uncertain about what to offer.

Also look for failures that inflate staff time without changing the final result. Staff may repeatedly check a record because the status is unclear, copy information into a second log, or contact the customer to confirm a detail already captured. These minutes may be operationally necessary during the pilot, but they should be visible. If the process requires this checking in production, count it in the ongoing staff-time estimate.

Review late corrections only after the observation window has closed. A request that appears correct at the end of a shift may need correction the next day. Conversely, a customer’s later change of plan should not be attributed to the workflow unless the agreed correction definition says it should. Keep the reason code and supporting note so that the distinction can be audited internally.

When a result looks implausibly clean, inspect missingness first. Check for gaps by site, shift, request type and status. Compare the number of eligible requests with a separate operational count where one is available. Then inspect a small set of original records from successful, corrected and unresolved cases. A measurement process that misses difficult cases can create confidence without reducing operational risk.

Step 10: Apply a decision rule and preserve the evidence

Set the decision rule before the final review. It should state which measures must meet a threshold, what minimum data completeness is acceptable, and what site-level exception requires more investigation. Thresholds depend on the business’s starting point and the consequences of errors; do not borrow an attractive percentage from another operation without testing whether it fits your request mix and tolerance.

A usable rule might require: answer quality to meet a pre-agreed minimum in each request category; booking completion to improve or remain acceptable at every participating site; correction rate not to exceed a specified ceiling; staff minutes per eligible request to fall without reducing required checking; and missing or unclassified events to stay below a defined limit. The numbers should be set by the team that owns the service, based on baseline performance and the cost of failure, not selected after results are known.

Choose among three practical outcomes. Expand where evidence is sufficiently complete and site-level results meet the agreed criteria. Continue a bounded pilot where the signal is promising but the sample, period or exception pattern is inconclusive. Pause or stop a workflow where the result misses a critical threshold or the measurement cannot verify that customer requests were completed correctly. The ship-or-kill decision framework can help structure that broader decision; this guide supplies the operational measures to bring into it.

When deciding whether more implementation support is useful, describe the specific gap: event capture, site-level comparison, exception routing or result verification. A team assessing operations automation should bring its request definitions and acceptance criteria into that conversation so that proposed work can be evaluated against the actual operating process.

For a deeper discussion of booking questions, staffing and takings in a distinct connector context, see the Dayrun MCP connector article. Do not let that separate topic distract from this pilot’s acceptance evidence: accurate answers, verified booking completion, manageable corrections and measured staff effort.

Printable pilot measurement worksheet

Copy this worksheet into the team’s working document and complete it before the pilot begins. Keep one version for the baseline and one for the pilot; do not overwrite the baseline definitions after results arrive.

Scope and period

  • Participating sites and accountable site contacts are listed.
  • Baseline and pilot dates, staffed hours and relevant weekdays are recorded.
  • Eligible request is defined in one sentence, including treatment of duplicates and out-of-scope contacts.
  • Request categories and rules for mixed requests are documented.
  • Relevant operating changes, promotions, policy changes and outages will be recorded with dates.

Measures and evidence

  • Answer correctness rubric defines correct, incomplete, incorrect and unanswered.
  • Review sample method and sample size are fixed before outcomes are inspected.
  • Booking completion means a verified result in the authoritative booking record.
  • Correction definition, denominator and observation window are fixed.
  • Staff-time method includes review, escalation, follow-up and correction work.
  • Site-level numerators, denominators, dates and request mix will be shown beside percentages.

Operational exceptions and decision

  • Unresolved requests have a named staff route and a record of final disposition.
  • Exception codes are consistent across sites and have a short written definition.
  • Missing-record checks compare the pilot log with a separate operational count where available.
  • Acceptance thresholds and minimum data-completeness criteria are set in advance.
  • The team has defined what qualifies for expansion, continued measurement or pause.

One-page review summary

At the review, fill in these values for baseline and pilot: eligible requests; reviewed answers and correct answers; booking requests and verified completions; requests requiring correction and the completed observation window; total staff minutes and minutes per eligible request; missing or unclassified events; and the corresponding values for each site. Add three short notes: the largest observed failure pattern, the most important difference between sites, and the next action with an owner and review date.

Summary: make the outcome verifiable

A useful multi-site pilot measurement process starts with a stable definition of an eligible request and a comparable baseline. It records what happened at each site, checks answers against an agreed rubric, verifies bookings in the operational record, tracks corrections through a defined window and includes the staff time needed to resolve exceptions.

Show counts as well as rates, calculate each measure with its own denominator, and inspect site results before relying on a pooled average. Treat missing records and unverified outcomes as unresolved evidence, not as successes. Decide in advance what results justify expansion, what requires another measurement period and what calls for a pause.

The purpose is not to produce a flattering pilot score. It is to determine whether the workflow reliably answers eligible requests, completes the bookings it claims to complete, avoids excessive corrections and changes staff effort in a way the operation can verify.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call