SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Tutorial 23 mins

Benchmark Agent Tail Latency: Trace One Workflow End to End

A reproducible tutorial for measuring agent p50 and p95 workflow latency, attributing time to stages, and separating fast failures from successful completion.

The PADISO Team ·

Why benchmark a workflow rather than a model call

An agent’s user-visible completion time is the elapsed time from accepting a task to reaching a defined terminal outcome. That interval may include orchestration, model requests, tool calls, retries, queueing, validation, and waiting between stages. A benchmark that times only one component cannot tell you how long the task took, or where that time went.

This tutorial builds a repeatable measurement for one bounded workflow. You will define its start and finish, record one trace per task, calculate end-to-end p50 and p95, and produce a stage breakdown that helps explain changes. The goal is not to declare one system universally faster. It is to make a meaningful comparison between two configurations under a workload your team can reproduce.

A trace is a set of related spans, with context that allows operations to be correlated. That relationship helps connect work across stages; it does not establish that the task succeeded. Record and evaluate the business outcome separately from timing. OpenTelemetry’s overview of traces and spans describes this correlation model.

For the worked example, assume a hypothetical support agent that looks up an order, checks its delivery state, and drafts a response. The numbers later in the article are illustrative calculations, not observed benchmark results. The same protocol works for another bounded task if you replace the stages and define a verifiable outcome.

A useful benchmark answers four separate questions: how long did a successful task take; how long did unsuccessful tasks run before their disposition; which stages account for elapsed time; and did the tested configuration complete the intended work correctly? Keeping these questions distinct prevents a fast error from appearing to be a performance improvement.

Prerequisites and setup

Before collecting data, assemble a stable test harness and one versioned workload. You need a way to submit tasks through the same entry point used by the system under evaluation, capture timestamps at the workflow boundary, attach a correlation identifier to each task, and preserve terminal outcomes. Your instrumentation should also expose stage start and end events, including retry attempts.

Prepare two configurations only when there is a specific comparison to make, such as a current deployment and a proposed change. Keep the application workflow, task inputs, tool data, test period, and completion criteria fixed. If a configuration change alters any of those elements, document it; otherwise, a latency difference may have several competing explanations.

Create a workload file with a stable task identifier, a category, an input reference or controlled input, and the expected business outcome. Avoid using live customer requests as an uncontrolled benchmark set. For the order example, categories might distinguish an order with a normal delivery status, an order with no matching record, and a record that requires a second lookup. Give each category a fixed proportion in the workload rather than allowing whichever tasks happen to arrive to define the mix.

Specify your environment in a run manifest before testing. Record the application build, configuration label, tool-service version if known, workload revision, start and end times, concurrency target, timeout policy, and relevant dependency conditions. The manifest is not a claim that every dependency can be held constant. It is a way to identify what changed and to avoid comparing runs whose differences have been forgotten.

Use a test environment or controlled data where possible. If a workflow can trigger a real-world side effect, configure the benchmark path so it does not send a message, change an order, or otherwise act on a real account. This tutorial measures the workflow; it does not require executing irreversible actions to do so. When the actual production path must be observed, make a separate plan for limiting and verifying those effects.

One practical setup is a small append-only event store or a structured log export with one row per event. The storage choice matters less than preserving task identity, timestamps, outcome, and attempt number. Do not reduce the record to one final duration field before retaining the underlying stage events; the raw events let you recalculate attribution when a boundary rule changes.

Step 1: Define the unit of work and terminal states

Write a one-sentence task definition that can be applied consistently. For example: “A task begins when the service accepts a valid order-status request and ends when it returns a response that passes the specified order and response checks, or reaches a recorded terminal failure.” This definition gives the benchmark a boundary independent of internal implementation details.

Choose the start timestamp at the point where the system takes responsibility for the request. If the user-facing service queues work before the agent begins, start at acceptance rather than at the first model request; otherwise the measured duration omits queueing that the user experiences. Keep the clock source consistent across components, preferably by collecting boundary events in the same process or ensuring the participating systems have synchronized clocks. Record timestamps with sufficient precision for the expected workflow duration.

Define terminal states before running the workload. At minimum, distinguish successful completion, validation failure, tool failure, timeout, cancellation, and an unexpected system error. A response that looks plausible but refers to the wrong order is not a successful completion. Likewise, an agent that returns an error quickly has reached a terminal state, but has not demonstrated fast successful work.

Use two related latency distributions. First, calculate successful completion latency from accepted task to verified successful result. This directly answers how long correct completions take. Second, report time to terminal disposition for all tasks, with terminal outcome alongside each value. Keep failure categories visible rather than folding them into a single success-only statistic. A timeout’s elapsed time is measured to the timeout disposition, not treated as a successful completion.

Set the timeout policy as part of the protocol, not after seeing the results. If the service has a defined task deadline, use it. If the test harness imposes a deadline, declare its duration and count tasks that hit it as timed out. Report how many tasks were censored by the deadline and do not silently discard them: excluding slow unfinished work can make the observed successful p95 look better while concealing that the run failed to finish a substantial portion of the workload.

Specify success checks that do not depend on the agent’s own assertion that it succeeded. In the example, a checker can compare the returned order identifier and delivery status with the controlled order record, then verify that the response follows the required format. A benchmark can have correct timing and still be invalid as an evaluation if its outcome check is merely “the agent returned something.” For broader evaluation design, see AI Agents in Production: Agent Evaluation Frameworks.

Step 2: Map the workflow into timed stages

Choose stages that correspond to operationally distinct waits or work, rather than every internal function. A practical initial map for the example is: admission and queue wait; orchestration and decision work; external tool wait; response validation; and final delivery. Keep the map small enough that operators can interpret it, but detailed enough to reveal whether a change moved time from one stage to another.

Separate elapsed time from active processing time. If a tool call takes 700 milliseconds from dispatch to response, the benchmark can attribute that interval to tool wait whether the tool spent time processing, queued internally, or crossing a network. Unless you have reliable instrumentation inside that dependency, do not label it “tool compute.” The measurement is the elapsed interval visible to your workflow.

Retries must remain visible. Record each attempt as its own event with an attempt number and a parent task identifier. The task’s end-to-end duration includes the time spent on failed attempts, backoff, and subsequent attempts. A report may show both total retry time and the number of retrying tasks, but it must not subtract retries from the user’s elapsed time.

Avoid adding nested span durations as if they were independent pieces of elapsed time. A parent orchestration span can contain a tool span, so summing both counts the same interval twice. For a stage breakdown that adds up to end-to-end elapsed time, assign each instant to one mutually exclusive bucket. One workable rule is to classify the wall-clock timeline in this priority order: external wait, queue wait, validation, and remaining orchestration. Document the priority rule so a reader can reproduce the allocation.

Some workflows execute work in parallel. In that case, the sum of the individual branch durations can exceed the overall task duration. Preserve the branch spans for diagnosis, but derive the top-level stage attribution from the critical path or from mutually exclusive wall-clock intervals. State which interpretation is used. A table headed “stage duration” without that distinction invites readers to add numbers that are not additive.

The following flow describes the measurement path. A retry returns to the tool-call stage; a task without a tool call can proceed from orchestration to outcome verification. The diagram does not assert that any component is successful merely because its span ended.

flowchart TD
accTitle: Workflow latency measurement path
accDescr: A task enters orchestration, may make a tool call and retry it, then reaches outcome verification, terminal classification, and event recording.
    A["Accept task"] --> B["Run orchestration"]
    B --> C["Call tool"]
    C --> D["Retry decision"]
    D -->|"Retry"| C
    D -->|"Stop"| E["Verify outcome"]
    B --> E
    E --> F["Classify terminal state"]
    F --> G["Record timings"]

Step 3: Define the event dataset

Before implementation, agree on a compact event schema. One record per stage event is more flexible than one record per task because retries and repeated validations can be represented without adding special columns for every possible attempt. Keep a task-level summary as a derived view, not as the sole retained dataset.

FieldPurposeExample or rule
run_idGroups one benchmark runStable label from the run manifest
task_idCorrelates stages for a taskRandom or fixture-based identifier
config_idIdentifies the tested configurationA label, not a free-form description
event_nameNames a stage or boundarytask.accepted, tool.wait, task.terminal
start_msEvent start on a consistent clockMonotonic time when available within a process
end_msEvent endOmit only for an explicitly open event
attemptDistinguishes repeated workZero for non-retry stages; increment tool attempts
outcomeRecords stage or task dispositionKeep task terminal outcome distinct from stage status
categorySupports workload stratificationA predeclared workload category
trace_idConnects related operationsOne trace identifier per task where possible

The example names are a suggested schema, not a requirement of a particular tracing product. If your instrumentation already uses other field names, keep them and maintain a mapping in the run manifest. What matters is that each task can be reconstructed, each interval has a defined meaning, and the terminal outcome can be joined to the duration calculation.

Use elapsed-time clocks carefully. Wall-clock timestamps are useful for aligning events across processes, while monotonic clocks are preferable for measuring durations within one process because clock adjustments can make calendar time jump. When stage events cross process boundaries, preserve the source timestamp and clock context you actually have; do not imply sub-millisecond accuracy if synchronization does not support it. If clock offsets might materially affect a stage allocation, call that out as a measurement limitation.

Add a data-quality check before calculating percentiles. Every admitted task should have one accepted event, one terminal disposition or an explicit still-open status, and no negative stage duration. Flag duplicated terminal records, missing identifiers, overlapping exclusive buckets, and events whose end precedes their start. Do not “repair” anomalies silently. Preserve the raw row, document the correction rule, and report the count of affected tasks.

Step 4: Instrument one task from boundary to outcome

Add instrumentation at the task entry point and at each chosen stage boundary. Generate or accept a task correlation identifier at admission, carry it through orchestration and tool calls, and attach it to each event. The identifier should connect events without embedding the task’s full sensitive input in timing logs. Keep input storage and timing records separate if the measurement does not need their contents.

At task acceptance, record run_id, task_id, config_id, workload category, and the start timestamp. At each stage, record a start and end event, stage name, and attempt number. At terminal disposition, record the end timestamp and outcome. Then run the independent outcome checker and attach its result. If the checker itself adds latency to the user-facing workflow, instrument it as part of the measured path; if it runs only in the benchmark harness after the response, keep it outside the user-facing duration and label it as evaluation time.

For illustration, the following pseudocode shows the intended boundaries, not a drop-in instrumentation library. clock() must be replaced with a suitable clock for the runtime, and emit() with the team’s event sink. The snippet leaves out application-specific exception handling and makes no claim about a vendor API.

def measure_task(request, run_id, config_id, clock, emit, run_agent, verify):
    task_id = new_test_task_id()
    accepted = clock()
    emit("task.accepted", task_id, run_id, config_id, accepted)

    try:
        agent_start = clock()
        result = run_agent(request, task_id=task_id, emit=emit)
        agent_end = clock()
        emit("agent.finished", task_id, run_id, config_id,
             agent_start, agent_end)

        check_start = clock()
        success = verify(request, result)
        check_end = clock()
        emit("outcome.checked", task_id, run_id, config_id,
             check_start, check_end, outcome=success)

        terminal = "success" if success else "validation_failure"
    except TimeoutError:
        terminal = "timeout"
    except Exception:
        terminal = "system_error"
        raise
    finally:
        ended = clock()
        emit("task.terminal", task_id, run_id, config_id,
             accepted, ended, outcome=terminal)

This sketch illustrates two important limits. First, a real implementation must ensure terminal is initialized for every control-flow path and that exception handling records failures without accidentally changing application behavior; the simplified snippet needs adaptation before use. Second, if a failure is re-raised, the harness still needs to persist the terminal record reliably. Test instrumentation with controlled cases for success, validation failure, timeout, and unexpected exception before treating the resulting data as complete.

Expected output from this step is a reconstructable sequence of events for each task. For a successful task with one retry, it should show one acceptance, the orchestration interval, two tool attempts, any wait between attempts, verification, and one terminal success. The elapsed task duration should run from acceptance through the terminal boundary and include all those waits that occurred before the result was returned.

Step 5: Build a fixed workload and run protocol

Create a dataset specification that another engineer can use without guessing what “the same workload” means. Include the task categories, number of tasks per category, fixed input fixtures or input-generation rules, expected outcome checks, task ordering or randomization rule, concurrency schedule, timeout, warm-up policy, and number of independent runs. Version the specification and retain the exact input set or a stable reference to it.

A benchmark’s workload mix is part of its result. If one configuration receives mostly straightforward lookups and another receives a larger share of missing-record cases, their overall p95 values are not directly comparable. Report both the aggregate distribution and category-level distributions. When the intended operational mix is known, set the category weights in advance and use the same weights for each configuration.

Choose concurrency to represent the decision you are making. A low-concurrency run can help isolate per-task behavior; a load-shaped run can expose queuing and contention. Do not describe one as a substitute for the other. If you run several concurrency levels, give each its own result and manifest. Record the offered rate or concurrency target and the number of tasks actually admitted; a target that the harness did not achieve is not the condition that was tested.

For a controlled comparison, interleave configurations or alternate run order where practical. Running all of configuration A in the morning and all of configuration B later can confound the result with changes in dependent services or background load. Interleaving does not remove every external influence, but it makes a time-dependent shift less likely to align perfectly with only one configuration. Preserve run-level results rather than merging all tasks into one large pool with no indication of when they occurred.

Decide whether to include warm-up tasks. If the production question concerns steady-state service, use a declared warm-up period and exclude only those designated warm-up tasks from the primary statistic. Still retain their data separately. If cold starts or first-use behavior are part of the user experience, measure and report them as their own condition instead of excluding them. Do not select a warm-up cutoff after inspecting which samples make a candidate look better.

A reproducible run record should state: dataset revision; configuration identifiers; task count and category counts; concurrency and achieved admissions; start and end times; timeout; exclusions and their rules; and the software build. If any of these are unknown, mark them unknown. Reproducibility means a future reader can reconstruct the method and its limits, not that every external system can be frozen.

Step 6: Calculate p50, p95, and stage attribution

For a defined list of durations sorted from smallest to largest, p50 is the median position and p95 describes the upper tail near the 95th percentile. State the percentile convention. For a simple reproducible choice, use the nearest-rank rule: for percentile p and n values, select the sorted value at one-based rank ceil(p × n). Other quantile conventions can produce slightly different values, especially in small samples, so do not compare reports that use different conventions as if they were identical.

Here is a small standard-library calculation for a list of successful completion durations. It assumes the input list has already been filtered using the outcome and cohort rules in the protocol. It does not determine which records are valid, estimate uncertainty, or decide what belongs in a stage.

import math

def nearest_rank(values, percentile):
    if not values:
        raise ValueError("percentile requires at least one value")
    if not 0 < percentile <= 1:
        raise ValueError("percentile must be in (0, 1]")
    ordered = sorted(values)
    rank = math.ceil(percentile * len(ordered))
    return ordered[rank - 1]

successful_ms = [
    820, 910, 1040, 1100, 1270, 1390, 1540, 1810, 2250, 3100
]
print("successful p50 ms:", nearest_rank(successful_ms, 0.50))
print("successful p95 ms:", nearest_rank(successful_ms, 0.95))

The expected output for this illustrative input is successful p50 ms: 1270 and successful p95 ms: 3100. These values demonstrate the calculation only. Ten observations are not a recommendation for a benchmark sample size, and the example says nothing about any real agent’s latency. Select a sample size suited to the variability and decision risk, and report the count supporting each percentile.

Calculate p50 and p95 separately for successful completion and for each important terminal failure category. Also report the success rate and the count that reached the deadline. A configuration with a lower successful p95 but a materially lower success rate may be a worse choice. Likewise, failure durations should not be mixed into successful latency: an early validation failure can pull a combined percentile down and create the appearance of speed.

For stage attribution, calculate a task-level duration for each mutually exclusive stage bucket using the same task boundaries. Summarize each bucket with its own distribution, but do not assume that the p95 of each stage adds up to the end-to-end p95: percentile values may come from different tasks. To explain an individual slow task, inspect that task’s timeline. To describe the population, report stage distributions and the proportion of tasks in which each stage was the largest contributor.

Where stages are sequential and exclusive, a task’s stage durations should reconcile with its end-to-end duration within the known timestamp precision. Establish a tolerance in advance. Where work overlaps, provide a separate critical-path view and state that branch-time sums are not additive. For a p95 cohort, it can be useful to show the stage composition of the slowest successful tasks, but define how that cohort is selected and retain their outcome checks.

A compact report table can show, by configuration and workload category: accepted task count; successful count; successful p50 and p95; timeout count; terminal failure counts; and median and p95 duration for major stages. Include units in the headings. Publish the dataset revision and quantile convention beside the table so the numbers remain interpretable when copied elsewhere.

Step 7: Work through a hypothetical comparison

Suppose a team tests two hypothetical configurations against the same fixed order-status workload. Each run contains 200 admitted tasks across three categories, uses the same timeout, and applies the same independent order and response checks. For this worked example, assume the runs are sufficiently controlled for the team to compare them, while recognizing that the numbers are invented solely to demonstrate analysis.

Imagine configuration A has successful p50 of 1.4 seconds and successful p95 of 4.8 seconds, with 190 successes and 10 terminal failures. Configuration B has successful p50 of 1.2 seconds and successful p95 of 4.1 seconds, with 172 successes and 28 terminal failures. Those fictional values make B look faster among successful tasks, but its success count is lower. The appropriate conclusion is not “B is faster” without qualification; it is that B’s successful subset completed sooner in this illustrative run while fewer tasks passed the outcome checks.

Now suppose stage attribution shows that the slow successful tasks in A spend most of their interval waiting on a second order lookup, while B’s slow successful tasks spend more time in response validation and fewer include a second lookup. That pattern could be consistent with several causes: different task paths, changed retry behavior, measurement error, or an actual stage change. The breakdown directs the next investigation; it does not establish causality by itself.

The team should first verify that the category counts and task order are comparable, then inspect traces for the long-tail tasks, retry counts, and terminal reasons. It should check whether B’s failures were concentrated in a category or dependency, and whether the stage boundaries were recorded consistently in both configurations. Only after these checks should it consider a follow-up run targeted at the suspected bottleneck. This is the value of a trace-based benchmark: it turns a summary difference into a set of testable explanations.

An illustrative calculation can show the importance of weighting. If a workload has 80% category X and 20% category Y, an aggregate percentile is calculated from the task-level durations under that mix, not by taking 80% of category X’s p95 and adding 20% of category Y’s p95. Percentiles are rank statistics, not weighted averages. To compare under another expected mix, construct the same declared task-level composition for both configurations or report category results separately.

A useful interpretation note for this hypothetical run would read: “Configuration B has a lower successful p50 and p95 in this sample, but also fewer verified successes. Do not select it on latency alone. Investigate the 28 failures, compare category-level distributions, and repeat the run if the failure pattern can be explained or corrected.” That statement preserves what the data supports without claiming a cause or promising a future result.

Step 8: Diagnose misleading results and operational failures

The most common analytical trap is survivorship bias. If a report includes only tasks that returned successful results, it can omit the slowest tasks that timed out or were still open when collection stopped. Preserve admission counts, terminal status, and deadline outcomes. A high successful p95 with many timeouts tells a different story from the same p95 with nearly all tasks successful.

A second trap is a shifting boundary. One run may start timing at HTTP acceptance, while another begins after queueing; one may end at the agent’s response, while another includes verification. Even correct timestamps then answer different questions. Store the boundary definition in the manifest and validate it with a controlled event sequence whose expected duration and stage order are known.

Missing spans can make an apparent stage improvement. If a tool call is not recorded during one run, its time may be attributed to orchestration or disappear from the stage table even though the end-to-end duration remains high. Compare the count of accepted tasks with the count of terminal records and expected stage events. Review trace completeness by configuration and category, not only in a few examples selected because they look unusual.

Retries can distort both tail latency and interpretation. An attempt-level report may contain many short events and hide that a task retried repeatedly. A task-level report should retain total elapsed time and retry count. Compare the distribution of retries across configurations, and inspect whether a change in timeout or retry policy altered the number of tasks that could finish. Do not infer that a stage is slow merely because it appears often; frequency and duration are separate measurements.

Clock issues matter when stages run across processes. A negative interval, impossible event order, or stage sum that differs dramatically from task duration may indicate clock skew, duplicated events, or an incorrect parent-child relationship. Use same-process monotonic durations where appropriate and cross-process timing only to the accuracy supported by the instrumentation. If precise attribution is not possible, report a coarser stage or mark the allocation uncertain rather than manufacturing precision.

Traffic conditions can change during a run. A dependency slowdown, queue buildup, or background workload can affect the tail without a code change. Capture the run window and any known dependency condition. Interleaving runs and preserving run-level distributions help identify temporal patterns. They cannot guarantee that external conditions were identical, so a single observed difference should not automatically be treated as a stable effect.

A final counterexample: suppose a new configuration quickly returns a generic response whenever an order lookup exceeds a short internal threshold. Its p50 and p95 for returned responses might fall because many tasks now end early. If the outcome checker counts those generic responses as failures, the faster values are not evidence of better task completion. If the checker mistakenly accepts them, the evaluation itself is defective. Review response validity and failure categories alongside latency, and keep the completion contract independent of the implementation being measured.

When a benchmark result is surprising, follow a fixed triage sequence: verify cohort and boundaries; check event completeness and clock plausibility; segment by workload category and terminal outcome; inspect task timelines and retries; compare run conditions; then design a targeted rerun. This ordering avoids prematurely tuning a stage based on a summary statistic that may be an artifact of changed inputs or missing events.

Step 9: Turn the result into an engineering decision

Decide in advance what latency change would matter to the workflow, and what success-rate or failure change would make a latency improvement unacceptable. These are product and operational thresholds, not universal constants. Derive them from the task’s user-facing needs and the cost of an incorrect or incomplete result. Record the decision rule before inspecting the comparison so the team is less tempted to redefine “good enough” around an attractive number.

Use p50 to understand the typical successful task and p95 to understand a slower part of the successful distribution. Neither statistic describes every task. Add the sample count, success rate, deadline count, and category breakdown so readers can see how much evidence supports each figure. If the decision is important and the observed difference is close to ordinary run-to-run variation, collect additional runs or use a suitable uncertainty analysis rather than treating a small numerical gap as decisive.

Time allocation can suggest where to investigate, but it does not automatically identify the highest-value fix. A stage that consumes a large share of elapsed time may be difficult to change or may be necessary for correctness. Conversely, a short stage that causes many failures can deserve priority despite adding little latency. Combine the timing evidence with outcome failures and the workflow’s requirements when choosing the next experiment.

Keep model-selection questions separate from this workflow measurement. If a particular stage appears to dominate, first establish whether the delay is actually in that stage and whether the outcome remains correct. Decisions about routing across models or upgrading a model have distinct tradeoffs; see Model Routing for Enterprise Agents: Cheap First, Escalate on Evidence and When to Upgrade an Agent’s Model—and When to Fix Its Tools. The benchmark here supplies workflow evidence, not a routing or model recommendation.

Also distinguish this controlled workflow test from a general-purpose task benchmark. A benchmark suite can reveal behavior on its own dataset and scoring rules, but it may not reproduce your tool waits, queueing, validation, or workload mix. For context on interpreting task benchmarks, see What SWE-bench, Terminal-Bench and OSWorld Can—and Cannot—Tell You. Use external results as a reason to formulate a question, not as a substitute for measuring your own defined workflow.

Step 10: Reproducible protocol and printable worksheet

Use the following worksheet in a run ticket or benchmark report. Fill it before the first run, then attach the raw event export and derived tables. A blank field is a useful warning: if the team cannot specify the boundary, workload, or terminal outcome, it is not ready to make a strong latency comparison.

Protocol worksheet

  • Decision being tested: State the concrete engineering decision this benchmark will inform.
  • Workflow and start boundary: Name the task and the exact event that starts user-visible elapsed time.
  • Terminal definition: List success and each failure, timeout, cancellation, or still-open state.
  • Independent success check: Specify how correctness is checked without relying only on the agent’s claim.
  • Stage map: Name each exclusive timing bucket and describe how retries and parallel work are attributed.
  • Dataset revision: Record task categories, fixture references, category counts, and the fixed workload mix.
  • Run conditions: Record configuration, application build, concurrency target and achieved admissions, run order, and test window.
  • Deadline and warm-up: State the timeout and which tasks, if any, are classified as warm-up.
  • Data-quality rules: List missing-event, duplicate, clock, and exclusion checks; preserve excluded records with reasons.
  • Statistics: State the percentile convention and report successful p50/p95, terminal outcomes, counts, and category breakdowns.
  • Decision rule: Specify what latency difference matters and what outcome degradation blocks adoption.
  • Reproduction record: Save the manifest, event schema, analysis code version, raw events, and report date.

A completed report should show the end-to-end successful p50 and p95 beside the number of successful tasks, not as isolated headline values. It should also show timeouts and other failures, category counts, major stage distributions, and any limitation in event coverage. If it reports only the best run or only the aggregate, retain that choice as a separate view rather than presenting it as the full result.

For teams establishing this method across several workflows, an evaluation plan can help align task definitions, success checks, and operational questions before implementation. AI evaluation and strategy is one relevant next step when the work needs broader evaluation planning. Keep the engagement focused on the decisions and measurements your team needs; the worksheet is intended to be usable without an external service.

The output of this tutorial is a protocol, a traceable event dataset, and a decision-ready report—not a universal ranking. Start with one bounded workflow, validate its boundaries and outcome checks, and repeat the same defined workload across the configurations that matter. When the p95 moves, the trace should make it possible to say which task paths changed, how stage time shifted, and whether verified completion improved along with the clock.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call