Table of contents
- Purpose and boundaries
- The 90-day delivery model
- Set the calendar and ownership
- Days 1–15: make the workflow measurable
- Days 16–30: build the smallest credible system
- Days 31–50: evaluate behavior and failure
- Days 51–65: prepare operations and controlled use
- Days 66–80: run a bounded production trial
- Days 81–90: decide, document and hand over
- Worked example: invoice exception triage
- Failure patterns and schedule recovery
- Printable delivery worksheet
- Summary and next steps
Purpose and boundaries
An AI pilot becomes a delivery project when a team can name the users, define the work the system may perform, measure acceptable behavior, and assign people to operate it. The transition is not achieved by extending a demo or adding a production label to a prototype. It requires evidence gathered in a sequence: first that the workflow is worth addressing, then that a bounded implementation behaves acceptably, and finally that the organization can support the resulting service.
This guide lays out that sequence over 90 calendar days. It is a planning baseline, not a promise that every system can be production-ready in three months. A workflow with difficult data access, high-impact decisions, complex integrations or unresolved ownership may need a longer plan—or may not be suitable for this schedule. The useful outcome at day 90 can be a controlled launch, a narrowly scoped continuation, or a documented stop. Spending the full period does not make launch the correct answer.
The scope here is delivery: who does what, when, and what evidence should exist before the next stage begins. It does not choose your organization’s first AI workflow or replace a separate decision framework for whether to ship or stop. If those decisions are still open, use a framework for selecting the first workflow to resolve the decision criteria before committing the delivery calendar.
A small delivery team still needs distinct responsibilities. A business owner is accountable for the workflow outcome; a technical lead owns the implementation; an operational owner will run the service; and a reviewer with authority to pause the work evaluates evidence at gates. One person may cover more than one role in a small organization, but the responsibilities should remain explicit. When a role has no named person, record it as an unresolved dependency rather than assuming the project team will absorb it indefinitely.
Treat the dates as working ranges. Days 1–30 establish scope and build a testable slice. Days 31–65 develop confidence in behavior and operations. Days 66–90 test the proposed service under bounded real use, then make and record a decision. A gate is a decision point, not a ceremonial meeting: evidence is reviewed, gaps are assigned, and work either proceeds, narrows, pauses or ends.
The 90-day delivery model
A useful delivery plan connects each phase to an artifact and a decision. A meeting without an artifact can create the impression of progress while leaving the next team unable to verify what was agreed. The artifacts in this plan should be concise enough to maintain: a one-page workflow definition, an evaluation set, a release and rollback plan, an operating record, and a final decision memo.
| Period | Main purpose | Required evidence | Gate decision |
|---|---|---|---|
| Days 1–15 | Bound the workflow and define measurement | Workflow map, baseline, exclusions, owners | Is the workflow sufficiently clear and valuable to prototype? |
| Days 16–30 | Build a narrow, testable slice | Data and integration map, working path, initial test set | Can the system be evaluated against representative cases? |
| Days 31–50 | Evaluate behavior and failure modes | Results by case type, defect log, revised design | Are remaining risks understood and addressable? |
| Days 51–65 | Prepare service operations | Monitoring plan, runbook, access and rollback design | Is the service supportable in a controlled trial? |
| Days 66–80 | Observe bounded real use | Trial log, user feedback, service and outcome measures | Is evidence strong enough for a launch decision? |
| Days 81–90 | Decide and transfer responsibility | Decision memo, named operating owner, follow-up plan | Launch narrowly, continue with conditions, or stop? |
The calendar deliberately separates model or system behavior from business outcome. A response that appears plausible is not proof that the user’s task was completed correctly. Measure both: whether the system produced an acceptable result under defined conditions, and whether the workflow changed in a way that matters to the organization. NIST describes its AI Risk Management Framework as voluntary, rather than a certification or guarantee; use it as a way to structure risk conversations, not as proof that a system is safe or ready. NIST AI Risk Management Framework
The diagram summarizes the gates. A failed gate should lead to a specific action: narrow the workflow, repair a known defect, gather missing evidence, or stop. The route back is intentional; it is not an invitation to repeat the same experiment without changing the design or acceptance criteria.
flowchart TD
accTitle: AI pilot delivery gates across ninety days
accDescr: A bounded workflow moves through definition, prototype, evaluation, operational readiness, and a controlled trial. A failed gate returns to scoped repair or evidence gathering; the final gate chooses a narrow launch or stop.
A["Define workflow and measures"] --> B["Build testable slice"]
B --> C["Evaluate cases and failures"]
C -->|"Evidence adequate"| D["Prepare operations"]
C -->|"Gap found"| B
D -->|"Ready for bounded use"| E["Run controlled trial"]
D -->|"Not ready"| C
E --> F["Launch narrowly or stop"]
The loop from evaluation to the build stage means the team changes something material—such as scope, instructions, data handling or a human review point—and then retests. Returning from operational readiness to evaluation means that an operational concern has changed the risk picture and needs evidence, not just a runbook edit. The final node includes two different outcomes; a production decision should not be framed as launch versus project failure.
Set the calendar and ownership
Before day one, reserve recurring working time for the people who can make decisions and access the workflow. A practical cadence is a short weekly delivery review, a more detailed gate review at phase boundaries, and a lightweight operational check during any live trial. The business owner should be able to attend the gates. If a key reviewer can only attend monthly, move the gate dates to match that constraint instead of building a plan around unavailable approvals.
Write down what each role can decide. The business owner defines the workflow boundary and confirms whether the outcome is useful. The technical lead can recommend an architecture and owns technical evidence. The operational owner confirms support procedures and capacity. The gate reviewer can reject evidence or require a narrower trial. A privacy, security, legal or domain specialist should be brought in when the actual data, users or consequences make their expertise relevant; this is not a reason to invent universal approval steps detached from the system being built.
Name a decision-maker for each gate and agree how decisions are recorded. A useful record has the decision, evidence reviewed, unresolved risks, conditions, owner and due date. If reviewers disagree, record the disagreement and the person authorized to resolve it. Do not convert an unresolved concern into a green status merely to preserve the date.
A plan also needs explicit capacity. Estimate the time people can actually commit, not just the ideal calendar. For instance, if a workflow owner can provide two hours a week, schedule observation and review around those two hours. Do not plan a 30-case labeling exercise that requires a full workday from that person without negotiating capacity. A compressed schedule that depends on invisible labor is a schedule risk, not a productivity win.
Finally, capture external dependencies in a short list: data access, a system owner, a test environment, vendor or procurement decisions, and any required operational support. Give each dependency an owner and a latest useful date. If a dependency misses that date, decide whether to use a safe substitute for evaluation, narrow the scope, or move the gate. Avoid using production data or access as a workaround for a missing test path without evaluating the consequences.
Days 1–15: make the workflow measurable
Start with the work as people perform it now. Observe a small set of representative cases, interview the people who handle exceptions, and draw the current path from incoming request to completed outcome. The objective is not to document every organizational variation. It is to identify where the proposed system would enter, what information it would receive, what it may return or change, and where a person remains responsible.
The workflow definition should specify the input, expected output, user, downstream action and exclusions. For an internal support process, the input might be a request with a category and message; the output might be a suggested routing queue and a short rationale. An explicit exclusion could be that the system does not close a case, send a customer-facing message, or make a decision about eligibility. Exclusions are useful because they constrain what the team must build and test in the first 90 days.
Establish a baseline before choosing a success threshold. Depending on the workflow, the baseline may include handling time, rework rate, queue age, escalation frequency or the proportion of cases routed correctly. Define how each measure is counted: which cases are included, what event starts and stops the clock, and what period is sampled. A before-and-after comparison is weak if the populations or measurement rules differ.
Choose a small set of representative cases and label what a good outcome looks like. Include routine inputs, incomplete information, ambiguous cases and known exceptions. For each case, record the expected result and the reason a human might disagree. If labels are subjective, ask two domain reviewers to label a sample independently and resolve disagreements before treating the set as a scorecard.
At the day-15 gate, ask whether the team can explain the work, the baseline, the exclusions and the measurement method in plain language. If the workflow is still changing in ways that alter the input or desired outcome, do not pretend a prototype can answer the original question. Narrow to a stable slice or use the next two weeks to settle the operating definition.
Gate artifact: a one-page workflow brief. Include the user, current process, proposed intervention, input and output, exclusions, baseline measures, representative case plan, business owner, technical lead, operational owner, and the day-15 decision. This document should be useful to someone joining later without requiring them to infer the project’s purpose from chat history.
Days 16–30: build the smallest credible system
Build only what is needed to produce evidence against the agreed cases. That usually means a narrow path from an authorized input to a reviewable output, with enough logging to connect a test case to its result. Resist building a complete interface, broad workflow automation or generalized agent behavior before you know which behavior needs to be evaluated.
Document the data path. List the source fields, transformations, model or service boundary, output fields, storage points and downstream actions. Mark which fields are required and how missing or malformed values are handled. This reveals whether a bad result comes from the model, the input preparation, a stale source record or an integration assumption. It also helps the team keep sensitive or irrelevant data out of the prototype where possible.
Define a stable test interface for the evaluation set. A case record might include an internal case identifier, the input text, relevant structured fields, expected routing, an acceptable alternative, and a reason for escalation. Do not use a customer name or other identifying detail when a synthetic or redacted case can test the same behavior. Keep the mapping to any live record separately controlled if such a mapping is necessary.
Keep external effects out of the early prototype. A returned suggestion should not silently change a production record, send a message, or trigger a payment. If a later trial requires an action, create a separate design for human review, exact action confirmation, failure handling and auditability. A model response is not itself authorization to perform an external action.
At the day-30 gate, demonstrate that the system can process the agreed test cases and produce outputs that can be inspected. The demonstration is not a claim of accuracy or readiness. It is a check that the evaluation can now proceed. If the team cannot reproduce a result or trace it to an input and configuration, fix the testability problem before drawing conclusions from output quality.
Gate artifact: a working test path, versioned case set and data-flow sketch. Record the implementation version and the configuration used for each evaluation run. Without that context, comparisons across revisions may mix different inputs or settings and create a misleading story about improvement.
Days 31–50: evaluate behavior and failure
Evaluate by meaningful case category rather than a single overall score. A system that handles routine cases well may still fail on missing fields, unusual wording or high-consequence exceptions. For each category, report the number of cases reviewed, the acceptable outcomes, the unacceptable outcomes and the uncertain judgments. Small samples need visible counts; a percentage can imply more precision than the data supports.
Use a structured defect log. Each entry should include the case identifier, observed result, expected result, severity, likely cause, proposed remedy, owner and retest status. Distinguish defects that change the intended outcome from presentation issues, and distinguish a model-behavior problem from a workflow or source-data problem. This helps direct effort to the part of the system that can actually fix the issue.
Test adversarial and awkward but plausible inputs that come from the workflow: contradictory fields, a blank description, unusual abbreviations, a request outside the supported categories, or source data that has changed since a previous step. The goal is not to produce an impressive collection of edge cases. It is to identify the inputs for which the system should abstain, ask for clarification or route to a person rather than invent a confident answer.
Set acceptance criteria before reviewing the final results. Define critical errors that block progression, minimum acceptable behavior for the bounded use, and the evidence needed to justify exceptions. For example, a routing suggestion might be allowed to have an alternative queue if both queues lead to the same trained review team, while a suggestion that bypasses a required human review would be a blocking error. The acceptable threshold must be tied to the workflow consequence, not selected because the current system happens to meet it.
Measure human review burden as well as output quality. If every result requires a person to inspect all source material and reconstruct the recommendation, the system may add work even when the output looks good. Ask reviewers to record whether they accepted the suggestion, changed it, or could not assess it, and why. Sample review should reflect how the proposed workflow will actually be used; a relaxed review process during evaluation can hide the cost of a more demanding real process.
At the day-50 gate, the team should be able to explain what fails, for whom, and under which input conditions. If a critical defect repeats after a design change, stop the progression to operational readiness until the cause is understood or the workflow is narrowed to exclude the problematic cases. The purpose of this gate is not to eliminate every possible error. It is to prevent known, material failure modes from being treated as an acceptable surprise.
Gate artifact: an evaluation summary by case category, a defect log and a list of open risks with owners and dates. Keep the raw case-level results available to reviewers. A polished summary without traceable cases is insufficient evidence for a meaningful decision.
Days 51–65: prepare operations and controlled use
A technically functional prototype is not an operated service. Before real use, identify who notices a failure, who can pause the workflow, how the user reports a problem, and who restores the prior process. The operational owner should review these responsibilities against actual working hours and support capacity. If nobody can respond during the period the workflow is used, reduce the trial window or keep the system in shadow mode.
Define observable service behavior in terms users experience. For example: the proportion of eligible requests that receive a result within the agreed window, and the proportion that are available for review without a processing error. Specify the denominator and measurement window. Google’s SRE guidance describes service-level objectives as measures of user-relevant service behavior and emphasizes defining the denominator and window. Implementing SLOs
Choose indicators that reveal a meaningful operational problem: processing failures, delayed results, a rise in manual corrections, a change in escalation volume, or a growing queue of cases awaiting review. Define thresholds that trigger investigation or pause, and identify who receives the signal. Do not add a dashboard merely to create the appearance of monitoring; every indicator needs an interpretation and an action.
Write a short runbook for the trial. It should say how to identify an affected case, where to find the relevant result and configuration, how to stop new cases entering the process, how to return to the previous workflow, and whom to contact. Include steps for a partial outage and for a result that was produced but not acted on. Specify what should be recorded before a case is retried, so a retry does not create duplicate work or an unintended external effect.
Plan the release boundary. State which users, case types, hours and downstream actions are included in the trial. If the system is only advisory, make the user action explicit and preserve the existing route for cases that the system cannot assess. If traffic is being shifted between versions, do not assume a hosted model endpoint provides a traffic-splitting control unless that has been verified for the chosen setup. A proposed canary may require external routing and a tested fallback.
The day-65 gate is a readiness review, not a launch announcement. Confirm that the operating owner has accepted the runbook, the pause mechanism has been demonstrated in a safe setting, the trial population is defined, and open risks have an explicit disposition. A missing support owner or uncertain rollback path should block live use, even if the evaluation results look promising.
Days 66–80: run a bounded production trial
Begin with a limited group of users or case types and a clearly stated observation period. The objective is to learn how the system behaves in the actual workflow, not to maximize volume. Preserve the prior process so that a person can continue handling work if the system is paused. Tell trial users what the system is and is not intended to do, what to review, and how to report a problem.
Where practical, start in shadow mode: generate results that are logged for comparison but do not affect the decision or action taken. Shadowing can reveal input differences and operational defects without exposing the workflow to the system’s output. It does not establish that users will understand or appropriately rely on an interactive recommendation, so it cannot substitute for a controlled test of the human interaction when that interaction is part of the proposed service.
Review a sample of cases on a regular schedule during the trial. Include routine and exception cases, and record the input, output, human decision, eventual outcome where available, processing delay and any correction. Avoid reviewing only cases users reported as problematic; reports are valuable but are not a representative sample. Set a sampling plan that fits the workload and the consequence of error, and reduce scope if the team cannot review enough to support the decision.
Separate immediate safety or workflow signals from slower business measures. A failed processing step may be visible the same day; a change in rework or completion time may need a longer observation period. Compare trial results to the baseline using comparable case types and measurement rules. Note changes in staffing, demand, policy or seasonality that may affect the comparison, instead of crediting every difference to the new system.
Use a pause rule written before the trial. Examples include a critical misroute, repeated processing failures, an unexplained change in the input population, or the loss of the person responsible for review. The exact rules depend on the workflow. When triggered, stop or narrow the relevant path, preserve records needed to investigate, and use the prior process. Avoid improvising a threshold after a result arrives simply because the team wants to keep the trial on schedule.
The day-80 gate reviews both evidence and practical burden. Do users understand when to rely on an output? Is review work manageable? Are errors caught before they cause downstream effects? Are service indicators within the agreed bounds? If results are mixed, identify whether the issue is the workflow boundary, system behavior, training, operations or sample size. A vague instruction to improve accuracy is not a useful continuation plan.
Days 81–90: decide, document and hand over
Reserve the final ten days for a decision and transfer, not an unplanned expansion of the trial. The decision memo should compare observed behavior with criteria set before the trial, explain the limits of the evidence, summarize operational incidents, and state the proposed next scope. Include the cases or measures that did not meet expectations; do not bury them in an appendix that decision-makers are unlikely to read.
Choose among three practical dispositions. A narrow launch means the evidence supports a defined workflow and the operating team accepts responsibility for it. A conditional continuation means there is a specific unresolved gap, a bounded next experiment, a named owner and a date for review. A stop means the expected value, behavior or operating burden does not justify more work now. These dispositions are more useful than a binary label that implies a pilot either succeeded or failed.
For a launch, record the allowed users and case types, the actions the system may support, the review requirement, the pause trigger, the service indicators and the accountable operational owner. State what changes require reevaluation, such as a material change to input data, workflow, downstream action or model configuration. This is a concise release boundary, not a substitute for the organization’s normal operational controls.
For a conditional continuation, define the smallest experiment that could resolve the gap. If the evidence is weak because a rare category was not sampled, collect more relevant cases. If reviewers disagree on what counts as correct, revise the rubric and relabel a defined sample. If queue delays are the issue, test the workflow under the actual operating schedule. Do not extend the project by default with an open-ended list of improvements.
For a stop, preserve the useful work: the workflow map, test cases, defect findings, integration notes and decision rationale. Remove or retire trial access according to the organization’s practices, and tell affected users what will happen next. A documented stop can prevent another team from rebuilding the same prototype without knowing why the earlier attempt was unsuitable.
Hand over operational ownership explicitly. The team that developed the system may not be the team that responds to a failure after launch. Name the person who maintains the runbook and service measures, the person who can pause the process, and the channel for users to raise issues. A deeper treatment of the ongoing operating role is available in the discussion of AI-agent ownership after launch; this 90-day plan only establishes the handoff point.
The final gate record should state the disposition, decision-maker, evidence date range, material limitations, outstanding risks, follow-up owner and review date. If changing model versions is part of a later plan, treat that as a new evaluation question rather than assuming a different version preserves the behavior observed in this trial. Questions to ask before approving a new frontier model addresses that separate decision.
Worked example: invoice exception triage
The following scenario is hypothetical and illustrates how the calendar can constrain a real workflow. A mid-sized distributor receives supplier invoices through a shared operations queue. Staff compare each invoice to purchase-order information, identify discrepancies and route exceptions to the appropriate team. The proposed system would extract a small set of fields and suggest a discrepancy category. It would not approve invoices, change payment details, or send messages to suppliers.
The business owner chooses a bounded question: can a suggestion help trained staff route eligible exceptions without increasing correction work or causing urgent cases to be missed? The first version excludes invoices with missing purchase-order references, suspected fraud, changed bank details and cases requiring a policy judgment. Those exceptions remain in the existing process. This is deliberately narrower than automating invoice processing end to end.
In days 1–15, the team observes how staff classify cases and finds that the current queue records arrival and assignment times but not consistently the reason for reassignment. It establishes a baseline from a defined recent period using cases with complete timestamps. The team records case volume, time to first assignment, the percentage reassigned, and the existing queue categories. It labels a sample of historical cases with the expected category and notes acceptable alternatives where two queues share responsibility.
During days 16–30, the technical lead builds a path that takes an eligible invoice record and returns a proposed category plus the extracted fields used to support it. The output is visible to an internal reviewer, but no queue or payment record changes automatically. The test set includes routine invoices, low-quality scans, missing references and inconsistent supplier identifiers. Each case has an internal identifier so reviewers can connect a result to the source without putting supplier names into a wider report.
In days 31–50, results are reviewed by category. The team discovers that two categories are often confused because staff use a local abbreviation that is absent from the reference data. Rather than raising a global accuracy target, the team decides whether that category can be excluded, the reference data can be corrected, or the user interface can request confirmation. It retests the changed design against the same cases and a separately held set, recording both the initial and revised results.
During days 51–65, the operational owner confirms that two trained staff members can review trial suggestions during agreed weekday hours. The runbook says to route unreadable scans and unsupported categories to the usual queue, stop the suggestion path if the result service is unavailable, and leave the invoice record unchanged until a human selects a category. The trial does not run overnight because no reviewer is available to examine exceptions then.
For days 66–80, the team first observes suggestions in shadow mode, then makes them visible to a small group of trained reviewers for a limited category set. Reviewers record acceptance, correction or escalation and the reason. The trial measures time to first assignment and reassignment using the same definitions as the baseline, while also checking whether staff have to repeat the extraction work. The business owner reviews a sample of cases each week rather than relying only on voluntary incident reports.
The hypothetical day-80 review finds that routine categories are easier to route, but one class of exception still causes inconsistent suggestions and additional review. The team does not describe the trial as a success simply because the common cases look good. It proposes a narrow next step: keep the supported categories, exclude the inconsistent class, and extend observation only long enough to determine whether the revised boundary reduces reassignments without shifting work to another team.
At day 90, the decision memo can support a limited launch only if the operating owner accepts the scope and the excluded cases still have a reliable route. Otherwise, the team records a conditional continuation with a specific evidence gap. This example illustrates why a 90-day plan is valuable even when it does not end in broad automation: it produces a bounded decision and leaves the existing business process intact when the evidence is not yet adequate.
Failure patterns and schedule recovery
A demonstration arrives before the problem is defined. The team builds an attractive interface in the first two weeks, then discovers that staff disagree about the desired category. Recover by pausing feature work, defining the labels with domain reviewers, and revising the day-15 gate. A late agreement about what counts as correct cannot be repaired with more model iterations alone.
The evaluation set mirrors easy work. Routine examples dominate because they are easiest to collect, while the cases that cause costly rework are missing. Recover by comparing the test set with the workflow’s actual case distribution and deliberately collecting important exception types. If the team cannot evaluate a material category, exclude it from the proposed trial rather than silently treating its absence as evidence of success.
A passing aggregate hides a consequential defect. Overall performance can conceal a failure concentrated in a rare but important class. Recover by reporting counts and outcomes by category, identifying blocking errors, and changing the workflow boundary or human review point. Do not choose a threshold solely to achieve a pass on a blended score.
The live trial changes more than one thing. New staffing, a changed queue policy and a system rollout occur together, so outcome changes cannot be attributed confidently. Recover by recording the changes, separating the observation period where possible, and avoiding causal claims the design cannot support. If separation is impossible, the final decision can still use direct case review and operational evidence, but should state the attribution limit.
A human review step becomes ceremonial. Users accept suggestions quickly because a queue is busy, even though the intended design assumes careful review. Recover by checking actual correction patterns and user behavior, clarifying which fields must be verified, and reducing the workflow scope if meaningful review is not feasible. Do not claim human oversight merely because a person can technically override an output.
The team reaches day 70 without an operational owner. This is a delivery failure, not a documentation gap. Pause the live trial, identify a person with time and authority to run the workflow, or keep the system out of production use. If the organization needs help establishing the delivery leadership and decision cadence, consider whether fractional CTO leadership is appropriate for the need; it is one possible support route, not a prerequisite for following this plan.
When a dependency slips, preserve the evidence sequence rather than compressing every remaining phase. For example, if data access is delayed by a week, use the time to refine the rubric or build a clearly labeled synthetic test path if that can test the same behavior. Do not use fabricated cases to claim real-world performance. If the delay prevents representative evaluation, move the gate and record why.
If the schedule is too long for the organization’s decision window, narrow the workflow before removing gates. Limit case types, users, operating hours or external actions. Cutting the evaluation period while retaining the same scope often reduces confidence precisely where the system’s consequences are largest. A smaller decision supported by adequate evidence is more useful than a broad decision supported by hurried evidence.
Printable delivery worksheet
Use the following worksheet in the kickoff and update it at each gate. The checkboxes are not a substitute for evidence: attach or link the named artifact, record the person responsible, and mark a gap openly. Keep the completed version with the decision record so that a new owner can understand the delivery state.
Scope and baseline
- Workflow and user: Write one sentence naming the person doing the work and the specific step where the system enters.
- Input and output: List fields or content entering the system and the exact result returned to the user.
- Excluded cases and actions: Name what the system must not handle or change during this delivery period.
- Current process: Attach a simple workflow map showing handoffs, exceptions and the existing route when work cannot proceed.
- Baseline: Record the measure, case population, time period, inclusion rules and data source.
- Decision roles: Name the business owner, technical lead, operational owner and gate decision-maker, including any combined roles.
Build and evaluation
- Data path: Record source fields, transformations, service boundary, stored outputs and downstream effects.
- Case set: List the case categories, expected outcomes, acceptable alternatives and how cases were selected.
- Acceptance criteria: Set critical errors, minimum behavior and escalation conditions before reviewing final results.
- Defect log: Record each material failure, cause hypothesis, assigned owner, remedy and retest result.
- Human review: Describe what the reviewer checks, what they can change, and when they must escalate.
- Version record: Record the implementation and configuration associated with each test result.
Trial and decision
- Trial boundary: Specify users, case types, hours, duration and permitted downstream actions.
- Operating owner: Name who monitors results, handles user reports and can pause the workflow.
- Service measures: Define user-relevant measures with their denominator, measurement window and response threshold.
- Fallback: Describe how work continues when the system is unavailable, uncertain or paused.
- Pause conditions: Record concrete triggers and the person authorized to act on them.
- Evidence summary: Compare observed results with the pre-agreed criteria and state evidence limitations.
- Disposition: Select narrow launch, conditional continuation or stop; name the decision-maker and rationale.
- Handoff: Record the operating owner, runbook location, unresolved risks, follow-up actions and review date.
For a printable one-page summary, copy the following fields into your project record: Workflow: ___; scope and exclusions: ___; business owner: ___; technical lead: ___; operational owner: ___; day-15 decision: ___; day-30 evidence: ___; day-50 evidence: ___; day-65 readiness decision: ___; trial boundary: ___; pause trigger: ___; day-90 disposition: ___; unresolved risk and owner: ___; next review date: ___. A blank field is a visible delivery gap, not a reason to infer agreement.
Summary and next steps
A useful 90-day AI delivery plan advances through evidence, not enthusiasm. It begins by defining the task and its baseline, builds only enough to test the behavior, evaluates errors by case type, prepares an operating path, and trials a bounded workflow with a real pause mechanism. It ends with a documented decision that can be narrow, conditional or a stop.
Start by naming the workflow owner and writing the exclusions. Then set the first gate date, reserve reviewer capacity, and agree what evidence must exist before building. Use the worksheet to expose missing ownership and dependencies early. At each gate, change scope or timing when the evidence requires it; do not treat the 90th day as a launch deadline.
If your team has already chosen a workflow but lacks the technical or delivery leadership to coordinate the work, clarify the role you need before adding people or tools. A guide to choosing between an advisor and a delivery partner can help frame that distinction. The next practical action is smaller: book the kickoff, name the four delivery responsibilities, and leave the first meeting with a measurable workflow brief and a date for the day-15 decision.