SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Explainer 25 mins

What Good AI Delivery Evidence Looks Like in a Case Study

A credible AI case study shows the baseline, timeframe, sample, failures and client-approved outcomes—not just a compelling result or model claim.

The PADISO Team ·

An AI case study is useful when a reader can tell what changed, how the change was measured, which cases were counted, and what did not work. Without those details, a strong-sounding result may be impossible to interpret. A percentage without a denominator, a comparison without a baseline, or a success story without its failures can persuade without helping anyone make a sound delivery decision.

This is not a demand to publish confidential data or every internal implementation detail. It is a way to make a public account specific enough to evaluate while respecting client boundaries. The evidence should connect an operational claim to a defined population, a period of observation, a credible comparison and an outcome the client has approved.

The distinction matters to buyers as well as delivery teams. A buyer needs to know whether the reported result resembles the work under consideration. A delivery team needs a record that can survive handover, internal review and later questions about how the conclusion was reached. A case study that meets both needs is more than promotional copy: it is a compact, inspectable account of a decision and its evidence.

1. What a case study is claiming

A case study usually makes several claims at once, even if its prose presents only one headline. It may imply that a system performed a task, that it improved a process, that the improvement was caused by the system, and that a similar organization could expect a comparable result. Those are different claims, and the evidence needed for each is different.

A model claim concerns system behavior: for example, whether an assistant classified a request correctly under specified conditions. A workflow claim concerns what happened in the operating process: perhaps staff spent less time triaging a defined queue. A business outcome claim connects that workflow change to a result that matters to the organization, such as shorter response time or fewer items requiring rework. A transfer claim suggests the result may apply beyond the observed setting.

Evidence becomes weaker as a case study moves from observed behavior to broad conclusions without adding support. If a sample of test prompts receives a high score, that does not by itself establish that a live workflow became faster. If a workflow became faster during a pilot, that does not by itself establish a durable business benefit. And a result in one team does not establish that another team will see the same effect.

A practical case study makes its central claim narrow enough to support. Instead of saying, “AI transformed customer support,” it might say, “During a six-week pilot, a draft-generation workflow reduced median handling time for one defined category of email, while a human agent reviewed every draft before sending.” The narrower statement identifies a task, a period, a measure and a control point. It also leaves room to state what was not measured.

Use a useful test: could a skeptical reader identify what data or record would confirm or contradict the headline? If not, the headline is probably too broad, the measure too vague, or the evidence trail too thin. The purpose is not to remove interpretation. It is to show the boundary between observation and interpretation.

2. Begin with a baseline that means something

A baseline is the measured state of the process before the change being evaluated. It gives the later result a reference point. “The team was slow” is not a baseline. “For this request category, the median elapsed time from assignment to first response was 18 hours during the four weeks before the pilot” is closer, provided the timestamps and category definition are reliable.

The baseline should describe the same process, population and measure that the case study later reports. If the pre-pilot figure covers all requests but the post-pilot figure covers only easy requests, the comparison is not like-for-like. If the baseline measures elapsed time and the outcome measures active staff time, they may both be useful, but they must not be presented as though they were the same measure.

Record how the baseline was assembled. Note the source system or record type, the relevant dates, the inclusion and exclusion rules, and any known data gaps. A reader need not receive confidential records, but the case study author should be able to explain how the figure was derived. When a number comes from staff estimates rather than timestamps, label it as an estimate and explain the method.

A baseline also needs context about variation. A monthly average can move because the workload changed, not because the workflow improved. Consider whether the process has a weekly cycle, a seasonal peak, a staffing change, a policy change or a change in the mix of cases. The case study does not need to model every possible influence, but it should disclose changes that materially affect interpretation.

Choose measures that reflect the actual job. For a review workflow, elapsed time alone may conceal whether reviewers had more rework. For a customer response process, speed alone may conceal whether answers were complete or whether customers had to follow up. For a document workflow, “documents processed” can be misleading if the system receives a different share of difficult documents. A useful measure pairs throughput or time with an indicator of quality, rework or unresolved cases when those factors affect the claim.

Where a clean baseline is unavailable, say so. A case study can still document a pilot’s observed behavior, describe the current measurement limitations and avoid asserting improvement. That is more credible than manufacturing a comparison from recollection. It can also help a buyer see that measurement design should be part of the delivery plan rather than an afterthought.

3. Define the sample and the denominator

A sample is the set of cases included in the reported analysis. The denominator is the number of eligible cases against which a rate or percentage is calculated. Both should be clear. “The assistant was accurate 92% of the time” leaves basic questions unanswered: accurate on which task, among how many cases, under what conditions, and according to whose judgment?

For a case study, define the unit being counted. It might be one incoming email, one document, one completed workflow, one human-reviewed draft or one customer interaction. Those units are not interchangeable. If a single interaction can generate several model calls or several staff actions, say what counts as one case so that totals are not inflated or misunderstood.

Describe the population the sample represents. A sample may include all eligible work during a stated period, a random subset, a set of scripted evaluation examples, or cases selected by staff. Each approach supports a different inference. A random subset can help estimate behavior in a larger defined population if selection and exclusions are handled appropriately. A hand-selected set of representative examples can help explain a workflow, but it should not be presented as a population-wide performance rate.

Selection rules deserve plain language. State whether the sample excluded incomplete records, unusual cases, duplicate submissions, unsupported languages or items routed around the new workflow. If the system was used only for a narrow category, say how staff identified that category. If difficult cases were escalated before reaching the system, that is important context: the measured sample may describe the routine queue, not all work arriving at the organization.

A small sample is not automatically useless. It may reveal a workflow defect, expose a recurring failure mode or support a decision to continue learning. But a small or selectively assembled sample should not carry a broad claim about expected performance. Report the count beside the result, distinguish exploratory observations from stable operating evidence, and avoid false precision. A result such as “46 of 52 reviewed drafts met the team’s acceptance criteria” is easier to inspect than “88.46% accurate,” especially when the evaluation itself involves judgment.

When sample composition changes over time, report that too. A system may appear to improve because the later period contains easier cases, or appear to worsen because staff began routing edge cases into the workflow. A case study should not treat a shifting population as a stable test set without explanation. If separate segments matter, report them separately when the sample supports doing so; otherwise, describe the limitation instead of implying segment-level certainty.

4. Make the timeframe visible

A timeframe is the period over which the baseline and observed outcome were collected. It is not merely a date range for the project. A case study should distinguish the period used to measure the pre-change process, the pilot or operating period, and the date through which outcomes were reviewed. This allows readers to see whether the result reflects an initial trial, a sustained workflow or a snapshot.

The right duration depends on the work. A short period may capture setup effects, staff learning or a temporary surge. A longer period may capture more variation, but it can also include changes in staffing, policy or workload that need to be accounted for. There is no universally correct duration. The important point is to explain why the observation window is relevant to the claim and what it does not capture.

Separate calendar time from exposure. A pilot that ran for six weeks may have processed only a small number of eligible cases because staff used it intermittently. Another six-week pilot may have covered most of the queue. State how many eligible cases entered the workflow and, where useful, what portion of the eligible work was actually handled through it. Do not let a long calendar period imply broad operational coverage if exposure was limited.

Record meaningful events within the period. A prompt or workflow change, a new review instruction, a staff training session, an outage, a change in incoming volume or a policy adjustment can affect the observed result. The case study need not provide a day-by-day log, but a short chronology helps readers understand whether the result came from one stable configuration or several iterations.

A concise evidence timeline can be expressed in prose: “The team collected baseline records from 1 March through 28 March, introduced the workflow on 1 April, and reviewed eligible cases through 12 May. The first week was treated as setup and excluded from the main time comparison; the exclusion rule was applied to both the count and the outcome calculation.” The dates and rule in such an example must, of course, come from the actual project. The point is to make the timeline and exclusions inspectable.

When a result is early, call it early. Do not turn an initial operating period into a claim of long-term stability. That distinction is especially important for work affected by changing demand, staff turnover or evolving processes. A later review may support a different conclusion, but the first case study should describe only what its observation window can show.

5. Separate system performance from business outcome

System performance and business outcome are related, but one is not a substitute for the other. System performance describes how the system behaves on defined tasks. Business outcome describes what changed for the organization or the people it serves. A case study should show the link between them rather than treating a model score as proof of organizational value.

For example, a high proportion of drafts accepted by reviewers may indicate that the drafts were useful under a stated review rule. It does not automatically show that the process became cheaper, faster or better for customers. Review time may have shifted rather than disappeared. Staff may have spent less time writing but more time checking. The end-to-end process may have stayed the same length because another queue became the bottleneck.

Trace the workflow from input to outcome. Identify the event that begins measurement, the actions the system performs, the human decisions that remain, and the event that marks completion. Then state which part changed. If the headline is about time, specify whether it means active handling time, elapsed time or time to a defined service milestone. If the headline is about quality, state the acceptance rule and who applied it.

Separate directly observed measures from calculated or inferred measures. A timestamp difference may be directly calculated from records, though the records can still be incomplete. A claim about “hours saved” may require assumptions about how many cases were processed, how much staff time each case previously took, and whether review time is included. Make those assumptions visible. If a calculation is illustrative rather than observed, label it as such and do not present it as a realized outcome.

External influences matter. Demand may have dropped, staffing may have increased, or a policy change may have altered the work. Those changes do not invalidate a result, but they affect what can be attributed to the system. If the project had no comparison group or controlled design, use appropriately bounded language: “The measured time was lower during the pilot period” is not the same as “the system caused the reduction.” Describe the evidence and the plausible interpretation separately.

A framework can help teams organize their risk-management work, but it should not be mistaken for an endorsement or outcome credential. NIST describes its AI RMF as a voluntary framework, not a certification or guarantee (NIST AI Risk Management Framework). For case-study writing, the practical lesson is to report the actual evidence and its limits rather than implying that a framework name proves a result.

6. Report failures as part of the result

A failure is any observed behavior or process outcome that prevents the workflow from meeting its stated acceptance criteria. It may be a wrong output, an incomplete output, a refusal where the task required completion, an escalation, a timeout, a formatting defect, a review burden or a human correction. A case study that reports only visible model errors can miss operational failures that matter just as much.

Define failure before counting it. “Incorrect” is not a reproducible category unless the team specifies what counts as incorrect for the task. For a draft, does a minor wording edit count as failure, or only a factual correction? For a classification, which labels are acceptable and how are ambiguous cases handled? An acceptance rule should be specific enough that two reviewers have a reasonable chance of applying it consistently.

Report both the number and the disposition when practical. “Seven cases failed” is less useful than “Seven of 120 cases did not meet the review rule: three were corrected before use, two were routed to the existing manual process, and two were excluded after a source-record issue was found.” The categories reveal the operational consequence. They also keep a single total from hiding failures with very different severity.

A failure rate is only meaningful when its denominator is clear. If 8 of 100 eligible cases failed, readers should know whether the 100 includes items that never reached the system, cases stopped by a reviewer, and cases with missing records. A failure that was intercepted before reaching a customer still matters, but it should be distinguished from a failure that passed through the control and caused an external effect.

Look for failures in the surrounding process, not only in the system output. Staff may bypass the workflow, paste incomplete information, misunderstand a review instruction or accept an output without checking a critical field. Those are not necessarily model failures, but they are delivery evidence. They can reveal that the designed process was difficult to use, that its controls were unclear or that the measurement missed important behavior.

A useful account describes what the team did when a failure occurred. Was the case corrected, escalated, excluded, retried or left unresolved? Did the response change the measured time or count? Was the acceptance rule revised during the observation period? If criteria changed, distinguish results before and after the change rather than combining them as though evaluation conditions remained constant.

Do not conceal an inconvenient failure by moving it into a footnote or treating it as an exception without a reason. Equally, do not imply that every anomaly makes the entire result meaningless. State its frequency, consequence and relevance to the claim. This gives a reader enough context to judge whether the remaining evidence is useful for the proposed decision.

7. A worked hypothetical example

Consider a hypothetical regional service company evaluating AI-assisted drafting for a narrow category of routine customer email. The team’s initial question is not whether AI can write an email. It is whether a draft-and-review workflow can reduce staff effort on eligible cases without increasing corrections or delaying replies. This is an illustrative scenario, not a report of an actual client result.

The team first defines an eligible case: an incoming email in one category, with a complete record and no exception that already requires specialist handling. One email thread counts as one case. The baseline covers four weeks of those cases. For each, the team records active staff handling time, elapsed time to first response, whether the response needed a correction, and whether it was escalated. Before collecting outcome data, the team writes down what counts as a material correction and what makes a case ineligible.

Suppose the hypothetical baseline includes 240 eligible cases. The median active handling time is 11 minutes, and 18 cases require a material correction under the agreed review rule. These values are illustrative assumptions. They are not benchmark figures and do not predict another organization’s results. Their purpose is to show how a reader can interpret a denominator and a measure together.

The pilot runs for five weeks, but the first week is used to train staff and resolve workflow issues. The evaluation period therefore covers four weeks, with the same eligibility rule applied. The team records the number of eligible cases, the number actually processed through the drafting workflow, the number reviewed, the number accepted with no material change, the number corrected, the number escalated and the handling time through final response. If staff choose not to use the workflow for an eligible case, that case remains visible in the eligible-work count rather than disappearing from the account.

Assume, for illustration, that 150 of 230 eligible pilot cases use the drafting workflow. The other 80 remain in the established process. Within the 150, reviewers identify 12 cases requiring a material correction, and 9 are escalated under the existing rule. The case study should not report only the 150 processed cases and imply that the workflow covered the full queue. It should show both 230 eligible cases and 150 workflow cases, then explain why the remaining cases were not processed through it.

Suppose measured median active handling time for the workflow cases is 8 minutes, including review and correction, while the comparable baseline median was 11 minutes. That observed difference is potentially useful, but it does not on its own establish a causal effect. The pilot group may differ from the baseline in case mix, staff, demand or other conditions. The case study should say which cases were compared, how medians were calculated, and which contextual changes the team knows about. If it cannot make the populations comparable, it should describe the result as an observed difference, not a proven saving.

The team should also account for effort outside the measured case. If staff spent additional time learning the workflow, maintaining instructions or resolving setup problems, that effort may matter to a broader business decision. A case study might report the per-case handling measure and separately disclose that setup and training time were not included. That is not a reason to discard the per-case finding; it is a boundary on what the finding means.

The failures tell an equally important story. If 12 of 150 cases needed material correction, the case study can describe the corrections by type, provided the categories are supported by records and the client approves disclosure. If several errors involved an omitted qualifier, the team might state that reviewers caught them before sending and that the workflow was restricted to reviewed drafts. It should not claim the problem is permanently solved merely because no known customer-facing error occurred during the short observation window.

The result could support a limited conclusion: the pilot processed a defined portion of one email category, recorded a lower median handling time for its measured cases, and found a specified number of corrections under a stated review rule. It may support a decision to continue measurement or adjust the workflow. It would not, by itself, prove company-wide savings, demonstrate performance on other categories, or establish that the same result will persist over a year.

That is what a useful case study does. It gives the reader enough detail to see the path from initial question to observed result, while preserving the distinction between the measured finding and the larger business decision.

8. A practical evidence path

The following flow shows a simple sequence for turning delivery records into a publishable claim. “Revise” can mean narrowing the claim, collecting additional evidence or deciding not to publish the result. It does not mean changing a result to make it more attractive.

flowchart TD
  accTitle: "From delivery evidence to a case-study claim"
  accDescr: "Define the claim, establish a baseline, select and count the sample, measure the timeframe, record failures, obtain client review, and publish only if approved; otherwise revise the claim or evidence."
  A["Define the outcome"] --> B["Freeze the baseline"]
  B --> C["Count the sample"]
  C --> D["Measure the timeframe"]
  D --> E["Record failures"]
  E --> F["Client reviews the claim"]
  F -->|"Approved"| G["Publish bounded result"]
  F -->|"Not approved"| B

Each step answers a different question. Defining the outcome prevents a model score from silently becoming a business result. Freezing the baseline makes the comparison explicit before the team knows which figure will look favorable. Counting the sample clarifies who or what is represented. Measuring the timeframe establishes when the observation applies. Recording failures keeps the account from becoming a success-only sample.

Client review belongs before publication, but review is not a substitute for sound measurement. The client should be able to verify that names, dates, process descriptions and outcome claims are accurate and that disclosure is permitted. If the client does not approve a specific number or description, remove it or replace it with an approved, still-accurate statement. Do not infer permission from participation in a project.

The return path in the diagram is deliberate. A review objection may reveal a factual error, an unclear definition, a confidentiality concern or a claim that exceeds what the records support. The right response may be to correct the evidence, narrow the language or omit the claim. Publication approval should not require the client to endorse an interpretation that the evidence cannot establish.

9. A counterexample: polished, but not evaluable

Imagine a case study that says an AI assistant “cut processing time by 60% and improved accuracy.” It offers no baseline dates, no case count, no definition of processing time, no description of how accuracy was assessed, and no account of cases routed away from the assistant. The client has approved the wording, but the reader cannot tell what was measured or whether the two periods involved comparable work.

The sentence may be factually intended, yet it leaves key questions unresolved. Was the 60% a reduction in active work or elapsed time? Did the measurement include review and rework? Was accuracy assessed on live cases, a selected evaluation set or staff opinion? Were difficult cases excluded? Did the team compare the same workflow before and after? Was the result observed for a week or a year?

The remedy is not necessarily to publish a dense technical report. A more defensible account could say: “For the defined document category, the team compared a specified pre-pilot period with a specified pilot period. It measured elapsed time from receipt to completed review across stated case counts. Some documents were excluded under a published-in-the-case-study rule; the pilot also recorded a stated number of corrections. The result describes those periods and does not establish performance for other document categories.” The actual case study should replace every general phrase with verified project details or omit the claim.

If the underlying records cannot answer the questions, the team should not fill gaps with confident prose. It may publish a qualitative description of the workflow change, explain that comparative outcomes were not measured, and state what the next evaluation would need to capture. That can still help a buyer understand the work, but it is a different kind of evidence from a quantified outcome claim.

A polished counterexample also shows why client approval alone is insufficient. Approval establishes that the client has reviewed and permitted the account; it does not transform an undefined measure into a reliable one. Conversely, careful measurement does not grant permission to publish confidential details. Measurement quality and publication approval are separate conditions, and both matter.

10. The buyer’s evidence worksheet

Use this worksheet when reading a case study or preparing one for review. It is designed to expose gaps in the claim, not to force every project into a single measurement method. A “not available” answer can be informative if it leads to a narrower conclusion.

Evidence fieldWhat to write downDecision it supports
Primary claimOne sentence stating exactly what changedCan the headline be evaluated?
Process boundaryStart event, end event and task includedAre the compared workflows the same?
BaselineMeasure, dates, source and calculation methodIs there a meaningful reference point?
Eligible populationWho or what could enter the workflowWhat work is the claim about?
Sample and denominatorUnit counted, total eligible, total measuredWhat do reported rates actually represent?
TimeframeBaseline period, pilot period and review cutoffHow long and under what conditions was it observed?
ExclusionsRule, count and reason for each material exclusionIs selection hiding difficult or incomplete cases?
Outcome measureDefinition, source and whether observed or calculatedIs this a system, workflow or business outcome?
FailuresCount, categories, disposition and consequenceWhat did not work, and how did the process respond?
Other changesStaffing, demand, policy or workflow changesWhat else may explain the observed difference?
LimitationsUnmeasured effects, small sample or weak comparisonHow far can the conclusion reasonably extend?
Client approvalApproved facts, wording, attribution and scopeIs this version permitted to be published?

Start with the claim and work down the worksheet. If the headline cannot be connected to a measure, reconsider it before debating how to phrase the result. If the measure has no clear denominator, establish the count. If the comparison is weak, narrow the claim rather than hiding the weakness in a general limitations paragraph.

For a buyer, the worksheet can also guide diligence conversations. Ask the delivery team to walk through one reported result from source record to published sentence. The purpose is not to demand access to another client’s confidential files. It is to understand the definitions, counting rules, review process and boundaries that make the claim interpretable. A credible team should be able to explain those choices without presenting private client information.

A buyer may also ask whether the case study’s workflow resembles the proposed work. Compare the inputs, decision points, human review, exception handling and outcome measure—not just the industry label or model category. A similar organization can have a materially different queue, data quality or risk tolerance. Similarity should be demonstrated in the operational details relevant to the claim.

For a delivery team, assign an owner to each evidence field at the start of the work. The project lead may define the claim and process boundary; an operations owner may confirm the baseline and case categories; analysts may document counting and calculations; and the client contact may review factual accuracy and publication scope. The roles need not match these labels, but the work should not be left to a writer assembling numbers at the end.

Maintain a compact evidence record alongside the draft: the measure definitions, date ranges, sample counts, exclusions, failure categories, calculation notes and approved wording. This does not need to expose customer data in the published article. It makes it possible to answer later questions without relying on memory or reconstructing the analysis from a polished paragraph.

11. Turning evidence into a publishable account

Draft the case study in the same order a reader needs to evaluate it: define the work, establish the baseline, state the sample and period, report the observed result, disclose material failures and limits, then explain the client-approved outcome. This sequence is not a required template, but it prevents the conclusion from arriving before the reader knows what was counted.

Prefer concrete labels to promotional shorthand. Replace “accuracy improved” with the task, review rule, sample and observed count. Replace “saved significant time” with the measured time definition and comparison. Replace “automated the process” with a precise account of which steps the system handled and which remained with people. Precision makes the claim easier to understand even when the result is modest.

Use ranges or rounded figures when exact values would imply more certainty than the method supports or expose sensitive information. But do not round in a way that changes the story. If a percentage is calculated from a small denominator, report the count as well. If a client-approved range is used, make clear that it is a range and do not reverse-engineer or imply a more exact figure.

Keep the distinction between observed and interpreted language visible. “The median recorded time was lower in the pilot period” reports an observation. “The workflow caused the reduction” interprets the observation and needs stronger support. “The result suggests a reason to test broader coverage” is a decision proposal, not a measured outcome. A reader benefits when these statements are not blended into a single confident claim.

Use limitations that change interpretation, not generic disclaimers. “The period was short and covered one category” tells a reader what the result does not establish. “Results may vary” is much less informative. A useful limitation names the boundary and its consequence: the sample cannot support a claim about rare cases; the before-and-after comparison cannot isolate causation; or the measure excludes setup and training effort.

When a case study supports a next decision rather than a broad success claim, say so. A team may have enough evidence to continue a bounded pilot, revise the review process or collect a better baseline, but not enough to expand to other tasks. Readers can assess that decision more effectively when the case study shows what evidence triggered it and what remains unknown. For the separate question of how a pilot should lead to a deployment decision, see The Ship-or-Kill Framework for AI Pilots.

12. Evidence as a delivery decision tool

The quality of a case study begins before anyone writes it. During planning, decide what outcome matters, what records can measure it, which cases qualify and how failures will be categorized. These choices are cheaper to make before the workflow changes. Once a pilot is under way, teams may find that the old process has no reliable timestamps or that “completed” means different things to different staff.

This does not mean every project requires elaborate measurement infrastructure. A narrow workflow may need only a clearly defined sample, a small number of well-chosen measures and an agreed review method. The appropriate effort depends on the consequence of the decision and the strength of the claim. A descriptive account of a small experiment can be useful with modest evidence; a broad claim about business impact needs more support.

The evidence plan should also fit the decision the organization expects to make. If the question is whether staff can use a drafting tool safely under review, record review outcomes and failure disposition. If the question is whether the process is faster, measure the relevant time boundary, including review and correction. If the question is whether customers receive better service, include an outcome that captures the customer-facing result rather than assuming speed is a proxy for quality.

If a team needs help defining these boundaries, measurement responsibilities and decision criteria, fractional CTO leadership may be relevant. The practical next step is to agree on the claim and evidence plan before drafting a success narrative, so that project work produces records suited to the decision the organization actually needs to make.

A well-supported case study does not promise that another organization will reproduce the result. It shows what was observed, under which conditions, for which cases, over what period, with what failures and with whose approval. That level of detail gives buyers a basis for judging relevance, gives delivery teams a record they can defend, and gives clients control over what is attributed to them.

The strongest conclusion may be a bounded one: a workflow improved on a specific measure for a defined sample, while other outcomes remain unmeasured. It may be that the result is mixed, that a failure changed the plan, or that the evidence supports further evaluation rather than expansion. When the account makes those distinctions clear, it is useful precisely because it does not ask the reader to confuse a compelling story with proof.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call