Start with the decision, not a sample-size rule
There is no universal number of test cases that makes an AI model comparison trustworthy. A set of 50 cases may be useful for finding obvious failures in a narrow workflow, yet far too small to distinguish models whose performance is close. A set of 5,000 may still mislead if it mostly repeats one easy task while omitting the situations that matter to the business.
The useful question is not “How many cases should we run?” in isolation. It is: “How much uncertainty can we tolerate for this decision, and which kinds of work must the test represent?” The answer depends on the decision at stake, the size of the expected difference, the variability of outcomes, and the distribution of tasks.
A test set is a sample from the work a model may encounter. Its score estimates performance on that work; it does not reveal an exact, permanent capability. A sample result therefore needs an uncertainty estimate and a description of what the sample represents. Without both, a precise-looking percentage can disguise weak evidence.
This article develops a practical way to size and structure a comparison. It focuses on confidence intervals, uncertainty, and stratified task coverage—not on choosing a particular model or designing a full benchmark. For release evaluation more broadly, see Benchmarks That Actually Matter for New Model Releases. For a worked approach to tool-use reliability, see Tool-Use Reliability Across Frontier Releases: A Benchmark Methodology.
1. Define what a “case” measures
A test case is one defined opportunity for a system to perform a task under specified conditions. It should include the input, relevant context, expected outcome or scoring rubric, and any constraints that affect evaluation. A case could be a customer request, a document to classify, a coding task, or a multi-step action request. The count is meaningful only when the unit is clear.
If a model receives one prompt and produces one answer, that may be a case. If the prompt requires five independent decisions, counting it as one case can hide variation among those decisions. Conversely, splitting one conversation into ten turns does not necessarily create ten independent cases: later turns depend on earlier model behavior, and all ten may share the same underlying scenario.
A case is also not automatically a real-world user, transaction, or deployment event. A curated test collection may have more unusual cases, more detailed instructions, or less noisy inputs than production. Its score describes performance on the defined evaluation population, not every possible use of the system.
Before counting, write down the outcome being estimated. For example: “The proportion of eligible support requests for which the model produces a response that meets the approved rubric without a critical factual error.” This makes the denominator, success definition, and intended population explicit. “Answer quality” alone is too broad to size a test responsibly.
If a workflow has more than one important outcome, do not compress them into a single pass rate prematurely. Track separate measures such as task completion, factual correctness, required-format compliance, and critical-error rate. A model that improves average writing quality while increasing a rare but serious failure may not be preferable. The number of cases needed can differ for each measure.
2. Understand uncertainty in a pass rate
Suppose each case has a binary outcome: pass or fail according to a fixed rule. The observed pass rate is the number of passes divided by the number of evaluated cases. If 80 of 100 cases pass, the observed rate is 80%. That is a description of the sample, not a guarantee that the underlying task population has an 80% success rate.
A confidence interval expresses sampling uncertainty around an estimated proportion under stated assumptions. It gives a range of values consistent with the observed data at a chosen confidence level; it does not mean that a particular future result is guaranteed to fall inside the interval. The interval’s width depends chiefly on the number of observations and the observed proportion, alongside the interval method. Small samples can leave wide uncertainty, especially when estimating a proportion near a boundary. NIST’s discussion of binomial proportion intervals explains the underlying interval problem and why the assumptions matter.
The simplest intuition is that more independent cases generally narrow uncertainty, but the improvement is not linear. Roughly, halving an interval’s width requires about four times as many independent observations when other conditions are comparable. This is a planning approximation, not a substitute for calculating an interval with an appropriate method.
For a quick illustration, imagine a sample with an observed pass rate near 80%. A result based on a few dozen cases should not be reported with the same confidence as the same percentage based on several hundred genuinely varied cases. The point estimate may match, but the first estimate is less stable. Reporting “80%” without the denominator and an interval suppresses that distinction.
Use an interval method suitable for binomial proportions rather than relying reflexively on a simple normal approximation, particularly with small samples or rates near zero or one. The calculation still cannot repair a biased sample, inconsistent scoring, duplicated examples, or dependence among cases. Statistical precision is not the same as representativeness.
For a model comparison, uncertainty concerns the difference between models, not merely the separate uncertainty around each model’s score. If both models answer the same cases, their outcomes are paired: a case where both succeed contributes different information from one where one succeeds and the other fails. A comparison should preserve case-level results and estimate uncertainty in the paired difference, rather than treating two sets of scores as unrelated samples.
3. Pick a meaningful precision target
A useful sample-size plan begins with the smallest difference that would change the decision. Call this the decision-relevant difference. If a team would not change its choice for a one-percentage-point improvement, spending effort to resolve a difference that small may not be worthwhile. If a small change in critical-error rate would alter whether a workflow can be launched, that outcome needs more sensitive measurement.
Set the target before seeing the comparison results. Otherwise, teams can unconsciously choose a precision threshold that makes the preferred model appear decisive. Record the outcome, minimum meaningful difference, acceptable uncertainty, and the action that follows each possible result.
A confidence interval can guide this decision. If the interval for the paired performance difference is narrow and lies entirely on one side of the decision threshold, the evidence may support a choice. If it spans both meaningful improvement and meaningful harm, the comparison is inconclusive for that decision—even if one model has the higher point estimate.
“Not statistically distinguishable” does not mean the models are equivalent. Equivalence requires defining a margin within which differences are practically unimportant and collecting evidence precise enough to support that claim. Similarly, a result that is statistically distinguishable may still be too small to justify migration, added operational complexity, or a change in user experience.
A planning calculation should use assumptions about the baseline rate, expected difference, desired confidence level, and desired power or precision. For paired binary outcomes, planning also depends on how often the models disagree on the same cases. If their answers usually succeed and fail together, a comparison may need a different number of cases than if outcomes frequently diverge. Estimate these inputs from a pilot or prior, relevant evaluation data where available; label them as assumptions rather than facts.
Do not treat a convenient calculator’s output as an order to stop at an exact number. It is conditional on the inputs and statistical design. If the assumed difference is too optimistic, the actual comparison may be underpowered. If the sample differs from deployment work, a narrow interval can still answer the wrong question.
4. Stratify cases so coverage matches the work
Stratification means dividing the evaluation population into meaningful groups and ensuring each group is represented deliberately. For a support assistant, groups might distinguish billing questions, account access, product troubleshooting, and policy-sensitive requests. For document processing, groups might reflect document type, scan quality, language, and whether key information is missing.
The strata should come from the task and decision, not from categories that are easy to count. A useful stratum changes the expected difficulty, consequence of error, or appropriate response. Avoid creating dozens of tiny categories with no practical interpretation; they fragment the sample and make estimates unstable.
There are two different goals that are often confused:
- Estimate average performance in the expected workload. Sample strata in proportions that resemble the workload, or calculate a weighted estimate using credible workload shares.
- Check performance in important subgroups. Allocate enough cases to each subgroup to examine it, even when that means oversampling rare but consequential work.
These goals can coexist, but the reporting must distinguish them. If rare account-recovery cases are deliberately oversampled, an unweighted overall pass rate no longer estimates the natural workload average. Report subgroup results and apply documented workload weights for the overall estimate, when those weights are defensible.
A compact coverage plan might include common routine work, less common complex work, and low-frequency high-consequence work. The labels are not a substitute for case definitions. Specify what qualifies, how cases are sampled or authored, and how borderline cases are assigned. Keep the categories mutually interpretable; overlapping labels can make counts and denominators ambiguous.
| Coverage view | What it answers | Main risk if ignored |
|---|---|---|
| Overall workload mix | How might average performance look across the expected task distribution? | Averages can hide weak performance in a small but important group. |
| Task family | Does the model handle distinct kinds of work consistently? | A dominant easy category can overwhelm smaller categories. |
| Difficulty or ambiguity | Does quality change when inputs are incomplete, conflicting, or unusually phrased? | Clean examples can exaggerate readiness. |
| Consequence of error | Are high-impact mistakes visible and measured separately? | A low overall error rate can obscure concentrated risk. |
| Input condition | Does performance depend on language, format, or source quality? | The test may not resemble actual incoming material. |
Stratification does not automatically make a sample representative. The cases within each group also need a defensible selection process. A set of handpicked “interesting” examples can be useful for discovering failure modes, but it should not be presented as a random estimate of the group’s pass rate.
5. Worked example: a support-response comparison
Consider a hypothetical mid-market software company comparing two models for drafting support responses. This is an illustrative design, not a report of measured results. The decision is whether one candidate is sufficiently better for a limited workflow to justify a controlled next step. The primary outcome is a binary rubric: a response passes only if it answers the request correctly, follows the required policy, and contains no critical factual error. The team separately records which rubric condition failed.
The company’s initial sample has 120 requests: 84 routine billing or account questions, 30 troubleshooting cases, and 6 policy-sensitive cases. The overall observed pass rate is easy to compute, but the sample leaves only six observations for the most consequential group. Even if both models pass all six, that is weak evidence about the true failure rate in policy-sensitive work. A perfect observed subgroup score does not establish rare-event safety.
The team first decides that the evaluation must answer two distinct questions. One is the workload-weighted average, based on its current estimate of incoming request mix. The other is whether either model shows a material weakness in the policy-sensitive group. The first needs a weighted estimate; the second needs deliberate subgroup coverage and a separate uncertainty statement.
Suppose the team’s current workload estimates are 70% routine, 25% troubleshooting, and 5% policy-sensitive. These are planning assumptions that must be checked against operational records; they are not measurements produced by the benchmark. A proportionate sample of 200 cases would yield about 140, 50, and 10 cases respectively. That improves coverage but still gives limited information about rare failures in the last group.
The team therefore chooses an illustrative design of 300 cases: 150 routine, 90 troubleshooting, and 60 policy-sensitive. This deliberately oversamples the rare group. To estimate the workload average, the team would weight each group’s result by the assumed 70/25/5 distribution. To judge the policy-sensitive subgroup, it would report its own result and interval, rather than letting its oversampling distort the overall figure.
Each request is presented to both models under the same fixed conditions, and reviewers score outputs without seeing which model produced them where practical. The comparison retains a row per case and per model, plus paired outcomes. The core comparison is the difference in pass rates on matched cases. For example, cases where both models pass do not demonstrate a win for either; cases where one passes and the other fails are especially informative about the difference.
The team does not declare a winner merely because one point estimate is higher. It sets a decision-relevant threshold in advance—for instance, a minimum improvement needed to justify moving to the next evaluation stage—and checks whether uncertainty around the paired difference is narrow enough to distinguish that threshold. The exact threshold depends on business consequences and is intentionally not universal.
If the overall weighted estimate favors one model but the policy-sensitive interval remains broad, a sensible decision may be “promising for routine work; insufficient evidence for the sensitive group.” That is more informative than forcing a single winner. The next action could be to collect additional policy-sensitive cases, narrow the proposed use, or revise the rubric if reviewers cannot score those cases consistently.
This design also exposes the cost of the decision: more cases in a rare stratum improve what can be learned about that stratum, but they do not represent its natural frequency. The team must keep the two purposes separate in its analysis and in any executive summary. Otherwise, an intentionally balanced test may be mistaken for a workload-weighted estimate.
6. Compare models on the same cases—and account for dependence
When two candidates are evaluated on different case sets, differences in case difficulty can masquerade as model differences. A matched comparison reduces that problem: both candidates see the same underlying cases, and the analyst compares their outcomes case by case. Matching does not eliminate all bias, but it makes the comparison more efficient when cases vary substantially in difficulty.
Preserve the paired data. For every case, record each model’s outcome, the task stratum, and any relevant scoring details. Then report how many cases both passed, both failed, and were discordant. The discordant cases help explain where the models differ; inspecting them can also reveal a rubric defect or an unexpected failure pattern. Do not silently discard disagreements that are inconvenient to the headline score.
Independence deserves special attention. Ten prompts derived from one source document may share the same vocabulary and facts. Ten turns in one conversation share context. One hundred near-duplicates generated from a template can create the appearance of a large sample while providing little new evidence. The effective information may be closer to the number of distinct source situations than to the raw row count.
Where cases are clustered, identify the cluster unit—such as customer, document, conversation, or source template—and keep related examples together when splitting development and evaluation data. For uncertainty calculations, use an approach that reflects clustering when it is material. A simple interval that assumes every row is independent can be too narrow when correlated cases are counted separately.
Repeated model runs on the same input raise another design choice. If model output varies across runs, one run per case measures a mixture of case difficulty and run variability. Repeated runs can help estimate variability, but they do not create new independent task scenarios. Report how many distinct cases and how many runs per case were used, and decide whether the target is typical single-run behavior or behavior averaged over repeated attempts.
Scoring variation is another source of uncertainty. If human reviewers disagree, the observed pass rate may reflect rubric interpretation as well as model quality. Calibrate the rubric on a sample, clarify ambiguous criteria, and preserve disagreement records. A large test set scored inconsistently can produce a very precise estimate of an unstable label.
7. Know when a rare failure needs a different plan
Average pass rates are poor evidence about very rare events unless the test is very large or the evaluation uses additional justified evidence. If a critical failure is absent from a small sample, the correct interpretation is that none was observed in that sample—not that the risk is zero. The binomial interval still leaves uncertainty about the underlying rate, and the observed sample may omit conditions that trigger the failure.
This creates a practical mismatch: teams may need confidence about a low-frequency, high-consequence event, yet ordinary-sized test sets may be incapable of establishing a sufficiently low rate. The answer is not to relabel a small sample as proof. Options include gathering more relevant cases, restricting the model’s permitted scope, requiring an independent check for specific actions, or collecting complementary evidence that addresses the failure mechanism. Each option answers a different question and should be described accordingly.
A targeted stress set can be valuable for finding whether a known weakness is present. Its hit rate should not be confused with the natural frequency of that weakness in everyday work. Conversely, a representative sample may contain too few rare cases to diagnose a failure mode. Use a representative or weighted sample for workload estimates and a separately identified challenge set for focused probing.
This distinction also applies to changing conditions. If a new document format or policy creates a new task subgroup, old cases may no longer support the same estimate. Adding a handful of examples can reveal obvious regressions but may not provide precise subgroup rates. Treat changes in the evaluation population as a reason to review both coverage and uncertainty, not simply append cases and keep the old interpretation.
8. Use a stopping rule, not repeated peeking
A fixed-size evaluation chooses the sample size or stopping point before examining the comparison outcome. This is straightforward to explain: run the planned cases, score them under the defined protocol, and analyze once. It reduces the temptation to stop when an attractive result appears or continue only when an unfavorable result needs to be overturned.
If the team expects to inspect results repeatedly and add cases as needed, it needs a sequential design with an appropriate analysis plan. Repeatedly checking an ordinary confidence interval and stopping as soon as it excludes zero can inflate the chance of a false-positive finding. “We will keep testing until the result is clear” is not a statistically neutral rule.
A practical review cadence can still be useful for operations. For example, teams may inspect early outputs for broken prompts, scoring failures, or safety issues, while keeping those exploratory checks separate from the final confirmatory comparison. If the evaluation is paused because of a serious issue, document the stopping reason and do not present the partial score as if it came from the original fixed plan.
Define what happens when the result is inconclusive. Possible outcomes include extending the sample according to a preplanned rule, limiting the decision to a supported subgroup, or choosing based on other operational criteria while acknowledging that the performance evidence does not distinguish candidates. A forced binary winner is not always the most honest decision artifact.
9. A reproducible sizing and analysis protocol
Use this protocol before collecting the final comparison set. It is designed to make the number of cases defensible and the result interpretable; it is not a claim that any one sample size fits every organization.
- Write the decision and target population. State what choice the comparison will inform and which incoming tasks the estimate is meant to represent. Exclude unsupported use cases explicitly.
- Define the unit and outcome. Specify one case, the pass criteria, critical failures, and any secondary measures. Decide whether outcomes are binary, graded, or both.
- Set strata and sampling intent. List meaningful task groups. Mark whether the aim is workload estimation, subgroup assurance, failure discovery, or a combination. State which groups will be oversampled.
- Choose the decision-relevant difference. Document the smallest improvement or harm that would change the decision. Select the uncertainty target and confidence approach before looking at final results.
- Estimate planning inputs. Use relevant prior data or a clearly labeled pilot to estimate baseline performance and, for paired comparisons, likely disagreement between candidates. Record assumptions and how sensitive the planned sample is to them.
- Check dependence and scoring. Identify repeated sources, conversations, or templates; decide how clusters will be handled. Calibrate reviewers and define how ambiguous cases are resolved.
- Set the stopping and analysis rules. State the planned sample, any permitted extension rule, subgroup and weighted estimates, interval method, and treatment of missing or invalid cases.
- Preserve case-level records. Store case ID, stratum, source or cluster ID, model outcome, rubric components, reviewer, and exclusion reason. Retain paired outcomes so the comparison can be reproduced.
- Report estimates with limits. Show denominators, intervals, subgroup counts, workload weights, and the number of discordant paired cases. Separate exploratory findings from the preplanned primary comparison.
- Tie evidence to action. State which decisions the evidence supports, which remain uncertain, and what additional cases or controls would change that boundary.
A simple flow summarizes the logic:
accTitle: Choosing a test-set plan accDescr: Define the decision and task mix, then check whether the planned cases can estimate the decision-relevant difference. If not, improve coverage or narrow the claim before collecting and reporting results.
flowchart TD
accDescr: Workflow stages and decisions: Define decision, Map task strata, Set precision target, Check sample plan, Revise design, Run and report. The adjacent text explains the conditions and exceptions.
accTitle: How Many Test Cases Are Enough for an AI Model Comparison? workflow
A["Define decision"] --> B["Map task strata"]
B --> C["Set precision target"]
C --> D["Check sample plan"]
D -->|"Coverage or precision weak"| E["Revise design"]
E --> C
D -->|"Adequate for decision"| F["Run and report"]
“Check sample plan” means examine both uncertainty and whether each important task group has enough cases for the intended claim. If either is weak, revise the design or narrow the claim; do not proceed as though a large overall count fixes a missing subgroup. “Adequate” means adequate for the specified decision and assumptions, not universally sufficient.
10. Read results without overstating them
A good comparison report puts the point estimate beside its uncertainty and denominator. It identifies whether the estimate is workload-weighted, presents important subgroup results, and explains how cases were selected. It reports the paired nature of the comparison when both models saw the same inputs. Readers should be able to distinguish observed sample performance from an estimate of expected performance.
Avoid statements such as “Model A is 4% better” when the calculation is actually a four-percentage-point difference in this sample, the interval is broad, and the test set oversampled difficult cases. Prefer a precise description: “On this defined set, the observed pass-rate difference was four percentage points; the interval includes both a practically meaningful advantage and little difference. The set oversampled troubleshooting cases, so the unweighted total is not a workload estimate.”
A result can be clear for one stratum and unresolved overall. It can also show no meaningful average difference while revealing complementary strengths or a serious failure concentrated in one group. Those are not statistical nuisances; they are information about where the system may or may not be suitable. Keep the interpretation aligned to the level at which evidence was collected.
For teams implementing a broader model-evaluation program, AI evaluation and strategy is a relevant next step. If the immediate question is how to conduct a particular model comparison, Astra vs Sonnet 5.5 vs Opus 5.5: How to Run Your Own Comparison covers that distinct exercise. A migration decision needs a regression-focused design, as discussed in Sonnet 5.5 Migration: Build a Regression Suite Before Switching; task routing is a separate decision addressed in Sonnet 5.5 vs Opus 5.5: Which Work Should Your Team Route to Each?.
Printable comparison worksheet
Complete this worksheet before the final evaluation. Keep the answers with the dataset and analysis so another team member can understand what the numbers mean.
- Decision: What action will this comparison inform? What evidence would change that action?
- Target population: Which tasks are included, and which are outside the claim?
- Case unit: What counts as one independent task opportunity? Which examples share a source, conversation, or template?
- Primary outcome: What exact criteria determine pass or failure? Which failures are critical?
- Task strata: What groups matter for task type, difficulty, input condition, or consequence?
- Sampling purpose: Is the sample intended to estimate workload performance, inspect subgroups, discover failures, or do more than one of these?
- Weights: If groups are oversampled, what evidence supports the workload weights used for any overall estimate?
- Comparison design: Will both models receive the same cases? How will paired outcomes be retained and analyzed?
- Precision target: What difference would matter in practice, and what uncertainty is acceptable for this decision?
- Planning assumptions: What baseline rate, paired disagreement, cluster structure, and scoring consistency are assumed?
- Stopping rule: Is the sample fixed? If it can expand, what preplanned rule governs the extension?
- Reporting: Will the report show denominators, intervals, subgroup results, weighting, exclusions, and limitations?
- Decision boundary: What will the team do if the interval remains inconclusive or a subgroup is too small?
Printable summary
Decision: ________________________________________________
Target task population: ____________________________________
Primary outcome and critical failures: ______________________
Strata and planned cases per stratum: _______________________
Meaningful difference and uncertainty target: ______________
Sampling, pairing, and dependence notes: ___________________
Stopping rule and analysis method: _________________________
What the evidence will not establish: _______________________
The defensible test-set size is the size that can answer a stated decision with an acceptable level of uncertainty across the task groups that matter. A raw count cannot do that alone. Define the population, preserve meaningful strata, compare candidates on matched cases where appropriate, and report what remains uncertain. When the planned data cannot support the decision, the honest response is to improve coverage, collect more relevant cases, narrow the intended use, or defer the claim—not to mistake a small interval-looking percentage for certainty.