SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Explainer 19 mins

Pass@1 vs Best-of-N: Which AI Benchmark Matches Your Workflow?

Pass@1 vs Best-of-N: Which AI Benchmark Matches Your Workflow?. Practical examples, tradeoffs and implementation guidance for technology leaders.

The PADISO Team ·

The short answer

Pass@1 and Best-of-N answer related but different questions. Pass@1 asks whether one attempt succeeds. Best-of-N generates several candidates and selects one using a defined scoring procedure. Oracle pass@N instead reports whether at least one candidate passed, regardless of whether the deployed system could identify it. The second result depends not only on generation, but also on how candidates are selected and what each attempt costs.

For a workflow that gives a user one answer and has no dependable way to judge alternatives, Pass@1 is usually the more representative starting point. For a workflow that can inspect, test, rank, or verify several candidates before acting, a Best-of-N experiment can help determine whether sampling improves the final outcome enough to justify its expense and delay.

The key is not to treat the larger number as proof that a model or workflow is better. An idealized pass@N score can count a candidate that an actual system would never recognize or choose. A practical evaluation therefore reports one-shot performance, the precise N-candidate procedure, and the result after the real selection step.

This distinction matters for engineering decisions. A team choosing a model for a single-response support workflow needs evidence about its first response. A team designing a code-generation pipeline with tests may also care about whether generating and checking several alternatives improves the accepted result. Neither metric is universally superior; each corresponds to a different operating design.

Define the terms before comparing scores

Pass@1 is the proportion or probability of tasks for which one generated attempt meets a defined success criterion. “One” means one opportunity for the system to produce the answer or action being evaluated. It does not mean the model has only one internal reasoning step, nor does it describe how many examples were in the evaluation set.

Pass@N commonly describes whether at least one of N generated candidates passes a criterion. If a task has a binary check—such as whether a proposed calculation matches a known answer—an evaluator can inspect all candidates and ask whether any passed. This is an existential result: it says a passing candidate appeared somewhere in the set. It does not, by itself, say that a deployed system could find that candidate.

Best-of-N describes a procedure in which N candidates are generated and a selection rule chooses one, or rejects them all. That rule might be a deterministic test, a ranking model, a human reviewer, or another explicitly defined mechanism. The score that matters operationally is whether the selected candidate succeeds, not merely whether a successful candidate existed among those generated.

Teams sometimes use “Best-of-N” as shorthand for the idealized pass@N result. That convention is understandable in a benchmark with an oracle that can identify the correct answer, but it can obscure a consequential gap. If the evaluation gives the selector access to the answer key while production does not, the reported number describes candidate availability under privileged inspection—not production selection quality.

Sampling means producing multiple candidate outputs under a defined generation procedure. Selection means deciding which candidate, if any, becomes the system’s answer or action. Selector is the component or rule that makes that decision. These are separate stages and should be measured separately when the product relies on both.

An agent evaluation needs an outcome, not just a transcript. One useful distinction is that task evaluation scores what happened, while the agent harness and evaluation harness are different parts of the setup. Anthropic’s discussion of agent evaluations provides context for that distinction. A plausible-looking transcript is not proof that the task succeeded; a result should be judged against the outcome the workflow actually requires.

The basic probability—and its assumptions

Suppose a single candidate has probability p of passing a binary success check. If candidate outcomes are independent and identically distributed, the probability that at least one of N candidates passes is:

Pass@N = 1 − (1 − p)ᴺ

For example, if p is 0.70 and N is 4, the idealized probability of seeing at least one passing candidate is 1 − 0.30⁴, or 0.9919. This is a mathematical illustration, not a benchmark result. It says that under the stated assumptions a passing candidate is very likely to appear; it says nothing about whether the selector can identify it, whether the task distribution is representative, or whether four attempts are worth the added resources.

Independence is a strong assumption. Candidates from the same model, prompt, and task may repeat the same error. Four nearly identical wrong answers do not provide four independent chances to recover. If failures are correlated, the simple formula overstates the benefit of additional samples. The only reliable way to know how much repetition helps on the target workflow is to evaluate repeated candidates under that workflow’s actual generation and selection rules.

The same caution applies to p. A single overall success rate can conceal different task classes. One candidate may perform well on routine requests and poorly on ambiguous ones, while another may show the reverse. The average can be useful for a defined workload, but it does not establish that every important task type benefits from N attempts.

Even when the “at least one passes” probability rises with N, operational quality need not rise at the same rate. The selector can choose a weaker answer, misread a test result, or fail to reject a candidate that should not be used. A benchmark should keep the candidate-generation result and the selected-result outcome distinct so a high oracle score cannot mask a poor selection stage.

What changes when N candidates are used

A single-attempt system has one generation path. It produces an output, and the workflow either accepts it, checks it, or asks a person to intervene. Its evaluation is comparatively direct: define the task, run one attempt under fixed conditions, and score the outcome.

A Best-of-N system adds a candidate pool and a decision point. Its end-to-end path is more like this:

flowchart TD
  accDescr: Workflow stages and decisions: Define task and success check, One candidate or N?, Run one candidate, Score outcome, cost, and latency, Generate N candidates, Valid selection rule?, Select or reject a candidate, Report oracle pass@N only. The adjacent text explains the conditions and exceptions.
  accTitle: Pass@1 vs Best-of-N — Which AI Benchmark Matches Your Workflow? workflow
    A["Define task and success check"] --> B{"One candidate or N?"}
    B -->|"One"| C["Run one candidate"]
    C --> G["Score outcome, cost, and latency"]
    B -->|"N candidates"| D["Generate N candidates"]
    D --> E{"Valid selection rule?"}
    E -->|"Yes"| F["Select or reject a candidate"]
    F --> G
    E -->|"No"| H["Report oracle pass@N only"]

accTitle: Comparing single-attempt and multi-candidate evaluation accDescr: Define a task and success check, then choose one candidate or N. One candidate is scored directly. N candidates require a valid selection rule for an operational score; without one, report only the oracle pass-at-N result.

The “valid selection rule?” branch is the core decision. A rule is valid for the evaluation when it uses information the operating system will actually have and its output is scored using the intended task outcome. If the benchmark checks every candidate against a hidden answer key and reports whether any one passed, it can still provide an oracle pass@N measure. It should not label that result as the deployed system’s selected-answer reliability.

A selector can be a rule rather than another model. For example, if candidates are alternative code changes and a specified test suite distinguishes acceptable from unacceptable changes, the test outcome can guide selection. If the task is to draft a response whose factual correctness cannot be mechanically checked, simply choosing the most fluent candidate does not establish correctness. The selection method has to fit the success criterion.

Selection can also introduce new failure modes. A judge can prefer confident wording over accurate content; a ranking rule can systematically favor one candidate style; a tie can be resolved inconsistently; or all candidates can fail while the system still chooses one. A sound procedure specifies what happens in each case, including whether the system abstains, retries, or asks for human review. The evaluation should count the resulting behavior, not silently discard inconvenient cases.

A hypothetical worked comparison

Consider a hypothetical internal workflow that drafts a short answer from a fixed record. Assume, solely for illustration, that a candidate passes the team’s binary factual check 70% of the time. A one-candidate setup would then pass 70% under this assumption. Four independent samples would give an idealized pass@4 of 99.19% by the formula above.

Now assume the workflow cannot use the answer key, but has a selector that identifies a passing candidate only 85% of the time when at least one passing candidate exists. For a deliberately simplified illustration, assume it never selects a passing candidate when all four candidates fail. The estimated selected-result success is then 0.9919 × 0.85, or about 84.3%. The improvement over the assumed 70% one-shot rate is about 14.3 percentage points—not the 29.19-point lift suggested by comparing 99.19% directly with 70%.

Those numbers are chosen to demonstrate the distinction; they are not measured model performance. Real selectors may sometimes succeed even when another candidate passes only imperfectly, or may select a passing candidate through a different process. They can also accept an incorrect candidate. A real evaluation must directly score the chosen output and include the rule for rejection or abstention rather than relying on this simplified calculation.

The candidate count also changes resource use. If one attempt consumes one unit of generation effort, four candidates require roughly four units before accounting for selection. A selector may add its own processing, and parallel generation can still increase completion time through queueing or coordination. The practical comparison should use the workflow’s measured usage and elapsed time, not assume that parallel calls are free or that each candidate has the same resource profile.

Suppose the one-shot result already meets the product’s acceptance threshold and the additional candidates rarely change which answer is selected. Then a better oracle pass@N number may not justify the extra work. Conversely, if failures are costly and an available test can reliably identify a good candidate, sampling could be worth evaluating even if it adds delay. These are decisions about the whole path, not about a benchmark label.

For a fuller treatment of how task-level evaluation translates into operating expense, see the worked model for cost per successful agent task. This article’s narrower point is that the success numerator must match what the deployed selection process can actually deliver.

Choose the metric that represents the workflow

Start with the production interaction, not the metric. Ask how many candidates the system can produce before a user, downstream service, or deadline needs an answer. If the system exposes the first response and no later check changes it, Pass@1 is the central reliability measure. A hidden pool of alternatives does not improve the user’s experience unless the system can select or validate one in time.

If the system generates alternatives and uses a real test, report the selected-result success rate under that test. Also report oracle pass@N where it helps answer the narrower diagnostic question, “Did generation produce at least one passing candidate?” The gap between the two figures is useful: it indicates whether the bottleneck is finding a good candidate or recognizing and using it.

If a human chooses the candidate, the human is part of the system being assessed. Record the review conditions that matter, such as what information was visible, whether the reviewer saw candidate order, and whether they could reject every option. A human-selection result should not be presented as model-only performance. It may be a perfectly appropriate workflow score, provided the measured arrangement reflects the intended operation.

If candidates are produced concurrently but the first completed answer is returned, the relevant metric may be first-result quality rather than best-of-N quality. If the workflow waits for all candidates, the relevant latency includes that wait. “N” is not a complete description of the budget: the stopping rule and timing policy also affect which candidates can influence the outcome.

A simple comparison table helps keep the claims aligned:

MeasureWhat it answersWhat it leaves outBest use
Pass@1Does one attempt meet the success criterion?Benefits or costs of additional candidatesSingle-response or no-selection workflows
Oracle pass@NDid at least one of N candidates pass?Whether the system can identify that candidateDiagnosing candidate-generation headroom
Selected-result successDid the actual selection procedure produce a passing result?Other budgets or task distributions not testedComparing deployable multi-candidate designs
Cost- or latency-constrained successHow often does the workflow succeed within a specified limit?Limits and assumptions outside the tested conditionsChoosing among designs with explicit operating constraints

The final row is not a substitute for the other measurements. It requires the team to define the limit and count the actual process that reaches an answer. When comparing systems, state whether costs and elapsed time include failed attempts, selection, and any work discarded before the final output.

Design a comparison that can be reproduced

A useful comparison begins with one task definition and one success rule shared across the alternatives. Write down what counts as a pass before running the evaluation. For a factual answer, that may require specific supported claims. For a tool-mediated task, it may require the intended external state change to be present and correct. Do not treat a model’s claim that it completed the task as independent proof of completion.

The evaluation should make the candidate budget explicit: N, whether candidates are generated sequentially or concurrently, when generation stops, and what happens if an attempt times out or returns an unusable result. Record the prompts and material inputs, model and deployment identifiers as available, generation settings, and selector version. If a random seed is exposed, record it; do not assume that every environment offers one or that recording it guarantees identical behavior.

Keep the task set fixed when comparing Pass@1 with an N-candidate procedure, so the difference in results is not simply a difference in task mix. If multiple runs are needed to capture variability, retain the task-level outcomes and run conditions. Do not count several candidates from one task as if they were several independent tasks when the question is how reliably the workflow handles tasks.

The evaluation harness should execute the same meaningful steps as the intended operation: provide the input, capture candidates, apply the selection rule, and score the resulting outcome. If the harness has a shortcut—such as access to labels that the production selector would not see—mark that result as oracle analysis, not selected-result performance. Agent evaluations should score task outcomes as well as inspect transcripts; the transcript can help diagnose a failure, but it is not the outcome itself.

A compact run record might include task ID, attempt index, candidate output reference, pass/fail result, selected candidate ID, selector decision, abstention status, elapsed time, and measured resource use. Store enough to reproduce how the final score was produced, while limiting sensitive input and output retention to what the evaluation requires. The point is traceability of the comparison, not collection of every possible field.

Sampling and holdout design are a separate, important problem: task selection can make a comparison look better or worse regardless of N. For that deeper treatment, see a practical approach to a private agent evaluation set. Here, keep the same appropriately chosen tasks across the one-shot and multi-candidate arms, and be clear about which workload the conclusion covers.

Read the result as an estimate, not a guarantee

A benchmark score is an estimate based on a finite set of tasks. A measured 90% success rate does not mean the next ten production tasks will contain exactly one failure, nor does one strong result establish that rare failures are acceptably uncommon. The estimate depends on which tasks were sampled, how outcomes were scored, and whether the evaluation conditions match the intended workflow.

A binomial proportion confidence interval expresses sampling uncertainty under its assumptions; a small sample does not establish the absence of rare failures. The NIST guidance on confidence intervals for proportions is a useful reference for interpreting that uncertainty. Keep the interval and the assumptions visible when the decision depends on a modest difference between scores.

For example, if a hypothetical test observes zero failures in 100 independent, representative trials, that still does not prove that the true failure rate is zero. A rough 95% upper bound using the familiar “rule of three” is about 3% for the failure rate. This is an illustrative calculation, not a safety conclusion. If observations are correlated, the sample is unrepresentative, or important task classes are missing, that simple interpretation is not justified.

For Best-of-N comparisons, uncertainty attaches to more than the candidate-pass rate. The selected-result rate depends on how often candidates pass and how often the selector chooses them correctly. With limited tasks, apparent selector performance may be noisy too. Report the observed outcomes for each stage, and avoid claiming a precise improvement when the difference is small relative to the evaluation’s uncertainty.

Repeated candidates for a task add information about sampling behavior, but they do not automatically add the same amount of information as new independent tasks. A practical report can show task-level Pass@1, oracle pass@N, and selected-result outcomes, then show how results vary across tasks and repeated runs. That layout helps reveal whether the gain is broad or concentrated in a few easy-to-rescue cases.

Counterexample: a higher oracle score, a worse product

Imagine an evaluation of a drafting workflow where four candidates are produced for every request. The benchmark’s oracle checks all four against a hidden reference and records a pass whenever any candidate is correct. The result is substantially higher than the single-attempt score. The team might reasonably infer that the candidate generator often produces a useful answer.

Now consider the actual product. It has no reliable way to compare factual support, always returns the candidate with the strongest tone, and waits for all four generations before responding. If confident but incorrect answers tend to rank highly, the selected-result rate may be lower than the one-shot rate. The user also waits longer, while the oracle’s hidden reference never exists in the live path.

This is not a paradox. The oracle measured candidate availability; the product needs candidate selection. The benchmark is useful as a diagnostic of potential generation quality, but it does not answer whether the deployed workflow is more reliable. A benchmark report that calls both figures “Best-of-4 accuracy” without explaining the selector makes the result difficult to act on.

A related failure occurs when all attempts share a blind spot. Increasing N can generate more versions of the same wrong assumption. If the task is ambiguous and every candidate interprets it the same way, sampling may produce variety in wording but no improvement in correctness. In that situation, an explicit clarification step or better input may matter more than a larger candidate pool; the evaluation should record the task outcome that reveals the problem.

Operational failures to watch for

Candidates are not comparable. If one arm gets more context, a different prompt, or a different stopping allowance, the result does not isolate the effect of N. Freeze the conditions that should be held constant, and document deliberate differences such as the selector or the total generation budget.

A selector sees information it will not have later. A test harness may score every candidate using an answer key and then choose the passing one. That is valid for oracle pass@N, not for a production-selection claim. Put the selector’s actual inputs on record and score its chosen output independently against the task criterion.

The evaluation counts attempts but misses the terminal state. For a workflow that changes an external system, a candidate that says “done” may not have produced the intended change. Score the business result through the available evidence for that task. Conversely, do not attribute a downstream failure to candidate quality if the evaluation cannot distinguish generation from execution; preserve enough stage-level data to locate where the failure occurred.

Failed or slow attempts disappear from the report. Excluding timeouts, empty responses, and rejected candidate sets can make an N-candidate design appear more reliable than it is. Define in advance how each is scored. If a product can abstain safely, count an abstention distinctly from an incorrect action, and report it as part of the workflow’s outcome rather than removing it from the denominator.

The comparison mixes task types or run conditions. A set with a different proportion of easy tasks can shift the overall score. Changes to model configuration, selector logic, or evaluation rules between runs can also make a small difference uninterpretable. Use a stable task set for the direct comparison, preserve run identifiers, and investigate disagreements at the task level before summarizing them.

The reported improvement ignores the operating limit. A measured gain may depend on a candidate count or waiting period the product cannot afford. Set the maximum N and latency target to plausible workflow constraints before interpreting results. A useful score is the success delivered under that stated budget, not the best figure observed after unlimited attempts.

A decision artifact for your next evaluation

Use this worksheet to define the comparison before anyone runs it. Its purpose is to prevent an oracle result, a selector result, and a production result from being collapsed into one number.

  • Workflow: Name the user or downstream process that receives the final answer, and state whether it sees one candidate, a selected candidate, or an abstention.
  • Success criterion: Describe the observable outcome that counts as a pass. Specify who or what scores it, and how ambiguous or partial results are handled.
  • One-shot arm: Record the fixed task inputs, generation conditions, and whether the workflow may reject the single candidate. Calculate Pass@1 from the task outcomes.
  • N-candidate arm: State N, the generation order or concurrency policy, stopping rule, and handling of timeouts and unusable candidates. Keep its task inputs aligned with the one-shot arm.
  • Selection rule: Name the exact signal used to choose or reject a candidate. State whether the rule can inspect a hidden label. If no realistic selection signal exists, label the any-candidate score “oracle pass@N.”
  • Outcome and budget: Record selected-result success, abstentions, elapsed time, and measured resource use. Define which generation and selection work is included.
  • Evidence and uncertainty: Report the number of tasks, task-level variation, repeated-run conditions, and an appropriate uncertainty interval for proportions. State the workload represented and important limitations.
  • Decision: Choose Pass@1, a defined N-candidate design, or a further evaluation. Explain the decision using the selected-result outcome and operating constraints, not the oracle score alone.

A concise printable summary can fit on one page: task and success rule; Pass@1; oracle pass@N, if measured; selected-result success; N and selector; time and resource use; uncertainty; and the decision with its reason. Keep raw task-level records available to the people responsible for interpreting the result. A polished headline without that trace is difficult to reproduce or diagnose.

Make the comparison serve a specific decision

Use Pass@1 when the product’s actual contract is one response per request, or when no practical mechanism can recognize the best candidate. It is also the right baseline for any multi-candidate experiment: without the single-attempt result on comparable tasks, the team cannot tell whether extra sampling helped.

Use oracle pass@N as a diagnostic when the question is whether the generator can produce a passing candidate at all. Pair it with an operational selector evaluation before treating it as a deployable reliability estimate. If the oracle-to-selected gap is large, invest in understanding selection rather than assuming that generating still more candidates will solve the problem.

Use selected-result performance when the intended workflow genuinely generates and chooses among candidates. Report the selector, rejection behavior, and budget alongside the score. If the practical choice is between more samples and more reasoning effort, evaluate that as a separate design decision; this discussion of reasoning effort, latency, and user-facing applications addresses that distinct tradeoff.

Finally, treat a benchmark as a decision instrument rather than a ranking contest. A higher score matters only if its success definition, selection procedure, task mix, and resource limits correspond to the workflow being built. For teams turning that comparison into an evaluation plan, AI evaluation and strategy is an appropriate next step. The immediate engineering discipline is simpler: measure one attempt, measure the actual selected result, and never call the best hidden candidate the system’s answer.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call