My thesis is simple: an AI benchmark is credible only when another team can reconstruct what was tested, what it cost, what failed, and which exact model and settings produced the reported result. A ranking without that record may be interesting, but it is not a sound basis for a purchasing, deployment or migration decision.
I would treat a benchmark report as a compact research artifact, not a marketing chart. Its job is to make the comparison inspectable: define the task, identify the data and system versions, expose relevant operating conditions, show how outcomes were judged, and preserve exceptions instead of smoothing them away. The proposal below is a practical reporting standard for teams publishing model comparisons, including internal evaluations whose decisions affect customers or budgets.
This is an editorial proposal, not a claim that one template fits every evaluation. It focuses on dataset, exact versions, settings, costs, failures and reproducibility. It does not prescribe how many cases every study needs or which model a team should choose. Those are related decisions with their own context; the point here is to ensure readers can understand what a reported comparison actually means.
A benchmark should make its claims inspectable
A benchmark result is a claim about performance under a particular set of conditions. If those conditions are hidden, a reader cannot tell whether the result applies to their workflow. “Model A scored higher” is incomplete unless the report explains what counted as success, which cases were included, what execution path was used and how an unsuccessful run affected the score.
That standard applies even when a benchmark is internal. A slide circulated to leadership can become the basis for a renewal, a migration or a production change. If the underlying evidence is difficult to recover, the organization may mistake a one-time observation for a stable property of a system. A detailed record does not make a result universal; it makes the boundaries of the result visible.
The most useful benchmark reports separate three things that are often collapsed. The first is the measurement design: cases, scoring rules and planned settings. The second is the execution record: what actually ran, including retries, errors and deviations. The third is the interpretation: what the observed results support, and what they do not. A polished summary can then sit on top of those layers without replacing them.
For teams developing tool-using systems, this distinction is especially important. An evaluation may need to assess outcomes as well as the interaction transcript, and the agent harness that coordinates work is distinct from the evaluation harness that assesses it. Those are separate parts of an evaluation design, not interchangeable labels (overview of agent evaluations). A report should identify both when both are relevant.
Define the question before assembling the dataset
A dataset is not just a folder of prompts. It is a set of defined cases intended to represent a particular decision or task. Before selecting examples, write a one-sentence decision question: for instance, “Can the candidate system complete these support-triage tasks within the stated quality and operating constraints?” That wording is illustrative; the actual question should reflect the workflow being evaluated.
Then describe the population the cases are meant to represent. Are they new requests, difficult escalations, routine structured inputs, or a deliberately balanced mix? Give readers enough information to judge whether the cases resemble the intended work. If examples were filtered, transformed, deduplicated or grouped, record those operations. If a dataset combines cases from different sources, identify the sources at a level that can be disclosed and explain how each contributes to the sample.
A practical case record needs a stable identifier and enough fields to trace the case from input to adjudication. The following fields are a proposed minimum; teams can add domain-specific fields without obscuring the common record.
| Field | What the report should record |
|---|---|
case_id | Stable, non-sensitive identifier used in logs and result files |
case_version | Version or revision of the input and expected outcome |
task_family | Defined group used to organize cases and summarize performance |
input_reference | Recoverable input or controlled reference to it |
expected_outcome | Scoring target, acceptable range or adjudication rule |
case_origin | Source category and relevant transformation history |
inclusion_status | Included, excluded or held out, with a reason where applicable |
sensitivity_note | Any restriction affecting sharing or reproduction |
The table is an artifact proposal, not a prescription for publishing sensitive inputs. A public report can provide stable case identifiers and a carefully described access path without disclosing private or confidential content. The key is not to imply that an undisclosed input is fully reproducible. State which parts can be independently inspected and which require controlled access, and explain the limitation plainly.
Include exclusions in the record rather than silently deleting them. An exclusion may be justified—for example, an input may be malformed relative to the defined task—but the reason and timing matter. If a team removes a case after seeing a model’s answer, readers need to know that. A rule written before execution is easier to defend than a post hoc decision that happens to improve one system’s apparent standing.
The dataset specification should also distinguish case-level success from aggregate success. A case may pass, partially pass, fail, or be unscorable according to defined rules. If the scoring scheme allows partial credit, explain the anchors or rubric that separate scores. If judgment requires a human reviewer, record the reviewer procedure and how disagreements are resolved. Without this information, two readers may apply the same label to different outcomes.
A related discussion of sample size has a narrower focus than this reporting proposal. For guidance on the distinct question of how many test cases to include, see How Many Test Cases Are Enough for an AI Model Comparison?. Here, the priority is to define and preserve each case well enough that the reported evaluation can be understood and rerun.
Identify the exact system that produced the result
“Model name” is not a complete version identifier. A report should state the exact model or endpoint identifier used at execution time, along with the date or period of the run. If a provider exposes a version or snapshot identifier, record it as observed. If it does not, say that the underlying version could not be independently fixed, and avoid wording that implies an immutable version was tested.
The system under evaluation may include more than a model. Record the application build or commit, prompt and instruction versions, tool definitions, retrieval configuration, data snapshot, and any middleware that can influence the output. List the evaluation harness and the agent or application harness separately where both exist. The purpose is not to publish every line of implementation; it is to capture every component that could plausibly change the result.
For each execution, retain a run identifier that ties together the case set, system configuration, timestamp, output, scoring decision and operational events. When a setting changes, start a new configuration identifier. Do not overwrite the earlier run record with the latest value. If a report combines multiple configurations, show which result belongs to which one rather than treating them as a single undifferentiated experiment.
A useful test is whether a technically capable colleague could answer, from the report and its retained records: “What exactly did this system receive, what was it allowed to do, and which version of the surrounding application handled the request?” If the answer requires relying on someone’s memory, the report has a versioning gap.
This matters in model comparisons because a change attributed to a model may actually come from a changed prompt, a different tool schema, a retrieval update or a revised scoring rule. In a migration evaluation, preserve the old and proposed configurations as separate, named records. The distinct implementation question of comparing named model options is explored in Astra vs Sonnet 5.5 vs Opus 5.5: How to Run Your Own Comparison; this article’s concern is the reporting record that lets readers interpret whichever comparison they conduct.
Publish settings that could change the outcome
A benchmark report should disclose the settings that govern the tested execution, not merely say that “default settings” were used. Defaults can change, and the phrase does not tell a reader which values were in force. Record explicit values when available, and identify the source and date of any setting that cannot be fixed or observed.
For a text-generation task, relevant settings may include the sampling configuration, output limit, stop conditions, system and user instructions, and any structured-output constraints. For a tool-using task, also describe the available tools, their schemas, tool-selection rules, maximum action or turn limits, and what happens when a tool returns an error. Include a setting only when it can affect the task or its measurement; a long inventory of irrelevant configuration is not a substitute for a coherent protocol.
Execution conditions deserve the same care. State whether cases were run sequentially or concurrently, whether the system received prior context, how a timeout was handled, and what retry policy applied. If cases were randomized or their order could affect later inputs, explain the ordering. If human reviewers could see system identity while scoring, disclose that, because it may affect interpretation of subjective outcomes.
Record deviations as deviations. Suppose an output limit was changed midway to prevent truncation. The final report should show which cases used each limit, why it changed, and whether the scores were recomputed under a consistent rule. A sentence such as “minor settings were adjusted during testing” is too vague to tell readers whether the comparison remained fair.
These details should appear in a machine-readable configuration file as well as the human-readable report when practical. A short JSON or YAML record can make exact values easier to compare, but it should not be treated as self-explanatory. Include a schema or field descriptions, and ensure that a reference to a prompt or dataset resolves to a specific version rather than a mutable filename such as latest.
Report costs as observed, bounded measurements
Cost reporting is not a universal price quote. It is an account of what the evaluation consumed under specified conditions. A useful report separates the provider charge, if available to the evaluator, from the broader cost of operating the run. That broader account may include tool calls, retrieval, infrastructure, human review and reruns. It should distinguish measured charges from estimates and state the accounting period and unit.
A reader should be able to tell whether the reported number is per case, per completed task, per attempt or for the entire run. Include the denominator and explain how failed or repeated attempts were counted. Reporting a cost per successful case while omitting the number of unsuccessful attempts can make an expensive, failure-prone process appear efficient. Where provider billing records are unavailable or aggregated, say so rather than presenting a modeled estimate as an invoice-backed total.
Consider this explicitly illustrative calculation. Assume a team runs 120 cases. It makes 120 initial attempts, 18 of those attempts are retried under a predeclared retry rule, and 9 cases still fail. If the team reports only “cost per successful case,” it must define whether the denominator is 111 completed cases, 120 assigned cases, or another quantity. It should also show the cost of all 138 attempts and identify any human review or tool charges included. The arithmetic is straightforward; choosing the denominator is a reporting decision that changes the interpretation.
When comparing systems, keep units and accounting boundaries consistent. If one result includes review labor and another does not, do not place both figures side by side without qualification. If the evaluation changes a setting that alters output length or the number of tool calls, describe that alongside cost. The goal is not to force every organization into one accounting model. It is to stop a reader from mistaking unlike measurements for a direct comparison.
Avoid extrapolating a small evaluation into an asserted production bill unless the assumptions are stated. A benchmark’s observed cost may not capture the request mix, concurrency, operational overhead or failure rate of a live workflow. A projection can still be useful if clearly labeled as a scenario: specify its assumed volume, case mix, retry behavior and included charges, and show how the estimate changes if those assumptions move.
Make failure part of the evidence
A benchmark report should account for failed executions and failed tasks separately. An execution failure means the planned run did not complete as specified—for example, an infrastructure error interrupted the record. A task failure means the system completed an attempt but did not meet the task’s success criterion. A malformed output, an incorrect tool action and an unanswered request may be different task outcomes even if all count as unsuccessful at an aggregate level.
Keep the raw status and the scoring interpretation. A concise failure taxonomy might distinguish timeout, provider or application error, invalid format, unavailable tool, incorrect result, partial result and unscorable case. Define categories before analyzing the comparison where practical. If the initial taxonomy proves inadequate, retain the original label and record the revision rather than retroactively making unlike events appear consistent.
The report must explain retries. Was a retry automatic, manually initiated or prohibited? Did it use the same configuration? Did the system receive the prior attempt’s output? Was a new external action possible? These details can change the task, not just the number of attempts. Preserve each attempt under its own identifier, then link it to the case and run. Do not replace the failed attempt with the successful retry in a way that hides how often recovery was needed.
For systems that can trigger consequential external effects, the evaluation design should prevent a benchmark retry from being mistaken for a safe replay. A team can test against a controlled environment or use a non-executing substitute, but the report must describe the setup accurately. A log entry saying “retry” does not establish that an external action happened exactly once. Track the attempted operation, the confirmation received and the resulting state separately when those distinctions matter to the task.
Failure analysis also needs a timeline. Record when the failure occurred, what the system had already done, whether the issue was detected automatically, and how it was resolved. That makes it possible to distinguish a recoverable interruption from a silent failure that produced a plausible-looking but wrong answer. When a reviewer reclassifies a result, preserve the initial decision and the reason for the change.
A benchmark that reports only its cleanest completed runs is not reporting the behavior the team evaluated. Failures are not embarrassing debris to remove; they are part of the observed system behavior. The relevant editorial choice is how to classify them and how clearly to show their effect on the claim.
Separate observed results from uncertainty
A score should travel with its denominator, scoring rule and uncertainty. If 87 of 100 cases pass, report 87/100 and define “pass”; do not provide only “87%.” If the sample is divided into task families, show the group sizes and avoid implying that a high aggregate score means every group performed similarly. A compact table can show overall and subgroup results while keeping the case-level record available for inspection.
For a proportion, an interval can help describe statistical uncertainty under the method’s assumptions; it is not a guarantee about future performance. NIST’s discussion of binomial proportion intervals describes uncertainty in that setting. For this reporting standard, do not treat a small sample as evidence that rare failures are acceptably unlikely (NIST discussion of binomial proportion confidence intervals). State the interval method and assumptions used. Do not imply that an interval corrects for a biased dataset, changed configuration or poorly defined outcome.
The distinction matters because a benchmark can be precise about the wrong target. If the cases omit a relevant task family, a narrow interval around the measured average does not establish performance on that missing work. If scoring rules are inconsistent, more cases can make a flawed measurement look more authoritative without making it more valid.
Use language calibrated to the scope. “On this version of the dataset, under the recorded settings, the system passed 87 of 100 cases” is a bounded observation. “The system is 87% reliable” is a much broader statement and may not be supported by the same evidence. A report should make the narrow claim easy to quote and the important limitations hard to miss.
For a model migration, comparisons should also account for paired cases: the same cases run through both configurations under stated conditions. The resulting case-level differences can explain whether one system’s gains coincide with losses elsewhere. A regression suite helps preserve those comparisons over time; for a separate migration-focused treatment, see Sonnet 5.5 Migration: Build a Regression Suite Before Switching.
Worked example: a hypothetical support-triage evaluation
Consider a hypothetical mid-market software company evaluating two configurations for drafting support-ticket classifications. The purpose is not to assert a result about any real system; it is to show how the proposed report would work. The decision is whether one configuration is suitable for a supervised triage workflow, not whether it can independently resolve customer issues.
The team defines 120 cases across three task families: routine routing, ambiguous ownership and incomplete requests. It assigns stable identifiers such as R-001, A-001 and I-001, stores a versioned input reference for each, and writes an expected routing label plus an allowed “request clarification” outcome where appropriate. Ten cases are reserved for rubric calibration and are not included in the main score. The report gives the reason for that separation and does not quietly add them later.
The two configurations receive the same case inputs and the same allowed output schema. The record names the model or endpoint identifier as observed, the application commit, prompt revision, tool schema revision, evaluation harness revision and run date. A setting file captures output limits, timeouts and retry policy. If the provider does not expose a fixed underlying snapshot, that limitation appears in the version statement rather than being covered by a confident product label.
For this hypothetical study, the team defines success as a correct routing label or a permitted clarification response when the case lacks enough information. A human reviewer scores each output against the rubric, with system identity hidden during scoring. A disagreement is sent for adjudication, and both the initial labels and final decision are retained. These are proposed design choices for the example, not claims about a tested process.
Suppose the illustrative records show that Configuration A passes 96 of 120 cases and Configuration B passes 100 of 120. Those numbers are invented solely to demonstrate reporting. The team should not publish “B is better” without examining case-level differences, subgroup counts, adjudications and the operating record. If B’s additional passes are concentrated in routine routing while it introduces errors in ambiguous ownership, that tradeoff may matter more than the four-case aggregate difference.
Now suppose the illustrative run also contains 18 retries and 9 cases that remain unsuccessful. The report lists the initial attempt, retry reason, configuration and final outcome for each affected case. It totals the observed charges for both attempts and states whether reviewer time is included. If the retries were triggered only after timeouts, that rule is written down; if an operator chose them case by case, the report says so and explains how that discretion might influence comparability.
A counterexample clarifies why this record matters. Imagine a summary that reports B at “83% accuracy” and A at “80%,” but B used a revised prompt on the final quarter of cases, retries are counted only when they fail, and five ambiguous cases were excluded after scoring. The visible ranking may be mathematically correct for the filtered records, yet it does not establish a fair comparison of the two named configurations. A reader cannot identify whether the difference came from the system, prompt changes, retry selection or exclusions.
The remedy is not a more persuasive chart. It is a corrected record: split the run by configuration, identify the prompt boundary, restore or justify the excluded cases under a rule applied consistently, and report retries under a common accounting rule. If the resulting comparison is no longer balanced, label it exploratory and schedule a new run rather than retroactively presenting it as a controlled head-to-head result.
A reproducible run protocol
A report should include a short protocol that another team can follow. The protocol need not reproduce an entire production stack, but it must identify the records and decisions that affect the evaluation. I recommend the following sequence because it catches common reporting failures before the summary is written.
-
Freeze the question and scoring rule. Name the decision, task families, inclusion criteria and outcome rubric. Record what counts as pass, partial pass, failure and unscorable. Decide how adjudication and exclusions work before examining comparative results.
-
Version the dataset and systems. Assign a dataset version and stable case IDs. Record exact model or endpoint identifiers as available, application and harness revisions, prompts, tools, relevant data snapshots and material settings. Capture what cannot be fixed or observed as an explicit limitation.
-
Create a run manifest. Give the planned execution a unique ID. List the configuration, date, case range, execution order, retry policy, timeout behavior, reviewer procedure and cost-accounting boundary. Where values are in a separate file, preserve a stable reference to that file.
-
Execute without overwriting. Write every initial attempt and retry as a distinct record. Link outputs, operational events and scores to their case and run IDs. Retain the original response when a correction or adjudication changes the final label.
-
Validate the records. Check that every planned case has a terminal status or an explained exception, every score has a defined basis, and every retry is counted. Compare actual settings with the manifest and log deviations before calculating the summary.
-
Calculate and publish bounded claims. Report numerators and denominators, subgroup results where relevant, observed costs and failure categories. State uncertainty methods and limitations. Archive the report with its manifest, dataset specification and an access description for any restricted material.
flowchart TD
A["Freeze protocol and versions"] --> B["Execute cases and retain attempts"]
B --> C["Validate records against manifest"]
C --> D{"Records complete and consistent?"}
D -->|"Yes"| E["Calculate and publish bounded results"]
D -->|"No"| F["Quarantine run and investigate"]
F --> G["Correct protocol, create new run ID"]
G --> B
accTitle: Benchmark execution and reporting flow
accDescr: Freeze the protocol, execute and retain attempts, then validate records. Complete and consistent records proceed to reporting. Incomplete or inconsistent records are quarantined, investigated, corrected under a new run identifier, and executed again.
The decision node is about record quality, not whether the result is favorable. A low score with complete records can be reported; a favorable score with missing attempts or unexplained configuration changes should not be treated as a clean result. Quarantining an inconsistent run preserves its history while preventing it from being silently merged into a comparable result. If a correction changes the protocol, a new run identifier makes the boundary visible.
The final archive should contain, or point to, the report, dataset specification, configuration manifest, case-level outcomes, attempt log, scoring rubric and cost ledger. For restricted data, include a description of what is withheld and a feasible route for authorized inspection if one exists. Test the archive by asking someone who did not run the evaluation to locate a particular case, recover its configuration and trace its final score to the underlying outcome.
The counterargument: this may be too much for a quick decision
The strongest objection is practical. Teams often need a fast directional answer, and a full case-level record can take longer to prepare than a lightweight comparison. Some input data cannot be published, some model identifiers may not expose a stable snapshot, and a highly detailed report can overwhelm a reader who only needs a decision brief.
Those are real constraints, but they argue for a clearly labeled tier of evidence—not for presenting a limited test as a definitive benchmark. A quick screen can publish a concise summary with its dataset size, task definition, versions, major settings, cost boundary and known failures. It can say that the result is preliminary and identify what evidence is missing before a consequential decision. The detailed record can remain internal or access-controlled where disclosure is inappropriate.
The reporting standard should be proportional to the claim. A small exploratory run need not imitate a formal research publication. But if a team uses its results to claim that a system is safer, more reliable or materially cheaper in a business workflow, then the relevant evidence must be available for scrutiny. A short report can still be honest; a short report that hides the conditions behind a confident generalization cannot.
There is also a legitimate concern that rigid templates can reward documentation rather than good measurement. A team could fill every field and still choose a biased dataset or weak scoring rule. I agree. This proposal is not a certification rubric, and completeness is not validity. Its narrower purpose is to expose the assumptions and execution history so that readers can challenge the design instead of guessing what happened.
A printable reporting worksheet
The following worksheet is intended to be usable inside a report or run archive. Mark an item “not applicable” only when the reason is clear. “Unknown” is a meaningful result: it tells a reader that a claim may have a limit. The worksheet does not replace the case-level records, but it gives reviewers a consistent path through them.
Study definition
- Decision question: Is the decision stated in one sentence, with the workflow and intended use bounded?
- Task population: Are the task families, intended case population and known omissions described?
- Scoring rule: Are pass, partial pass, failure and unscorable outcomes defined before comparison?
- Inclusion and exclusion: Are the criteria recorded, with identifiers and reasons for excluded or held-out cases?
- Dataset revision: Can the reader identify the exact case set and the version of each case used?
System and execution
- Model identity: Are exact model or endpoint identifiers recorded as observed, with any snapshot limitation stated?
- Surrounding system: Are relevant application, prompt, tool, retrieval, data and harness revisions named?
- Settings: Are material generation, output, timeout, concurrency and tool-use conditions specified?
- Run identity: Does each execution have a stable run ID and date or time period?
- Deviations: Are actual conditions compared with the planned manifest, with changes documented rather than overwritten?
Outcomes, failures and cost
- Case-level trace: Can each reported score be traced to a case, attempt, output and scoring decision?
- Retries: Are retry triggers, attempt counts, changed conditions and final outcomes visible?
- Failure categories: Are execution errors distinguished from task failures and unscorable cases?
- Aggregate denominators: Are numerators and denominators shown for overall and relevant subgroup results?
- Cost boundary: Does the report say what charges and labor are included, what is estimated and which unit is used?
- Uncertainty: Is any interval or uncertainty statement accompanied by its method and limitations?
Reproduction and publication
- Protocol: Could a qualified colleague reconstruct the sequence of steps and evaluation decisions?
- Access limits: Are restricted inputs or records identified, with the limitation on independent reproduction stated?
- Archive: Are the manifest, rubric, attempt records and cost ledger retained at stable references?
- Claim scope: Does the conclusion describe observed performance under the recorded conditions rather than imply universal behavior?
- Review: Has someone other than the run owner checked that summary figures reconcile with the retained records?
For a printable summary, the report owner can place this compact record at the front of the archive: Decision: [workflow and choice]. Dataset: [version, case count and task families]. Systems: [exact identifiers and application revisions]. Settings: [manifest reference and material conditions]. Results: [numerators, denominators and scoring rule]. Failures: [categories, retries and unresolved cases]. Cost: [unit, included components and observed total or estimate]. Limits: [important unknowns and access restrictions]. Reproduction: [archive reference and run ID]. A completed summary should point to evidence, not substitute for it.
My recommendation
I would make this reporting standard the default for any model comparison likely to influence an operating decision. Keep the public summary readable, but preserve the records that let a reviewer reconstruct the result. Treat dataset identity, system version, settings, costs and failures as part of the result itself—not supplementary notes that can be dropped when a chart is copied into a presentation.
For a team designing an evaluation program, the useful next step is to establish a reporting protocol before the next consequential comparison. If the team needs help translating a business decision into an evaluation plan and a defensible reporting process, explore AI evaluation and strategy. The core principle remains independent of any service or tool: publish enough of the method and execution history that the claim can be examined, bounded and, where possible, reproduced.