Why a model comparison can be a harness comparison
A benchmark result is not produced by a model in isolation. It comes from a model operating inside a defined setup: instructions, tools, context, execution rules, stopping conditions, and a way to judge the result. Change any of those and the measured outcome may change, even when the underlying model stays fixed.
A harness is the surrounding system that presents a task to a model and manages its work. In an agent evaluation, the agent harness may decide which tools are available, how tool calls are executed, what information is returned, and when the run ends. An evaluation harness is the separate machinery that selects cases, records runs, scores evidence, and summarizes results. Agent evaluations should assess outcomes as well as transcripts; the agent harness and evaluation harness are distinct components. (Anthropic’s discussion of agent evaluations)
A model effect is a change attributable to swapping the model while holding the relevant harness conditions stable. A harness effect is a change attributable to altering the surrounding setup while holding the model stable. These are experimental distinctions, not claims that a model or harness acts independently of everything else. In practice, the effect of a model can depend on the harness: a model that benefits from a longer context or a particular tool interface may not show the same advantage in another setup.
That interaction is why a leaderboard rank alone is weak evidence for a procurement or deployment decision. If Model B scores higher than Model A in a different harness, the result does not tell you whether B is better for your task, whether its harness was more favorable, or whether both factors contributed. A controlled comparison must create runs where those explanations can be separated.
This article focuses on designing that controlled experiment. It does not select a universal benchmark, estimate the total cost of an agent task, or decide whether repeated attempts are an appropriate product behavior. For broader benchmark selection, see Benchmarks That Actually Matter for New Model Releases. For repeated tool-use measurement, the distinct tool-use reliability methodology covers that problem in more depth.
The comparison you are trying to make
Suppose an organization wants to compare two models on a support-operations task. One system receives a case, searches a knowledge store, drafts a proposed response, and may update a case record. A second system uses a different model, but also has a revised prompt, a different search tool, and a more permissive execution loop. If the second system completes more cases, the result is a useful observation about those two complete systems. It is not yet evidence that the second model caused the difference.
The experimental unit should therefore be defined before any runs begin. It might be one software issue, one customer case, or one request to create a structured report. Each unit needs a stable task input and a rule for deciding whether the requested business outcome occurred. For issue resolution, for example, the unit is not merely a generated patch; the intended outcome could include satisfying specified acceptance conditions in a controlled environment.
A case is one task instance. A run is one execution of one case under a specified configuration. A condition is the configuration being compared, such as the same model with harness A versus harness B. A replicate is another run under the same condition, useful when outcomes can vary across runs. Keeping these terms explicit prevents a common counting error: treating many runs of a few cases as if they were many independent task examples.
There are two basic comparisons. First, hold the model constant and compare harnesses. This estimates how much the chosen harness changes measured performance for that model. Second, hold the harness constant and compare models. This estimates model differences under that particular harness. To evaluate both, run a crossed design: each model is tested in each harness, on the same eligible cases.
| Question | What stays fixed | What changes | What the result supports |
|---|---|---|---|
| Harness comparison | Model, task cases, scoring rule | Harness configuration | Difference between harness conditions for the tested model |
| Model comparison | Harness, task cases, scoring rule | Model | Difference between models under the tested harness |
| Crossed comparison | Task cases, scoring rule | Model and harness as separate factors | Model effects, harness effects, and evidence of interaction |
The table describes the intended contrasts, not a guarantee of causal certainty. A crossed design can still be compromised by inconsistent data, hidden configuration differences, or a scoring process that changes between runs. The protocol must make those risks visible.
What belongs in a harness specification
A harness specification is a versioned description of the parts of the run that can influence what the model sees or does. It should be detailed enough that another operator can reconstruct the condition, not just recognize its name. Calling a setup “agent v2” is not enough if the prompt, tool schema, retrieval behavior, or stopping rule changed during the evaluation.
Record at least the following configuration groups:
- Task presentation: system and task instructions, formatting requirements, context assembly, and any examples supplied to the model.
- Available actions: tool names and descriptions, argument schemas, permissions in the test environment, and the data each tool can return.
- Execution policy: how tool calls are validated and handled, what happens after an error, whether results are truncated, and how many model turns are permitted.
- State and context: initial state, memory or history supplied, retrieval inputs, and whether information persists between steps or runs.
- Completion behavior: stop conditions, handoff rules, response format, and what counts as a completed run versus a timeout or failure.
- Run environment: software versions and settings that materially affect the task, including the environment in which proposed changes are checked.
This list is not a demand to freeze every operational detail forever. It is a way to distinguish intended experimental factors from incidental drift. For example, if a harness comparison is meant to test whether a revised tool-error policy improves completion, then the error policy is the planned difference. If the revised version also changes the instructions, retrieval index, and task timeout, interpretation becomes ambiguous.
The model specification also needs to be concrete. Record the exact model identifier available to the experiment, the endpoint or deployment configuration used, and any inference settings that can be controlled. If a vendor changes a model behind a stable label, the label alone may not establish that two runs used the same model version. Keep dated run records and avoid claiming that a configuration was identical where you cannot verify it.
Some details cannot be made identical across models. A model may require a different input format, or a tool interface may need a model-specific adapter. Decide in advance whether the question is “Which model performs better with one common harness?” or “Which model performs better in a reasonable harness tailored to it?” The first prioritizes comparability; the second evaluates complete, model-specific systems. Both can be legitimate, but they answer different questions.
If an adapter is necessary, treat it as an explicit part of the condition. Document its transformations, validation, and error handling. Do not silently give one model extra context, more retries, or a more forgiving parser and then describe the result as a pure model comparison. Conversely, forcing a common interface that breaks one model’s intended use may answer an artificial question. State the tradeoff and, where the decision matters, run both a common-harness comparison and a separately labeled tailored-system comparison.
Build a task set that supports attribution
A good task set is not simply a pile of examples. It has a declared purpose, a sampling rule, a fixed representation, and a scoring method defined before results are examined. Start by describing the production decision the evaluation will inform. “Choose a model for issue triage” is too broad to guide case construction. A more useful scope identifies the input, allowed actions, required outcome, and boundary conditions, such as issues that can be resolved from repository context versus issues that need clarification.
Write a dataset specification before collecting or freezing cases. Include the case identifier, source or construction method, task category, input fields, expected state or target, scoring evidence, and exclusions. Keep sensitive or mutable source data out of the run record where it is not needed; use stable case references and controlled copies where appropriate. The objective is repeatability, not copying every operational record into an evaluation system.
For software issue resolution, dataset and harness choices are especially consequential. SWE-bench evaluates software issue resolution, so a meaningful score comparison should identify the dataset version and harness used. (SWE-bench) A case set, repository snapshot, test procedure, and patch-handling method all shape what “resolved” means. A score without those identifiers is difficult to interpret as evidence about a particular implementation decision.
Separate development cases from held-out evaluation cases when teams will use results to tune prompts or tools. Development cases help reveal defects and improve the experiment. Held-out cases reduce the temptation to optimize directly for the same examples used to report success. Keep the boundary operational: record who can inspect the held-out labels, how changes are proposed, and when the set is considered consumed. Repeatedly checking the same held-out set can turn it into a de facto tuning set.
Before runs, test whether each case can be scored consistently. Define acceptable outcomes, partial credit if relevant, and failure categories. For example, a task might be judged on whether the requested field was set correctly, whether an unsupported action occurred, and whether the final response accurately reflects the resulting state. Those dimensions should be grounded in the task’s actual requirements, not selected after one system’s behavior is visible.
A transcript is evidence, but it is not automatically the outcome. A polished answer can claim success without a state change, while a correct state change may be accompanied by a clumsy explanation. Score observable task results separately from process traces when both matter. Preserve enough trace information to diagnose why a run failed, but do not use the transcript as a substitute for checking the target state.
A reproducible controlled protocol
The following protocol is designed for a team comparing two harnesses and two models. Reduce it to one factor at a time if the immediate question is narrower. The key discipline is to define conditions, cases, and scoring before reading comparative results.
-
Write the decision statement. Specify what decision the evaluation can inform, who will use it, and what it cannot establish. Example: “Compare the current and revised tool-handling policies for this case set, using one fixed model configuration.” Avoid framing the study as a search for the universally best model.
-
Name the factors and levels. List each model condition and harness condition using stable identifiers. Mark which factor is intentionally varied in each contrast. Record every planned adapter or exception. If a setting cannot be held constant, label it as a possible confound rather than hiding it in implementation notes.
-
Freeze the task set and scoring rules. Assign stable case IDs, preserve the relevant initial state, and define success evidence and failure categories. Write scoring instructions that a reviewer can apply without knowing which condition produced a run. If automated checks are used, describe what they inspect and how ambiguous outcomes are handled.
-
Validate each condition on a small, separate set. The purpose is to catch configuration and logging defects, not to choose a winner. Confirm that each condition receives the intended inputs, has the intended tools, records failures, and reaches the declared stopping condition. Do not use this validation set as the reported held-out comparison if it informed changes.
-
Plan pairing, order, and repetitions. Where possible, run every condition on the same cases so case difficulty is paired rather than left to chance. Randomize or balance run order to reduce effects from service changes, shared-resource contention, or operator sequencing. Decide repetition policy before results are visible. If repeated runs are too expensive or operationally impractical, say so and limit the claim accordingly.
-
Execute with immutable run identifiers. Each record should link case ID, condition ID, model identifier, harness version, run order, timestamp, initial state reference, and outcome evidence. Store the actual rendered input and relevant tool events, subject to the organization’s data-handling rules. A configuration name without the effective run details is a weak audit trail.
-
Score without condition cues where practical. Remove model and harness labels from review material when those labels could bias judgment. Apply the predeclared rubric, record disagreements, and resolve them using a documented rule. If a reviewer must inspect the transcript to understand an outcome, keep that review separate from any automated outcome check.
-
Check protocol integrity before calculating a comparison. Look for missing traces, mismatched inputs, unexpected tool behavior, changes to the task state, and scoring exceptions. Mark invalid runs and apply the predefined rerun rule. Do not selectively rerun a poor result just because it looks anomalous; that practice can favor the condition that receives more chances.
-
Analyze the planned contrasts. Compare harnesses within each model and models within each harness. Report the number of cases and valid runs, the outcome definition, exclusions, and uncertainty or variability appropriate to the design. Show category-level patterns where they help explain the aggregate, while avoiding a long search for a favorable slice after the fact.
-
Make a bounded decision. State what the data supports for the tested cases and configurations, what remains uncertain, and what evidence would change the decision. A measured advantage in one harness does not automatically transfer to another model, task distribution, or deployment environment.
This sequence is deliberately more demanding than running two configurations and comparing a single percentage. The overhead buys interpretability. If the experiment is intended only as an early screening exercise, label it as such and keep the claim proportionate to the controls actually implemented.
Worked hypothetical: separating a tool-policy change
Consider a hypothetical engineering team deciding whether to adopt a revised harness for repository maintenance tasks. The current harness exposes a read-only search tool and a patch-submission action. The proposed harness changes how tool errors are returned and allows a limited retry after a recoverable error. The team also wants to know whether a second model should be considered. This is a design example, not a report of measured results.
The team first defines a task unit: one issue paired with a fixed repository snapshot and a written acceptance condition. It selects cases from several issue categories, such as a localized defect, a test adjustment, and a change that requires understanding multiple files. The dataset record stores the issue text, snapshot reference, starting state, category, expected evidence, and any exclusion reason. It does not declare success merely because a patch was produced.
The first experimental contrast uses one fixed model configuration with the current and revised harness. That isolates the planned tool-policy change more cleanly. The second contrast runs two model conditions within each harness. The crossed design reveals whether the revised policy helps both models similarly or whether one model interacts differently with the changed error handling. A model-only comparison under the current harness remains available for an immediate decision, but it is not confused with the harness comparison.
The team freezes the task text, tool schemas, repository snapshots, and scoring rubric. The harness versions differ only in the error-return and retry behavior that the experiment is meant to test. If a model-specific adapter is required, its exact transformation is documented and reviewed before the evaluation. The team does not quietly increase the retry allowance for one model or modify prompts after seeing the first comparative outcomes.
For each run, the record captures whether the task’s acceptance condition was met, whether an action failed, whether a retry occurred, whether the run stopped normally, and whether the final response accurately describes the result. These are separate fields rather than a single opaque pass/fail label. A reviewer can then distinguish a tool-recovery issue from an incorrect patch, a timeout, or a claim that overstates what happened.
The team pairs conditions on the same case IDs and balances execution order. It decides beforehand what makes a run invalid—for example, a missing initial snapshot reference or a harness version that does not match the planned condition. A valid but unsuccessful run remains in the comparison. An invalid run is handled under the stated rerun rule and remains visible in the audit record.
After scoring, the team could find that the revised policy changes outcomes for one category but not another, or that a model difference appears only under one harness. No particular result is assumed here. The important interpretive point is that a category pattern can guide the next experiment, while the aggregate alone may conceal it. The team should report the actual case mix and avoid describing a result on this sample as a general property of all repository tasks.
A follow-up question might concern whether the system should make multiple attempts in normal operation. That requires a separate decision about the intended sampling and selection procedure; it should not be smuggled into a harness comparison. The distinct Pass@1 versus Best-of-N discussion addresses that comparison. Likewise, task economics belong in a separate cost model rather than being inferred from completion rates; see how to model cost per successful agent task.
Reading effects, interactions and uncertainty
A main effect is the average difference associated with one factor across the levels of another factor in the tested design. In a crossed model-and-harness comparison, a model main effect summarizes the model contrast across tested harnesses; a harness main effect summarizes the harness contrast across tested models. These summaries can be useful, but they may conceal a meaningful interaction.
An interaction occurs when the effect of one factor changes depending on the other factor. If a revised harness helps one model but not another, reporting only an average harness improvement can mislead. Show the within-model harness contrasts and within-harness model contrasts, then discuss whether the pattern is consistent enough to support a deployment choice. Do not interpret a noisy difference in a small subset as a stable interaction without adequate evidence.
Pairing cases matters because tasks differ in difficulty. If every condition sees the same cases, analysis can compare outcomes case by case and account for that shared difficulty. If conditions receive unrelated samples, an apparent difference may reflect sample composition. When the same case is run repeatedly, remember that those repetitions are not new independent cases. Report both the case count and run count so readers can see the design’s actual breadth.
Randomizing run order helps guard against time-linked changes, but it does not cure every source of variation. Service conditions, background load, data drift, and external state may still differ. Record relevant changes and define when a run block should be paused. If conditions cannot be interleaved, balance their order across blocks and treat time as a possible explanation, not a footnote.
Scores should be accompanied by uncertainty appropriate to the unit of analysis. For a small task set, a difference of a few cases can be unstable even when the displayed percentages look precise. Avoid decimal precision that the sample design cannot support. Report case-level counts and category patterns, explain the method used to summarize variability, and distinguish exploratory observations from a predeclared primary comparison.
A useful practical distinction is between measurement uncertainty and scope uncertainty. Measurement uncertainty concerns how much the estimate might vary under the evaluation’s sampling and execution conditions. Scope uncertainty concerns whether the selected cases and setup represent the future work where the result will be used. More runs on the same narrow set may reduce some execution noise, but they do not establish that the set represents a different task distribution.
Failure modes that invalidate an attribution
A controlled experiment can fail in ways that still produce neat tables. The following failure timeline illustrates how a plausible result can lose its meaning. Before the first run, a team freezes the task set but not the prompt template. During execution, an operator edits the instructions to fix a formatting issue in one condition. Later, a tool schema is updated in both environments, but only one run log records the change. At scoring time, reviewers know which condition produced each transcript. The final comparison may be internally consistent as a spreadsheet, yet it no longer isolates the intended harness difference.
Prevent this with versioned artifacts and explicit run invalidation rules. Store a content identifier or controlled version reference for task inputs, prompts, tool schemas, and scoring instructions. If a change is necessary, stop the affected block, record the change, and decide whether to restart or analyze the conditions as separate versions. Do not silently merge records from two configurations under one condition label.
A second failure is selective repair. One condition encounters an error, so an operator reruns it, while a successful run in another condition is retained. The extra chance can improve the apparent outcome of the repaired condition. Predeclare which failures qualify for rerun, whether the first attempt remains in the operational metric, and how a replacement is labeled. For causal analysis, preserve the original event even if an invalid technical run is excluded from a specific calculation.
A third failure is score leakage. A reviewer sees the model name or polished response and unconsciously applies a more favorable interpretation. Blinding is not always possible, but the scoring process can separate objective state checks from judgment calls, hide labels where practical, and record ambiguous cases. If human judgment is central, use a written rubric and have a second reviewer examine a defined subset rather than resolving differences informally after the aggregate appears.
A fourth failure is outcome substitution. The team intended to assess successful task completion, but later reports the rate of syntactically valid outputs because that measure is easier to collect. That can be a useful diagnostic, but it is not the original business outcome. Keep each metric tied to a specific question, and make clear when the decision metric differs from an intermediate quality signal.
Finally, a correct result can still be used beyond its scope. A harness may help on the tested task set and fail on work with different data, constraints, or tools. That is not proof the experiment was worthless; it is a reason to make the generalization boundary explicit and select follow-up cases based on the deployment decision.
Flow of a defensible comparison
The diagram summarizes the protocol. “Freeze factors” means documenting the model, harness, cases, scoring, and planned differences before comparative runs. The validity gate checks whether each run followed its assigned condition and retained the required evidence. An invalid run is repaired or excluded only under the predeclared rule; it does not simply disappear. The final comparison estimates the planned contrasts rather than treating every observed difference as a model effect.
flowchart TD
accTitle: Controlled model and harness comparison
accDescr: Define tasks, freeze experimental factors, execute assigned runs, check protocol validity, repair invalid records under a preset rule, then estimate the planned effects.
A["Define task set"] --> B["Freeze factors and scoring"]
B --> C["Run assigned conditions"]
C --> D["Check protocol validity"]
D -->|"Invalid"| E["Repair by preset rule"]
E --> C
D -->|"Valid"| F["Estimate planned effects"]
The repair loop is not permission to keep rerunning until a preferred result appears. It applies only to predefined technical invalidity, such as a missing input record or an incorrectly assigned configuration. A valid run that performs poorly is data, not a defect to erase. The experiment record should show the original run, the reason for any replacement, and how each is treated in analysis.
A printable experiment worksheet
Use this worksheet as a compact protocol record. Complete it before comparative runs, then attach the detailed configuration and case specifications. If a field does not apply, write “not applicable” with a brief reason rather than leaving it ambiguous.
Decision and scope
- Decision: State the operational choice this evaluation may inform.
- Primary question: Write one sentence naming the factor contrast, such as harness A versus harness B for one fixed model.
- Boundary: Name the tasks, users, environment, and cases outside the study’s intended claim.
- Success evidence: Specify the observable outcome that represents task completion; keep intermediate signals separate.
Conditions and task set
- Model conditions: Record exact identifiers and controllable settings for each model condition.
- Harness conditions: Version prompts, tools, adapters, context assembly, execution rules, and stopping behavior.
- Planned differences: List the intended changes and identify unavoidable differences that could confound interpretation.
- Case specification: Define case IDs, selection method, starting state, categories, expected evidence, and exclusions.
- Held-out boundary: Record which cases can be used for tuning and which are reserved for the reported comparison.
Execution and scoring
- Pairing and order: State whether each condition sees the same cases and how run order is balanced or randomized.
- Repetition policy: Define repetitions, retry rules, invalid-run criteria, and the treatment of original attempts.
- Run record: Capture condition, case, model, harness version, order, timestamp, relevant trace, and outcome evidence.
- Scoring rubric: Define success, partial outcomes if applicable, failure categories, reviewer process, and ambiguity handling.
- Protocol check: Identify who verifies assignment, input consistency, trace completeness, and scoring exceptions.
Reporting and decision
- Planned contrasts: Report within-model harness comparisons and within-harness model comparisons where the design supports them.
- Counts and exclusions: Show case counts, run counts, invalid runs, reruns, and reasons for exclusions.
- Uncertainty: Explain what variation the analysis captures and what the sample cannot establish.
- Interaction check: Describe whether the harness difference appears consistent across models or depends on the model condition.
- Bounded recommendation: State the tested configuration, supported decision, unresolved uncertainty, and the next evidence needed.
For a quick review before publication or an internal decision meeting, check that a reader can answer four questions without asking the experiment owner: What changed? What stayed fixed? What counted as success? Which records were excluded, and why? If any answer depends on undocumented operator judgment, the comparison is not yet ready to support a strong attribution.
Choosing the next experiment
A single controlled study should narrow uncertainty, not be expected to settle every model and system question. If the harness contrast is decisive for the intended workflow, the next step may be a representative operational evaluation under the selected configuration. If the result depends on model condition, investigate the interaction with cases designed to distinguish likely causes, such as tool-error recovery or context demands. If scoring disagreement dominates, improve outcome definitions before spending effort on more runs.
When the implementation contains several changed components, use staged comparisons. First isolate the highest-priority component with a narrow contrast. Then test combinations that matter to the product decision. A full factorial design can become expensive as factors multiply; staged testing is often more tractable, provided each stage is described honestly and later interactions are not ruled out merely because they were not measured.
Teams may also need support translating an operational decision into an evaluation plan. AI evaluation and strategy is the relevant next step when the decision, task definition, or comparison design needs structured review. The goal is a protocol that answers the organization’s question—not a larger score table detached from the system it is meant to inform.
The central discipline is simple to state and demanding to maintain: keep the task and scoring stable, make intended changes explicit, and preserve enough evidence to reconstruct every condition. With that discipline, a model comparison can distinguish model effects from harness effects, reveal where those effects interact, and state the limits of what the result supports. Without it, even a precise-looking benchmark may compare two different systems while attributing the difference to only one.