SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Comparison 20 mins

An Agent Harness Is Not an Evaluation Harness: Here’s the Difference

An Agent Harness Is Not an Evaluation Harness: Here’s the Difference. Practical examples, tradeoffs and implementation guidance for technology leaders.

The PADISO Team ·

Overview: two harnesses, two different jobs

An agent harness is the machinery that lets an agent attempt work. It assembles the model call, tools, run state, limits, and effects on external systems. An evaluation harness is the machinery that determines whether an attempt met a defined standard. It supplies test cases, captures evidence, applies graders, and reports results.

They are easy to confuse because both can invoke an agent, collect a transcript, and produce a status. But the status answers different questions. An execution harness asks, “What should happen next so this task can proceed?” An evaluation harness asks, “Given the result and evidence, how well did this attempt satisfy the requirement?”

This distinction matters when a team is deciding what to build, debugging inconsistent results, or using test outcomes to decide whether a change is safe to release. A reliable executor does not prove that its outputs are good. A convincing evaluation report does not make a fragile executor reliable.

Agent evaluations should grade outcomes, not transcripts alone. Anthropic’s discussion of agent evaluations makes that distinction useful: a fluent interaction can still fail the task, while a result may meet its requirement despite a different path through the conversation.

DimensionAgent harnessEvaluation harness
Primary purposeExecute and coordinate a taskJudge task performance against defined criteria
Main inputA request, operating context, tools, and constraintsA test case, expected properties, evidence, and grading rules
Main outputA result, state transition, or controlled failureA score, pass/fail decision, diagnosis, or comparison
Typical ownerProduct or platform engineeringProduct, quality, research, or evaluation engineering
Main riskUnsafe, incomplete, duplicated, or untraceable executionMisleading grades, weak coverage, or conclusions unsupported by evidence

A team may build both in one repository or use separate services. The boundary is conceptual, not a requirement for two products. The important design choice is to keep execution responsibility distinct from judgment, even when code is shared.

Start with the execution harness

The execution harness owns the live attempt. At minimum, it needs to know which task is being run, what input it received, which model and instructions are in scope, what tools are available, and what state must survive between steps. It also needs a way to stop: a tool error, a deadline, a usage limit, an invalid response, or a business rule can all make continuation inappropriate.

An execution loop is more than repeatedly asking a model what to do. It has to validate a proposed action, decide whether that action is permitted, call the appropriate tool, record what actually happened, and feed a truthful result back into the next decision. When the attempt ends, it should preserve enough information to explain whether it finished, stopped, or needs a person or another system to resume it.

That record should distinguish the agent’s claim from observed facts. If an agent says it updated an account, the harness should not treat the claim as proof that the account changed. The executor should record the tool result or other available confirmation, and preserve uncertainty if the result cannot be observed. This is a design recommendation, not a guarantee that every external system offers a reliable confirmation path.

For a broader component map, see the anatomy of an AI agent. This article focuses on the narrower boundary between running an attempt and evaluating it, rather than repeating a general agent architecture tour.

Execution strengths and tradeoffs

A purpose-built execution harness can make limits and failure paths explicit. It can carry task identifiers through tool calls, save state before a long pause, prevent an invalid action from reaching a downstream service, and return a useful failure classification instead of an opaque exception. Those are operational capabilities: they help a system behave predictably while work is underway.

The cost is that execution infrastructure has to deal with changing reality. A network request can time out after the remote system has processed it. A model response can be malformed. A task can outlive a process. A retry can repeat an external effect. The harness needs policies for these conditions, and those policies are part of the application rather than incidental plumbing.

The harness also cannot decide by itself that a task result is correct just because the run ended normally. A tool may have returned success for an incomplete update. A response may meet a format constraint while omitting a material fact. A clean end state means only that the execution process reached its defined stopping point.

Evaluation strengths and tradeoffs

An evaluation harness makes quality claims inspectable. It can apply fixed checks to structured fields, compare outputs against task-specific expectations, ask a judge to assess a bounded criterion, or combine multiple forms of evidence. It can also rerun a stable set of cases after a change and show where behavior differs.

Its principal weakness is that a score can look more authoritative than its measurement. A test set may omit an important class of inputs. A grader may reward a plausible explanation rather than a correct outcome. A binary threshold can conceal a serious failure in one case behind acceptable results elsewhere. An evaluation needs a stated target, evidence, and known limits before a number is useful.

It also does not own the live task. A failing grade should inform a decision; it should not, by itself, retry a customer transaction or reverse an external change. Those are execution and product decisions. A test runner may trigger a controlled test attempt, but it should not be mistaken for the component that safely handles production work.

Compare the responsibilities feature by feature

1. Inputs and boundaries

The execution harness begins with a task that must be carried through a sequence of decisions. Its input includes operating parameters: request content, relevant state, tool definitions, and limits. It may receive new observations as work proceeds. Its boundary ends when the task reaches a defined terminal state, pauses for a valid handoff, or fails under an explicit stopping rule.

The evaluation harness begins with a case designed to test a claim. A case needs a prompt or starting state, the conditions that matter, and a rule for interpreting the result. It may run an agent as part of the test, but the case is not itself the production task. Its boundary ends when evidence has been collected and assessed, with any uncertainty recorded.

The distinction affects repeatability. A live request may depend on changing records or external services. An evaluation case should control, simulate, or record those dependencies sufficiently to make comparisons meaningful. If the case changes silently between runs, an apparent model improvement or regression may actually reflect different input conditions.

Execution harness — pros: it can work with current task state and perform the next allowed operation. Cons: that same connection to changing systems makes exact replay difficult and increases the consequence of a bad action.

Evaluation harness — pros: it can define a stable set of cases and compare candidate behavior against a common target. Cons: too much isolation can hide problems that only appear with real system state, timing, or tool behavior.

2. State, continuity, and recovery

Execution state answers operational questions: what has already happened, what remains, what must not be repeated, and where to resume after a pause. A run identifier, step record, last confirmed result, and explicit terminal status are more useful than storing a long conversation with no interpretation. State should capture facts needed for recovery, not merely preserve every token.

For work that spans context windows, progress artifacts and structured handoffs can help preserve continuity; they do not guarantee completion. Anthropic’s treatment of long-running agent harnesses is relevant to this narrow point. The practical design implication is to make resumed work inspectable: a later process should be able to tell what was confirmed, what is pending, and which assumptions need checking before continuing.

Evaluation state serves a different purpose. It records which case ran, which candidate or configuration was assessed, what evidence was collected, and how each criterion was graded. It should make a result reproducible enough to investigate. An evaluation record is not automatically a recovery checkpoint for a production task, and a production checkpoint is not an evaluation result.

Execution harness — pros: it can resume work from recorded operational state rather than asking the agent to infer what happened. Cons: stale or ambiguous state can cause a resumed attempt to repeat an action or trust an outdated assumption.

Evaluation harness — pros: it can retain case-level history and make changes over time visible. Cons: recording a run without preserving the tested inputs and grading basis makes the history hard to interpret.

3. Tools, effects, and observation

The executor owns the decision to invoke a tool and the handling of its result. That does not mean every executor should expose unrestricted capabilities. It means the execution path must validate the requested operation, supply only the needed inputs, handle errors, and record what was returned. The specific contract between an agent and a tool deserves its own design treatment; see how to design an agent tool contract.

A tool call has at least three meaningful states: not attempted, attempted with a confirmed result, and attempted with an uncertain result. A timeout after submission belongs in the third category unless the remote system can confirm whether it processed the request. Automatically treating uncertainty as failure and retrying may duplicate an effect. Automatically treating it as success may leave the task incomplete. The harness should surface the uncertainty and follow a defined recovery policy.

An evaluation harness may invoke tools while exercising a case, but tool execution is a means of producing evidence, not the grading decision. For example, a test can inspect a resulting record or compare a captured action with an expected property. If it grades only the model’s description of what it supposedly did, it risks confusing a claim with an observed outcome.

Execution harness — pros: it can coordinate real operations and return their actual responses to the next step. Cons: it has to manage side effects, uncertain outcomes, and downstream failures that a test-only environment may never encounter.

Evaluation harness — pros: it can examine outputs and recorded effects against a stated criterion. Cons: its verdict is only as sound as the evidence it can observe; a test that cannot see the relevant business state must not imply that it verified that state.

4. Grading and evidence

An evaluator needs to answer a narrow question, such as whether required fields are present, a constraint was respected, or a business result matches the case’s expected condition. Deterministic checks are suitable for objective properties. A model-based grader may help assess less mechanical qualities, but its judgment should be bounded by a rubric and supported by the specific evidence it is asked to inspect.

A useful report keeps the evidence close to the finding. Instead of saying only “pass,” it can record the case identifier, criterion, observed value, expected property, and any grader rationale needed for review. Instead of a single aggregate score, it can preserve per-case results and group failures by kind. This makes the report a diagnostic instrument rather than a decorative metric.

The executor should not quietly absorb the evaluator’s role by declaring success based on a model’s final message. Conversely, the evaluator should not modify the run record to make a result appear complete. If an evaluation identifies a concern, that finding can lead to a separate decision: reject a release, inspect a case, revise a prompt, or change a tool. The path from finding to action should be explicit.

Execution harness — pros: it can expose the operational facts needed for a later judgment. Cons: its completion status is not a quality grade and should not be presented as one.

Evaluation harness — pros: it can turn stated requirements into repeatable checks and reveal weaknesses that anecdotal review misses. Cons: a grader can be wrong, incomplete, or poorly aligned with the actual business requirement.

5. Failure handling and stopping

For an execution harness, failure handling is part of the task lifecycle. It needs to decide what to do when input is invalid, a tool refuses an operation, an external call times out, an output cannot be parsed, or a run reaches its limit. Continuing is not always the safest response. A controlled stop with a precise reason can be better than another model turn that increases uncertainty.

The evaluation harness needs its own failure taxonomy. A case may fail because the agent produced a wrong result, because an evaluator could not assess the evidence, because a dependency was unavailable, or because the test setup itself was invalid. Combining these into one failed status makes the score difficult to act on. In particular, an infrastructure failure should not be reported as proof that the agent failed the task.

The two systems should preserve linked but distinct statuses. The execution record can say that a run stopped after an uncertain tool result. The evaluation record can say that the case is not gradable because the intended outcome could not be observed. That is more honest than forcing either system to declare success or failure without sufficient evidence.

Execution harness — pros: it can stop at the point where continuing would exceed a limit or rely on an unsafe assumption. Cons: conservative stops can leave legitimate work unfinished and require a well-defined human or automated recovery path.

Evaluation harness — pros: it can separate agent-quality failures from test-system failures. Cons: maintaining those distinctions requires operational discipline; a simple pass/fail dashboard may erase them.

Worked example: reviewing an invoice exception

The following is a hypothetical design for a mid-market operations team. An agent receives an invoice exception and is expected to compare the invoice with a purchase record, identify a discrepancy, and prepare a recommendation. It must not approve payment. The case is useful because the desired outcome is not just a polished explanation: the review must identify the relevant discrepancy and leave an inspectable record for a person to act on.

The execution harness receives a task identifier, invoice data, a purchase record reference, an allowed set of read operations, and a deadline. It fetches the relevant record, asks the agent to classify the discrepancy, validates the proposed structured result, and stores a recommendation. It records whether the read succeeded, which values were compared, and whether the recommendation was written. A failed read ends the attempt as incomplete; it does not invite the agent to invent a matching purchase record.

The evaluation harness receives a fixed test case with known invoice and purchase-record values. Its checks can verify that the agent identifies a deliberately mismatched quantity, cites the correct field values in its structured output, avoids claiming approval, and produces the expected recommendation status. If the case does not provide enough evidence to decide, the expected behavior can be a request for review rather than a guessed conclusion.

Here is a compact execution flow. It shows a single allowed read and a controlled stop when that read fails. The evaluation harness is deliberately outside the operational path: it assesses a recorded test attempt after execution rather than authorizing the invoice action.

flowchart TD
%%{init: {"flowchart": {"htmlLabels": false}}}%%
accTitle: Invoice review execution and separate evaluation
accDescr: The executor validates a task, reads records, handles read failure, records a recommendation, and then an evaluator grades the evidence.
    A["Receive task"] --> B["Validate inputs"]
    B --> C["Read records"]
    C -->|"Read fails"| D["Stop and record failure"]
    C -->|"Read succeeds"| E["Record recommendation"]
    E --> F["Evaluate test evidence"]

The branch matters. A failed read leads to a recorded stop, not a fabricated result or a hidden retry. A successful read permits a recommendation to be recorded, but does not mean payment has been approved. The final evaluation step is meaningful for a controlled test case; in a live workflow, the evaluator is not a substitute for the organization’s decision process.

A practical execution record might contain these fields: task_id, case_id when applicable, input_revision, run_status, read_status, observed_invoice_quantity, observed_order_quantity, recommendation, write_status, failure_code, and finished_at. The production schema should use the application’s actual field types and data-handling rules. The field names here are illustrative, not a claim about a product API.

An evaluation record could separately contain evaluation_id, task_id, test_case_id, candidate_revision, criterion_results, evidence_references, grader_status, and overall_disposition. Keeping separate records allows the team to ask whether execution completed and whether the test met its acceptance criteria without collapsing those into one ambiguous status.

For a small test set, the case might include three scenarios: matching quantities, a quantity mismatch, and a missing purchase record. The expected outcomes differ. The first should produce a no-discrepancy recommendation; the second should identify the mismatch; the third should stop or request review because the evidence is incomplete. These are not three interchangeable examples to inflate coverage: each tests a distinct decision boundary.

The evaluator should check observable properties, not reward the agent for using a particular sequence of words. For the mismatch case, it can compare the cited quantities with the test fixture, verify the classification, and ensure the output does not state that payment was approved. If a free-form explanation is also assessed, its rubric should distinguish factual support from writing quality.

A run that gives the correct recommendation but fails to store it is an execution failure, even if the model’s response text looks right. A run that stores a recommendation but names the wrong quantity is an outcome failure. An evaluation setup that cannot inspect the stored record has an evidence gap. Those distinctions help the team choose the right fix: repair persistence, improve task behavior, or improve observation in the test.

Operational failure analysis: where the boundary is tested

Consider a timeline in which the harness submits a recommendation write and then loses its connection before receiving a response. The executor cannot infer from the timeout alone whether the remote system saved the record. Marking the action as definitely failed and immediately repeating it may create a duplicate. Marking it successful without checking may hide a missing recommendation.

A safer design records the write as uncertain, retains the request identifier and the exact intended payload, and uses an available read or reconciliation path before deciding whether to resume. If there is no reliable way to establish the outcome, the run should remain unresolved and surface that limitation. The design must not promise exactly-once external effects simply because it has retries or a unique local task identifier.

The evaluation harness should preserve the uncertainty too. If the expected evidence is absent because the write result could not be confirmed, the case may be ungradable rather than a clean agent failure. If the record exists but contains the wrong recommendation, that is different evidence and should be graded accordingly. The grading logic should not convert transport ambiguity into a confident quality claim.

Another common failure occurs when a long task resumes from a summary that omits what has already been confirmed. The resumed executor may repeat a read or overwrite a prior recommendation. A progress artifact should identify completed steps, their observed results, pending work, and any unresolved action. For continuity across context windows, progress artifacts and structured handoffs can help, but they do not make completion certain. The evaluation should separately test whether the final result meets requirements after the resume, not assume that successful resumption proves quality.

A third failure is a grader that treats the agent’s own assertion as evidence. Suppose the output says, “I found a five-unit discrepancy,” but the fixture differs by two. A transcript-only grade could reward a confident explanation. A result-oriented check compares the claim with the controlled values and marks the discrepancy. Where the evidence is not accessible to the grader, the report should say that verification was unavailable rather than claim a verified pass.

A fourth failure is a mismatch between the case and the live contract. A test may expect a field that the current application no longer records, or use sample data that does not reflect a relevant boundary condition. The evaluator can then report a regression that is actually a stale test, or miss a defect because its cases are too narrow. Test ownership includes reviewing assumptions and updating cases when requirements change; it does not mean editing expected results until every candidate passes.

Operational traces and evaluation reports should therefore be related by stable identifiers but retain their own meaning. A useful case report can point to the execution record, the relevant evidence, the candidate revision, and the criterion that failed. An execution record can point back to the originating test case when it was part of a controlled run. Neither record should be used as a substitute for the other.

If a change affects tool inputs or outputs, an evaluation may detect behavioral effects, while a contract test can check the tool boundary itself. The distinct treatment of that boundary is covered in adding contract tests for agent tools; this comparison does not attempt to reproduce it. Likewise, when a live task must wait for a person, the execution flow needs a safe pause and resumption design. See human approval in an agent harness for that separate concern.

Counterexample: a green evaluation and a bad production run

Imagine an invoice-review agent that passes every case in a small offline evaluation. The cases use complete records, stable values, and successful reads. The evaluator verifies the right discrepancy classifications and confirms that the response avoids claiming approval. That is useful evidence about the tested behavior, not proof that production execution is safe under every condition.

In production, a record lookup may return stale data, an external service may time out, or the task may resume after its state has changed. If the executor treats missing data as zero, retries an uncertain write without reconciliation, or resumes from an outdated snapshot, the resulting operation can still be wrong. The evaluation harness did not fail merely because this happened; it answered a narrower question about the cases it ran.

Now reverse the example. An execution harness may safely handle timeouts, preserve state, and stop on missing records. That is strong operational design. It does not establish that the agent consistently identifies the correct discrepancy or provides a useful recommendation. Safe failure is valuable, but it is not a quality grade for successful outputs.

The practical lesson is not to make every evaluation reproduce all of production. It is to state what each test covers, observe meaningful outcomes, and use operational data to identify cases the current evaluation misses. When a new failure reveals a real requirement, add a representative case and decide whether the executor, the agent behavior, or the grading method needs to change.

Decision worksheet: what belongs in which harness?

Use the following questions during design review. The point is to assign each responsibility to the component that can answer it with evidence, not to require separate teams or services.

  • Does this component decide what happens next in a live task? If it selects a tool, validates an action, handles a timeout, records progress, or determines whether to stop or resume, it belongs to execution. Write down the relevant state transition and its failure behavior.
  • Does this component decide whether a result met a stated criterion? If it compares a result with expected properties, applies a rubric, or reports performance across cases, it belongs to evaluation. Name the criterion and the evidence needed to judge it.
  • Can the result be independently observed? Identify the record, tool response, or other evidence that supports the claim. If the only evidence is the agent saying it succeeded, label the conclusion as unverified.
  • Are operational errors distinct from quality failures? Keep dependency failures, invalid test setup, ungradable cases, and incorrect task results distinguishable. Define what each status means before aggregating results.
  • Can an uncertain external effect be reconciled? Specify what the executor does after a timeout, including whether it can check the remote state. If it cannot establish the outcome, preserve uncertainty and choose a safe stopping path rather than claiming certainty.
  • Can the evaluation be interpreted later? Preserve the test input or revision, candidate identifier, criteria, evidence references, and per-case findings. A score without its basis is not a useful decision record.
  • Is a passing result being used for a claim beyond its coverage? State the tested scenario boundaries. Do not use a small set of controlled cases to imply guarantees about unobserved inputs or production conditions.

A printable summary for a design review can be kept to four lines: Executor owns: actions, state, limits, and operational failures. Evaluator owns: cases, evidence, criteria, and grades. Shared link: stable task and test identifiers. Unresolved boundary: any claim that neither component can verify. If the last line is nonempty, describe the uncertainty instead of hiding it inside a green status.

Verdict: choose by the decision you need to make

Build or improve an execution harness when the immediate problem is unreliable task progression: tools are called inconsistently, failures are difficult to recover from, state is lost between steps, or the system cannot distinguish confirmed work from uncertain effects. Start by defining the task lifecycle and the smallest set of state fields needed to resume or stop honestly.

Build or improve an evaluation harness when the immediate problem is uncertainty about behavior quality: teams disagree about whether an output is good, releases are judged from anecdotes, or a change appears to help without repeatable evidence. Start with a decision that the evaluation must inform, then define representative cases and observable criteria. Do not begin with a single aggregate score unless you can explain what it measures.

Build both when an agent performs consequential multi-step work and the team must make repeatable decisions about its quality. Keep run records and grading records distinct, connect them with identifiers, and inspect the gap between a safe execution and a correct outcome. A combined codebase can still have separate responsibilities; two separately deployed systems can still be confused if their statuses are blended.

If the team is deciding between adopting an agent framework and building orchestration itself, that is an adjacent implementation choice, not the same comparison. When the Claude Agent SDK may suit an agent build can inform that discussion. Whichever implementation route is selected, the team still needs to decide how execution state, evidence, and evaluation criteria fit together.

For an implementation review that needs to connect these design decisions to a working agent system, AI agent engineering is an appropriate next step. The core decision remains yours: assign live operation to the execution harness, assign performance judgment to the evaluation harness, and make every claim traceable to the evidence that supports it.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call