An AI model can generate text, select a proposed action, or request a tool. That alone does not explain how a working agent receives information, executes an action, handles an error, or decides when to stop. The software coordinating those steps is commonly called an AI harness.
A harness matters because the same model can behave very differently depending on the context it receives, the tools it can request, the rules that govern execution, and the state carried between steps. Evaluating an agent therefore means evaluating the surrounding system as well as the model.
This explainer builds from the basic definition to the execution loop, then examines context, tools, state, controls, and practical ways to assess examples such as Codex, Claude Code, and Deep Agents. It stays focused on how a harness works; separate questions about runtimes, long-running workflow architecture, and run-manifest design are linked where they become relevant.
What an AI harness is—and what it is not
An AI harness is the software that organizes a model’s participation in a task. It supplies task-specific input, makes relevant context available, exposes permitted tools, processes tool requests, tracks what has happened, and applies rules for continuing or stopping. In short, the model proposes; the harness coordinates what happens around that proposal.
A model is the component that processes its input and produces an output. A tool is an operation the surrounding system can make available, such as reading a particular record or running a bounded calculation. A harness connects the model’s output to those tools and to the rest of the application. These are useful conceptual distinctions even when a product packages several components together.
The harness is not simply a long prompt. Instructions can shape what the model is asked to do, but the execution system determines which tool requests are actually run, what results return, what happens after an error, and whether the task can be resumed. A prompt that says “do not send the report” is not an adequate substitute for a system that has no sending action available in that workflow.
Nor is a harness the same thing as a model. If two systems use the same model but provide different context, tools, and stopping rules, they are different working systems. Conversely, a team can change the selected model without necessarily changing every part of the surrounding execution design. Keeping those concepts separate helps engineering leaders identify what they are actually choosing or building.
The term is broad, so product boundaries do not always line up neatly with the concept. A product may provide a harness as a managed service, a software package may provide some of the building blocks, or an organization may write its own coordination code. For a more detailed distinction among those architectural categories, see Agent Harness vs Agent Framework vs Runtime: A Practical Guide. Here, the practical question is narrower: what work must the harness perform for an agent task to behave predictably?
The execution loop, one decision at a time
A useful way to understand a harness is to follow a single task through its execution loop. The loop starts with a bounded request, supplies information needed for the next decision, receives a model output, and handles that output according to its type. If the output requests a tool, the harness checks and runs the request; if it is a final response, the harness may return it. Tool results then feed the next decision or a defined stop path.
flowchart TD
accTitle: A controlled AI harness execution loop
accDescr: The harness receives a task and limits, loads scoped context, and asks a model for an output. A tool request is checked and run, then its result is inspected and the loop continues or stops. A final response returns to the caller, while a denied request or error follows a stop path.
A["Receive task and limits"] --> B["Load scoped context"]
B --> C["Model proposes next output"]
C -->|"Tool request"| D["Check and run tool"]
C -->|"Final response"| F["Return result"]
D -->|"Result"| E["Inspect result"]
E -->|"Continue"| C
E -->|"Done"| F
D -->|"Denied or error"| G["Stop and report"]
E -->|"Unsafe or unusable"| G
The diagram represents a control pattern, not a promise about any named product. The important branch is between a proposed tool action and its execution. A request should not become an action merely because it appears in a model response. The harness needs to interpret the request, apply the workflow’s rules, and either run it or take a defined refusal path.
The second important decision comes after a tool returns. “The tool ran” and “the task succeeded” are different statements. A query can complete but return no matching records; a file operation can finish but affect the wrong destination; a service can return an error that requires a different next step. The harness should make the result available for inspection rather than treating every completed call as proof of business success.
The loop also needs a way to end. A final answer, a deliberate stop, a request for human attention, and a checkpoint for later continuation are not interchangeable outcomes. Decide which end states matter for the particular workflow and make them visible to the system that called the harness. For workflows that must persist through interruptions or take much longer than one interaction, see Durable Execution for Agents That Run for Hours; this article does not try to design that broader execution architecture.
Context: give the model what this step needs
Context is the information supplied or made available to the model for a particular decision. It can include the user’s request, task-specific instructions, selected records, prior tool results, and relevant state. It is not necessarily the whole history of a system or every file a team owns.
A practical context design begins with a question: what facts must be available for the next action to be sensible? For a report-drafting task, that might mean the reporting period, approved source names, the definitions of key measures, and the data returned by a specific query. It may not require unrelated team conversations, every past report, or credentials for systems the task does not need.
Too little context can make a run ambiguous. If a request says “compare this week with the prior period” but the reporting period and comparison rule are missing, the model must either guess or ask. Too much context can obscure the relevant facts, increase the chance of mixing unrelated information, and make it harder to tell why a decision was made. The goal is not maximum context; it is adequate, selected context.
Separate instructions from evidence. “Use the approved reporting period” is a rule. “The selected period is 1–7 September” is task data. “The query returned 1,248 records” is a tool result. Keeping those kinds of information distinguishable makes it easier to detect a missing input or a conflict. It also helps a reviewer see whether a final statement came from a tool result, an instruction, or an unsupported assumption.
Context should also be appropriate to the current step. A model deciding which permitted data query to request may need the available field names and the user’s reporting objective. After the query, the result and the definition of the measure may be more useful. Reusing one giant context package for every stage is convenient to implement but does not automatically make it clearer or safer.
When work spans sessions, a transcript is not necessarily a good substitute for a concise record of progress. A later run needs enough information to understand what has been completed and what still needs attention, without treating an old conversation as proof that an action took place. For a related discussion of durable, file-backed context, see Memory Patterns for Multi-Session Agents: File-Backed Context.
Tools: turn a proposal into a bounded action
A tool is a callable operation exposed to an agent workflow. Examples might include retrieving a specific record, calculating a total from supplied values, or saving a draft to a designated location. These examples describe common design choices, not capabilities guaranteed by any particular product. A tool’s real boundary comes from what its implementation accepts and does.
A useful tool definition is narrower than a general instruction such as “manage the reporting system.” It identifies the operation, the inputs it accepts, the result it returns, and the consequences of running it. A hypothetical read_weekly_metrics operation might accept a start date and end date, return a structured result with named measures, and have no write effect. A separate draft-saving action would have a different input and consequence. Keeping the operations distinct makes it easier to reason about what each request can do.
Validation belongs between the proposed request and the underlying operation. The harness can check that required fields are present, values use an expected format, and the requested operation is available for this task. It can reject a date range outside the run’s allowed period, for example, instead of passing an ambiguous request directly to another system. The precise checks depend on the operation; validation should be designed around its actual effects.
Tool results need a consistent shape. If one result says “success” while another uses an empty object to represent a failed lookup, downstream logic becomes difficult to interpret. A useful result can distinguish a completed operation, a domain-level result such as “no records found,” and an execution error. The model may help interpret those results, but the harness should preserve the underlying distinction rather than flattening every outcome into a sentence.
Tool access should be assessed by task, not by the number of tools available. A workflow that only drafts a report may need selected read operations and a draft destination; adding an unrelated action broadens the system without improving that task. Conversely, an absent operation may cause the model to produce instructions for a person rather than complete the action itself. The design decision is to expose the smallest useful set of operations, then test whether it is sufficient for the defined task.
State: remember what has happened
State is the information the harness retains about a run beyond the model’s current output. It can include the task status, a record of tool requests and results, a count of attempts, and a pointer to an outstanding review. State matters because a sequence of actions cannot be understood solely from the latest message if earlier steps affect what should happen next.
A simple state record might distinguish started, in_progress, blocked, and completed. It could also include the current step, the last tool outcome, and the reason for any stop. Those names are an illustrative design, not a standard that every system must adopt. The useful property is that a person or calling process can determine what the run believes has happened and what outcome it reached.
State must not confuse an intention with an effect. “The agent requested a draft save” is not the same as “the draft was saved.” “The model reported success” is not independent verification that an external system accepted the operation. When a tool returns a result, record that result and its meaning; when a business outcome needs an additional check, represent that check separately.
This distinction is especially important when an interruption happens near the boundary between a tool call and a state update. Suppose a workflow asks a tool to save a draft, but the process stops before recording the returned result. On resumption, blindly repeating the action could produce an unintended duplicate, while blindly marking it complete could misrepresent what happened. The correct recovery depends on the tool and available evidence; the harness should surface uncertainty rather than silently choose whichever answer is convenient.
For a small one-turn task, state may be limited to the current request and its result. For work that crosses sessions or needs durable recovery, the required design expands considerably. The distinction helps teams avoid both extremes: ignoring state in a multi-step workflow, or building a long-term persistence system when a bounded request does not need one. A dedicated design artifact can help formalize the run record; see Build an AI Harness Run Manifest: State, Tools and Stop Conditions.
Controls: define what can continue and what must stop
Controls are the rules the harness applies to the workflow. They may limit which tools can run, validate inputs, cap attempts, define a stop condition, or route a particular outcome to a person. Controls turn broad intent into decisions the execution system can apply at the moment they matter.
A retry rule is a useful example. If a read operation returns a temporary error, a workflow may allow a limited retry. If the same operation keeps returning the same error, continuing indefinitely is not a useful form of persistence. The harness should count attempts and turn the repeated failure into a defined outcome, such as stopping with a reason or returning control to an operator. The appropriate retry policy depends on the operation and consequences; it should not be inferred from a model’s willingness to try again.
A stop condition can be tied to the task’s meaning. The run might stop when it has assembled all required inputs, when an input is missing, when a permitted attempt limit is reached, or when a requested action falls outside the workflow’s scope. These are examples to adapt, not a universal list. A useful stop condition is explicit enough that a reviewer can tell why the system ended instead of having to infer it from a vague final message.
A harness can also distinguish between a decision it can make and one that needs a human. For instance, a missing reporting period could prompt a request for clarification; a read result with an unexpected shape could stop the run for inspection. The point is not to route every uncertainty to a person. It is to choose which uncertainties the system may resolve under the defined rules and which ones should interrupt automation.
Keep model confidence language separate from execution evidence. A fluent statement that a task is complete is still a model output. Completion should be based on the workflow’s relevant evidence: required operations ran, returned usable results, and any required outcome check passed. If that evidence is absent, the harness should be able to report an incomplete or unverified status rather than forcing a confident-sounding conclusion.
What Codex, Claude Code, and Deep Agents illustrate
The examples are helpful as reference points, but they do not imply that every product uses the same design or provides the same boundaries. Compare what is documented for the particular offering and version under consideration, then test it against your task. The word “harness” names a role in a system; it is not a guarantee that two systems with that label behave alike.
OpenAI’s Agents API overview describes a hosted Codex harness and distinguishes the harness from the selected model. (OpenAI Agents API overview) For an evaluation, ask which behavior belongs to the hosted harness, which depends on model selection, and which choices remain with your application. That separation is useful even when an offering presents the pieces together, because it clarifies what you are assessing when the task or model changes.
Anthropic’s guidance for long-running agents discusses progress artifacts and structured handoffs across context windows, while noting that these do not guarantee completion. (Anthropic guidance on long-running agents) For a team considering Claude Code or the Claude Agent SDK, this is a reminder to inspect how progress is represented and what a later session would need to continue. It is not a reason to assume that every handoff will succeed or that a progress note proves a task was completed. A more implementation-specific comparison is available in Claude Agent SDK: When It Beats Your Own Orchestration.
Deep Agents describes a harness that combines context and tool execution; planning is optional. (Deep Agents overview) That description invites a concrete evaluation: which context and tool behaviors does a particular workflow need, and does planning solve a real problem for it? Optional planning should not be treated as a requirement for every task, just as the presence of a capability is not proof that it fits a particular operating process.
Across these examples, avoid comparing labels alone. Write down the behavior you need: where the task input comes from, what actions can be requested, how tool results are inspected, what state is available after a stop, and how the system represents incomplete work. Then compare the implementation against those requirements. This produces a more useful decision than asking whether a product is “an agent” in the abstract.
Worked hypothetical: prepare, check, and route a weekly report
Consider a hypothetical operations team that wants an agent to gather a small set of weekly metrics and prepare a report draft. The task is deliberately bounded: retrieve figures for a named reporting period, calculate one agreed comparison, write a draft to a designated location, and return its status. The example is illustrative; it does not describe a PADISO deployment or a tested product workflow.
First, define the request fields. The run needs a reporting period, a list of approved measures, a destination for the draft, and a rule for the comparison period. If any required field is missing, the harness should return a request for clarification rather than letting the model invent a date range or select a destination. This turns an ambiguous natural-language task into inputs that the rest of the loop can inspect.
Next, provide only the context needed for that run. The model might receive the measure definitions, the supplied period, and the permitted operations. It does not need to be handed a broad instruction to “manage reporting.” Separating the task’s definitions from values retrieved later also makes it possible to see whether a result came from the source operation or from the model’s own wording.
Then expose a read operation for the approved measures and a separate operation for saving a draft. The harness checks each proposed request against the task fields and operation boundaries. If the model asks for a measure that is not on the approved list, the harness rejects that request and records why. If the read operation returns an error, the error becomes an execution outcome; it is not converted into a zero value to keep the report moving.
After retrieval, inspect the result. The harness can check that required measures are present and that the reporting period matches the requested period. If one measure is absent, the workflow should mark the input as incomplete and take its chosen stop or review path. It should not let a plausible-looking draft conceal a missing value. If all required results are present, the comparison calculation can be performed or checked according to the system’s defined approach.
When the draft is prepared, save it only to the designated destination and retain the operation’s returned status. If the save reports an error, the run should not mark the draft as saved. If the process stops before the result is recorded, the state should show that the outcome is uncertain, not quietly claim completion. How to resolve an uncertain external action depends on the operation’s behavior and available records; a harness should not pretend that repeating an action is always harmless.
Finally, return a concise status that distinguishes completed, blocked, and incomplete work. A completed report can identify the period and the draft location, subject to whatever evidence the workflow actually received. A blocked run can identify the missing input or failed operation. That explicit result gives an operator a next step without asking them to reconstruct the entire model conversation.
Failure analysis: follow the consequence, not just the error message
A useful failure review traces what happened before and after an error. If a tool request is rejected, did the harness prevent the operation from running? If a tool returned an error, did the run stop, retry within its rule, or incorrectly continue with missing data? If an interruption occurred after a side effect, does the recorded state distinguish a confirmed result from an uncertain one? These questions test the execution design rather than the persuasiveness of a final response.
Consider a malformed input. The model proposes a request with a missing date. A weak design may pass it to a downstream operation and accept whatever default the operation applies. A stronger design validates the field before execution and returns a specific input error. The key difference is not a more elaborate prompt; it is where the boundary sits between a proposed action and an actual operation.
Now consider an empty result. An empty result can mean there were no records, the query was too narrow, or the operation failed to retrieve data, depending on the tool’s contract. If the harness collapses these cases into “nothing found,” a later draft may present the wrong interpretation. Define the result meanings, preserve them in state, and ensure the next step uses the correct branch.
Repeated failure exposes another weakness: a model may continue to request variations without making progress. The harness should be able to detect a relevant repeated condition and stop according to a bounded rule. That does not require treating every repeated request as malicious or every retry as wrong; it requires deciding in advance what persistence means for this operation and what evidence would justify another attempt.
A final failure pattern is completion by assertion. The model says “done,” but a required tool result is missing or the result does not satisfy the task. In that case, the harness should report what it can establish and identify what remains unverified. This protects business operators from mistaking a polished narrative for a confirmed outcome, while giving engineering teams a concrete signal to investigate.
A practical harness evaluation worksheet
Use this worksheet to evaluate a specific task, not to award a product a general score. Fill it in with the people who own the workflow and the people who will operate the software. Each answer should refer to observable behavior, such as an input field, a tool result, or a stop status, rather than broad statements like “the agent is reliable.”
| Design question | Record a concrete answer | Acceptance evidence to look for |
|---|---|---|
| Task boundary | What exact outcome is the run meant to produce, and what is outside its scope? | A sample request can be accepted or rejected without guessing the intended task. |
| Required context | Which instructions, records, and prior results are needed for each decision? | A reviewer can identify the source and purpose of each important input. |
| Available tools | What operations may be requested, and what does each accept and return? | An invalid or out-of-scope request is denied before the operation runs. |
| Result interpretation | What is the difference between a valid empty result, an error, and a usable result? | The workflow follows distinct paths for outcomes that have different meanings. |
| State | What should be retained after each significant step or stop? | A later inspection can distinguish a proposed action from a confirmed result. |
| Retry and stop rules | Which failures allow another attempt, and what condition ends the run? | Repeated failure reaches a defined status instead of an open-ended loop. |
| Completion evidence | What observable evidence is required before the task is marked complete? | A model’s unsupported completion statement cannot substitute for missing evidence. |
| Operator handoff | What should a person see when the run is blocked or uncertain? | The handoff names the issue and next decision without implying that incomplete work succeeded. |
A compact printable summary can be derived from the worksheet:
- State the task outcome and its boundary in one or two specific sentences.
- List required inputs and say what happens when each is missing.
- Name each available operation, its accepted fields, and the meaning of its result.
- Describe the state retained after a tool request, result, failure, and stop.
- Set a task-specific retry rule and identify the condition that ends the run.
- Define the evidence needed for “complete” separately from a model’s claim.
- Run at least one representative success path and the failure paths that could change the business outcome.
The worksheet is deliberately about execution behavior. It does not settle how a team should structure every component, store all long-term workflow state, or govern every business process. Those are related design questions, and they deserve their own treatment rather than being compressed into a generic definition of “harness.”
Choosing what to adopt or build
Start by describing one real task at the level of inputs, permitted actions, results, and stop outcomes. If an available harness covers those needs, evaluate the behavior against the worksheet and the failure cases that matter to the workflow. If a key boundary is missing, determine whether it belongs in the product configuration, an application layer, or a different implementation approach before adding more tools or broader access.
Keep the first evaluation small enough to inspect. A task with one or two bounded operations can reveal whether context is supplied appropriately, whether tool results are interpretable, and whether incomplete work remains visible. Expanding to more steps before those basics are clear makes it harder to identify which part of the system caused an unexpected result.
A custom harness may be justified when the workflow needs control over a particular tool boundary, result interpretation, or stop condition. It also creates responsibility for maintaining that coordination logic and testing it as the workflow changes. An existing offering may reduce the amount of implementation a team needs to own, but the team still needs to understand what it does for the task and what outcomes it cannot establish on its own.
The decision is therefore not simply “use an agent” versus “do not use an agent.” It is whether a model-plus-harness design can perform a bounded task with useful evidence and understandable failure paths, compared with the non-agent process it would replace or assist. If your team needs help designing or implementing such a workflow, explore AI agent engineering.
An AI harness is best understood through the decisions it makes around a model: what information is supplied, what actions can be requested, how results are interpreted, what state is retained, and when execution stops. Those decisions—not the label attached to a product—determine whether an agent workflow is understandable enough to evaluate and controlled enough to operate.