SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Opinion 19 mins

When Should You Build Your Own Agent Harness?

When Should You Build Your Own Agent Harness?. Practical examples, tradeoffs and implementation guidance for technology leaders.

The PADISO Team ·

My thesis: build only when control is worth owning

My position is deliberately conservative: build your own agent harness only when a material workflow requirement cannot be met with a suitable existing harness, and when your organization is prepared to operate the difference. A bespoke execution loop is not a badge of technical maturity. It is a product and operations decision that adds a component your team must understand, test, observe, change and recover when something goes wrong.

An agent harness is the runtime around an agent’s model calls and tools: it assembles the information needed for a run, decides how execution proceeds, invokes approved actions, and handles the result. It is distinct from the business workflow itself. It is also distinct from the model. A team can choose a model and tools while using a framework to coordinate them; building a harness means taking responsibility for some or all of that coordination logic.

That distinction matters because customization is easy to value and operational ownership is easy to discount. A tailored loop can make a real difference when it enforces a business-specific sequence, produces a required audit record, or supports recovery from partial work. But writing a loop that calls a model and dispatches tools is the beginning of the obligation, not the end of the project.

I would not make the decision by asking whether a team can build a harness. Many capable teams can. I would ask whether the expected benefit is specific, important, and durable enough to justify owning runtime behavior that an existing option would otherwise provide. If the answer is still vague after the workflow has been written down, the decision is not ready.

This is an argument about build versus buy, including the option to adopt and extend an existing framework. It is not an argument that one framework fits every workflow. For the surrounding system concepts, see The Anatomy of an AI Agent: What Each Component Does. The decision here is narrower: who should own execution behavior, and why?

First separate the workflow from the harness

Teams often reach for custom orchestration before they have established which behavior actually needs customization. Start by describing the business workflow in ordinary operational language. Name the request, the inputs that must be present, the decisions being made, the actions that can change an external system, and the evidence that tells an operator the work is complete. Do not start with a diagram of model calls.

Then mark which parts are genuinely agentic. A workflow may contain a model-assisted classification step followed by deterministic validation and a human or system action. That does not automatically call for a general-purpose agent loop. Conversely, a workflow that requires iterative tool use may still fit an existing harness if its execution boundaries, context needs and failure behavior can be expressed safely there.

A harness becomes a meaningful build-versus-buy choice when it owns runtime decisions: what information to load, how a run may continue, which tool actions are permitted, how outcomes are checked, and what happens when a step fails. The more your custom code determines these behaviors, the more your organization owns the consequences. A thin adapter around an existing framework and a homegrown execution runtime are not equivalent investments, even if both are described internally as “our harness.”

The supplier’s packaging is not the decision criterion either. A framework can be adopted without surrendering every application-specific policy. A custom component can still depend heavily on third-party libraries and model interfaces. The useful question is where the boundary sits: which behaviors do you need to control directly, which can you configure, and which should remain the responsibility of a maintained dependency?

For a focused comparison of one SDK-based approach with custom orchestration, read Claude Agent SDK: When It Beats Your Own Orchestration. Treat that as a separate implementation decision. This article’s recommendation is to evaluate the operational boundary before selecting a particular technology.

The four conditions that can justify a custom harness

A custom harness has a credible case when a specific requirement crosses a threshold that configuration, extension or a different existing option cannot meet. The requirement should be concrete enough to test. “We want more flexibility” is not. “Every external action must be checked against an immutable, run-specific authorization record before dispatch” is a testable design need, although the team must still establish whether an existing option can meet it.

The first condition is workflow-specific control. The workflow may require a strict sequence of state transitions, a specialized context assembly policy, or a narrowly defined execution boundary that cannot be represented clearly in an available harness. Custom code is more defensible when that control protects a real business invariant, not merely when it makes an architecture diagram look cleaner.

The second is a durable integration or operational constraint. For example, an organization might have an established internal job system, trace format or recovery process that an agent runtime must participate in. That is not automatically a reason to build the whole runtime. It may justify a small adapter or extension. The team should identify the smallest boundary that must be custom rather than allowing one integration requirement to expand into a complete reinvention.

The third is a meaningful failure or audit requirement. If the business needs to know exactly which request, inputs, tool intent, authorization decision and result belong to a run, the design must make those relationships explicit. A bespoke harness may help if it gives the team the necessary control and an existing option cannot. But custom code does not create reliable records by itself; the data model, persistence, operational access and retention decisions still need owners.

The fourth is a capable, funded operating team. A harness has to be maintained after the original builders move to other work. Someone must own upgrades, incident response, test coverage, deployment, logging, recovery and deprecation. If no team is assigned those responsibilities, the case for custom control weakens sharply. A component without an owner is not a strategic asset; it is an unattended dependency that happens to live in your repository.

Use a scorecard that counts ownership, not just features

Score each criterion from 0 to 2. A score of 0 means the requirement is absent or an existing option appears sufficient. A score of 1 means the requirement exists but its importance or fit is uncertain. A score of 2 means there is a documented need, an identifiable gap in existing options, and an acceptance test that could demonstrate the gap. Record the evidence for each score; do not treat the sum as a mechanical verdict.

CriterionWhat earns a high scoreWhat to write down
Workflow fitRequired execution behavior cannot be expressed adequately through an available optionThe exact invariant and a failing example
Control boundaryThe team needs direct control over a consequential runtime decisionThe decision, its owner and its permitted values
Failure handlingRequired recovery behavior is not adequately supported by an available designFailure state, operator action and recovery test
Integration fitA durable internal constraint needs more than a small adapterThe integration boundary and why extension is insufficient
ObservabilityOperators need run-level evidence the proposed design can provide and useRequired fields, retention owner and incident question answered
Maintenance capacityA named team can operate and upgrade the componentPrimary owner, backup owner and recurring work allocation
Exit and change costThe team can change dependencies or retire the workflow without a hidden rewriteMigration seam, stored state and decommission plan

The scores are a conversation prompt. A total cannot rescue a critical gap. A high score for workflow fit does not compensate for no operational owner. A promising custom prototype does not prove that an existing framework fails a production requirement. Conversely, a low feature gap does not prove that a framework is suitable if the team cannot operate the surrounding workflow safely.

Add two explicit columns to your internal version of the table: evidence and owner. Evidence might be a written workflow, a review of candidate options, a failure-injection plan, or an agreed acceptance criterion. The owner is the person or team accountable for verifying it. If either column is blank, mark the score uncertain rather than filling the gap with optimism.

As a practical interpretation, a low or uncertain total points toward adopting an existing option or running a bounded evaluation. Strong workflow and failure-control evidence can justify a custom component only if maintenance capacity is also credible. A mixed result often supports a hybrid: use a maintained harness for common runtime behavior and write a narrow adapter or policy layer for the requirement that is genuinely distinctive.

A hypothetical worked decision: processing supplier changes

Consider a clearly hypothetical mid-market company that receives supplier-change requests and wants to automate the preparation of updates. The workflow has three source inputs: a submitted request, an internal supplier record, and a change proposal. The system may draft a proposed update, but the business requires a designated employee to verify it before a change is sent to the supplier system. This scenario is illustrative, not a claim about a real organization or deployment.

Suppose the team’s first proposal is to build a custom harness because it wants “full control.” That phrase is too broad to justify a runtime. The team should instead turn the need into specific requirements: every run must have a stable run identifier; the submitted request must be linked to the supplier record used; the action payload must be visible before dispatch; and the system must distinguish a generated proposal from a confirmed business result. Those are requirements to assess, not presumed properties of any product.

The team could compare a suitable existing harness with a small custom layer against those requirements. It might discover that the existing option can support the workflow if the application supplies its own validation and persistence. In that case, a custom adapter for record lookup and payload checks could be the narrower solution. If the option cannot make the required execution boundary explicit, and that limitation is demonstrated through a reproducible acceptance test, a custom harness becomes more plausible.

The hypothetical scorecard might read as follows. Workflow fit receives 2 only if the team documents a requirement the candidate option cannot represent. Failure handling receives 2 only if the proposed custom design demonstrates how it prevents an ambiguous or partial result from being reported as complete. Maintenance capacity receives 0 if nobody has accepted responsibility for upgrades and incidents. Under those conditions, even two strong technical scores do not justify proceeding to production: the missing owner is a decision blocker.

The critical design question is not whether a model can draft the update. It is how the application establishes that a permitted action was prepared, authorized and completed. Model output is not proof that a downstream business change occurred. The system needs an independent way to inspect the resulting state or otherwise record the action’s outcome. The exact verification method depends on the external system and must be designed for that environment; it should not be invented from the model’s response.

A rough, explicitly illustrative calculation can make hidden effort visible. Assume, solely for planning discussion, that a custom implementation needs 10 engineering days initially and 2 engineering days each month for maintenance and operational work. Assume an existing option requires 4 initial days and 1 day per month for integration and upkeep. Across a 12-month planning horizon, those assumptions produce 34 days for the custom path and 16 for the existing-option path. These are invented planning inputs, not cost estimates, benchmarks or expected outcomes. The point is to make recurring ownership visible and replace the assumed values with the team’s own estimates.

That calculation still omits the cost of a serious incident, migration, staffing changes and evaluation. It also omits any measurable business value from a capability the existing option cannot provide. The right decision is not the one with the smallest invented total. It is the one whose material requirements and ongoing obligations are explicit enough for accountable leaders to accept.

Draw the execution boundary before writing the runtime

A useful artifact is a small execution diagram that makes the stop conditions visible. The diagram below is illustrative. It describes a conservative pattern for a workflow with an external action; it is not a product capability claim or a complete deployment architecture.

flowchart TD
  accTitle: Controlled agent execution path
  accDescr: A request is validated, context is loaded, an action is executed only after authorization, the result is checked, and failures stop for recovery.
  A["Request received"] --> B["Validate and persist"]
  B --> C["Load bounded context"]
  C --> D["Prepare exact action payload"]
  D --> E["Authorize, then execute"]
  E --> F["Verify external result"]
  F --> G["Record result or stop for recovery"]

The sequence separates preparation from execution. A generated proposal is not yet an authorized action. When approval is required, it must occur before execution, refer to the exact payload that will be sent, and expire so an old approval cannot silently authorize a later, changed action. If the payload changes after authorization, the old approval no longer describes the action being taken. This is a design recommendation for a consequential workflow, not a claim that a particular framework enforces it.

The result check is also separate from the model’s account of what happened. A model response may explain the intended action, but it cannot establish that an external system accepted or persisted it. The application needs an appropriate confirmation signal, or it must report that the outcome remains unknown. The diagram’s final node deliberately allows a stop for recovery rather than requiring the runtime to pretend every run has a clean answer.

Failures can occur at each boundary. Validation may reject incomplete input. Context loading may fail or return a record that does not match the request. Payload preparation may produce a value that fails business validation. Authorization may be absent, expired or tied to a different payload. The external action may time out after the receiving system has acted, leaving the sender uncertain whether retrying is safe. Verification may be unavailable even though the action was accepted.

The correct response to uncertainty is not automatically to retry. The system should retain the run identifier, the action intent, the exact payload version, the authorization reference, the attempted operation and whatever result evidence exists. Then it should stop automatic progression when it cannot establish whether a repeat could create a second effect. Exactly-once external effects should never be promised on the strength of a harness design alone; the external system’s behavior and the integration protocol matter.

For a practical data artifact, define a run record before implementing the loop. One possible illustrative schema is: run_id; workflow_version; input_reference; context_references; state; action_name; payload_version; payload_digest; approval_reference; attempt_reference; external_result_reference; failure_code; created_at; and updated_at. Those field names are a proposed design, not a prescribed standard. Store only information the workflow needs, and decide how each field is protected, retained and made available to operators.

A state model should be just as explicit: received, validated, prepared, awaiting_authorization, authorized, executing, result_unknown, verified, rejected or recovery_required. The important property is not the vocabulary. It is that an operator can distinguish an action that was never sent from one that may have succeeded but whose result has not been verified. That distinction drives safe recovery.

Count operational ownership as part of the architecture

A build-versus-buy comparison is incomplete unless it names who owns the runtime after release. For each responsibility, name a team rather than assuming the original author will remain available. The owner does not need to be one person, but accountability should not dissolve between application engineering, platform engineering and operations.

Runtime and dependency changes: someone must decide when to update the harness and its dependencies, review behavior changes, and maintain a path for rollback or migration. A dependency update is not just a package change if it can alter execution behavior that the business workflow depends on. The owner should know which acceptance tests must pass before an update is promoted.

Run visibility: operators need enough information to answer practical questions: which request is this; which workflow version ran; which external action was attempted; what outcome is known; and what should happen next? A dashboard is not a substitute for meaningful run records. Conversely, detailed records without an operational procedure can leave staff unable to act. Define who can inspect them and who is responsible for resolving ambiguous states.

Recovery and incident response: specify who handles stuck runs, failed validations, unavailable dependencies, duplicate submissions and uncertain external outcomes. Recovery instructions should describe the evidence needed before an operator resumes, rejects or escalates a run. Do not design an automatic replay rule around a hopeful assumption that a failed response means no external action occurred.

Testing and release: someone must maintain tests for the business invariants and integration boundaries that make the custom harness worthwhile. The purpose is not to prove that a model always behaves identically. It is to check that the application handles required inputs, invalid actions, failures and state transitions as intended. Tool contracts can change independently of the model loop; Tool Changes Can Break Agents: Add Contract Tests to CI explores that distinct testing concern.

Retirement and migration: a workflow may stop being useful, or the organization may choose a different runtime. The owner should know where run state lives, which interfaces are depended on, and how pending work would be completed or safely closed. A custom harness that cannot be changed or retired without losing operational history creates a continuing constraint. Planning an exit is part of the build decision, not an admission that the design will fail.

These obligations are not arguments against custom software. They are the price of deciding that custom control matters. If the team can identify owners, allocate time and define the operational evidence it needs, custom code may be a responsible choice. If not, choosing a maintained option can be a way to keep scarce engineering attention focused on the workflow rather than on a runtime the business did not ask to own.

The counterargument: frameworks can constrain the work

The strongest counterargument is that an existing harness may impose abstractions that fit the common case but make an unusual workflow awkward. A team might need a specific execution boundary, a specialized state model, or integration behavior that an available option does not expose in a maintainable way. Wrapping every decision in adapters can create complexity of its own. In that situation, a custom runtime may be easier to understand than a stack of workarounds.

That argument is valid when it is demonstrated against a real workflow and a defined acceptance test. It is weaker when the team has evaluated only one option, has not tried a narrow extension, or is optimizing for developer preference rather than a business requirement. A prototype can show that a custom loop is feasible; it does not show that its maintenance burden is acceptable or that the candidate framework is unsuitable.

There is also a legitimate concern about dependence on an external project. A team may want control over release timing or behavior. But owning source code is not the same as having sustainable control. If the team cannot maintain the component, test its behavior and manage changes, a custom implementation may replace visible dependency risk with less visible internal dependency risk.

A sound response is to test the smallest disputed capability. Write an acceptance case that includes normal execution, a failure, and recovery. Compare an existing option, a narrow extension and the proposed custom component against the same case. Record what each requires the application team to implement. If the difference is merely that one route feels more elegant, defer the decision. If the existing option cannot satisfy a material invariant without fragile workarounds, the evidence for custom control is stronger.

The relevant product fact should remain modest: LangChain’s DeepAgents documentation describes a harness that combines context and tool execution, with planning optional (DeepAgents overview). That is one example of how a harness can be packaged, not proof that it fits a particular workflow. Evaluate an option against the requirements and ownership model you have actually documented.

A counterexample: a one-off internal drafting assistant

Imagine a different, hypothetical case: an internal team wants a low-stakes assistant that drafts summaries from a small set of approved inputs. A person reviews every draft, and the system does not take external actions. The team has no established requirement for a custom state machine, unusual execution policy or special recovery behavior. Its immediate uncertainty is whether the task is useful and whether the inputs produce acceptable drafts.

Building a bespoke harness first is difficult to defend here. The team has not identified a material gap that needs custom control. It risks spending time on runtime behavior before it has learned whether the workflow merits continued investment. A suitable existing option, a simple bounded prototype, or an ordinary application workflow may be enough to evaluate the task. The choice depends on the actual requirements, but custom orchestration has not earned its additional ownership burden.

The counterexample is not an argument that low-stakes systems need no engineering care. The team should still define the input boundary, review process, retention expectations and a way to stop the experiment. The point is that those responsibilities do not automatically require a new general-purpose harness. Scope the implementation to the unresolved decision, then revisit the runtime choice if the workflow acquires requirements that change the scorecard.

That is why I resist “we will need it later” as a standalone justification. A credible future need can be recorded as an option, with a trigger for reassessment. It should not be treated as a present requirement unless someone can say what event would make it real. Prematurely building for an undefined future workflow adds current maintenance without proving future value.

A printable decision worksheet

Use this worksheet in an architecture review. It is intended to fit into a decision record; it is not a downloadable template. Check a box only when the team can point to concrete evidence, not because the statement sounds desirable.

Workflow and requirement

  • We have written the workflow from request to verified business outcome, including the points where an external effect can occur.
  • We have named the specific runtime behavior that an existing option cannot meet, and separated it from preferences such as language, style or familiarity.
  • We have an acceptance case that demonstrates the gap, including at least one failure and the expected recovery state.
  • We have compared a suitable existing option, a narrow extension and a custom implementation against the same requirement.

Ownership and operation

  • A named team owns upgrades, release decisions, run visibility, incident response and eventual retirement.
  • Operators can distinguish work that was not attempted from work whose external result is uncertain.
  • The run record identifies the workflow version and connects inputs, action intent, authorization and outcome evidence.
  • Any authorization required for an external action occurs before execution, applies to the exact payload and expires.
  • The team has described how it will respond to an ambiguous timeout without assuming a repeated request is harmless.

Decision and review

  • The scorecard records evidence and an owner for every material criterion; uncertain claims are labeled uncertain.
  • The cost discussion includes continuing maintenance and incident work, not only initial implementation.
  • The design has a boundary that allows the team to replace an existing dependency or retire the custom component.
  • We have set a review trigger, such as a demonstrated workflow gap, a sustained operating burden, or a change in the business requirement.

Printable summary: Build a custom harness only when a documented, material workflow requirement cannot be met cleanly by an available option, a test demonstrates the gap, and an accountable team accepts the ongoing operational work. Otherwise, adopt a suitable option, extend it narrowly, or defer the runtime decision until the need is clearer.

My recommendation

My recommendation is to make custom ownership the exception that must be justified by workflow evidence, not the default reward for having an capable engineering team. Choose a custom harness when it buys control over a consequential requirement that alternatives cannot meet cleanly, and when the people who will operate it are part of the decision before implementation begins.

Choose an existing option or a narrow extension when the important requirement can be satisfied without taking ownership of a new runtime. Keep an evaluation bounded when the workflow itself is still uncertain. Revisit the decision when a concrete constraint emerges, not when a vague expectation of future complexity becomes uncomfortable.

If your scorecard shows a real gap but the team has not assigned ownership, do not treat that as a minor project-management detail. It changes the architecture decision. Work out who will maintain the component and what operational capacity they can commit, or choose a design that asks less of the organization. Where a team wants help evaluating that boundary or shaping an implementation plan, AI agent engineering is an appropriate next step.

The standard I would use is simple: every custom runtime decision should have a named business reason, an observable acceptance test and an owner for the failure path. If one is missing, the argument for building is incomplete. If all three are clear, the team can make a deliberate choice—and can explain not only why it built the harness, but what it has agreed to operate.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call