SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Tutorial 18 mins

Build an AI Harness Run Manifest: State, Tools and Stop Conditions

A hands-on tutorial for designing an AI harness run manifest, with explicit state transitions, approval-bound actions, stop conditions and failure handling.

The PADISO Team ·

What you will build

A run manifest is the explicit record of what one agent run is allowed to do, what it has done, and what must happen before it stops. It is not the agent’s conversation transcript, and it is not a prompt with extra headings. It is a control artifact that lets the surrounding application distinguish a plausible model response from a valid next action.

This walkthrough builds a small, application-owned manifest and a state machine around it. The example is a hypothetical customer-support workflow: an agent reviews a refund request, gathers order evidence, drafts a refund action, waits for approval of that exact action, executes it, and checks the business result. The scenario illustrates a design; it is not a claim about a deployed system or a guarantee that external payment effects can be made exactly once.

If you need the component-level orientation first, start with The Anatomy of an AI Agent: What Each Component Does. This tutorial assumes you already have an application that can call a model and expose selected tools. It focuses on the run record and its transitions, not on choosing a framework or designing durable multi-hour execution. For those separate decisions, see Agent Harness vs Agent Framework vs Runtime: A Practical Guide and Durable Execution for Agents That Run for Hours.

The target is modest but useful: every run has a unique identity, bounded inputs, named state, recorded evidence, an explicit action payload, a clear approval gate, and a stopping condition. When a run fails, an operator should be able to tell whether to retry a read, request a new approval, reconcile an uncertain external result, or stop for a person.

Prerequisites and setup

You need an application service capable of persisting one run record, a model-call boundary, and a small set of application-controlled tools. For the tutorial, assume three conceptual operations: retrieve an order, prepare a refund request, and inspect the refund status. These are illustrative names, not vendor API names. Adapt them to your own application and confirm each operation’s actual behavior before exposing it to an agent.

Keep the model outside the authority boundary. It may suggest which evidence to collect and propose a structured action. Your application decides whether the request is valid, whether the run can transition, whether approval is required, and whether an external operation may be attempted. A schema can reject malformed data, but valid shape alone does not confer authority to execute a business action. JSON Schema’s object reference describes properties, required fields, and additional-property handling.

Choose a persistence mechanism that can atomically save a run’s current state and revision. A relational row with a version number is one option; a transactional document store can also work. The example uses ordinary JSON objects so the design is readable. It does not prescribe a database, queue, model provider, payment processor, or approval product.

Before coding, decide which fields are authoritative. In this design, the application supplies the run identifier, tenant or account scope, request identifier, allowed tools, state, revision, timestamps, and approval record. The model may contribute a proposed rationale and action fields, but those fields are untrusted until checked against application-owned order data and policy.

Set a clock source for timestamps and a stable serialization format for the action payload. Stable serialization matters because approval must refer to the exact proposed action, not merely a similar description. Record a digest of canonical payload bytes, or store the exact payload and compare it byte-for-byte after canonicalization. If your language or database changes object-key ordering, do not rely on raw, incidental serialization.

Finally, decide what constitutes an observed business result. For this hypothetical flow, a successful tool response is not enough: the application makes a separate status read and accepts completion only when the returned refund record matches the approved order, amount, and currency. That distinction prevents a model’s claim, or an ambiguous network response, from becoming the run’s completion signal.

Step 1: Define the manifest and its invariants

Start with the smallest record that can explain and constrain a run. Avoid putting the entire model transcript into this object. Store a transcript separately if your product needs it; the manifest should remain a compact control record that an operator and application can inspect quickly.

The example manifest has identity, input, state, evidence, action, approval, and outcome areas. revision increases whenever the authoritative record changes. state is selected by the application from a finite set. attempt is an application-managed counter for bounded retries. stop_reason is present when the run is terminal. Fields such as tenant_id and request_id scope lookups so that an action cannot silently migrate to another customer or request.

The central invariants are more important than the field names. A run cannot execute before the evidence has been checked. A run cannot execute a payload different from the one approved. An approval has an expiration time and is rejected after that time. A rejected, expired, or invalid approval does not fall through to execution. A terminal run cannot resume by simply changing its state field; a deliberate new run or controlled recovery path is required.

An illustrative initial record might look like this:

{
  "run_id": "run_7f31",
  "tenant_id": "shop_north",
  "request_id": "case_2841",
  "created_at": "2026-09-30T14:00:00Z",
  "revision": 1,
  "state": "queued",
  "attempt": 0,
  "input": {
    "order_id": "order_913",
    "requested_amount_minor": 4200,
    "currency": "USD",
    "reason": "item arrived damaged"
  },
  "allowed_operations": ["read_order", "propose_refund", "read_refund_status"],
  "evidence": [],
  "action": null,
  "approval": null,
  "outcome": null,
  "stop_reason": null
}

The amount is represented in minor currency units to avoid floating-point ambiguity. That is a choice for this example, not a universal payment convention; use the unit and currency rules of the system that owns the transaction. An input check should reject missing currency, negative amounts, unexpected fields, invalid identifiers, or values outside the request’s allowed range before any model call.

The allowed_operations list is not a substitute for server-side authorization. It is an auditable declaration of the intended tool surface for this run. The dispatcher must still check the authenticated principal, tenant scope, operation-specific rules, and current run state at the point of use. Never accept a model-provided operation list as the source of permission.

Step 2: Make transitions explicit

Write down the states before connecting the model. This example uses queued, preflight, collecting, proposed, awaiting_approval, executing, verifying, and terminal states completed, rejected, expired, failed, and needs_reconciliation. The names are illustrative; the important property is that every transition has a defined trigger and guard.

A queued run may enter preflight only after the application has loaded the request in its proper account scope. Preflight checks the input shape, request status, amount bounds, and availability of required read operations. If a guard fails, the run stops with a specific reason rather than asking the model to repair an authoritative identity or permission failure.

After preflight, the application retrieves the order and stores relevant facts as evidence: source record identifier, observed amount and currency, order status, retrieval time, and any version or update marker the source provides. Do not store a bare model summary as if it were the source record. A summary can help the next reasoning step, but the manifest should retain enough structured provenance to revisit the decision.

The model can then propose an action. The application checks that the proposed order matches the input order, the currency matches verified evidence, the requested amount is within the permitted amount, and the reason maps to an allowed policy outcome. If validation fails, do not quietly edit the proposal and proceed. Reject it, ask for a new proposal where appropriate, or stop for review. Any changed payload must pass the same checks and receive a fresh approval.

Approval is a state transition, not a sentence in the transcript. The record should identify the approver or approval mechanism, decision time, expiration time, and digest of the exact action payload. If any bound field changes—order, amount, currency, destination, or operation—the old approval no longer authorizes the new payload. If no approval arrives before expiry, the run moves to expired; it does not remain eligible for a delayed execution callback.

Execution requires a fresh read of the authoritative manifest revision, an unexpired approval, a matching action digest, and an allowed transition. The dispatcher should enforce these guards immediately before the external call. A previous successful check is not sufficient if another process could have modified the run in the meantime.

The flow below shows the main path and the stop branches. A stop is an intentional outcome: preflight rejection and approval expiry must not connect to the execution node.

flowchart TD
    accTitle: Run manifest state flow
    accDescr: A run passes preflight, evidence collection, action proposal, approval, execution and verification. Failed preflight or approval stops before execution; an unverified execution result stops for reconciliation.
    Q["Queued manifest"] --> P["Preflight checks"]
    P -->|"valid"| E["Collect evidence"]
    P -->|"invalid"| S["Stop with reason"]
    E --> D["Draft exact action"]
    D --> A["Approval valid and current?"]
    A -->|"no"| S
    A -->|"yes"| X["Execute approved payload"]
    X --> V["Verify business result"]

The diagram omits implementation details such as retries and reconciliation queues to stay readable. Adjacent code and rules supply those details. In particular, a timeout after the external call is not equivalent to a confirmed failure: the external system may have accepted the action even when the application did not receive its response.

Step 3: Create a bounded action proposal

For the hypothetical case, suppose the order evidence confirms the customer’s order, a refundable balance of 4,200 minor units, and USD currency. The model proposes refunding 4,200 units to the original payment method. The application, not the model, fills or verifies the account scope, order reference, currency, permitted operation, and policy result.

Store the proposal as a typed object with only fields the executor requires. For example, operation might be refund_order; order_id, amount_minor, and currency identify the requested effect; evidence_refs point to the records supporting it. The application can also retain a short explanation for the approver. Do not let free-form explanation fields control the dispatcher.

Before requesting approval, create a canonical representation of the action and compute its digest using a standard cryptographic hash implementation in your application language. Save both the action and digest in the same revision. The approval interface should display the material action fields clearly and bind the approval record to that digest. A general “approve this case” button is insufficient if the payload could change after the decision.

A useful approval record contains decision, approver_id, decided_at, expires_at, and action_digest. It should also include the run revision or another concurrency token. On execution, verify the decision is affirmative, the expiry is still in the future according to the application clock, the digest matches, and the relevant revision has not changed. Expired approvals require a new decision, even if the action appears unchanged.

This design intentionally requires a person before the modeled refund action executes. Whether a particular organization should use human approval, a lower-risk automated limit, or a different workflow is a business policy decision. The engineering invariant remains: the policy is enforced by code, and a model cannot declare its own approval or bypass a missing decision.

Step 4: Implement transition guards

The following Python-like example uses only the standard library for the state transition logic. It is illustrative rather than a complete service: persistence, identity checks, tool calls, canonicalization, and approval capture belong in your application. In particular, do not paste this into production without adding atomic storage updates and your actual policy checks.

from datetime import datetime, timezone

TERMINAL = {"completed", "rejected", "expired", "failed", "needs_reconciliation"}


def require(condition, message):
    if not condition:
        raise ValueError(message)


def approve_for_execution(run, now, approved_digest):
    require(run["state"] == "awaiting_approval", "wrong state")
    approval = run.get("approval")
    action = run.get("action")
    require(approval is not None and action is not None, "missing approval or action")
    require(approval["decision"] == "approved", "approval not affirmative")
    require(approval["expires_at"] > now, "approval expired")
    require(approval["action_digest"] == approved_digest, "approval digest mismatch")
    require(action["digest"] == approved_digest, "action changed after approval")
    run["state"] = "executing"
    run["revision"] += 1
    return run


def stop(run, state, reason):
    require(state in TERMINAL, "not a terminal state")
    run["state"] = state
    run["stop_reason"] = reason
    run["revision"] += 1
    return run


now = datetime.now(timezone.utc).isoformat()

The final now line only demonstrates obtaining a UTC timestamp; production code should compare parsed, timezone-aware instants rather than lexicographic strings. The sample guard accepts a digest argument only to make the check visible. In a real dispatcher, retrieve the approved digest and stored action from the authoritative persisted record; do not let a caller supply the value that makes its own check pass.

The TERMINAL set makes the state boundary visible, but a production transition function should define every allowed source-to-destination pair. For example, awaiting_approval can become executing, rejected, or expired; it cannot jump to completed. A run in needs_reconciliation should enter a controlled reconciliation procedure, not restart execution automatically.

Persist transitions with a compare-and-swap condition on revision, or an equivalent transaction. If two workers load revision 8, both must not independently commit different next states as revision 9. One update should win; the other should reload and reevaluate. This is a concurrency guard, not a claim that the external action itself is exactly-once.

Step 5: Execute and verify as separate events

When the state moves to executing, write an execution-attempt record before calling the external system. Include a unique attempt identifier, run revision, action digest, start time, and status such as started. If the external system supports a documented idempotency mechanism, use it according to its actual contract; do not invent one or assume a repeated call is harmless.

On a clear success response, record the response identifier and transition to verifying. Then read the business record independently and compare its order, amount, currency, and status with the approved payload. Only after that comparison succeeds should the run become completed. Save the observed record reference and verification time as outcome evidence.

If the call returns a clear validation rejection, record the response and stop as failed or route for review, depending on your policy. Do not change the action and retry under the old approval. A corrected action is a new proposal and needs a new digest and, where required, a new approval.

If the call times out, the network connection drops, or the process crashes after sending the request, mark the result as uncertain and move to needs_reconciliation. First inspect the external business record using a read operation. If the approved effect exists, record it and complete after matching the fields. If it does not exist, decide whether a retry is safe under the external system’s documented semantics. If the result cannot be established, keep the run stopped for a person rather than guessing.

This separation addresses a common counterexample: an agent says “refund submitted,” the application receives a timeout, and a retry issues a second refund. The model’s narrative is not evidence of submission. Nor does the absence of a response prove that the first request failed. The manifest must preserve that uncertainty explicitly until a read or authoritative response resolves it.

Step 6: Walk through the hypothetical run

At 14:00 UTC, the application creates run_7f31 in queued at revision 1. Preflight loads the case under shop_north, confirms that the request is open, validates the input amount and currency, and checks the operation policy. It records the result and moves to collecting. An invalid account scope or closed request would stop here with a specific reason, before the model sees any action tool.

The read operation returns order facts. The application stores the source reference, retrieval time, order status, refundable balance, and currency. It may pass a compact, scoped summary to the model to obtain a proposed disposition. The model’s text is not copied into the evidence fields as truth. The application compares the structured proposal against its retrieved record and policy.

Suppose the proposal is 4,200 USD for order_913, with a concise reason tied to the damaged-item request. The application verifies the amount is within the recorded refundable balance and that the operation is permitted. It stores the exact action and digest, moves to awaiting_approval, and presents those fields to the approver. The approver accepts before the expiry time, and the system stores the decision and the same digest.

A worker reloads the current run, confirms its revision and state, checks expiry and digest, and records the attempt. It submits the action once. The response provides a reference, so the run advances to verifying; the application reads the refund record and checks the approved order, amount, currency, and status. A matching result produces completed with the observed record reference. This expected output means the business state was observed, not merely that the model returned a completion message.

Now consider a changed-amount counterexample. After approval for 4,200 USD, a later model turn proposes 3,900 USD because it reinterprets the request. The old approval digest does not match. The dispatcher refuses execution and returns the run to a proposal or review path according to policy. It must not treat the new amount as a harmless correction. The changed payload requires fresh validation and a new approval.

A second counterexample is an approval arriving after expiry. Even if the approver clicked an old notification and the action is otherwise unchanged, the system marks the approval unusable. It requests a fresh decision or stops the run. This avoids delayed callbacks turning a stale decision into present authority.

Step 7: Add operational stop conditions

Define stop reasons as stable machine-readable codes with a short operator-facing explanation. Useful categories for this workflow include invalid_input, scope_mismatch, evidence_conflict, policy_denied, proposal_invalid, approval_rejected, approval_expired, attempt_limit, and external_result_uncertain. Avoid a single error value that forces an operator to reconstruct every failure from logs.

Set a maximum number of model proposal cycles and read retries. The right limits depend on latency, the cost and semantics of your reads, and the operational response path; choose them explicitly rather than allowing an agent to loop until a time budget is exhausted. Increment counters in application state, and stop when a limit is reached. Do not let the model reset counters by requesting a new turn.

Define a deadline for the run and separate deadlines for approval and an individual tool call. A whole-run deadline should lead to a terminal or recovery state with a reason. A tool-call timeout should be classified by operation: a failed read may be retryable, while a timed-out write may have an uncertain effect and require reconciliation. Treating both as the same generic retry loses important semantics.

Store enough event history to reconstruct transitions: prior state, next state, revision, triggering event, actor or worker identity, timestamp, and a reference to relevant evidence. Protect this history from being overwritten by a later summary. The run manifest can hold the current snapshot; an append-only event record or equivalent history makes debugging and audit review more intelligible.

Useful operational signals are transition counts and age by state. For example, a growing population of runs awaiting approval may indicate an overloaded review queue; repeated needs_reconciliation outcomes may indicate an external interface or response-handling problem. These are diagnostic signals, not proof of a root cause. Review them alongside representative run records before changing thresholds or allowing more automation.

Troubleshooting common failures

The run is stuck in collecting. Check whether the read operation timed out, whether the worker recorded an attempt, and whether the state update was committed. Retry only if the operation is read-only and the retry budget remains. If the worker died after fetching evidence but before saving it, repeat the read and record a fresh observation rather than pretending an unsaved response is durable.

Approval appears valid but execution is blocked. Compare the current action digest with the approved digest, inspect expiry using the application’s UTC clock, and confirm the run revision did not change. A mismatch is a correct stop, not a reason to weaken the guard. If a field changed, create a newly validated proposal and obtain a decision over that exact payload.

The external call timed out. Do not immediately repeat the write. Mark the attempt uncertain, use an authoritative read to look for the effect, and follow the external system’s documented retry rules. If evidence is inconclusive, route to reconciliation. This may take longer than an automatic retry, but it avoids making an unknown state worse.

A tool returns success but verification fails. Keep the run out of completed. Record the response and verification evidence, identify which approved field differs, and stop for review. A success response may mean that a request was accepted without proving the final business result you require.

State updates conflict between workers. Use a revision check in the storage transaction. The losing worker reloads current state and discards any transition that no longer applies. For a write that may already have been sent, a storage conflict does not cancel the external side effect; check the attempt record and reconcile before another execution attempt.

Worksheet: review a manifest before deployment

Use this compact worksheet for one workflow, not as a substitute for testing its actual integrations. Write a concrete answer beside each item. If an answer depends on a model behaving correctly, move that control into application code or a verified external boundary.

Run identity and inputs

  • Run identity: Which stable identifier ties the request, manifest, events, and external record together? Confirm that concurrent requests cannot reuse it accidentally.
  • Scope: Which account, tenant, or business unit owns the request? Identify where the dispatcher checks that scope, not only where it is stored.
  • Input limits: Which fields are required, which values are bounded, and which unexpected fields are rejected? Document currency units and identifier formats.
  • Evidence provenance: For each fact that affects the action, record its source reference, observation time, and any useful version marker. Separate source facts from model summaries.

State and action

  • Allowed transitions: List every permitted state change and its guard. Confirm no path reaches execution directly from a model response or from a terminal state.
  • Action payload: Name every field that can change the business effect. Specify how the application validates it and how it creates a stable digest.
  • Tool boundary: Identify the exact application dispatcher that checks state, scope, and operation policy immediately before a call. Do not treat an advertised tool list as authorization.
  • Approval binding: Record who or what can approve, which exact payload is displayed, how the decision binds to its digest, and when it expires. Define the result of rejection or expiry.

Failure and completion

  • Retry policy: Set separate retry limits for reads and writes. Define how you determine whether a write’s effect is known after a timeout.
  • Stop reasons: Give each non-successful path a stable code, an operator explanation, and a safe next action. Avoid automatic retries for uncertain effects.
  • Verification: Name the authoritative read that proves the desired business result and the fields it must match. Specify what happens when a result is partial or contradictory.
  • Concurrency: Define how revision conflicts are detected and what a worker must do after losing a write race.
  • Recovery: Decide which terminal states require a new run, which can enter controlled reconciliation, and who may make that transition.

For a printable review, copy this worksheet into the team’s design record and fill it in for one real workflow. The useful output is not a checked box count; it is a set of implementable guards, named owners, and explicit failure paths. If the team cannot say how it knows a write succeeded after a timeout, keep that operation out of unattended execution until the uncertainty has a designed resolution.

Setup completion and next step

A first implementation is ready for a controlled rollout when an invalid input stops before model reasoning, every action is checked against recorded evidence, approval is tied to an unexpired exact payload, and the executor rechecks state immediately before acting. It should also distinguish a confirmed failure from an unknown external outcome, and it should reach completed only after a separate business-result check.

Start with a read-only workflow or a narrowly scoped action in a non-production environment. Exercise the transitions deliberately: malformed input, conflicting evidence, rejected approval, expired approval, changed payload, duplicate worker, clear tool rejection, timeout after request submission, and verification mismatch. These are proposed acceptance cases; results depend on your actual tools and application, so record what your implementation observes rather than assuming this illustrative design has been tested for you.

Once the run record is stable, connect it to your own persistence and operational workflow. Keep the manifest focused on state and action control rather than expanding it into a general framework, long-running workflow engine, or memory system. For teams that need help designing or implementing that boundary, AI agent engineering is the relevant next step.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call