SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Tutorial 24 mins

Design an Agent Tool Contract with JSON Schema and Failure Codes

Design an agent tool contract with JSON Schema, scoped permissions, approval-bound execution, stable failure codes and validation cases.

The PADISO Team ·

What you will build

This tutorial develops a concrete tool contract for an AI agent that can help process customer refunds. The example is hypothetical: it is not a PADISO implementation, a tested integration, or a claim about any particular payment system. Its purpose is to show how to turn a vague instruction such as “refund the customer” into a bounded interface with explicit inputs, permissions, approval rules, and observable outcomes.

The central design choice is to separate proposing a refund from executing one. A model may prepare a request, but preparation does not move money. A trusted application records a human approval against the exact proposal, and a separate execution tool checks that approval before asking the business system to perform the action. This boundary makes it easier to reject changed, expired, duplicated, or unauthorized requests.

The worked example is deliberately narrower than a general guide to agent tools. It focuses on the contract that sits between an agent and a business action: which fields are valid, which actor may call each operation, what conditions must hold, and what the caller should do when something fails. For a broader view of how tools fit into the agent runtime, see The Anatomy of an AI Agent: What Each Component Does.

By the end, you will have a pair of illustrative JSON Schemas, a permission matrix, validation cases, stable failure codes, and a small execution flow. Treat these as a design artifact to adapt, not as a drop-in refund system. Your own order model, approval policy, identity system, and payment processor determine the production details.

Prerequisites and setup

You need a clear owner for the business rule being encoded. For the example, that owner must decide which orders can be refunded, which reasons are acceptable, which amounts require human approval, and how to confirm the order’s current state. An engineering team should not infer those rules from model behavior or bury them in prompt text.

You also need an application layer that can authenticate the caller, load authoritative order data, persist a proposal, record an approval, and invoke the existing business system. The agent must not receive credentials that let it bypass this layer. The model supplies a request; trusted server-side code decides whether the request is allowed.

The sample schemas use JSON Schema to describe the shape of tool arguments. JSON Schema can define object properties, required fields, and whether additional properties are permitted. Validation establishes whether data conforms to the schema; it does not authorize a refund or prove that an order is eligible. Keep those checks in application logic. JSON Schema’s object reference describes these object constraints.

For local schema validation, choose a JSON Schema validator compatible with the schema version you adopt. This article uses Python-style examples for readability; they are illustrative and have not been executed here. Pin your chosen validator in the application’s dependency management, validate the schema itself during development, and run application tests against the same version used in deployment.

Before implementation, write down the identity model. At minimum, distinguish the agent’s tool-call identity, the employee or service that approved a proposal, and the identity used by the application to perform the external action. Those identities may be represented differently in your system, but collapsing them into a single “agent is authorized” flag makes later audits and revocation harder.

Step 1: Define the business boundary before the schema

Start by stating what each operation is allowed to do in plain language. In this design, prepare_refund creates a proposed refund record and has no external financial effect. execute_refund attempts to carry out a previously approved proposal. It cannot accept a new amount or order identifier from the agent; it receives identifiers for a proposal and approval, then reloads the authoritative records itself.

That split is more important than the names. If a single tool both interprets a free-form request and performs the refund, it combines intent parsing, policy enforcement, approval handling, and external execution at one boundary. Separate operations let the application insert review between preparation and execution, and let each operation have its own permission rule and failure behavior.

Define the invariant that must hold at execution time: the approved proposal must still exist, be unexpired, remain unexecuted, and match the exact action being requested. The application should compare the stored proposal with the approval record, including the order, amount, currency, reason, and a canonical representation or digest of the approved payload. Approval of one proposal must not authorize a modified proposal.

Do not rely on a prompt instruction such as “ask before issuing a large refund” as the enforcement mechanism. Prompts can help the agent choose an appropriate next step, but policy must be checked where the action is authorized. Likewise, a model-generated statement that a customer is eligible is not evidence. The application should fetch current order and refund state from its authoritative sources.

For this tutorial, assume a fictional order service accepts refunds in integer minor units, such as cents, and that the application has an internal reason-code list. Those are illustrative assumptions, not universal payment conventions. Confirm the units, currencies, rounding rules, and supported reason codes for your own systems before using any comparable design.

Step 2: Specify the proposal tool input

A proposal should contain only the information needed to request an action. Avoid accepting a customer name, email address, raw payment data, or a prose justification if the application can resolve or record those details elsewhere. Extra fields increase the number of ways a model can accidentally disclose data or influence a decision outside the intended contract.

Here is an illustrative schema for prepare_refund. The dollar limit is intentionally not embedded as a universal policy: the schema checks types and allowed values, while application policy decides the permitted amount for the particular order. The example uses Draft 2020-12 syntax, which you should keep consistent with your selected validator.

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "PrepareRefundInput",
  "type": "object",
  "properties": {
    "order_id": {
      "type": "string",
      "minLength": 1,
      "maxLength": 80
    },
    "amount_minor": {
      "type": "integer",
      "minimum": 1
    },
    "currency": {
      "type": "string",
      "pattern": "^[A-Z]{3}$"
    },
    "reason_code": {
      "type": "string",
      "enum": ["duplicate_charge", "returned_item", "service_issue"]
    }
  },
  "required": ["order_id", "amount_minor", "currency", "reason_code"],
  "additionalProperties": false
}

The schema rejects missing required values, wrong primitive types, unknown reason codes, and undeclared keys. The additionalProperties setting is especially useful for a narrow tool because it prevents an input such as {"override_policy": true} from being silently accepted as an extra field by a permissive validator. It does not stop an application bug from ignoring validation, and it does not make the values truthful.

The amount being an integer avoids floating-point ambiguity in this illustrative contract, but it does not settle currency-specific rules. A business service still needs to check that the currency is supported, that the order is denominated in that currency, that the amount does not exceed the refundable balance, and that the value is within any policy limit. Those checks require trusted order data and policy, not additional confidence from the model.

An identifier’s schema length is a defensive shape constraint, not proof that the identifier exists or belongs to the current customer. The server must resolve order_id in the authenticated business context and apply its access policy. If the conversation already has a verified order context, consider binding the request to that context instead of allowing the model to select an arbitrary order identifier.

A useful review question is whether each property is necessary for the model to choose the action. If currency can be read from the order, the tool could omit it and prevent a model from supplying a conflicting value. If the reason is selected from a fixed internal list, retain the enum. If the application has no meaningful use for a field, remove it rather than adding it “for flexibility.”

Step 3: Define the execution tool separately

The execution tool should accept references to a proposal and an approval, not a second copy of the refund’s substantive fields. If the caller can change amount_minor during execution, the approval can become detached from the actual action. Fetching the stored proposal by identifier lets the application compare and execute the same record that the approver reviewed.

An illustrative execution schema is deliberately small:

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "ExecuteApprovedRefundInput",
  "type": "object",
  "properties": {
    "proposal_id": {
      "type": "string",
      "minLength": 1,
      "maxLength": 80
    },
    "approval_id": {
      "type": "string",
      "minLength": 1,
      "maxLength": 80
    }
  },
  "required": ["proposal_id", "approval_id"],
  "additionalProperties": false
}

The schema is not a credential system. An agent that knows an approval identifier must not automatically gain permission to use it. The application must verify that the approval exists, was created by an authorized approver, applies to this proposal, is still valid, and has not already been consumed. It must also verify the agent or runtime is permitted to invoke execution in the current workflow.

Create the approval only after displaying the exact proposal data to the authorized reviewer. Store the approver identity, proposal identifier, payload digest or equivalent immutable binding, decision time, and expiry. The runtime should refuse execution if the proposal changes after approval, even if the changed value seems harmless. Approval should expire according to a policy chosen for the business process; this article does not prescribe a duration.

If your workflow uses an approval token rather than an approval identifier, treat it as a sensitive capability. Do not place it in ordinary conversation context, logs that are broadly accessible, or reusable agent memory. The safest contract is often one where the model can request execution but the application—not the model—holds and validates the approval record.

Step 4: Assign permissions to each operation

Write permissions in terms of actors, operations, and conditions. “The agent can refund orders” is too broad to review. A matrix makes differences between proposing, approving, and executing visible before implementation.

ActorPrepare proposalApprove proposalExecute approved proposalKey restriction
Agent runtimeAllowed within assigned customer/order contextNot allowedAllowed only through guarded toolCannot set approval or bypass policy
Authorized reviewerMay request a proposalAllowed under business policyNot by defaultApproval binds to exact proposal
Refund service identityNot applicableNot applicablePerforms external request after checksCredentials remain server-side
Unauthenticated callerNot allowedNot allowedNot allowedMust not reach tool execution

This is an example permission model, not a recommendation that every organization grant agents execution access. Some teams may require a reviewer to trigger the execution step themselves. Others may allow the agent runtime to submit an already approved proposal while a service identity performs the external operation. Decide based on the consequence of a mistaken action, reversibility, operational controls, and your organization’s authority model.

Implement authorization on the server for every call. Tool visibility in a model configuration is not a substitute for checking identity and scope at invocation time. The server should establish the caller’s current permissions, load the target records, and reject actions outside the assigned context. Do not infer authorization from the model’s explanation, the user’s claim, or an identifier supplied in the request.

Keep the permission decision separate from schema validation. A request can be perfectly valid JSON and still be unauthorized. Conversely, an authorized actor can submit malformed data. These are different failure classes, should be tested independently, and should not be collapsed into one vague “tool failed” response.

A practical review artifact is a per-tool row that names the actor, allowed scope, preconditions, denied cases, and audit event. Include who can change each rule and how a change is reviewed. If the same developer can change the schema, the authorization check, and the approval policy without an independent review, the contract may be syntactically strict but operationally weak.

Step 5: Validate shape, policy, and approval in separate layers

Treat validation as a sequence of gates. First parse the tool arguments as JSON and validate their shape against the schema. Then authenticate and authorize the caller. Next load the order and check business eligibility. For execution, retrieve the proposal and approval, verify their binding and expiry, and only then call the external business system.

A compact illustrative validator could look like this:

from jsonschema import Draft202012Validator

prepare_validator = Draft202012Validator(prepare_refund_schema)

errors = sorted(
    prepare_validator.iter_errors(tool_arguments),
    key=lambda error: list(error.path),
)

if errors:
    return {
        "ok": False,
        "code": "INVALID_ARGUMENTS",
        "retryable": False,
        "message": "The refund request does not match the tool contract."
    }

# Continue with authenticated caller checks and business-policy checks.

This snippet assumes that prepare_refund_schema and tool_arguments have already been loaded by the application. It illustrates where schema validation belongs; it does not implement authentication, order lookup, approval, persistence, or an external refund call. A production handler must ensure parsing failures are also handled and that internal validator details are not automatically returned to the agent.

Expected result for a valid argument object is that the shape check passes and the request proceeds to authorization and business eligibility checks. For an object with a missing currency or an unexpected override_policy key, the shape check returns INVALID_ARGUMENTS, and no proposal is created. A passing schema is not an expected refund success; later gates can still deny the request.

For the proposal operation, application-level checks should include whether the order exists in the caller’s permitted scope, whether the requested currency agrees with the order, whether the amount is positive and within the currently refundable balance, whether the order state allows a refund, and whether the reason is compatible with the action. The exact rules are business-specific. Query authoritative state at decision time rather than trusting a summary supplied in the conversation.

For execution, re-read the proposal and approval immediately before acting. Verify exact payload binding, expiry, unused status, caller permission, and current business eligibility. If an intervening change means the order is no longer eligible, stop and return a clear denial or conflict. Do not silently alter the approved amount to make the operation pass; create a new proposal and obtain a new approval instead.

The execution flow below shows the application boundary rather than model reasoning. Each rejection branch stops before the external action. A successful check reaches the business system, but the application must still interpret the result and verify the resulting business state.

flowchart TD
  accDescr: Workflow stages and decisions: Agent submits tool call, Validate shape and caller, Return stable failure, Load proposal and approval, Check exact binding and expiry, Invoke business system, Verify and record outcome. The adjacent text explains the conditions and exceptions.
  accTitle: Design an Agent Tool Contract with JSON Schema and Failure Codes workflow
    A["Agent submits tool call"] --> B["Validate shape and caller"]
    B -->|"Rejected"| X["Return stable failure"]
    B -->|"Allowed"| C["Load proposal and approval"]
    C --> D["Check exact binding and expiry"]
    D -->|"Rejected"| X
    D -->|"Valid"| E["Invoke business system"]
    E --> F["Verify and record outcome"]

accTitle: Guarded refund tool execution accDescr: The application validates a tool call, loads the proposal and approval, checks their binding and expiry, invokes the business system only after those checks, then verifies and records the outcome. Any rejection returns a stable failure without advancing to execution.

The first node is an agent request, not an authorization decision. The shape and caller gate checks data validity and runtime identity; rejection returns a stable failure. Loading the records does not itself approve anything. The binding and expiry gate prevents stale or changed authorization from reaching the external system. Finally, invocation and outcome verification are separate: an accepted request is not proof that the refund completed as intended.

Step 6: Return stable, actionable failure codes

Define a small error vocabulary before wiring the tool into the agent. A stable code lets the application, model, operator, and monitoring system distinguish invalid input from a business denial or an uncertain external outcome. Keep the code predictable even if the user-facing message changes.

CodeMeaningSuggested agent behaviorExecution occurred?
INVALID_ARGUMENTSInput is malformed or violates the schemaCorrect fields only when the correction is clear; otherwise ask for clarificationNo
NOT_AUTHORIZEDCaller or scope is not permittedStop; do not retry with another identifierNo
ORDER_NOT_ELIGIBLECurrent order state or policy prevents the proposal/actionExplain the limitation without claiming completionNo
APPROVAL_REQUIREDProposal needs an authorized human decisionRoute to the approved review processNo
APPROVAL_INVALIDApproval is missing, expired, mismatched, or already consumedStop and request a new valid approval if policy permitsNo
CONFLICTState changed between proposal and executionRefresh state and prepare a new proposal if appropriateNo, unless outcome is separately confirmed
DEPENDENCY_UNAVAILABLEA required internal or external dependency cannot be reachedRetry only under bounded application policyUnknown until checked
OUTCOME_UNCONFIRMEDThe request may have reached the business system, but its final state is not knownReconcile before attempting another actionUnknown

The last two cases need careful treatment. A network timeout can occur before a request is sent, after it is accepted, or while the response is lost. The tool should not report a definitive failure if the external effect may have happened. Mark the outcome uncertain, retain the operation reference and correlation data available to the application, and reconcile against the business system before another attempt.

An idempotency key can help a downstream system recognize repeated requests when that system supports it, but it is not a promise of exactly-once effects. Its scope, retention, and behavior depend on the external system. The application should preserve a stable key for retries of the same intended operation, avoid reusing it for a changed proposal, and still verify the resulting business state.

Give the agent enough information to choose a safe next step, but do not return stack traces, credentials, internal network details, or raw payment-system responses by default. A response can include a stable code, a short explanation, a correlation identifier, and whether execution is known not to have occurred, confirmed, or still uncertain. Keep sensitive diagnostic detail in access-controlled operator logs.

The agent should not be encouraged to “try again” for every failure. INVALID_ARGUMENTS may be correctable, NOT_AUTHORIZED should stop, APPROVAL_REQUIRED should go to review, and OUTCOME_UNCONFIRMED should trigger reconciliation rather than a second refund attempt. Encode those distinctions in orchestration behavior as well as in documentation.

Step 7: Add validation cases before connecting the tool

Create tests around the contract’s boundaries, not just a happy-path JSON object. The table below is a starter artifact. Each case should assert both the response and the absence or presence of side effects. Use a fake business-system adapter in tests where appropriate; the examples do not claim that any test has been run.

CaseInput or stateExpected resultSide-effect expectation
Valid proposal shapeAll required fields; allowed reason codeShape validation passesNo external refund call
Missing fieldcurrency omittedINVALID_ARGUMENTSNo proposal persisted
Unexpected propertyoverride_policy includedINVALID_ARGUMENTSNo proposal persisted
Wrong amount typeDecimal string or fractional numberINVALID_ARGUMENTSNo proposal persisted
Unknown reasonReason not in the enumINVALID_ARGUMENTSNo proposal persisted
Unauthorized orderValid shape, order outside caller scopeNOT_AUTHORIZEDNo proposal or execution
Amount above refundable balanceValid shape, authoritative balance is lowerORDER_NOT_ELIGIBLENo external call
Changed proposalApproved record no longer matches stored proposalAPPROVAL_INVALIDNo external call
Expired approvalApproval expiry precedes execution timeAPPROVAL_INVALIDNo external call
Duplicate execution attemptProposal already marked consumedReject or return reconciled prior stateNo second external effect assumed
Timeout after submissionExternal result cannot be confirmedOUTCOME_UNCONFIRMEDReconcile before retry

For every negative case, test that no later step runs. For instance, a schema error should not trigger an order lookup that has side effects, and a failed approval check must not reach the refund adapter. In a unit test, assert calls to the relevant boundary are absent. In an integration test, use an isolated environment and controlled records appropriate to your system; do not use real customer transactions as a shortcut for exercising error paths.

Test boundary values chosen by your policy owner: the smallest permitted amount, the exact refundable balance, a value just above that balance, unsupported currency, and an order whose state changes after proposal creation. The schema can express some simple numeric bounds, but current balance and state are runtime facts. Keep their tests in the application-policy layer, where those values can be loaded from controlled fixtures.

Also test malformed JSON separately from a valid JSON object that violates the schema. The parser may fail before the validator runs, and both cases should produce a safe, understandable response. Verify that diagnostic output does not expose raw input unnecessarily; tool arguments can contain customer or operational data even when this example keeps fields minimal.

For execution, include a test where the reviewer approves a proposal and then one approved field is changed before execution. The action must be denied, even if the changed payload would pass the schema. Add separate cases for the approval belonging to a different proposal, an approval from a disallowed actor, an expired record, and a repeated call after a previously confirmed outcome.

Step 8: Walk through a hypothetical request

Suppose a fictional customer asks for a refund for a returned item. The agent has access to the current conversation and a narrow order context, but it cannot inspect payment credentials or grant itself permission. It proposes an amount and reason using prepare_refund. The application validates the shape, confirms that the caller’s context includes the order, loads the current refundable balance, and evaluates the organization’s policy.

If the request is eligible for review, the application creates a proposal record containing the authoritative order reference and the requested action. It returns a proposal status and identifier to the workflow. The authorized reviewer sees the exact amount, currency, reason, and relevant order context in the organization’s review interface. The reviewer’s decision is recorded against that proposal, not as a general permission for the agent to issue refunds.

When the workflow later requests execution, the agent supplies the proposal and approval identifiers. The server loads both records, confirms that the approval actor is authorized, checks that the approval is still valid and bound to the exact proposal, and verifies that the proposal has not already been consumed. It then checks current eligibility again. If the order changed after review, the application stops and asks for a fresh proposal and approval where policy allows.

If the business system returns a clear success response, that is an input to outcome verification, not a substitute for it. The application should confirm the relevant state using a trustworthy result or subsequent read appropriate to that system, then record the outcome and expose a concise status to the agent. If the response is ambiguous, mark the operation unconfirmed and reconcile before any retry. Never let the model turn an uncertain tool response into a confident claim that the customer has been refunded.

This sequence includes friction by design. A proposal may require a human decision; an external request may need reconciliation. Removing those steps can make a demo shorter, but it also erases the distinctions that let a team understand who approved what, what was actually attempted, and whether the intended business outcome occurred.

The example also shows why an approval flag supplied as a tool argument is not enough. A model could provide approved: true regardless of the real review process. A trusted approval record has an identified approver, a specific bound proposal, a valid time window, and a status that the application can verify. The model can ask to proceed; it cannot manufacture the authority to proceed.

Step 9: Handle operational failures deliberately

A tool contract continues to matter after deployment because failures occur between its gates. A policy denial is different from a dependency outage, and both differ from a request whose external result is unknown. If all three become “refund failed,” operators will have difficulty deciding whether to correct data, restore a dependency, or reconcile a possible completed action.

Log enough structured information to investigate without turning logs into a second source of sensitive customer data. A useful event can include the tool operation, stable failure code, internal correlation identifier, proposal reference where permitted, timestamp, and relevant state transition. Restrict access and retention according to your own data governance. The exact fields and retention rules are organization-specific and are not prescribed here.

Record the proposal, approval, execution attempt, and verified outcome as distinct events. This provides a trace of the decision path without pretending that a single “tool call succeeded” event captures the entire transaction. If the system supports append-only or otherwise controlled audit records, evaluate how they fit your existing operational controls rather than assuming a particular storage design.

Build bounded retry behavior around known semantics. Retrying a read may be safe under one system’s rules; retrying a financial side effect after an ambiguous response may not be. Use the external system’s documented behavior, stable operation identifiers where supported, and reconciliation. Do not claim exactly-once execution merely because the application sends an idempotency value.

A deployment should have a way to disable or narrow the execution permission without removing the agent’s ability to answer questions or prepare non-executing proposals. This can be an operational control in the application’s authorization layer. It is particularly useful when a dependency is degraded, an approval workflow is unavailable, or the team is investigating unexpected behavior.

If you use a managed agent framework, check how its tool lifecycle, approval handling, and server-side authorization fit your design before moving these rules into framework configuration. The implementation trade-offs are discussed in Claude Agent SDK: When It Beats Your Own Orchestration; the tool contract still needs to express your application’s own policy regardless of the orchestration choice.

Do not expand this implementation into a broad redesign of the agent’s memory or evaluation program. Those are separate concerns. For context handling, see Context Engineering for Agents: A Keep, Summarise or Retrieve Experiment. For the distinction between exercising an agent and measuring its quality, see An Agent Harness Is Not an Evaluation Harness: Here’s the Difference. If the workflow must resume across sessions, Why Long-Running Agents Lose Their Place—and How to Resume Them covers that separate design problem.

Step 10: Review the contract as a changeable interface

Treat a tool schema and its behavioral rules as a versioned interface. Changing a field name, enum value, permission condition, or failure code can affect the model configuration, application handler, test suite, operator playbook, and downstream reporting. Review changes as contract changes, not merely prompt edits.

Before adding a field, identify which layer owns its truth. Model-selected intent belongs in the request only when the model needs to choose it. Current order status belongs to the authoritative business system. Approval status belongs to the approval record. Caller scope belongs to authenticated application context. Putting authoritative values into the model’s input may create a second, potentially stale version of the truth.

Before relaxing additionalProperties, ask how unknown arguments will be interpreted and whether they can alter authority or business intent. Strict rejection is useful for a small, sensitive tool because it exposes mismatches instead of silently accepting them. If forward compatibility requires tolerant input, define exactly which fields are ignored, how they are logged, and why ignoring them cannot weaken policy.

Before adding a retry option, ask what happens if the previous request succeeded but its response was lost. Before adding a “force” field, identify who may set it and which independent control authorizes it. Before changing approval expiry, consider what can change between review and execution. These questions prevent convenience fields from becoming undocumented escape routes.

A useful release review compares the proposed schema, authorization behavior, failure-code semantics, test cases, and operator instructions together. Require a reviewer who can challenge the business policy, not only a reviewer who can spot invalid JSON. Keep the contract readable enough that product, operations, and engineering can all identify the boundary they are relying on.

Troubleshooting common problems

The schema accepts a field the team expected to reject. Check that the validator is using the intended schema version and that the object has additionalProperties set as designed. Confirm that the application actually invokes validation on the runtime input, rather than validating only a separate example or relying on the model’s advertised tool shape.

A valid tool call is still denied. That may be correct. Schema validation checks shape, while authorization and business policy check caller scope and current eligibility. Return a stable code that indicates which gate denied the action, and provide only enough information for the agent or operator to take an allowed next step.

The agent keeps retrying a denied action. Make failure semantics explicit to the orchestration layer. A permission denial should stop, an approval requirement should route to review, and an uncertain external outcome should trigger reconciliation rather than immediate resubmission. Avoid giving the model a generic retry instruction that erases these distinctions.

The reviewer approved one amount, but the tool attempts another. Do not pass editable action fields to the execution tool. Load the stored proposal and verify that the approval binds to its exact payload. Reject mismatches and require a new review for a changed proposal.

An execution timed out and nobody knows whether it worked. Do not translate the timeout into a definitive failure. Record an unconfirmed state, use the available operation reference to reconcile against the authoritative system, and prevent an unverified second action. The downstream service’s own semantics determine what retry strategy is safe.

The agent says the refund completed, but the business record does not show it. Separate tool invocation from verified outcome in the tool response. Have the application return a completion claim only after its defined verification step. If verification is delayed or unavailable, report a pending or unconfirmed status instead of asking the model to infer success from its own prior message.

The schema has become difficult to maintain. Revisit whether the tool combines distinct operations or accepts values that the application can retrieve itself. A smaller proposal interface and a separate guarded execution interface are often easier to reason about than a single broad tool with many optional flags. Keep the contract aligned with one meaningful business action at a time.

Implementation handoff

Use the following worksheet in a design review before exposing a consequential tool to an agent. It is a printable artifact within this article, not a downloadable template or a claim that completing it alone makes an implementation safe.

  • Name the action and its boundary. State what the tool does and, just as importantly, what it cannot do. Identify whether it reads, proposes, changes, or executes a business action.
  • List each input with its source of truth. Mark whether a field is model-selected, derived from authenticated context, or loaded from a business system. Remove fields the model does not need to provide.
  • Define schema constraints. Specify types, required fields, allowed values, length or numeric bounds where appropriate, and how unknown properties are handled. State the schema version and validator used by the application.
  • Write authorization rules separately. Identify the caller, permitted scope, denied cases, and the server-side check that enforces each rule. Do not treat schema conformance as permission.
  • Bind approval to an exact proposal. Record the approver, proposal identity, exact approved payload or its equivalent binding, expiry, and consumption state. Ensure execution rejects changed or stale approvals.
  • Define stable outcomes. Choose failure codes and state whether an external effect is known not to have occurred, confirmed, or uncertain. Give the agent a safe next step for each code.
  • Test refusals and ambiguity. Cover malformed input, extra fields, unauthorized scope, policy denial, expired or mismatched approval, duplicate requests, and timeouts. Assert that rejected requests do not reach the external action boundary.
  • Plan verification and reconciliation. State how the application confirms the resulting business state and what operators do when a response is ambiguous. Do not promise exactly-once effects without support from the actual system semantics.
  • Assign ownership and change review. Name the business-policy owner and technical owner. Review schema, permissions, approval behavior, error handling, and tests together when changing the contract.

The concise release criterion is this: someone reviewing the implementation should be able to tell what the agent may request, what the application must verify, who may approve, what exact action that approval covers, and how the system behaves when the outcome is uncertain. If any of those answers exist only in a prompt or an operator’s memory, the contract is not yet complete.

For teams that need help translating an agent workflow into a bounded tool interface and application-side controls, AI agent engineering is an appropriate next step. The concrete design remains specific to your systems, permissions, and business policy.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call