SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Tutorial 21 mins

Tool Changes Can Break Agents: Add Contract Tests to CI

Build CI contract tests for agent tools that catch schema, permission and error-behavior regressions before a changed tool reaches production.

The PADISO Team ·

A tool change can pass ordinary unit tests and still leave an agent unable to use it safely. A renamed field may make a formerly valid call unusable. A permission check may begin allowing an action for the wrong role. An error response may change shape, causing the agent to treat a failed operation as success or to retry an unsafe action.

Contract tests make those changes visible before deployment. They exercise the boundary between an agent and a tool: the input and output shapes the tool accepts, the conditions under which it permits an action, and the failure behavior callers can rely on. This tutorial builds a small contract-test suite and connects it to CI. The examples are illustrative; adapt them to the language, authorization model and tool adapter in your system.

The scope here is deliberately narrow. It is not a guide to agent orchestration or approval workflows. For resilience to retries and restarts, pair these tests with Idempotent Tools for Agents: Surviving Retries and Restarts. For the larger workflow context, see AI Agents in Production: Long-Running Workflows.

1. Define the contract you intend to protect

A contract is the observable agreement at a boundary, not an implementation detail. For an agent tool, write down what a caller may submit, what a successful response means, which caller identities may perform which actions, and how the tool reports a rejected or failed request. A test can then detect a change to that agreement even when the changed code still runs.

Keep three concerns distinct. A schema test checks the structure and basic constraints of data. A permission test checks whether a particular identity may perform a particular action on a particular resource. An error-behavior test checks what callers can observe when the request is invalid, unauthorized, unavailable or rejected for a business reason. These tests complement one another; passing one category does not establish the others.

For example, a request can be perfectly valid JSON and still be unauthorized. A denied request can return a well-formed error object but accidentally make a downstream change before returning it. A successful response can match its schema while describing an operation that did not actually complete. Contract tests should make these boundaries explicit rather than treating “valid JSON” as proof of correctness.

JSON Schema describes object properties, required fields and whether additional properties are allowed; validating a schema does not authorize a business action. JSON Schema’s object reference is useful when defining those structural rules. The authorization decision belongs in the tool’s business and access-control logic, and needs its own tests.

If the tool is exposed through a protocol boundary, test the tool semantics your application promises rather than assuming a transport-level check covers them. In MCP, the host/client/server architecture is distinct from the transport and authorization concerns; those layers do not define the business meaning of a tool action. The MCP architecture specification provides the relevant architectural context.

Before writing tests, name the consumer of each contract. It may be an agent framework, a tool adapter, a service calling the tool, or a human-operated process. If there is more than one consumer, record which ones must remain compatible. A schema change that is safe for a newly deployed adapter can still break an older worker that remains active during a rollout.

flowchart TD
  accTitle: Contract test gate for an agent tool release
  accDescr: The gate tests both accepted and rejected calls, including whether failures caused an external effect. A schema-only pass cannot establish behavior compatibility.
  N0["Proposed schema change"] --> N1["Replay accepted inputs"]
  N1["Replay accepted inputs"] --> N2["Replay rejected inputs"]
  N2["Replay rejected inputs"] --> N3["Check error and side-effect contract"]
  N3["Check error and side-effect contract"] --> N4["Compare prior version"]
  N4["Compare prior version"] --> N5["Approve compatible release"]

The gate tests both accepted and rejected calls, including whether failures caused an external effect. A schema-only pass cannot establish behavior compatibility.

2. Prerequisites and a minimal test fixture

You need a tool boundary that can be called in a test environment, a stable way to supply caller identity, and a replaceable downstream dependency. The dependency could be a fake repository, a local test database, or a stubbed service client. The crucial property is that permission-denied and malformed-input tests can prove that no real external action occurred.

The examples below use Python, pytest and the jsonschema package. They illustrate a test shape, not a claim that any particular tool or protocol uses these interfaces. Install the packages in your project’s normal development environment, or substitute equivalent libraries already in use. Keep the schema and tests under version control beside the tool implementation so a contract change can be reviewed with the code that caused it.

Set up a layout such as this, adjusted to your repository conventions:

agent_tools/
  update_record.py
  contracts/
    update_record.schema.json
  tests/
    test_update_record_contract.py
    test_update_record_permissions.py
    test_update_record_errors.py

In the test fixture, inject a fake record store and a controlled identity. Do not let a contract test silently connect to production systems. A fake can record requested operations, allowing a test to assert both the returned response and whether a side effect was attempted.

A useful minimum test fixture has four properties. It creates the tool using the same validation and authorization code as production. It supplies caller identity through the same boundary the tool actually uses. It gives the tool a deterministic dependency that can succeed or fail on command. And it exposes a count or log of attempted external operations for negative assertions.

Avoid creating a second, test-only implementation of the tool’s rules. If the test duplicates permission conditions instead of exercising the production decision path, the implementation and test can drift together unnoticed. Stubs should replace external effects, not replace the behavior under test.

3. Step one: write down request and response shapes

Start with one representative tool, not every tool in the estate. Choose a boundary where a change could produce a meaningful caller failure, such as updating a customer record or creating an internal work item. Record the fields that are required, the types and constraints on each field, and whether unknown fields are rejected or tolerated.

The following illustrative schema represents an update request. The names and rules are examples, not recommendations for a particular business domain. The schema intentionally makes record_id and status required and disallows unrecognized top-level fields. That choice is a compatibility decision: a caller that starts sending an extra field will receive a validation failure until the contract is deliberately revised.

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "record_id": { "type": "string", "minLength": 1 },
    "status": { "type": "string", "enum": ["open", "closed"] },
    "reason": { "type": "string", "maxLength": 500 }
  },
  "required": ["record_id", "status"],
  "additionalProperties": false
}

Decide whether omitted and explicit null mean different things. In many systems, omission means “leave unchanged,” whereas null means “clear this value”; in others, null is invalid. A test should reflect the intended behavior rather than inherit an accidental language default. The same goes for empty strings, whitespace-only values, case sensitivity and boundary lengths.

Write a response contract as well. A successful response should distinguish the outcome callers may safely rely on from incidental implementation details. For instance, a stable result might include an outcome identifier and a documented status. Do not make a model-facing tool response imply that an external business result has been verified if the tool has only accepted or queued a request. The contract should represent what the tool knows at that point in the workflow.

Use a validator directly in tests where that is the same validation mechanism used by the tool, or test through the tool boundary if it is not. A compact schema test can check both accepted and rejected examples:

import json
from pathlib import Path
from jsonschema import Draft202012Validator

SCHEMA = json.loads(
    Path("agent_tools/contracts/update_record.schema.json").read_text()
)
VALIDATOR = Draft202012Validator(SCHEMA)

def test_minimal_update_request_is_valid():
    request = {"record_id": "R-104", "status": "closed"}
    assert list(VALIDATOR.iter_errors(request)) == []

def test_unknown_request_field_is_rejected():
    request = {
        "record_id": "R-104",
        "status": "closed",
        "priority_override": "urgent",
    }
    assert list(VALIDATOR.iter_errors(request))

These examples check the schema’s observable rules. They do not prove the tool uses this schema at runtime, nor do they establish that the caller is permitted to close R-104. Add an integration-level contract test that calls the actual tool boundary with a valid and invalid request. That test catches a common gap: a correct schema file exists, but runtime validation was bypassed or applied to a different object.

4. Step two: test compatibility, not just the happy path

A contract suite needs both positive and negative examples. Positive cases show that intended callers still have a usable route. Negative cases show that malformed or unsupported inputs are rejected predictably. A suite containing only invalid requests can pass while every useful call is broken; a suite containing only the happy path can miss dangerous acceptance changes.

Build a small input matrix. Include a minimal valid request, a valid request with every optional field, a missing required field, an invalid enum value, a wrong type, a boundary-length value and an unexpected property. Include null where it is relevant to the contract. Each case should test one meaningful condition so a failure report identifies the changed rule instead of presenting an opaque collection of errors.

For each invalid request, assert the externally visible contract: whether validation rejects it, what stable error category is returned, and whether no downstream operation was attempted. Avoid assertions on full error prose or stack traces. Those details can change without changing the caller-facing behavior, and tests coupled to them create noisy CI failures.

A compatibility table helps reviewers reason about proposed changes before editing a schema. “Safe” here is not universal: it depends on which callers remain active, whether they tolerate extra fields, and how the deployment is sequenced. Treat the table as a prompt for a decision, not an automatic versioning policy.

Proposed changeTypical caller riskTest or review question
Add an optional response fieldOlder consumers may reject unknown fieldsDo active consumers tolerate extra response properties?
Add a required request fieldExisting callers may omit itIs there a default or staged migration for every caller?
Narrow an accepted enumPreviously valid calls may failAre all producers restricted to the remaining values?
Rename or remove a fieldCallers may send or read the old nameCan producer and consumer changes be sequenced safely?
Loosen accepted inputInvalid or unintended requests may passDoes the tool still enforce the business rule separately?

The table’s point is that compatibility has direction. An optional response field can be harmless for tolerant consumers and breaking for strict ones. An extra accepted request field may improve forward compatibility but can also conceal a misspelled field. Pick the rule that fits the callers and make it testable.

When a contract change is intentional, update the schema, example fixtures and expectations in the same change. Ask the reviewer to identify affected callers and the rollout order. A green test suite should mean that the changed contract is understood, not that the tests were weakened until the new implementation passed.

5. Step three: test permissions as decisions with observable effects

A permission contract is a decision over at least an identity, an action and a target. Make each dimension visible in the test data. A vague “authorized user” fixture cannot catch a role-to-action mix-up. Prefer named identities and actions that communicate the boundary, such as reader attempting update, or record_owner updating a record outside their permitted scope.

Write tests for both allowed and denied cases. An allowed case should show that the intended identity can perform the intended operation and that the expected downstream call is made. A denied case should show that the tool returns its defined denial outcome and that no protected action was attempted. If permissions also depend on attributes such as tenant, account or record ownership, vary those attributes in separate cases.

A useful test table might include a permitted editor updating an in-scope record, a read-only identity attempting the same update, and an editor attempting to update a record outside the editor’s scope. Add an unauthenticated or unrecognized identity if that state is meaningful at the tool boundary. Do not collapse all these cases into one test: they represent different regressions and should yield useful failure names.

Here is a deliberately simplified test shape. make_tool and the result fields stand for the interfaces in your application; substitute actual ones. The important assertions are that the permission decision is exercised and that denied calls do not reach the fake dependency.

def test_reader_cannot_update_record(fake_store, make_tool):
    tool = make_tool(store=fake_store, identity="reader")

    result = tool.update_record({"record_id": "R-104", "status": "closed"})

    assert result.category == "permission_denied"
    assert fake_store.update_calls == []

def test_editor_can_update_in_scope_record(fake_store, make_tool):
    tool = make_tool(store=fake_store, identity="editor")

    result = tool.update_record({"record_id": "R-104", "status": "closed"})

    assert result.category == "success"
    assert fake_store.update_calls == [("R-104", "closed")]

This is not a complete authorization test if the tool’s real policy also depends on scope, ownership or a resource attribute. Extend the fixture to represent those facts and include a denied case for each relevant boundary. Keep permissions outside the schema: a field such as role supplied by the agent must not become proof of the caller’s actual identity merely because it validates structurally.

Test the point at which authorization occurs. If the implementation validates input before checking permissions, malformed requests from an unauthorized caller may produce a validation response. If it checks permissions first, the caller may receive a denial. Either sequence can be deliberate, but the returned behavior should not accidentally reveal data or vary unpredictably after a refactor. Record the intended precedence as part of the error contract where it matters.

6. Step four: standardize and test error behavior

A tool error is part of its interface. Callers need to distinguish at least the conditions that require different handling: malformed input, denied action, a business-level rejection, and an unavailable dependency. Use the categories your application actually supports. Do not invent a protocol status code or pretend every error is safely retryable.

For each category, decide what the tool returns, whether the failure is retryable by the caller, and whether any side effect may already have occurred. The last question is essential. A timeout after a downstream request does not, by itself, prove that the request failed to take effect. A test should not teach an agent to retry blindly just because an operation returned an error-shaped response.

Define a stable envelope only if it fits the existing tool contract. An illustrative internal result might have a category, a caller-safe message, and a retryable indicator where the tool can determine that property. The implementation must not set retryable to true merely because an exception was caught. If the outcome is unknown, represent that uncertainty explicitly in a way the actual caller understands.

Test each error category with a controlled dependency. For example, configure the fake store to raise a known connection failure and assert the tool maps it to the expected category. Configure a business rule to reject an update and assert it is not reported as a transient outage. Ensure that an unexpected internal exception does not leak implementation details into the caller-facing message, while preserving enough server-side diagnostic context for the service’s normal observability.

The contract suite should also test that errors are not silently converted to success. A response with a valid shape but the wrong outcome category is a behavioral failure. Likewise, if a tool returns a success result before the downstream operation has completed, its contract must say that it represents acceptance rather than completion. Keep these states distinct in both names and test expectations.

Include a test for partial uncertainty where it is plausible. A downstream call might time out after the remote system accepted the request. The safe contract may be “outcome unknown; reconcile before attempting again,” rather than “operation failed.” How the tool resolves that state depends on its implementation and business requirements. For retry and restart design, see Idempotent Tools for Agents: Surviving Retries and Restarts; contract tests should verify the tool’s chosen observable behavior, not promise exactly-once effects.

7. Step five: connect the suite to CI and review contract changes

Run contract tests whenever a change can affect the tool boundary: schema, validation code, permission policy, error mapping, adapter behavior or a dependency wrapper. In a CI system, make the test command part of the required checks for the code that owns the tool. The exact command and workflow syntax depend on your repository and CI provider; the principle is to make a contract failure visible before the change is accepted for deployment.

A small project might run a command such as pytest agent_tools/tests after installing its development dependencies. Treat that as illustrative, not a universal invocation. Keep the command consistent with the project’s existing test discovery and environment setup. A test that passes only on one developer’s machine is not a reliable release boundary.

Add useful failure output. Identify the contract, input case and expected versus actual category. For schema validation, report the failing field path where practical. For permission tests, use descriptive names that make the identity, action and scope clear. For error tests, report the category and side-effect assertion. Avoid dumping secrets, real customer data or full sensitive payloads into CI logs.

Make contract changes visible in code review. A reviewer should be able to see what callers will now send or receive, which permission boundary changed, and what happens on failure. If the schema is generated, test the generated artifact or validate that generation is deterministic; do not review a handwritten source file while CI executes a stale generated schema.

During a rolling deployment, old and new consumers or tool workers can coexist. Plan compatibility around that overlap. For example, add a new optional field before requiring a caller to send it; deploy compatible consumers; then make the field mandatory only after old callers have been removed or otherwise accounted for. The correct sequence depends on the runtime and deployment model, so make the assumptions explicit rather than relying on a single all-at-once release.

A test suite can also flag a potentially breaking schema diff without deciding whether it is forbidden. That distinction matters: a strict block on every contract change can encourage teams to bypass or weaken the guard. A more useful workflow identifies the changed rule, requires an intentional update to examples and consumer expectations, and routes high-impact changes to the appropriate review.

When tools are exposed through a host/client/server arrangement, decide which layer owns each assertion. A protocol client may verify that a request can be transported, while the tool’s own tests verify action semantics and permission outcomes. Avoid duplicating the same test at every layer unless each copy protects a distinct boundary. For architectural responsibilities, consult the MCP architecture specification; keep this suite focused on the application’s tool contract.

8. Worked hypothetical: a permission regression during a field change

Consider a hypothetical support tool that updates the status of a record. Its contract accepts record_id and status; an editor can update records in an assigned account, while a reader can only inspect them. The implementation is changed to add a reason field. During the same refactor, the team reorganizes how identity and account scope reach the update function.

In the failure path, the new schema test passes because reason is optional and the original required fields remain valid. A happy-path permission test also passes because its editor fixture uses a record within the editor’s account. But a scope regression causes the refactored policy check to confirm only the editor role and omit the account comparison. The editor can now update an out-of-scope record. Without a denied cross-scope case, the suite remains green.

A second problem appears when the downstream store times out after receiving the update. The tool maps all exceptions to a generic temporary_failure response that tells the caller it may retry. That response fails to distinguish a connection failure before submission from an unknown outcome after submission. If an agent retries, the external effect may be repeated unless the underlying operation has appropriate safeguards. The contract suite cannot make the operation exactly once; it can catch an unsafe claim about what the error means.

A recovery trace makes the intended checks concrete:

EventExpected observationPass condition
Valid editor updates an in-scope recordSuccess category and one fake-store callResponse and call match the contract
Reader attempts the same updatePermission-denied category and no store callNo protected action is attempted
Editor targets an out-of-scope recordPermission-denied category and no store callScope is part of the tested decision
Request includes an unsupported fieldValidation rejection and no store callRuntime boundary enforces the request shape
Store rejects before accepting a requestDefined dependency-error categoryCaller is not told the update succeeded
Store outcome is uncertain after timeoutUnknown or reconciliation-required outcomeCaller is not told to retry as though failure were certain

In this hypothetical, CI should fail when the out-of-scope test observes a store call. The fix belongs in the production permission decision, not in a test-only condition. Then rerun the permission cases to show that the in-scope editor remains allowed while the reader and out-of-scope editor remain denied.

For the timeout case, change the error mapping only after deciding what the tool can actually know and how the wider workflow will recover. If the tool has a reliable way to look up the operation’s outcome, test that reconciliation path. If it does not, the caller-facing result should preserve uncertainty rather than assert a definite failure. The relevant recovery policy may involve workflow state and restart behavior, which is covered more broadly in AI Agents in Production: Long-Running Workflows.

This scenario illustrates why a schema-only suite is insufficient. The request remains structurally valid throughout the permission regression. It also shows why an error-shape assertion without a side-effect check can miss a dangerous change. A useful contract test observes the decision at the boundary and the attempted action behind it.

9. Step six: diagnose failures without weakening the guard

When a contract test fails in CI, first identify which agreement changed: request shape, response shape, authorization outcome, error category or side-effect behavior. Compare the test fixture and implementation at that boundary. Do not immediately update the expected value just to get a green build; determine whether the implementation changed intentionally and whether every affected caller can tolerate it.

If a schema test fails because the runtime accepts a field the schema rejects, verify which validator is authoritative. The test may be checking a stale schema, or the implementation may be bypassing validation. If runtime behavior is intentionally more permissive, make that policy explicit and test it through the actual boundary. Do not leave two conflicting definitions without a clear owner.

If a permission test fails, inspect the identity and resource scope supplied to the tool. Confirm that the test exercises the production authorization path and that the fake represents the relevant relationship accurately. A permission test that fails only because its fixture lacks required identity attributes is not evidence that the production policy is wrong; it is evidence that the contract assumptions need clarification.

If an error test fails after a dependency update, determine whether the underlying failure is genuinely different or simply wrapped differently. Map dependency-specific exceptions into the stable categories your callers use at the adapter boundary. Keep that mapping narrow and test it with representative failures. Do not expose arbitrary exception text as a tool result or label every unfamiliar exception as retryable.

If tests are flaky, inspect shared state, timestamps, network access and order dependence. Contract tests should be deterministic enough to serve as a release guard. Replace uncontrolled dependencies with fakes at the boundary, and reserve tests against real external services for a separate environment with its own operational purpose. Do not quietly skip a failing contract test in CI without assigning an owner and a resolution path.

A test that is too strict can also block harmless change. If a test asserts an error message word-for-word, replace that with assertions on stable categories and required caller-visible fields. If it asserts private implementation call order, ask whether that order is part of the contract. Preserve assertions that protect actual behavior, especially denial without side effects, rather than removing them because they are inconvenient.

10. Use this printable contract worksheet

Before merging a tool change, complete this worksheet in the pull request or the team’s existing design record. It is intended to be copied into a review, not downloaded from a separate location. Keep answers concise but specific enough that another engineer can write or inspect the corresponding tests.

Request and response

  • Tool boundary named: Identify the callable tool and the consumers that depend on it.
  • Required fields listed: Record field names, types, constraints and the meaning of omitted or null values.
  • Unknown fields addressed: State whether extra properties are rejected or accepted, and why that choice fits current callers.
  • Success meaning defined: Say whether success means accepted, completed or verified; do not blur those states.
  • Positive and negative examples added: Include at least one valid request and representative invalid requests tied to specific rules.

Permissions and effects

  • Identity and scope represented: Name the caller, action and target attributes used in the real authorization decision.
  • Allowed case tested: Show that a permitted caller can make the intended operation.
  • Denied cases tested: Include the relevant role, scope or identity failures as distinct cases.
  • No-effect behavior asserted: For a denied or invalid request, verify that the protected dependency was not called.
  • Trust boundary checked: Do not accept caller-supplied fields as proof of identity or permission unless the production design explicitly establishes that trust.

Errors and release

  • Failure categories defined: Distinguish the error cases callers need to handle differently.
  • Retry claims justified: Do not call an outcome retryable when the tool cannot tell whether an external effect occurred.
  • Error response tested: Assert stable categories and caller-visible fields, not incidental prose or stack traces.
  • CI runs the contract suite: Confirm the test command runs in the required checks for changes to the tool boundary.
  • Compatibility and rollout considered: Identify old or concurrent consumers that may be affected by the change.
  • Intentional change reviewed: Update fixtures and contract expectations alongside implementation changes, with the affected callers and deployment sequence understood.

A completed worksheet is not proof that the tool is safe in every circumstance. It is a compact way to make the contract reviewable and to connect each important promise to an executable test. If the boundaries or test fixtures are hard to define, that is useful information: the tool may have hidden dependencies or ambiguous outcomes that need to be clarified before a schema edit becomes a production change.

11. Keep the test boundary focused

Contract tests work best when their purpose is precise. They should tell the team whether callers can still submit the intended data, whether the right callers can perform the action, and whether failures remain distinguishable and honest. They should not become a substitute for testing an entire agent workflow or every downstream service in one brittle test.

Some questions sit next to this work but deserve their own design treatment. If multiple agents or components share responsibility, coordination costs and boundaries need separate analysis; see Single Agent or Multiple Agents? Start with the Coordination Cost. If a person must review an action before it runs, approval must be bound to the exact payload and the execution path; Human Approval in an AI Harness: Pause, Review and Resume Safely addresses that distinct problem. For deciding when an agent should stop and request help, see When an Agent Should Stop and Ask a Person.

For teams building the surrounding CI and deployment controls, production platform engineering may be a relevant next step. The immediate implementation decision, however, is small: select one consequential tool boundary, define its request, permission and error contracts, and make CI fail on a real regression rather than on incidental wording.

The reliable starting point is a few high-signal cases: a valid call, a malformed call, a permitted identity, a denied identity or scope, and a failure whose effect is known or uncertain. Run them through the actual tool boundary with controlled dependencies. Once those tests reveal a stable contract, expand only where another caller, permission edge or error state introduces a distinct risk.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call