SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Checklist 22 mins

Buying an AI Agent: Questions for Security, Procurement and Engineering

A practical buyer checklist for matching security, procurement and engineering evidence to an AI agent’s actual data, tools, autonomy and failure impact.

The PADISO Team ·

Buying an AI agent is not a single product decision. It is a decision about a particular system in a particular workflow: what information it can access, what actions it can take, who can intervene, and what happens when it gets something wrong. A convincing demonstration does not answer those questions. Neither does a long security questionnaire if its answers do not describe the proposed deployment.

Use this checklist to request evidence in proportion to the agent’s intended use. A read-only assistant with access to a small set of approved documents needs different proof from an agent that can change records, contact customers, or initiate transactions. The goal is not to collect the most documents. It is to resolve the uncertainties that could change the purchase, deployment boundary, or operating plan.

The checklist is organized around decisions. For each item, record the evidence received, the person responsible for assessing it, any unresolved gap, and the decision that follows. Ask for specific artifacts where possible: a system diagram, sample logs, a test result, a contract clause, or a named operational owner. Treat unsupported assurances as open questions rather than as evidence.

1. Write down what you are actually buying

  • Name the task, users, and business boundary. Describe the work the agent is expected to perform, who will use or be affected by it, and where the process begins and ends. “Automate customer operations” is too broad to evaluate. “Draft a response to a support ticket using the ticket and approved policy documents; a staff member sends it” is a scope that can be examined.

    Keep the first description in ordinary operational language. Include the current process, the intended change, the systems involved, and the human role before and after the agent acts. This helps procurement compare proposals against the same need and gives engineering a stable basis for testing. If the seller’s proposal quietly expands the task—for example, from drafting a response to sending it—record that as a scope change, not a minor configuration detail.

  • List the actions the agent may take, not just its advertised capabilities. For every action, write the target system, the kind of change, and whether it is read-only, preparatory, or externally effective. “Can access CRM” conceals important differences between reading a record, drafting an update, saving it, and triggering a downstream process.

    Ask the seller to distinguish what the system can technically do from what your proposed configuration permits it to do. A purchase decision should be based on the deployed boundary, not a broad feature list. If the boundary cannot be described clearly, use a narrower design or defer the decision until it can.

  • Identify who and what could be affected by an error. Record affected people, business records, customer commitments, financial activity, and downstream processes. Note whether an incorrect output is easy to spot before it matters, or could appear authoritative and move unnoticed into later work.

    This is a practical way to size evidence requests. A wrong internal summary that a trained employee reviews may justify a different acceptance threshold from an incorrect change that is automatically sent to a customer. Do not treat “human in the loop” as a complete risk description: specify what the person sees, what they must check, and whether they can stop the action in time.

  • Set the boundaries of the first deployment. State which teams, users, records, action types, and operating hours are included. State what is explicitly excluded. Ask whether a proposed trial or initial purchase can be kept within those boundaries without relying on informal promises.

    A narrow scope creates a meaningful basis for evaluation. It also makes it possible to distinguish a failure caused by the intended use from one caused by an unplanned expansion. If the vendor cannot explain how the proposed configuration corresponds to your boundaries, ask for a revised proposal before assessing the remaining evidence.

flowchart TD
  accTitle: Evidence gates for buying an agent
  accDescr: Commercial fit, security boundaries and engineering recovery evidence must all be reviewed. Unresolved material gaps return to the vendor before a purchase decision.
  A["Define purchase scope"] --> B["Review commercial evidence"]
  B --> C["Inspect access and data boundaries"]
  C --> D["Demonstrate failure recovery"]
  D --> E{"Material gaps resolved?"}
  E -->|"No"| F["Request evidence or decline"]
  E -->|"Yes"| G["Approve bounded engagement"]

The sequence gives each review function a concrete contribution. A convincing demonstration cannot replace missing contractual or access evidence, and a signed contract cannot prove the recovery path works. The final decision concerns the bounded engagement that was actually reviewed.

2. Ask for architecture evidence that answers your questions

  • Request a deployment-specific data and system diagram. It should show the agent’s relevant components, the information entering and leaving them, the systems it can reach, and the points where your organization or the supplier operates part of the service. Ask the supplier to identify any parts of the proposed flow that are not yet confirmed.

    A diagram is useful when it helps you ask concrete follow-up questions: Does a support ticket go to a supplier-operated service? Does the agent receive a full record or selected fields? Where is a proposed action reviewed? Avoid accepting a generic architecture illustration as a description of your configuration. Compare the diagram with the workflow you wrote in section one and mark every unexplained connection.

  • Request a field-level account of inputs and outputs. Ask which categories of information the system receives, which it can return, and which it can place into another system. Include prompts, retrieved material, generated content, tool requests, and results where those are part of the proposed design.

    Field-level detail matters because broad labels such as “business data” make it hard to decide whether a particular workflow is appropriate. Ask whether the agent needs each field for the task and whether a less sensitive substitute would work. If the supplier cannot provide a dependable account of data flows for the proposed use, treat that as a material deployment gap rather than filling it with assumptions.

  • Clarify the boundary between your organization and the supplier. Identify who configures the system, who controls connected accounts, who can change the workflow, and who is expected to respond when the service or an integration fails. Ask which parts of the arrangement depend on other providers or on components your team must operate.

    This is not an attempt to document every party in a complex supply chain. It is a way to establish who can answer specific questions and who can make a specific change. For each critical boundary, record the contact or role that owns the answer. If a supplier’s response depends on your team providing a control, make that dependency visible in the proposal and operating plan.

  • Compare the proposed configuration with the evidence offered. Check that diagrams, demonstrations, test reports, and contract descriptions address the same product arrangement and scope. Ask the supplier to identify any difference between the material reviewed and what you would buy.

    Evidence from a different deployment can still be informative, but it should not be mistaken for proof about your own setup. A feature may be available only in a different configuration; a test may omit a connected system; an operational explanation may describe a service your team is not purchasing. Ask what transfers and what does not, then keep the answer with the procurement record.

3. Examine data handling for the intended task

  • Ask what information is collected, retained, and used for the proposed service. Request a clear account of the relevant data categories and the supplier’s stated handling practices. Where a detail depends on a setting, contract, or product configuration, ask for the condition and how it will be verified.

    Keep the request tied to the workflow. A team evaluating a summarization task may need to know whether full source records are sent, what happens to submitted content, and who can access the resulting material. It may not need an abstract inventory of every data type the product could handle. Ask for enough specificity to decide whether the proposed inputs are appropriate and whether the arrangement matches internal requirements.

  • Check that access to information is no broader than the task needs. Ask which users, repositories, records, and fields are in scope, how access is assigned, and how the proposed arrangement changes when a user’s access changes. Request a configuration explanation or demonstration that answers those questions for the intended workflow.

    A useful review tests a simple boundary: could a user obtain information through the agent that they could not otherwise access in the intended process? If the answer is unclear, ask the supplier and your technical owner to trace a representative request from user identity through the information returned. Do not infer access behavior from a general claim that the system is secure.

  • Decide what the agent must not receive. Identify sensitive fields, restricted record types, or content that is unnecessary for the task. Ask whether the workflow can exclude them, and who verifies that exclusion when the connection or configuration changes.

    This is often more useful than asking for a sweeping assurance that all data is protected. A narrower input can reduce the consequences of an unexpected output or an access mistake, while also simplifying review. If the task cannot be completed without a category of information, document why it is needed and consider that fact when deciding whether the proposed use is acceptable.

  • Establish what happens to outputs and records of activity. Ask what is stored, where it is available to your team, and whether the information needed to investigate a disputed result will be available. Specify the records you need for the workflow, such as the input reference, proposed action, reviewer, approval, and outcome, rather than asking vaguely for “full auditability.”

    The right level of recordkeeping depends on how the output will be used. A draft that is discarded may need a different record from an action that changes a customer account. Ask how long relevant records remain available and what your team can retrieve, without assuming that a generated answer alone explains why a downstream change occurred.

4. Evaluate tool access and authority to act

  • Inventory every tool connection and its permitted operations. For each connected system, request the name or type of system, the data available, the operations permitted, and the intended purpose. Ask whether access is read-only or can change state, and whether the same connection could be used for actions outside the stated workflow.

    Judge the design by the authority it gives the agent, not by the marketing label attached to a connector. A connection used to look up a record has a different consequence from one that can update the record. If a proposed connection grants more authority than the task needs, ask whether the scope can be reduced or the action can be moved behind an explicit human step.

  • Define authorization boundaries for tool-connected work. For tool-connected deployments, assess token handling and confused-deputy risks explicitly; authorization boundaries matter, and tool output should not be trusted automatically. Model Context Protocol security best practices

    Ask the engineering owner to explain how identity, permissions, and the intended action are bound together in the proposed design. Check whether the agent can cause a system to act with authority that the requesting user should not have, and whether returned content can influence a later action without review. The answer should describe the actual boundary and the points where an action is constrained, rather than relying on a general statement that the system follows instructions.

  • Require review before consequential execution. If an action can affect a person, business record, or external process, specify who reviews it, what information they see, and what they must approve. Approval should precede execution, identify the exact payload to be sent, and expire if it is not used within the agreed window.

    A review that approves a general intention—“update the customer record”—does not necessarily approve a later payload with different fields or content. Make the reviewed action visible in a form the reviewer can understand, and bind approval to that action. If the agent can alter the payload after approval, the design has not preserved the approval boundary. For actions where timely review is impractical, consider a lower-impact mode or exclude the action from the first deployment.

  • Decide how untrusted content is handled. Identify whether the agent reads material that may contain instructions, such as user-submitted text or retrieved documents, and ask how the proposed workflow prevents that material from silently expanding what the agent is authorized to do.

    Do not accept “the model knows to ignore bad instructions” as a complete control description. Ask which actions are separately constrained and how a reviewer can tell what information informed a proposed action. This is particularly important when the system combines content from different sources with the ability to invoke tools. If the answer is unclear, keep the design read-only or require a human to move information into the action step.

5. Request evidence that tests the real workflow

  • Ask for an evaluation plan based on representative tasks. Request the task set, the inputs used, the expected outcomes, the kinds of failure being measured, and the person responsible for deciding whether results are acceptable. Ask how the proposed evaluation reflects ordinary cases as well as cases that could cause meaningful harm or rework.

    A polished demonstration usually shows selected interactions. It does not establish how the agent behaves across the variation your staff will encounter. Provide examples of real task patterns with sensitive details removed or replaced, where appropriate, and ask how the seller would test them. Your own technical and operational owners should agree on what counts as an acceptable result before reviewing results.

  • Require examples of failure, not only success. Ask what the system does when an input is incomplete, conflicting, out of scope, or unavailable. Request examples that show whether it stops, asks for clarification, returns an uncertain answer, or attempts an action despite missing information.

    The important question is not whether any system can avoid every error. It is whether predictable failure modes are visible and contained in the proposed workflow. Ask for the original input, the agent’s response, any tool action, and the human outcome for a small set of representative failures. If the supplier presents only successful examples, request the missing cases before using the demonstration as a purchase signal.

  • Test the handoff and stop conditions. Define what should happen when the agent cannot complete a task, a connected system is unavailable, a reviewer rejects a suggestion, or the output is not fit for use. Ask who receives the handoff and how the unfinished task is tracked.

    A workflow is not operationally complete merely because it can generate a response. Decide what status is recorded, how the human owner knows work remains, and whether a retry could create a duplicate external effect. Avoid assuming that repeated requests are harmless or that an external action happens exactly once. For consequential actions, design a visible reconciliation step that checks the resulting business state before the task is marked complete.

  • Define acceptance criteria before a trial or purchase decision. Specify task quality, review burden, acceptable failure classes, completion time if relevant, and the conditions that require a stop or scope change. State who measures each criterion and what evidence will be retained.

    Separate quality of generated content from the business result. A plausible response does not prove that the correct record changed, the right person received it, or the work was completed. Agree on outcome checks in the underlying process. For a broader discussion of integration and ongoing operations in an AI business case, see An AI Business Case That Includes Integration and Ongoing Operations; keep this purchase checklist focused on the evidence needed to judge the proposed agent.

The NIST AI Risk Management Framework is voluntary; it is not a certification or a guarantee of safe performance. NIST AI Risk Management Framework It can inform how an organization thinks about risk, but a reference to a framework should not replace evidence about this deployment. Ask what specific practice, artifact, or decision a supplier is pointing to, and whether it covers the workflow under review.

6. Examine operating readiness and failure handling

  • Name an operational owner and a technical owner. Identify who monitors the workflow, who can change its configuration, and who takes responsibility for incidents that involve both business process and technical behavior. Record a backup contact for each role.

    The owner does not need to be the person who built the system. They do need enough authority to pause the workflow, convene the right people, and decide whether it can resume. If the supplier expects your team to monitor performance or manage integrations, make that work explicit before signing. A responsibility that exists only in an informal conversation is easy to miss when the first incident occurs.

  • Agree on a pause and recovery procedure. Ask how the agent can be taken out of the workflow, what continues manually, and how the team identifies work that was incomplete or needs correction. Define what event triggers a pause and who can make that call.

    A useful procedure distinguishes three questions: how to stop new actions, how to account for actions already completed, and how to resume safely. Test the procedure as a tabletop exercise before expanding the scope. If turning off the agent also interrupts a process that has no manual fallback, that dependency should affect the purchase decision and rollout plan.

  • Specify incident evidence and response responsibilities. Decide what information the supplier and your team must be able to provide after a failure, who will investigate each part, and how affected business work will be reviewed. Request the supplier’s process for reporting relevant service issues and explaining changes that affect the deployed arrangement.

    Keep the request practical. The team needs a route to understand what happened, which tasks may be affected, and what action to take next. Do not accept an incident process that leaves your organization unable to identify the records or people requiring follow-up. Align the proposed response with the consequence of the workflow, rather than copying a process designed for a different service.

  • Set change-control expectations for material changes. Ask how you will learn about changes to the agent, its connections, or the workflow that could alter its behavior or evidence basis. Agree who assesses the change and whether re-evaluation is needed before the revised configuration is used.

    A purchase is not a one-time review of an unchanging object. A new connection, expanded access, altered task, or different review step can change what the organization is relying on. The contract and operating plan should give the right people a way to identify relevant changes. Not every update needs a new procurement cycle, but changes to action authority or affected populations should not pass without an owner’s review.

7. Make procurement evidence usable

  • Separate evidence from assurances and marketing claims. For each important claim, ask what artifact supports it, what scope it covers, and who can answer follow-up questions. Record claims that remain unsupported as unresolved items with an owner and decision date.

    An assurance can help locate the relevant evidence, but it should not close a material question by itself. For example, “the workflow is monitored” needs a description of what is monitored, who sees an alert, and what they do next. “The output is reviewed” needs a defined reviewer and a clear account of what approval permits. This makes supplier responses comparable and reduces the risk that different teams interpret the same phrase differently.

  • Check that the agreement matches the reviewed operating model. Compare the contract, proposal, configuration plan, and responsibility allocation. Look for differences in included services, supplier and customer responsibilities, change notification, access, support, data handling, and the scope of permitted use.

    Procurement should ask legal counsel to assess contractual language where appropriate; this checklist is not legal advice. The practical review is to avoid buying one arrangement while engineering evaluates another. If an important operating commitment appears only in a sales conversation, ask whether it can be reflected in an appropriate written document before relying on it.

  • Record dependencies and exclusions explicitly. List any controls, integrations, staffing, or manual steps that your organization must provide. Record features or use cases excluded from the proposal and what would require reassessment.

    This prevents an “included” label from obscuring work that has simply moved to your team. It also helps a buyer compare options fairly: a lower-complexity proposal may still require substantial internal effort, while a narrower proposal may be acceptable if the missing capability is not needed. Keep the list tied to the intended workflow rather than trying to settle every future possibility during the initial purchase.

  • Identify renewal and exit evidence. Ask what information your organization can retrieve for its operational records, what transition support is described, and which steps would be needed to stop using the agent. Consider whether the process can return to a manual method or another system without losing track of active work.

    This is not a demand to predict every future supplier change. It is a check that the team can end or pause the arrangement without leaving open tasks and unclear ownership. The exit plan should cover the practical workflow: stop new submissions, account for in-progress work, preserve needed records, and tell affected staff what process replaces the agent.

8. Worked example: a support-response agent

Hypothetical scenario: A mid-market company is considering an agent that reads incoming support tickets and approved policy material, drafts a reply, and offers it to a support representative. The representative—not the agent—sends the response. The company wants to evaluate the proposal for one support team and a defined ticket category before considering any wider use.

The buyer starts by writing the boundary in a single sentence: the agent may read the ticket and specified policy content, draft a response, and present it to an assigned representative; it may not send a message, change an account, or resolve a ticket. That description gives engineering something testable and gives procurement a concrete scope to compare with the proposal. If a supplier describes the same system as capable of handling tickets end to end, the buyer records that as a possible later scope, not as authority granted by the initial purchase.

The team then requests a diagram of the actual data path and a field-level description of ticket content and policy inputs. It asks which tool operations are enabled, how representatives see a draft, and what record shows whether a draft was accepted, changed, or discarded. It does not treat a sample response as proof of correct access or task completion. Those are separate questions requiring separate evidence.

For evaluation, the team selects representative ticket patterns and defines what a useful draft must do: address the request, stay within the approved policy material, avoid asserting facts absent from the ticket or policy, and make uncertainty visible enough for the representative to review. It also defines unacceptable outcomes, such as a draft that invents a customer commitment or includes information outside the intended record boundary. These are proposed acceptance criteria for this hypothetical design, not claims about any particular product’s performance.

The team checks the failure path. If policy material is missing, the system should not silently fill the gap with a confident answer. If the connected ticketing system is unavailable, the task must remain visible to staff rather than being treated as complete. A representative’s rejection should leave an understandable status, so the team can distinguish “draft not useful” from “draft never arrived.” Before any wider rollout, the owner would compare operational outcomes—not merely how fluent the drafts appear—with the agreed criteria.

Counterexample: Suppose the proposal changes during review so the agent can send replies automatically whenever its draft passes an internal confidence threshold. That is no longer the same purchase decision. The organization would need evidence about action authority, the exact execution boundary, failure handling, and how affected work is detected. If those questions remain unanswered, the sound decision is to keep the initial scope in draft-only mode or defer the expanded use. A confident-looking output is not a substitute for evidence about the consequences of sending it.

9. Printable buyer worksheet and responsibility matrix

Use this worksheet in a procurement meeting or print it as a record of the decision. Keep evidence references short but specific: a file name and section, a meeting date and owner, or a link to an approved internal record. Mark “not applicable” only with a reason. A blank field means the question is open, not that the risk is absent.

Decision areaEvidence or answer to recordOwnerStatus / next decision
Intended task and excluded uses
Users and affected parties
Data fields and sources in scope
Connected systems and permitted actions
Review step and exact approval boundary
Evaluation cases and acceptance criteria
Failure, pause, and recovery process
Supplier and internal responsibilities
Contract or configuration dependency
Unresolved gap and decision date

Complete the responsibility allocation in plain language. “Shared” is not enough unless the team can say who performs each step. The same person may hold more than one role in a smaller organization, but the role still needs a named owner and a backup.

ActivitySupplier contact or roleInternal ownerEvidence that closes the question
Explain the proposed data and system flow
Confirm access and action boundaries
Review workflow tests and failures
Pause the workflow and track open work
Assess a material configuration change
Review the purchase decision at renewal

Before the meeting ends, turn each unresolved item into a decision record. State the evidence still needed, who will obtain or assess it, the date it is due, and what decision follows if it remains unavailable. For example, the decision might be to narrow the use, add a manual review step, request a contract clarification, or decline the proposed scope. These are alternatives to consider, not automatic remedies; the right response depends on the consequence and the importance of the unanswered question.

10. Decide whether to proceed, narrow, or stop

  • Proceed only when the evidence matches the intended boundary. Confirm that the reviewed system, task, data, tool authority, human review, and operating responsibilities describe the same proposal. Record the acceptance criteria and the person accountable for checking them.

    “Proceed” can mean a limited evaluation or a defined operational deployment; state which one. A trial that uses different data, grants different permissions, or lacks the eventual review step may not tell you what you need to know about the intended design. Record those differences before treating a trial result as relevant to a purchase decision.

  • Narrow the scope when a specific gap can be avoided. Consider removing a data category, connection, user group, or consequential action while retaining the part of the workflow that can be evaluated responsibly. Ask the supplier and engineering owner to confirm the revised boundary and update the evidence record.

    Narrowing is a substantive design decision, not a way to relabel the original proposal. Confirm that the reduced scope actually removes the unresolved consequence. For instance, turning off an action may reduce what the agent can change, but it does not answer an unresolved question about information access. Recheck the remaining boundary rather than assuming one limitation resolves every gap.

  • Stop or defer when a material question cannot be answered. Do not treat an unanswered question as acceptable merely because the demonstration was persuasive, a deadline is near, or the contract is ready. State what would need to change for the decision to be reconsidered.

    A clear deferral is useful: it tells the supplier what evidence would matter and prevents internal teams from continuing on incompatible assumptions. If the uncertainty concerns a core part of the intended use, a smaller or different design may be more appropriate than proceeding with a promise to investigate later. Keep the decision proportionate, document the reason, and revisit it only when the missing evidence or proposed boundary has materially changed.

If your organization needs help translating a proposed use into a bounded technical design, evidence request, and operating decision, consider fractional CTO leadership. Where the main uncertainty is whether an AI initiative is ready to move beyond experimentation, the Ship-or-Kill Framework for AI Pilots is a related decision resource. For guidance on choosing between an advisor and a delivery partner, see Fractional CTO, AI Advisor, or Delivery Partner: Which One You Actually Need.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call