SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
How-to Guide 17 mins

AI Agent SLOs: Define Reliability in Terms of Completed Work

Set AI agent SLOs around verified completed work, with explicit latency budgets, intervention thresholds and a practical operational test.

The PADISO Team ·

Prerequisites

Before setting targets, identify the agent workflow, the people who rely on it, and the consequence of an incomplete or late result. Bring a representative sample of recent tasks, including ordinary cases, difficult cases and known failure cases. You do not need perfect historical telemetry to begin, but you do need a credible way to distinguish a finished task from a task that merely produced a plausible response.

Agree on the boundary of the service being measured. It might begin when an eligible request is accepted and end when a result is verified, an explicit exception is returned, or the request is abandoned. Choose one boundary for the first version and document it. Mixing submission, model response and business completion in the same metric makes the resulting target hard to interpret.

Confirm that you can record, at minimum, a task identifier, eligibility decision, start and end timestamps, final status, verification outcome and whether a person intervened. Keep the record proportionate to the workflow and apply your organization’s data-handling rules. The purpose is to measure service behavior, not to collect every prompt or reconstruct every conversation by default.

You should also know who can respond when the service misses its target. A threshold without an owner, an action and a communication path is a dashboard decoration. Assign an operational owner for the measured workflow and decide who can pause intake, switch to a manual route or accept a backlog while an incident is assessed.

Step 1: Define the unit of completed work

Start with the unit a user recognizes: one request, case, document, transaction or approved batch. Avoid using an agent turn, tool call, generated answer or successful API response as the unit unless that event itself is the value users need. A multi-step workflow can return several technically successful intermediate responses and still fail to complete one business task.

Write a completion statement with three parts: the required result, the evidence that verifies it, and the terminal states that are not success. For example, a reconciliation task might be complete when every eligible invoice is either matched to an identified purchase order with required fields recorded or routed to a named exception category. A message saying “matched” is not evidence if the record was not actually saved or the matching rule was not satisfied.

Separate success from a safe, useful stop. If the agent cannot resolve an ambiguous case but creates a correctly classified exception for a person, that may be a successful handoff for a workflow whose purpose includes triage. It should not be counted as an autonomous resolution. Track both outcomes so a rising handoff rate cannot disguise declining automation coverage.

Set eligibility before measuring performance. Exclude only cases that the service contract explicitly does not accept, and record exclusions with a reason. If difficult cases are silently omitted after processing starts, the denominator will become more favorable precisely when the system is struggling. A change to the eligibility rule should be versioned and reviewed alongside the SLO.

A service-level objective measures user-relevant behavior over a defined denominator and time window; decide those boundaries before interpreting the result (Google’s SLO implementation workbook). For an agent, that means writing down which accepted tasks count, how a completed task is verified, and whether the window is calendar-based or rolling. Treat the target as a service commitment, not as a claim about a model’s general intelligence.

Step 2: Specify success criteria that can be audited

Translate the completion statement into observable checks. A useful first version usually has a primary outcome and a small number of guardrail measures. The primary outcome answers, “Did an eligible task reach its defined terminal state?” Guardrails explain whether that apparent success was useful, timely and operationally sustainable.

For each task, define a status vocabulary that does not collapse different outcomes together. One practical scheme is verified_complete, verified_exception, pending, failed and cancelled. Keep a separate human_intervention flag or event. That preserves the distinction between “the service completed the work,” “the service safely routed the work,” and “a person had to finish it.” Choose names that fit your systems, but define them in plain language.

Verification should be tied to the business result, not to the agent’s description of what it did. Depending on the workflow, evidence might be a persisted record with expected fields, a reconciled total, a downstream acknowledgement, or a reviewer’s recorded decision. If a human review is the only reliable check, label the metric as human-verified and account for review delay. Do not treat a confident completion message as independent verification of that same completion.

Establish a rule for partial work. If a request contains ten line items and nine are correct, is the request complete, partially complete or failed? The answer depends on what the requester needs. One option is a task-level SLO that requires all required items to pass and a separate item-level measure for diagnosis. Another is a documented threshold for partial completion where the business process genuinely accepts it. Do not select the easier rule after seeing the results.

Make the acceptance test specific enough that two reviewers would usually classify the same task the same way. Record examples of pass, fail and handoff cases. If adjudicators disagree, resolve the rule before using the label to hold a team to a target. This prevents a numerical SLO from creating pressure to relabel uncertain outputs as successful.

Step 3: Choose a denominator and reporting window

The denominator is the population against which success is judged. A simple request-based measure might be: verified completed eligible tasks divided by eligible tasks accepted in the measurement cohort. State whether a timely human handoff counts as success for that measure. If it does, publish autonomous completion separately; if it does not, report successful handoff as a different service outcome.

Use an accepted-task cohort rather than only tasks that reached a final state. Otherwise, work that remains stuck or times out disappears from the calculation. Give unfinished work a defined treatment at the end of the observation period, such as counting it as not yet successful for the completion measure while also reporting its age. Explain how late completion is recorded: it can change the eventual outcome without erasing the fact that the original latency objective was missed.

Choose a window that matches the decision you will make. A rolling window can show recent deterioration without waiting for a month to close. A calendar window can align with operational reporting, but it can hide a bad final day inside an otherwise healthy month. Either can work if it is consistent, includes enough eligible work to be meaningful, and is accompanied by the actual count. A percentage without its numerator and denominator can make a handful of cases look like a stable trend.

For a new or low-volume workflow, do not create false precision by setting an ambitious percentage against a tiny sample. Report the count, review every miss and use a longer window or a clearly labeled provisional target until there is enough operating evidence. This is not a reason to ignore failures: a single high-impact error may require immediate action even when the aggregate percentage remains high.

Maintain distinct measures for completion, quality and timeliness. Combining them into one score makes tradeoffs invisible. A service might complete more tasks only because it accepts lower-quality matches, or meet a time objective by routing every case to a person. Separate measures let operators see the shift and decide whether it is acceptable.

Step 4: Set a latency objective users can feel

Measure elapsed time from the agreed start event to verified completion, not only the time to the first generated response. A useful operational view includes time to first useful progress, time to terminal status, and age of unresolved work. The headline objective should match the user promise; the supporting measures help locate delay.

Choose a percentile or another explicit rule that describes the experience you intend to protect. For example, a workflow might require a specified proportion of eligible tasks to reach a verified terminal outcome within its deadline over a rolling window. State the deadline and the population. Avoid reporting an average alone: a small group of extremely slow cases can disappear inside a favorable mean, while a percentile gives a clearer view of a portion of the distribution.

Set the deadline from the work’s actual operating context. A request needed before a scheduled downstream batch may have a different useful time limit from an interactive case that blocks an employee. Ask process owners when a result stops being useful, what fallback exists, and how long manual handling takes. Use those answers to define a service target, then validate whether the proposed system can meet it under representative load. Do not back into a user promise from a model’s best-case response time.

Break the end-to-end budget into stages for diagnosis, while keeping the end-to-end result authoritative. A sample budget might reserve time for intake and validation, retrieval or context preparation, agent execution, tool or system response, verification and a safety margin for scheduling variability. Those allocations are planning assumptions, not independent promises. If one stage consumes less time, the user has not received a faster completed task unless the end-to-end measure confirms it.

For multi-hop workflows, include the accumulated waiting and execution time across hops. Track where time is spent without making every internal component a separate SLO. Teams reviewing per-hop token and context constraints can use Token Budget Management Across Agent Hops as a related implementation resource; this article’s latency target remains the verified task outcome.

Decide how timeouts and pauses affect the clock. If a workflow waits for a person, report both end-to-end time and active service time if that distinction aids diagnosis, but do not remove the wait from the user-facing measure unless the service commitment explicitly excludes it. A pause can be operationally appropriate and still mean that the original time promise was not met.

Step 5: Build an intervention policy around the objective

An intervention threshold is a condition that prompts a defined operational response. It is not the same as the SLO target. Set thresholds for conditions that matter before the final reporting window shows a miss: a short-term rise in unresolved task age, a sudden increase in failed verification, a queue growing faster than it drains, or a spike in human handoffs. Pick signals that your team can observe and act on.

Assign each threshold an owner, an initial action and an escalation point. For example, an operator may inspect a sample of recent failures, stop admitting a non-urgent batch, or route new work to a documented manual path while the engineering owner investigates. The exact response depends on the workflow. The important design requirement is that the threshold does not simply generate an alert that nobody is expected to acknowledge.

Define thresholds using both rate and count when volume varies. A high failure percentage across a few requests may warrant case review but not the same response as a sustained failure rate across a large queue. Conversely, a low percentage can still represent an unacceptable number of delayed high-priority cases. Add a severity rule for specific outcomes that should trigger action regardless of aggregate performance.

Choose a response that preserves truthful measurement. Do not reset the window, delete the affected cohort or redefine the denominator to make an alert disappear. Record service changes and the time they took effect so operators can compare periods fairly. If intake is paused, distinguish requests accepted before the pause from requests not accepted by the service.

Operational context matters: observability practices can help teams connect task outcomes with traces and system behavior. For a deeper treatment of that implementation work, see AI Agents in Production: Agent Observability. Keep this SLO article focused on what counts as reliable completed work and what decision follows when the objective is at risk.

Step 6: Work a hypothetical service design

Consider a hypothetical invoice-reconciliation agent used by an accounts-payable team. Each accepted invoice is expected to be matched to a purchase order or placed into a review queue with a specific exception reason. The team’s operational requirement is to have routine invoices ready before its afternoon processing cutoff. These details are illustrative, not a report of a deployed system.

The first design decision is the unit: one eligible invoice. A verified autonomous success requires a persisted match with the required invoice identifier, purchase-order identifier and amount checks passing. A verified handoff requires a persisted exception category and a visible review item. A generated statement without the corresponding record is a failure. The team reports autonomous completion and successful handoff separately so that a shift toward manual review cannot be mistaken for improved automation.

Suppose, purely for illustration, the owners propose that at least 96% of eligible invoices achieve verified autonomous completion within 12 minutes over a rolling 28-day window. They also propose a separate handoff objective and a quality guardrail for incorrect persisted matches. Those figures are discussion inputs, not universal recommendations. The owners would need to validate them against actual request volumes, cutoff times, historical handling and the cost of a wrong match before adopting them.

For planning, they divide the 12-minute end-to-end allowance into intake and validation, context preparation, agent work and tool response, verification, and a residual buffer. The initial allocations add to the deadline, but they are not independent contractual limits. If the agent-work stage takes longer than its planning allowance, the team first checks the end-to-end completion time and task outcomes. It does not declare an SLO miss solely because an internal planning slice was exceeded.

The owners set an intervention signal when the unresolved queue’s oldest eligible invoice approaches the processing cutoff, or when verification failures rise above a reviewed operational threshold. The on-call operator inspects recent cases and can pause a non-urgent intake batch or invoke the existing manual route. The workflow’s service owner decides whether a broader change is needed. The team records the affected task cohort so a temporary fallback does not erase the original service performance.

This design forces useful choices into the open. A quick but incorrect match fails the quality guardrail. A correct match completed after the cutoff may pass eventual completion but miss the latency objective. A correctly classified exception can satisfy the handoff measure without inflating autonomous success. Each result points to a different operational decision.

Step 7: Trace a failure and recovery

Use a concrete failure timeline to test whether your definitions match operational reality. In the hypothetical invoice workflow, an invoice is accepted at 1:00 p.m. Context preparation completes, but a downstream record write does not return a clear confirmation. The agent reports a match, while the verification check cannot find the persisted record. At 1:12 p.m., the task has reached its latency deadline without verified completion.

At that point, the task must not be counted as a timely autonomous success. It remains unresolved or failed according to the state vocabulary, and its age continues to appear in the queue view. An operator checks the authoritative record and determines whether the write occurred. If the record is absent, the operator routes the invoice to the manual queue. The outcome becomes a verified handoff only when that queue entry and its exception reason are confirmed.

At 1:20 p.m., a reviewer completes the invoice manually. That later completion may update the final business outcome, but it does not retroactively satisfy the 12-minute objective. Keep the two facts: the work was eventually completed, and the service missed its timely completion commitment. This distinction lets the team improve recovery without turning late work into on-time work in the report.

The test should pass only if the recorded start time, deadline miss, failed verification, intervention, handoff and eventual disposition are all visible and classified consistently. It should fail if the agent’s response alone marks the task complete, if unresolved work vanishes from the denominator, or if a late manual finish erases the latency miss. It should also fail if an alert fires without an identified responder or next action.

This artifact tests measurement and response, not a particular recovery mechanism. Retry behavior and duplicate external effects require their own design treatment; see Retries Without Duplicate Actions: Idempotency for AI Agents for that separate concern. Here, record enough to determine whether the first attempt produced a verified result and what the service did when it did not.

Step 8: Draw the decision path and test its branches

The flow below shows the classification sequence for one accepted task. It distinguishes verification from a model’s own completion claim and routes unverified or late work to intervention rather than counting it as a success.

flowchart TD
accTitle: Classify an agent task against its SLO
accDescr: An accepted task is checked for verified completion. Verified work is then checked against its deadline. Unverified work goes to intervention, and late work is recorded as an SLO miss.
    A["Task accepted"] --> B["Check completion evidence"]
    B --> C{"Verified result?"}
    C -->|"Yes"| D{"Within deadline?"}
    C -->|"No"| E["Intervene and record unresolved"]
    D -->|"Yes"| F["Count timely success"]
    D -->|"No"| G["Record late outcome"]

The verification decision asks whether the agreed business evidence exists, not whether a response sounds complete. If the task is not verified, the operator follows the defined intervention route; the task remains visible while its final outcome is established. If evidence exists, the deadline decision is based on the recorded start and verified completion timestamps. A late but correct result contributes to eventual completion reporting but not to the timely-success measure.

Run the flow against a normal task, a missing-record failure, a valid exception handoff and a late completion. Include one case where the agent reports success but the evidence check fails. Review the resulting event record and dashboard classification with the people who will operate the workflow. If they disagree about what the diagram means, revise the state definitions before expanding the service.

Step 9: Review the target and preserve useful context

Review an SLO on a schedule tied to actual operational decisions, and after material changes to eligibility, workflow stages, verification rules or user deadlines. Compare outcomes by meaningful task class when one aggregate hides important differences. A change in case mix can make an unchanged system appear better or worse; record the composition of the measured population so the team can interpret a shift rather than merely react to it.

When an objective is missed, first establish which condition failed: completion, quality, timeliness or successful handoff. Then locate the affected cases, identify when the pattern began, and choose an action that addresses the observed failure. A slower context lookup calls for a different response than incorrect verification or an increase in ambiguous inputs. Avoid changing the target as the first response to a miss; revise it only when the user promise or service design has genuinely changed, and document the reason.

Use historical observations to calibrate provisional targets, but do not equate “what the system currently manages” with “what users need.” Compare the current distribution with the business deadline and available fallback. If the objective is consistently missed, determine whether to change the implementation, narrow the accepted workflow, improve the manual path or renegotiate the service promise. Each option has a different cost and should be an explicit decision.

Keep the record concise enough to maintain: the measured population, completion rule, verification source, denominator, window, latency deadline, quality guardrail, intervention thresholds, response owners and revision history. That record gives engineering and operations a shared definition when a graph looks healthy but users report unfinished work, or when an aggregate miss obscures a small number of urgent cases.

Practical worksheet and key takeaways

Use this worksheet in a design review or operating runbook. Fill it in with the people who own the workflow and the people who will respond to its failures. The example language is a prompt, not a preset target.

DecisionRecord for this workflow
Unit of workOne accepted request, case, document or transaction; define it precisely.
Eligibility ruleState which requests enter the measured population and why any are excluded.
Verified completionName the business result and the independent evidence that confirms it.
Safe handoffDefine what must be recorded before a human handoff counts as successful.
Primary measureSpecify numerator, denominator and whether handoffs are included.
Time windowChoose a rolling or calendar window and report task counts with the rate.
Latency objectiveSet the start event, verified end event and user-relevant deadline.
Quality guardrailIdentify an outcome that must not be traded away for speed or completion.
Intervention thresholdName the observable condition that prompts review or fallback.
ResponseAssign an owner, first action and escalation path.
Failure testCheck an unverified claim, unresolved task, valid handoff and late completion.
Review triggerSpecify when changes to workflow, population or user need require reassessment.

The most important design choice is to measure the work users need, not the activity the agent performed along the way. Define completion with evidence, preserve late and unresolved cases in the measurement, and report handoffs separately from autonomous results.

A useful latency objective starts and ends at events the business can recognize. Internal stage budgets support diagnosis, while intervention thresholds give operators time to act before a reporting window closes. Make each threshold actionable and test the response with a realistic failure timeline.

Treat targets as operational decisions that can be reviewed when the workflow or user need changes. For teams turning these definitions into service instrumentation and operating procedures, production platform engineering may be a relevant next step. Keep the SLO itself simple enough to explain, verify and use when a task does not go as planned.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call