SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Tutorial 20 mins

Cost per Successful Agent Task: A Worked Cost Model

A worked, reproducible model for calculating agent cost per successful task—including retries, tool calls, human review and failed outcomes.

The PADISO Team ·

What this calculation answers

An agent’s cost per successful task is the total cost of operating it over a defined period divided by the number of tasks that meet a defined business acceptance test during that period. It is not the model’s price per call, the average token bill, or the cost of a task the agent merely attempted. The numerator includes the work consumed by successes and failures; the denominator includes only verified successes.

That distinction matters whenever an agent can retry, call external tools, or send uncertain work to a person. A workflow that appears inexpensive when measured by its first model response can become costly once repeated attempts and review time are counted. Conversely, a higher model charge can be economically sensible if it reduces expensive rework while maintaining the required outcome quality.

This walkthrough builds a monthly calculation from an event log, using explicitly hypothetical inputs. The example is not a price quote, performance claim, or prediction. Replace each assumption with measured data from your own workflow before using the result for a purchase, staffing decision, or production forecast.

The method is deliberately narrow: it helps compare the operating cost of a defined agent task. It does not establish whether the task should be automated, determine the value of the resulting business outcome, or prove that one model is better. Those decisions require separate evidence. For a related approach to measuring whether an agent actually completes its work, see AI Agents in Production: Agent Evaluation Frameworks.

Prerequisites and setup

Before calculating a unit cost, define the unit. A useful task definition names the event that starts one unit of work, the required result, the time window for completion, and the evidence that qualifies as success. For example, one unit might be a single customer-submitted invoice packet that must be classified and entered into an internal queue. A batch of ten packets is ten tasks, not one task, if each packet can independently succeed or fail.

Write the acceptance rule in a form an evaluator can apply consistently. “The agent handled the packet” is not an acceptance rule. “The required supplier, invoice date, total, and currency match the source document, and the record is present in the correct queue” is closer. If a person must correct the result before it is usable, decide whether that corrected outcome counts as success and record the human work needed to achieve it.

Choose a measurement window that captures ordinary variation. A calendar month is convenient, but a month with a system outage, unusual seasonal volume, or a changed workflow may not represent normal operation. Keep the observed window visible in the worksheet. If the workflow changes materially, start a new measurement period rather than blending unlike versions into one average.

Prepare an event record with at least one row per task and enough detail to reconstruct its attempts. Useful fields include task ID, start and end timestamps, workflow version, outcome, failure reason, model attempts, model cost, tool calls, tool cost, review minutes, reviewer outcome, and any allocated fixed operating cost. Store counts and monetary amounts at sufficient precision; round only when presenting the final result.

You also need cost assumptions that are explicit and attributable. Use actual invoices or internal accounting data where available. If a shared service supports several workflows, document the allocation rule rather than assigning its entire cost to one task. Unit economics relates costs to business units of value; assumptions and indirect costs belong in the calculation. FinOps Foundation: Unit Economics

For comparisons, hold the task definition, acceptance rule, observation window, and cost boundaries constant. A comparison that includes human review for one option but omits it for another is not an apples-to-apples result. Likewise, do not compare one system’s successful tasks with another system’s completed attempts.

Step 1: Define success before counting it

Create a short acceptance specification before looking at the aggregate cost. The specification should say what must be true, how it will be checked, who or what performs the check, and when the check occurs. Separate properties that are mandatory from preferences. If a mandatory field is wrong, the task fails even if the response is fluent or the tool call returned successfully.

For the hypothetical invoice workflow, define success as a record appearing in the intended queue with four required fields matching the source: supplier, invoice date, total, and currency. A task that reaches the queue with a wrong total is not a successful task. A task that is corrected by a reviewer may count as a successful delivered outcome, but its review minutes and any correction actions remain costs in the numerator.

Record outcome categories that expose what happened rather than hiding it in a single pass/fail number. For example: accepted without correction, accepted after human correction, unresolved after review, and abandoned or technically failed. These categories make it possible to report both the number of usable outcomes and the amount of labor required to produce them.

Set a task boundary for retries. An attempt is one execution of the agent workflow that can consume model or tool resources. A retry is another execution associated with the same task ID. Do not count retries as new tasks. If a person restarts a failed task, retain the original ID and log the additional work. Otherwise the numerator may record the extra cost while the denominator accidentally counts the same underlying task more than once.

Decide how to treat late outcomes. If a task started in the measurement month but was accepted the next month, choose a consistent rule: assign both cost and outcome to the start cohort, or report it as pending until resolved. For a first operational calculation, a start-cohort view is often easier to audit. State the cutoff date and avoid treating still-pending tasks as successes.

Step 2: Build a cost ledger that follows the task

The ledger should capture variable costs generated by each attempt and the human effort associated with the final outcome. It should also include the attributable operating costs that continue even when task volume is low. The aim is not to allocate every corporate expense to an agent. It is to include the costs a decision-maker would reasonably need to operate the measured workflow and to explain how each amount was assigned.

For the illustrative monthly example, assume 1,000 tasks enter the workflow. The observed event log contains 1,000 initial attempts, 200 first retries, 40 second retries, and 8 third retries. Total attempts are therefore 1,248. The retry counts are hypothetical inputs, not a statement about typical agent behavior. They represent an explicit sequence: some tasks require another attempt, and a smaller number require further attempts.

Assume, for calculation only, an average model execution cost of $0.012 per attempt. The model line is 1,248 multiplied by $0.012, or $14.976. This simplified average is suitable only if the cost figure already accounts for the relevant variation in usage and rates. In a real ledger, calculate from recorded usage and the applicable invoice or internal chargeback method instead of assuming every attempt costs the same.

Next, assume 2.4 billable or internally costed tool calls per attempt on average, at $0.004 per call. The estimated call count is 1,248 multiplied by 2.4, or 2,995.2 calls; the fractional count is an average across tasks, not a literal partial call. At the assumed unit cost, tool cost is 2,995.2 multiplied by $0.004, or $11.9808. Keep the call count, rate, and calculated amount separate so a reviewer can replace any assumption.

Do not equate a tool call with a successful external action. A call can fail, time out, return unusable data, or require a second call. Count the charges and any internally allocated tool-service cost that actually arose. If there is a separate fee for a particular operation, use the recorded fee or an explicitly documented allocation. Do not invent a universal tool-call price.

For people, assume all 1,000 tasks receive a final human review. Suppose 930 are accepted after an average of two minutes of review each, while 70 unresolved tasks require six minutes each to inspect and route. That is 1,860 minutes plus 420 minutes, or 2,280 minutes: 38 hours. At an assumed loaded labor rate of $42 per hour, review cost is $1,596. This rate and the review durations are illustrative; replace them with your organization’s compensation and time-accounting assumptions.

Finally, assume $180 of monthly workflow operating cost is attributable to this process. This could represent an allocation of ongoing hosting, monitoring, or support work, but the example does not assert any actual provider charge. Document what the $180 includes and why the allocation is fair. If the amount combines one-time implementation work with recurring operations, separate those categories rather than silently treating them as the same kind of cost.

Step 3: Calculate the successful-task cost

Add the cost lines for the measurement window before dividing: $14.976 in model execution, $11.9808 in tool calls, $1,596 in human review, and $180 in allocated operations. The total is $1,802.9568. The example has 930 tasks accepted under the specified outcome rule. Divide $1,802.9568 by 930, giving approximately $1.9387, or $1.94 per successful task when rounded to cents.

The denominator is 930, not 1,000, because 70 tasks remain unresolved under the example’s acceptance rule. The cost of attempting those tasks still belongs in the numerator. Excluding their cost would make failure look free; counting them as successful would make the outcome rate appear better than it is. Show both the numerator and denominator beside the rounded unit cost so the result remains inspectable.

The calculation can be represented as:

cost_per_successful_task =
    (model_attempt_cost
     + tool_call_cost
     + human_review_cost
     + allocated_operating_cost
     + other_included_costs)
    / accepted_task_count

For a cohort that includes one-time setup expense, add that expense only if the decision under consideration calls for it, and label the treatment. A launch decision may need to see setup cost separately from steady-state operating cost. A mature workflow comparison may instead report recurring cost per successful task and show setup as a distinct line. Do not divide one-time cost by an assumed future volume without displaying that volume and the period over which it is spread.

A small numerator does not guarantee a useful unit cost. If the denominator is low, a few outcomes can move the ratio substantially. Alongside the ratio, report task volume, success count, unresolved count, retry distribution, review hours, observation dates, and the acceptance-rule version. Readers need these figures to judge whether the result is stable enough for the decision at hand.

A separate rate can help distinguish outcome quality from expense: successful tasks divided by initiated tasks. In this example that is 930 divided by 1,000, or 93%. Keep the rate separate from the dollar metric. A lower cost per successful task can result from a genuinely cheaper workflow, but it can also result from a changed acceptance rule or an incomplete cost ledger.

Step 4: Reproduce the arithmetic from a task log

A spreadsheet is sufficient for a first calculation if each formula can be traced to a recorded quantity. A compact monthly summary might contain the following columns: cost category, quantity, unit rate, extended cost, data source, and inclusion note. For model attempts, quantity is the number of attempts and the rate is the measured average cost per attempt. For review, quantity can be minutes and the rate can be cost per minute. For fixed operations, quantity can be one month and the rate the documented allocation.

Keep the task-level log separate from the summary. A task-level table should make it possible to answer, for a selected task, how many attempts occurred, whether a tool was called, what the final reviewer decided, and how much review time was recorded. The summary should be derived from those rows where possible. If an amount comes from an invoice or a shared-service allocation rather than a task event, identify the source and allocation method in the summary.

An illustrative pseudocode calculation is:

# Illustrative only: replace these values with reconciled cohort data.
model_cost = sum(task.model_cost for task in tasks)
tool_cost = sum(task.tool_cost for task in tasks)
review_cost = sum(task.review_minutes for task in tasks) * cost_per_minute
operating_cost = monthly_allocated_operating_cost

successful = sum(task.accepted for task in tasks)
total_cost = model_cost + tool_cost + review_cost + operating_cost

if successful == 0:
    cost_per_success = None
else:
    cost_per_success = total_cost / successful

The expected output for the hypothetical inputs is a total cost of $1,802.9568, 930 accepted tasks, and a cost per accepted task of about $1.94. The separate success rate is 93%. These are arithmetic outputs from the assumptions above, not observed results. If your code returns a different value, first check whether review minutes were converted to hours consistently and whether the unresolved tasks were excluded from the denominator but retained in cost.

A useful reconciliation is to compare the model and tool totals with billing or internal usage records for the same window. Then compare review totals with time records or a sampling method. Finally, verify that the number of task IDs in the cohort matches the count of started tasks, and that every task has exactly one terminal category for the selected cutoff. The word exactly here describes the reporting rule, not a guarantee about external side effects: your event design must determine how duplicates and late events are handled.

Step 5: Read the result as a decision aid, not a score

The $1.94 figure says what the specified workflow cost per accepted task under the stated month, definition, and assumptions. It does not say whether that cost is good. To judge it, compare against the cost of the existing process or a relevant alternative using the same unit and outcome standard. Include the labor and operating costs of the alternative too; comparing a fully loaded agent workflow with only the software line of a manual workflow will bias the choice.

Separate measured values from assumptions. In the example, the attempt counts, call average, review times, labor rate, and monthly allocation were chosen for demonstration. In an actual decision, some may be observed directly and others estimated. Mark each as measured, allocated, or assumed, and prioritize replacing the assumptions with higher-quality observations when they materially affect the decision.

Sensitivity analysis is a simple way to identify which assumptions deserve attention. Recalculate the unit cost with review taking one minute longer per accepted task, with more tasks requiring a second attempt, or with a different defensible labor rate. Change one input at a time first, then test a combined scenario. The purpose is not to predict an exact future bill; it is to see whether the decision changes under plausible variations in the inputs your team has identified.

In this example, human review is the largest listed cost line. That does not prove review should be removed. Review may catch costly errors, and the acceptance rule may require it. Instead, it points to a concrete measurement question: how much review time is spent confirming correct fields, correcting errors, investigating ambiguous cases, and routing unresolved work? Those activities have different causes and may call for different workflow changes.

Model and tool charges are smaller in the illustrative calculation, but they should not be dismissed. A more expensive execution pattern can accumulate with volume, and tool failures may generate both repeated charges and staff investigation. Track cost by attempt and by outcome category so a change in retry behavior does not disappear inside a monthly average.

For model selection, task-level unit cost is one operational input among quality, latency, and other requirements. Published benchmark scores do not automatically predict the cost of your task under your own workflow. See Benchmarks That Actually Matter for New Model Releases for a separate discussion of selecting evidence that matches a decision. Comparisons can also change with the surrounding workflow and evaluation setup; Same Model, Different Harness: Why Benchmark Results Change treats that distinct issue. This walkthrough does not estimate model performance from benchmark scores.

Step 6: Make the event flow auditable

Use one stable task ID from intake through final disposition. Record each attempt as a child event linked to that ID, rather than overwriting the prior attempt’s status. A minimal sequence stores task creation, attempt start and end, tool call outcomes, review start and end, acceptance decision, and any final correction or escalation. This makes it possible to distinguish one expensive task from several cheap tasks without relying on memory or a monthly estimate.

The flow below shows where the unit-cost calculation receives its denominator and where failed outcomes remain visible. A task reaches the successful-task count only after its business acceptance check passes. Both successful and unsuccessful paths contribute their incurred costs to the aggregate.

flowchart TD
    A["Task arrives"] --> B["Run attempts and tools"]
    B --> C["Verify business result"]
    C --> D{"Meets acceptance rule?"}
    D -->|"Yes"| E["Count successful task"]
    D -->|"No"| F["Record unresolved outcome"]
    E --> G["Aggregate costs and outcomes"]
    F --> G

    accTitle: Agent task cost measurement flow
    accDescr: A task is executed and its business result is checked. Accepted tasks enter the success denominator; unresolved tasks remain failures. Costs from both paths are aggregated.

The business-result check is intentionally distinct from the model’s claim that it finished. A successful API response, a nonempty answer, or a tool call returning without an error can be useful operational evidence, but it is not necessarily evidence that the requested business outcome occurred. For the invoice example, verify the actual queued record and required field values rather than treating the agent’s final message as proof.

When events arrive late or more than once, define how the ledger resolves them. Deduplicate using stable event identifiers where available, and use a documented rule for repeated status updates. Preserve the raw event history if practical so corrections can be traced. Do not silently discard a second event merely because it looks redundant: it might represent a legitimate retry or a correction that has cost and outcome implications.

If the workflow permits a human to approve an external action, log the review decision and the exact task state being approved. The accounting question is not only whether an approval occurred, but whether it created extra time or changed the final outcome. Record human intervention as a cost even when it is an intended control rather than a failure.

Step 7: Test the boundaries with counterexamples

Consider a report that divides the model invoice by all 1,000 initiated tasks and announces a tiny cost per task. It omits tool charges, review labor, allocated operations, and the 70 unresolved outcomes. That number answers a different question: model expense per initiated task under a partial ledger. It cannot be presented as cost per successful task. Labeling the denominator and cost boundary correctly is more important than adding decimal places.

A second counterexample is an overly permissive success definition. Suppose a workflow counts a task as successful whenever the agent produces a structured response, even if the total is wrong or the record never reaches the queue. The success count may rise and the calculated cost may fall, but the underlying business work is not being completed to the required standard. The remedy is not a more elaborate formula; it is an acceptance test tied to the intended outcome.

A third trap is averaging away a high-cost minority. Suppose most tasks finish quickly but a small group triggers several retries and extended investigation. A monthly average may be valid for forecasting aggregate spend, yet it can conceal a queue of difficult cases that consume disproportionate staff time. Report the distribution of attempts and review minutes, or at minimum show ordinary, high-effort, and unresolved categories alongside the mean.

A fourth trap is mixing different task types. A simple invoice with all fields visible and a damaged, ambiguous scan may have very different review and retry profiles. If those types are combined, a shift in the mix can change the unit cost even when the workflow itself has not changed. Add a task-type field and compare like with like when the categories are operationally meaningful and have enough observations to be useful.

A fifth trap is counting a human-corrected result as a success without counting the correction work. That may be a valid definition of a delivered outcome, but the associated labor belongs in the numerator. Report whether success means accepted as produced or accepted after intervention. Keeping these outcomes distinct makes the figure more informative than forcing every organization into one universal definition.

Step 8: Troubleshoot discrepancies and failure patterns

If the reconstructed cost does not match a finance or usage total, first check the time boundary. Invoices may be grouped by billing date while task logs are grouped by task start date. Usage can cross midnight, and a task may finish after the cohort closes. Write down whether cost follows the event date, invoice period, or start cohort, then reconcile on that same basis before investigating individual discrepancies.

If model cost seems too low, check whether retries were logged as new executions or overwritten under the original task. Compare recorded attempt counts with usage records for the same workflow and period. Confirm that cached, background, or recovery executions are included if they consume billable or internal resources. Do not infer an unrecorded unit rate to force agreement; identify the missing quantity or document the remaining difference.

If tool cost seems implausible, distinguish calls attempted from calls charged and from successful external operations. A retry may repeat a charge even when it does not complete the business action. Conversely, a call may be internally operated and have no separate invoice line while still consuming infrastructure or support. State which costs the calculation includes, and avoid counting both an allocated service cost and a chargeback for the same underlying expense.

If review cost is far below what operators report, check what the time field represents. Review time can exclude waiting, correction, escalation, and context gathering. Decide which labor belongs to the workflow boundary and apply the same rule across alternatives. Where direct timing is unavailable, conduct a documented sample and label the resulting estimate rather than treating it as a measured total.

If the success count changes after a reporting run, inspect acceptance decisions and late updates. A reviewer may revise a disposition after discovering a field error, or a previously pending task may later be resolved. Preserve the decision timestamp and the rule version used. Recompute the cohort when definitions change, and do not compare an updated series with an earlier one without explaining the change.

A retry loop deserves a separate incident review when it consumes cost without improving the chance of acceptance. Trace the timeline for representative failures: initial attempt, reason for retry, information available to the next attempt, tool result, reviewer action, and final disposition. A retry that simply repeats the same invalid input is different from one that receives corrected data. The ledger quantifies the cost; the event sequence helps explain its cause.

Step 9: Use the result to choose the next measurement

Once the first calculation is reproducible, choose a specific uncertainty to reduce. If review dominates the numerator, sample review work and categorize minutes by verification, correction, ambiguity, and escalation. If retries dominate variable execution costs, record the triggering condition and whether the retry changed the inputs or action. If allocated operations are the least certain line, improve the allocation rationale before arguing over small model-cost differences.

Set a comparison period and acceptance rule before evaluating a proposed workflow change. Define what would count as a material improvement for your organization, including any minimum outcome-quality condition. A lower cost is not a success if the task no longer meets its required business standard. Keep the previous workflow’s original ledger so the comparison can be recreated after assumptions are revised.

Do not treat a short observation window as proof of stable unit economics. Review how task volume, task type, reviewer mix, and failure modes vary over time. If the workflow has limited volume, publish the raw counts and avoid implying precision from a ratio calculated on very few accepted tasks. The useful conclusion may be that the current evidence is not yet sufficient to select an option.

Benchmark design can help answer different questions, such as whether a model handles representative task inputs or whether a sampling strategy matches a workflow. Those questions should remain distinct from the cost ledger. For an adjacent evaluation-design decision, Pass@1 vs Best-of-N: Which AI Benchmark Matches Your Workflow? discusses a separate comparison framework. Do not transfer a benchmark result into the cost numerator or assume it establishes the successful-task denominator.

For teams defining a measurement plan, acceptance criteria, or a like-for-like operating comparison, AI evaluation and strategy is a relevant next step. The useful output of that work should be an auditable definition of the task, a cost boundary, and a plan for collecting the unresolved inputs—not a promised savings figure.

Printable worksheet

Use this worksheet as a concise record for one workflow and one measurement period. Fill in the fields before calculating the ratio. Preserve the completed version alongside the underlying task log, invoice references, and allocation notes so another person can reproduce the result.

Scope and outcome

  • Workflow name and version: ______________________________
  • Measurement window and cohort rule: ______________________________
  • Unit of work, including task start event: ______________________________
  • Required business outcome: ______________________________
  • Acceptance test and evidence source: ______________________________
  • Treatment of human-corrected outcomes: ______________________________
  • Treatment of pending and late outcomes: ______________________________

Counts and costs

  • Tasks started: __________
  • Accepted tasks: __________
  • Unresolved or failed tasks: __________
  • Total attempts, including retries: __________
  • Model cost and source: __________
  • Tool cost, call count, and source: __________
  • Human minutes, rate, and cost: __________
  • Allocated operating cost and allocation rule: __________
  • Other included costs and exclusions: __________
  • Total included cost: __________

Result and review

  • Successful-task cost = total included cost ÷ accepted tasks: __________
  • Success rate = accepted tasks ÷ tasks started: __________
  • Largest uncertain input: __________
  • Cost or outcome category needing investigation: __________
  • Acceptance-rule or workflow changes since prior comparison: __________
  • Reconciliation completed against usage, labor, and finance records: __________

For the hypothetical example in this article, the worksheet would show $1,802.9568 total cost, 930 accepted outcomes, and approximately $1.94 per successful task. Its assumptions include 1,248 attempts, 2,995.2 average tool calls, 38 review hours at $42 per hour, and $180 in allocated monthly operations. Those entries are placeholders for demonstration. A real worksheet should retain the actual source, date, and owner of every measured or allocated value.

The calculation is most useful when its boundaries are visible. Define the accepted outcome first, retain costs from failed work, include human effort and defensible operating allocations, and keep assumptions beside the arithmetic. Then the resulting unit cost can support a specific comparison without being mistaken for a universal model price or an outcome guarantee.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call