Applied AI

What does an AI agent cost per successful task?

Calculate AI agent cost per successful task using all model calls, retries, tools and human review, then compare options without hiding failures.

All agent attempts and review costs flow into one cost-per-successful-task calculation.
All agent attempts and review costs flow into one cost-per-successful-task calculation.

The short version

Divide the full cost of a workflow by the number of tasks that meet an agreed acceptance rule. Count failed attempts, retries, tool use and human review in the numerator; do not call every model response a successful task.

Choose a business outcome before counting tokens

An agent can make three model calls and return a polished answer while the customer still has to reopen the ticket. If the job is resolving a support request, define success as an accepted resolution under a documented review rule, not merely a completed API call. For document intake, it might be a correct structured record that a reviewer can approve without rework. Different workflows need different units; do not compare their costs as though they were interchangeable.

Record an outcome for every eligible task: accepted, corrected by a person, escalated, abandoned or failed. Fix the observation window and the acceptance test before comparing two versions. The FinOps Foundation emphasizes cost per use case and unit of work because raw token totals do not say whether the spend produced the intended result.

Evidence: FinOps Foundation: FinOps for AI tools and services · FinOps Foundation: Forecast AI services costs

Count the whole workflow cost

For each task, add model input and output charges, retrieval and other tool charges, hosting or dedicated capacity, and the human time needed to review or repair the result. Allocate shared infrastructure on an explicit basis, such as active task time or measured usage. Keep one-time build costs separate from operating cost unless the decision genuinely needs a payback calculation.

Retries belong to the task that triggered them. A provider invoice may show inexpensive average calls, but a loop that repeats a retrieval and a model step three times can make the outcome expensive. Instrument the trace with task ID, model, token usage, tool calls, retry count, review time and final outcome. OpenTelemetry provides GenAI attributes for model and token usage; its documentation also warns that tool arguments and results can contain sensitive information, so collect only what the cost model needs.

Evidence: FinOps Foundation: FinOps for AI tools and services · OpenTelemetry: GenAI semantic convention attributes

Work through an illustrative calculation

Suppose a pilot receives 100 eligible tasks. Model calls and tools cost $48, review and correction time costs $20, and 80 tasks pass the acceptance rule. The operating cost per accepted task is ($48 + $20) / 80 = $0.85. The other 20 tasks still contributed cost; removing them from the numerator would make the workflow appear cheaper than it is. These figures are illustrative, not Dopstack client results or market benchmarks.

Now compare a cheaper model that cuts model charges but sends more tasks to review. Keep the same task mix, acceptance rule and time window. Compare cost per accepted task alongside quality, elapsed time and escalation rate. A lower token price is useful only if the total outcome improves under the constraints that matter to the business.

  • Use a stable, reviewable definition of success.
  • Include failed attempts and retries in total cost.
  • Price human correction time explicitly.
  • Compare quality and latency beside cost.

Evidence: FinOps Foundation: Forecast AI services costs

Turn the number into an engineering decision

Segment the result by use case, customer tier, task difficulty and model version before acting on an average. A few long documents or a bad retrieval path can dominate spend. If retries are the driver, fix timeouts or tool contracts before changing models. If review time dominates, improve evidence presentation and routing rather than compressing prompts blindly.

Recalculate after a material prompt, model, pricing or workflow change. Set a cost ceiling only with an associated quality floor and a plan for tasks the agent cannot complete. FinOps guidance frames unit economics as an ongoing measure of technology value; the number is most useful when it changes a specific operating choice.

Evidence: FinOps Foundation: FinOps for AI tools and services · FinOps Foundation: Forecast AI services costs

Sources & further reading

Written by Dopstack Technologies

We design and build software, cloud infrastructure and AI workflows. These notes explain engineering decisions; illustrative scenarios are not claims of client results.

Meet the team

Keep exploring.

Put the idea to work

Start with the problem in front of you.

Tell us where your system gets difficult. We’ll help you map the constraints and a practical first step.

Talk to an engineer