The short version
Divide the full cost of a workflow by the number of tasks that meet an agreed acceptance rule. Count failed attempts, retries, tool use and human review in the numerator; do not call every model response a successful task.
Choose a business outcome before counting tokens
An agent can make three model calls and return a polished answer while the customer still has to reopen the ticket. If the job is resolving a support request, define success as an accepted resolution under a documented review rule, not merely a completed API call. For document intake, it might be a correct structured record that a reviewer can approve without rework. Different workflows need different units; do not compare their costs as though they were interchangeable.
Record an outcome for every eligible task: accepted, corrected by a person, escalated, abandoned or failed. Fix the observation window and the acceptance test before comparing two versions. The FinOps Foundation emphasizes cost per use case and unit of work because raw token totals do not say whether the spend produced the intended result.
Evidence: FinOps Foundation: FinOps for AI tools and services · FinOps Foundation: Forecast AI services costs
Count the whole workflow cost
For each task, add model input and output charges, retrieval and other tool charges, hosting or dedicated capacity, and the human time needed to review or repair the result. Allocate shared infrastructure on an explicit basis, such as active task time or measured usage. Keep one-time build costs separate from operating cost unless the decision genuinely needs a payback calculation.
Retries belong to the task that triggered them. A provider invoice may show inexpensive average calls, but a loop that repeats a retrieval and a model step three times can make the outcome expensive. Instrument the trace with task ID, model, token usage, tool calls, retry count, review time and final outcome. OpenTelemetry provides GenAI attributes for model and token usage; its documentation also warns that tool arguments and results can contain sensitive information, so collect only what the cost model needs.
Evidence: FinOps Foundation: FinOps for AI tools and services · OpenTelemetry: GenAI semantic convention attributes
Work through an illustrative calculation
Suppose a pilot receives 100 eligible tasks. Model calls and tools cost $48, review and correction time costs $20, and 80 tasks pass the acceptance rule. The operating cost per accepted task is ($48 + $20) / 80 = $0.85. The other 20 tasks still contributed cost; removing them from the numerator would make the workflow appear cheaper than it is. These figures are illustrative, not Dopstack client results or market benchmarks.
Now compare a cheaper model that cuts model charges but sends more tasks to review. Keep the same task mix, acceptance rule and time window. Compare cost per accepted task alongside quality, elapsed time and escalation rate. A lower token price is useful only if the total outcome improves under the constraints that matter to the business.
- Use a stable, reviewable definition of success.
- Include failed attempts and retries in total cost.
- Price human correction time explicitly.
- Compare quality and latency beside cost.
Turn the number into an engineering decision
Segment the result by use case, customer tier, task difficulty and model version before acting on an average. A few long documents or a bad retrieval path can dominate spend. If retries are the driver, fix timeouts or tool contracts before changing models. If review time dominates, improve evidence presentation and routing rather than compressing prompts blindly.
Recalculate after a material prompt, model, pricing or workflow change. Set a cost ceiling only with an associated quality floor and a plan for tasks the agent cannot complete. FinOps guidance frames unit economics as an ongoing measure of technology value; the number is most useful when it changes a specific operating choice.
Evidence: FinOps Foundation: FinOps for AI tools and services · FinOps Foundation: Forecast AI services costs
Sources & further reading
- FinOps Foundation: FinOps for AI tools and services
Recommends cost per use case, token breakdown and retry visibility. Vendor examples on the page are not endorsements here.
- FinOps Foundation: Forecast AI services costs
Defines cost per unit of work and pairs cost with quality and latency; the arithmetic example is illustrative.
- OpenTelemetry: GenAI semantic convention attributes
Documents token and tool attributes plus sensitivity cautions. Instrumentation support varies by library and version.
Written by Dopstack Technologies
We design and build software, cloud infrastructure and AI workflows. These notes explain engineering decisions; illustrative scenarios are not claims of client results.
Meet the team