2026年10月7日 • Engineering • 6 分鐘閱讀

Measuring the Cost of a Complete AI Workflow

Measuring the Cost of a Complete AI Workflow

A model call is only one part of the cost of getting business work done. A workflow may retrieve records, call external tools, repeat a failed step, wait for review, and require someone to resolve an ambiguous result. A low price per call can coexist with an expensive completed task.

The useful unit is the outcome the team actually accepts. Define that outcome, measure the resources consumed while trying to reach it, and divide by the number of accepted results. This connects a cost estimate to the same work and quality conditions that justify using the system.

Choose a task and a measurement period

Start with a specific task: preparing a reviewed service summary, reconciling a defined set of records, or producing an evidence-backed answer for an employee. Specify what acceptance requires. A draft, a reviewed draft, and a delivered communication are different units of work.

Then choose a period and a currency. Include every attempt within that boundary, including rejected and unresolved outcomes. If a task continues into the next period, decide how its costs and eventual result will be attributed. Apply that rule consistently rather than moving expensive failures out of the calculation.

Use the same task definition for comparisons. A system that creates unreviewed summaries cannot be compared directly with a process that checks and delivers them. Changes to acceptance criteria should be reported as changes to the measurement, alongside any change in cost.

Collect the costs that follow the work

Model usage may be observable from provider records, but the rest needs deliberate accounting. Include retrieval and storage, external tool operations, retries, and the infrastructure supporting the workflow. Avoid counting the same shared expense in both an allocated infrastructure charge and a tool charge.

Measure human effort too. Review, correction, exception handling, and operational maintenance all consume time. Use an agreed planning rate and state whether it includes overhead. That rate is an estimate for this calculation, not a universal valuation of an employee's time.

Separate initial setup from ongoing operation. Mapping records, building integrations, preparing evaluation cases, and training reviewers may dominate the first period. Show those costs explicitly before choosing any allocation across later periods. Readers should be able to distinguish an observed operating expense from an assumption about future use.

Work through a transparent example

Suppose a team evaluates a hypothetical service-summary workflow for one week. It attempts 1,000 summaries and accepts 850 after review. All figures below are invented planning inputs in Canadian dollars. They are neither current model prices nor Accentrust performance results.

Weekly cost componentAssumed basisCost in CAD
Model processingAll attempts, including repeats120
Retrieval and tool operationsMeasured operations allocated to this workflow80
Review and correction30 hours at an assumed CAD 45 per hour1,350
Operational support5 hours at the same assumed rate225
Total ongoing costSum of the four components1,775

Dividing CAD 1,775 by 850 accepted summaries produces an ongoing cost of approximately CAD 2.09 per accepted summary. Dividing by all 1,000 attempts gives CAD 1.78 per attempt. Both are useful measurements, but they answer different questions.

If initial setup adds a hypothetical CAD 900, the first week's total is CAD 2,675, or approximately CAD 3.15 per accepted summary. That does not establish the cost of a typical later week. It makes the first period's additional expense visible.

The calculation should travel with its inputs: the task definition, period, attempt and acceptance counts, currency, included costs, and assumptions about human effort.

Keep unsuccessful work in the numerator

Retries, rejected drafts, and unresolved tool results still use resources. Excluding them understates what the team spent to produce the accepted outcomes. Track them as categories so their contribution can be explained and improved.

An unresolved request also needs a careful counting rule. If a tool call times out and its result has not been checked, it is premature to count the task as successfully completed. Conversely, a later verified result should not be counted twice because the original request was retried.

For tasks that correctly stop for human approval, record the system's disposition separately from business completion. A safe stop can be the desired behavior for an evaluation case while providing no accepted customer deliverable for a cost denominator. The accounting rule should follow the selected task boundary.

Separate cost from waiting and capacity

Price does not capture whether the workflow fits the working day. Measure elapsed time to acceptance and the effort required at each stage. A summary might take little active processing time but remain in a review queue until the next morning.

Inspect peak periods. If the same reviewer handles every exception, increasing throughput may increase waiting rather than accepted output. Treat reviewer availability and external service limits as operating constraints. Do not assume that doubling requests will double completed work at the same unit cost.

Waiting time should not be converted into a financial charge without an explicit reason and method. Show the timing measure directly when its business value is uncertain. This lets the responsible team assess a service delay without disguising an estimate as a recorded expense.

Test the assumptions that matter most

Use a sensitivity calculation to identify what could change the decision. In the hypothetical week, ten additional review hours add CAD 450. With all other inputs unchanged, ongoing cost becomes CAD 2,225, or approximately CAD 2.62 per accepted summary.

Alternatively, if acceptance falls to 700 while total ongoing cost remains CAD 1,775, the unit cost rises to approximately CAD 2.54. These isolated changes are exploratory calculations. A real reduction in acceptance might also change review effort, retries, and processing cost.

Testing such assumptions helps decide what to measure next. If review dominates the calculation, a small reduction in model expense may matter less than clearer evidence or fewer correction cycles. Quality and permission requirements still apply; reducing review by accepting unreliable results changes the task rather than improving its economics.

Give someone ownership of the estimate

Maintain a concise record with the task owner, period, input sources, calculation, exclusions, and the next review date. Update it when volumes, acceptance criteria, provider terms, or workflow steps change. Record a range when an important input remains uncertain.

Accentrust presents Signals as a platform capability concerned with operational understanding. A useful cost review turns that concern into concrete questions: what consumed resources, how much acceptable work resulted, and which assumption could alter the decision? Those questions support a better comparison than the price of an isolated call.

繼續閱讀

查看全部