A cheaper model can lower the cost of one call while increasing the number of calls needed to finish the task. Another model is more expensive per token but needs less correction. The product question is the cost of delivering a useful outcome under the required quality standard.

Per-token pricing is a component cost. It does not capture retrieval, tool fees, failed attempts, user waiting, human review, or the maintenance needed to operate the workflow.

Define the outcome before the denominator

For a synthetic document-extraction service, an accepted outcome might be a complete record whose required fields pass validation and review. A model response, a parseable JSON object, and a usable record are different boundaries.

State whether the measure is cost per attempted task, accepted record, or completed user workflow. Retain failure counts and unsupported tasks. A low cost per accepted record achieved by rejecting difficult inputs can conceal an expensive experience for the users who still need help.

Use an explicit accounting example

Suppose a synthetic batch has 1,000 attempts, €200 in model and tool charges, and 900 accepted records. Automated variable cost per accepted record is €200/900≈€0.222. If review consumed 120 minutes at an illustrative fully loaded rate of €30 per hour, add €60; the total becomes €260/900≈€0.289.

These are invented accounting inputs, not provider prices or results from a deployed system. The example also excludes fixed infrastructure, engineering, and incident costs. Include those when they matter to the decision, and state the allocation method rather than hiding it inside one precise number.

Keep the request-level trace

Aggregate all calls under the same user task, including unsuccessful candidates and retries. Separate cached responses, batch work, and live generation. A tool fee or a long-running external query may dominate the model charge.

Record actual billable usage where available. Input length, output length, reasoning computation, cache rules, and provider pricing differ. Avoid estimating a current bill from an old price table. Cost analysis should identify the pricing date and model/service version.

Compare quality and completion together

A candidate that lowers cost while missing a critical field has not completed the same job. Compare accepted outcomes, critical failures, coverage, correction effort, and turnaround under the same contract. Use task-level pairing when possible.

If a human reviewer is part of the workflow, measure their service time and errors. A model that shifts effort into review can look efficient in the API dashboard while increasing total delivery cost.

Optimize the expensive work, not every token

Remove redundant retrieval or unnecessary calls, shorten repeated context where evidence remains sufficient, cache stable authorized material, and use an appropriately capable model for bounded steps. Official latency guidance describes several such engineering levers; whether they help depends on the task.

Caching needs source version, user scope, expiry, and invalidation. Model routing needs its own evaluation. Neither is a free optimization if it introduces stale data or confident failure on an unrecognized difficult case.

Keep the decision robust to volume changes

Separate fixed and variable costs, then examine plausible traffic, failure, and review-rate ranges. A low-volume internal workflow may be dominated by maintenance. A high-volume workflow may be dominated by inference or review. Avoid assuming that one architecture is economical at every scale.

The useful cost report stays next to the quality report. It asks what it takes to deliver the promised result, including the tasks that do not succeed, rather than rewarding the cheapest isolated response.

Sources and further reading