Multi-cloudAI costs7 min read

Compare model quality and cost with a repeatable evaluation

Sources checked September 7, 2026Varies by scope
On this page
THE SHORT ANSWER

Select the least expensive model that passes the required quality and operational gates on representative tasks. Compare estimated scoped workflow cost per accepted result, not a single token rate.

Why this is worth a look

A lower token rate does not necessarily produce a lower application cost. Models may consume different token volumes, need different numbers of attempts, or involve priced tools, safeguards, modalities, tiers, or capacity arrangements. Evaluate with your own prompts, data, latency targets, and error tolerances rather than relying only on public benchmarks. Fixed tasks and explicit pass criteria make model changes comparable, while provider telemetry and applicable rates show the measured cost conditions.

Run this check

CHECKLIST

A read-only worksheet for one candidate model. Run it in your evaluation workspace, using application logs and provider telemetry for the stated window, then apply the candidate's applicable rate assumptions.

Repeatable model evaluation record
Evaluation ID and date: ____________________
Evaluation window and timezone: ____________________
Dataset version and immutable location: ____________________
Candidate provider, model, version, region, and deployment: ____________________
Prompt/template version: ____________________
Sampling and output limits: ____________________
Retry policy and maximum agent steps: ____________________

Cost scope and rate assumptions
Included components (model, tools, safeguards, retrieval, review, capacity): ____________________
Rate source, rate-card date, tier, region, and contract basis: ____________________
Excluded components and reason: ____________________

Required quality gates
[ ] task success definition documented
[ ] deterministic checks documented where possible
[ ] rubric and evaluator model/version documented
[ ] safety and policy checks documented
[ ] minimum passing score fixed before the run

Record per test case
[ ] test_case_id
[ ] pass or fail and reason
[ ] end-to-end latency and provider latency where available
[ ] time to first token where applicable
[ ] input and output tokens
[ ] cache token classes where the provider exposes them
[ ] model attempts, retry reasons, and tool calls from application logs
[ ] failed outputs and human-review outcome where used
[ ] estimated scoped cost under the stated rate assumptions

Aggregate for the evaluation window
pass rate: __________
p95 end-to-end latency: __________
model attempts per accepted result: __________
tool calls per accepted result: __________
tokens per accepted result: __________
estimated scoped workflow cost per accepted result: __________
failed cases requiring human review: __________

Decision, exceptions, and unresolved regressions: ____________________

How to confirm it

  1. 01

    Freeze representative tasks

    Build a versioned set of normal, difficult, and policy-sensitive tasks from approved patterns. Define expected results or scoring criteria before running candidates, and keep the dataset, prompts, tools, and retrieval fixtures identical for each comparison.

  2. 02

    Define gates before testing

    Set minimum task-success, safety, structured-output, and latency requirements. Use deterministic computation-based metrics where possible and fixed or documented rubrics where judgment is necessary. Record the evaluator model and version, and fix them for the run.

  3. 03

    Run identical configurations

    Use the same prompt templates, tools, retrieval fixtures, output limits, retry policy, and maximum agent steps. Record model version, region, deployment, and cache configuration. These controls keep the comparison interpretable across candidates.

  4. 04

    Capture available workflow usage

    Use provider telemetry and application logs for the evaluation window. Capture input and output tokens, cache token classes where exposed, provider latency and end-to-end latency, attempts, retry reasons, tool calls, and failed outputs. Provider metrics differ by cloud and deployment, so mark unavailable fields rather than estimating them silently.

  5. 05

    Compare scoped accepted-result cost

    Apply the stated contracted or public rates to measured units and included components, then divide estimated scoped workflow cost by accepted results. Reject candidates that miss a required gate even when their average cost is lower. Retain case-level results alongside aggregate scores so regressions can be reviewed.

Before making changes

Assumptions: the task set and configuration represent the intended workload, logs cover the stated evaluation window, and rates match the candidate's model, modality, region, tier, and capacity arrangement. This is estimated scoped workflow cost unless all relevant model, tool, safeguard, retrieval, review, and capacity charges are included. Re-run the suite before relevant changes reach production. Evaluators add cost, so validate the rubric on a reviewed sample.

Ignore this comparison when the model, configuration, workload, and applicable rates are fixed, and no alternative can be deployed or negotiated during the decision period.

Primary sources