DeepSeek V4 Flash completed these business tasks for $0.00028 per success
DeepSeek V4 Flash had the lowest observed inference cost per successful task: $0.00028, with 89 of 90 runs passing. We measured all 990 included scored runs. Every request in the 990-run comparison has a reconciled Gateway bill. Muse Spark is excluded from this comparison because one timeout charge is unresolved. Its 90 original runs remain in the downloadable source results. This is a result on 30 original synthetic tasks, not a claim about every workflow.
One invoice email says $5,000. Another email with the same invoice number says $4,500. A person reading the messages can see what happened: the second explicitly replaces the first after a pricing error. But a useful AI workflow has to do more than find a plausible number. It needs to tell you which record is current, what changed, and where the answer came from.
That is the kind of work we wanted to price. We gave 11 models the same synthetic business tasks, checked their answers, and compared their bills. We included record extraction, invoice reconciliation, and support requests that required applying a policy to a customer's records. Every task had a defined correct outcome.
Qwen 3.8 Flash, GPT-5.6 Sol, Gemini 3.8 Flash, GPT-6 Astra tied for the highest observed success count: 90/90 each. The 990 included scored runs cost $11.82081. Separate development/preflight and pilot runs cost $0.74671 and are excluded from the rankings. The main comparison includes spend on scored failures.
Every included model passed between 87 and 90 of its 90 runs. This small synthetic task set separates inference prices much more clearly than it establishes differences in reliability.
Two invoices, one decision
Our duplicate-invoice example is deliberately small enough to inspect. The original email, email-dup-1, contains a $5,000 server license renewal. Its replacement, email-dup-2, has a later issue date and a note explaining the pricing error. Both carry invoice number TSI-1001.
The task asks for both amounts, the discrepancy, and the corrected version. Returning only $4,500 would leave out part of the request. Returning $9,500 as money owed would confuse the history of the invoice with the amount to pay. The answer also needs references to the records that support it.
These are original synthetic records. The task distinguishes a plausible response from a complete, correct business result.
DeepSeek V4 Flash made 4 inference requests and 5 tool calls, then passed every substantive check. Its bill was $0.00030.
GPT-6 Astra made 4 inference requests and 3 tool calls, then passed every substantive check. Its bill was $0.04277.
These are the first trials on the same task. Using unrounded charges, their bills differed by 140.8 times. Read the exact submissions, tool steps, and generation bills.
The useful part of a trace is the sequence of decisions. Did the model read both records? Did it notice the correction? Did it submit every field the task needed? And how many model calls did it take to reach the result? Looking at the final paragraph alone hides much of that work.
Add up the whole bill
A failed answer still consumes inference. That is why our main metric includes the spend on every attempt:
Cost per successful task = total inference spend ÷ successful runs.
For a made-up example, 100 attempts costing $2 with 80 successes means 2.5 cents per success. Dividing by all 100 attempts gives two cents per attempt, a different metric.
We count a successful run only when the submitted business result passes the task's deterministic checks. We also show success rate separately. A low cost per success can be useful evidence, but it is not enough to choose a model for a workflow where an incorrect payment or refund has serious consequences.

| Model |
Passed |
Cost / success |
Total bill |
| DeepSeek V4 Flash |
89/90 |
$0.00028 |
$0.02464 |
| GLM-5.3 Flash |
89/90 |
$0.00078 |
$0.06914 |
| Qwen 3.8 Flash |
90/90 |
$0.00084 |
$0.07538 |
| GPT-5.6 Luna |
88/90 |
$0.00107 |
$0.09425 |
| DeepSeek V4 Pro |
89/90 |
$0.00646 |
$0.57467 |
| GPT-5.6 Sol |
90/90 |
$0.00929 |
$0.83610 |
| Gemini 3.8 Flash |
90/90 |
$0.01122 |
$1.01019 |
| Qwen 3.8 Max |
88/90 |
$0.01249 |
$1.09892 |
| Claude Sonnet 5 |
87/90 |
$0.01625 |
$1.41372 |
| Claude Opus 5 |
89/90 |
$0.03439 |
$3.06108 |
| GPT-6 Astra |
90/90 |
$0.03959 |
$3.56271 |
Extraction: DeepSeek V4 Flash had the lowest observed cost per success among fully billed groups at $0.00025, passing 30/30 runs. Reconciliation: DeepSeek V4 Flash had the lowest observed cost per success among fully billed groups at $0.00024, passing 30/30 runs. Support: DeepSeek V4 Flash had the lowest observed cost per success among fully billed groups at $0.00034, passing 29/30 runs. Whole-task bootstrap intervals are shown in the chart and study tables; these point estimates alone do not establish a statistically reliable ordering.
The workflow breakdown matters because applications do different jobs. A model's overall score blends extraction, reconciliation, and support equally in this study. Your app may spend nearly all its time on one of those. Start with the closest workflow, then test your own examples before treating a broad average as a routing rule.
What the token price leaves open
Token prices alone cannot tell you how much context a model will request, how long it will reason, how many tools it will call, or whether its result will pass your checks.
Our price comparison puts two different observations next to each other. One is the advertised charge for a fixed illustrative basket of input and output tokens. The other is the observed inference cost per successful task. The fixed basket makes the price-list comparison understandable; it is not a guess at how many tokens an agent ought to use.

The cheapest advertised basket belonged to DeepSeek V4 Flash, at $0.00023. The lowest observed cost per successful task belonged to DeepSeek V4 Flash. The cheapest token basket also led task economics here. The two dollar amounts buy different things and should not be divided into a savings claim.
Low token prices help when a model completes the work reliably. Higher inference costs may be worthwhile when a model solves cases a cheaper option misses. Test which applies to your workflow.
Follow the spend on failures

Across the included models, $0.14744 went to runs that failed the checks. The confirmed total is $11.82081. 979 of 990 included runs passed. Failed attempts remain in every reported cost-per-success figure.
Some failures point toward a model choice. Others may point toward the task, the tool interface, or a limit that cuts off otherwise useful work. The right next action depends on what happened. Switching models does not repair an ambiguous business policy, and increasing a token limit does not make a wrong calculation correct.
For example, on sup-008, Claude Opus 5 made 3 inference requests and 7 tool calls, then failed: submission.params.decision: required property missing. Its bill was $0.06478. The complete submitted result and grader explanation are in the trace appendix.
We kept failed attempts in the saved results so those explanations can be checked. A timeout does not quietly disappear from the success rate. A missing billing record does not become a zero-dollar call. The published economics require reconciled charges; unresolved billing is a reason to withhold a cost claim.
The same successful runs, side by side
The shared-success comparison contains 81 matched runs per model. We requested 85, but only 81 of the 90 aligned task/trial positions passed for every included model. We kept the existing runs and added none. This post-hoc subset compares costs conditional on shared success; it cannot estimate reliability or replace the full cost-per-success results above. Matched costs and selection.
Run a smaller version inside your app
You do not need a large benchmark to begin. Pick a few recurring tasks that your team recognizes. Include the awkward cases: duplicate records, missing information, partial payments, and exceptions to the usual policy. Write down what a correct outcome requires before looking at model responses.
Give candidate models the same inputs and tools, start each trial fresh, and preserve the outputs. Check the substantive result: a refund's decision, amount, recipient, and policy evidence, or a reconciliation's matches and remaining balances.
Count the whole inference bill and report success rate alongside cost per success and latency. Repeat the tasks. When models look close, gather more representative examples before declaring a winner.
Our comparison uses 30 synthetic tasks with three independent trials for each of 11 models. The code, fixtures, graders, traces, and detailed protocol are ungated. The agent harness uses the same tools and per-task limits across models. In the original ten-model cohort, concurrency rose from two to eight after 134 of its 900 scored runs, preserving every attempt. Astra and GLM ran later with eight workers; Muse Spark is excluded from the main comparison. These changes and the unresolved timeout charge are recorded in the study. The appendix explains how we handle routing, caching, billing, and uncertainty.
This is still one bounded test. It does not include the labor to review an answer, the cost of external tools, or the consequences of a wrong action. Those can dominate inference costs in a real business. Treat these results as evidence for what to test next in your application.
Find out what your own AI workflows cost
MeterGraph helps you see the calls behind a workflow and what they cost. Start by instrumenting one workflow you already run. With a supported Node client, install the MeterGraph SDK, set an ingest key, and wrap the client where it is created:
npm install 'metergraph@^0.3.0'
export METERGRAPH_APP_TOKEN='<your ingest key>'
import * as mg from "metergraph";
import OpenAI from "openai";
const client = mg.wrap(new OpenAI());
// Run your existing provider call inside this trace.
const result = await mg.trace("resolve-ticket", () =>
resolveTicket(client, ticket)
);
await mg.flush(); // Before a short-lived process exits.
Here resolveTicket is your application's existing workflow. The benchmark also offers optional OTLP export for its FX calls. You can run the full benchmark with only a Gateway key.
Find out what your own AI workflows cost.
Study and limitations · Complete result tables · Code and raw results · Trace examples
Prepared for review. Public repository and article URLs will be added at launch; no signup is needed for the files.