Measured results

Every request in the 990-run comparison has a reconciled Gateway bill.

Original synthetic data. All 990 included scored runs, 11 exact model IDs. The 12-model source contains 1,080 runs; Muse Spark is excluded only from the comparison because its billing is incomplete. To reproduce the reports and figures, follow the instructions in the downloadable benchmark archive.

Actual charges include Sol's promotional Gateway pricing. See the billed-versus-market comparison before treating the ordering as a comparison at undiscounted rates.

Overall

Model Passed Cost / success Total bill
DeepSeek V4 Flash 89/90 $0.00028 $0.02464
GLM-5.3 Flash 89/90 $0.00078 $0.06914
Qwen 3.8 Flash 90/90 $0.00084 $0.07538
GPT-5.6 Luna 88/90 $0.00107 $0.09425
DeepSeek V4 Pro 89/90 $0.00646 $0.57467
GPT-5.6 Sol 90/90 $0.00929 $0.83610
Gemini 3.8 Flash 90/90 $0.01122 $1.01019
Qwen 3.8 Max 88/90 $0.01249 $1.09892
Claude Sonnet 5 87/90 $0.01625 $1.41372
Claude Opus 5 89/90 $0.03439 $3.06108
GPT-6 Astra 90/90 $0.03959 $3.56271

Workflow estimates and uncertainty

Model Workflow Passed Cost / success 95% cost interval 95% success interval
GPT-5.6 Luna overall 88/90 $0.00107 $0.00101 to $0.00114 94.4% to 100.0%
GPT-5.6 Luna extraction 29/30 $0.00102 $0.00090 to $0.00120 90.0% to 100.0%
GPT-5.6 Luna reconciliation 30/30 $0.00103 $0.00098 to $0.00109 100.0% to 100.0%
GPT-5.6 Luna support 29/30 $0.00116 $0.00107 to $0.00128 90.0% to 100.0%
GPT-5.6 Sol overall 90/90 $0.00929 $0.00894 to $0.00969 100.0% to 100.0%
GPT-5.6 Sol extraction 30/30 $0.00874 $0.00813 to $0.00935 100.0% to 100.0%
GPT-5.6 Sol reconciliation 30/30 $0.00979 $0.00940 to $0.01020 100.0% to 100.0%
GPT-5.6 Sol support 30/30 $0.00934 $0.00857 to $0.01027 100.0% to 100.0%
Claude Opus 5 overall 89/90 $0.03439 $0.03195 to $0.03714 96.7% to 100.0%
Claude Opus 5 extraction 30/30 $0.02905 $0.02490 to $0.03281 100.0% to 100.0%
Claude Opus 5 reconciliation 30/30 $0.03071 $0.02692 to $0.03458 100.0% to 100.0%
Claude Opus 5 support 29/30 $0.04373 $0.03883 to $0.05162 90.0% to 100.0%
Claude Sonnet 5 overall 87/90 $0.01625 $0.01475 to $0.01804 91.1% to 100.0%
Claude Sonnet 5 extraction 30/30 $0.01358 $0.01209 to $0.01497 100.0% to 100.0%
Claude Sonnet 5 reconciliation 29/30 $0.01488 $0.01230 to $0.01892 90.0% to 100.0%
Claude Sonnet 5 support 28/30 $0.02053 $0.01730 to $0.02521 80.0% to 100.0%
DeepSeek V4 Flash overall 89/90 $0.00028 $0.00026 to $0.00030 96.7% to 100.0%
DeepSeek V4 Flash extraction 30/30 $0.00025 $0.00023 to $0.00027 100.0% to 100.0%
DeepSeek V4 Flash reconciliation 30/30 $0.00024 $0.00021 to $0.00027 100.0% to 100.0%
DeepSeek V4 Flash support 29/30 $0.00034 $0.00030 to $0.00039 90.0% to 100.0%
DeepSeek V4 Pro overall 89/90 $0.00646 $0.00601 to $0.00699 96.7% to 100.0%
DeepSeek V4 Pro extraction 30/30 $0.00566 $0.00515 to $0.00619 100.0% to 100.0%
DeepSeek V4 Pro reconciliation 29/30 $0.00615 $0.00543 to $0.00693 90.0% to 100.0%
DeepSeek V4 Pro support 30/30 $0.00755 $0.00657 to $0.00868 100.0% to 100.0%
Qwen 3.8 Flash overall 90/90 $0.00084 $0.00077 to $0.00091 100.0% to 100.0%
Qwen 3.8 Flash extraction 30/30 $0.00071 $0.00061 to $0.00080 100.0% to 100.0%
Qwen 3.8 Flash reconciliation 30/30 $0.00074 $0.00065 to $0.00083 100.0% to 100.0%
Qwen 3.8 Flash support 30/30 $0.00107 $0.00091 to $0.00123 100.0% to 100.0%
Qwen 3.8 Max overall 88/90 $0.01249 $0.01159 to $0.01343 94.4% to 100.0%
Qwen 3.8 Max extraction 28/30 $0.01101 $0.01012 to $0.01182 83.3% to 100.0%
Qwen 3.8 Max reconciliation 30/30 $0.01144 $0.01008 to $0.01284 100.0% to 100.0%
Qwen 3.8 Max support 30/30 $0.01491 $0.01290 to $0.01711 100.0% to 100.0%
Gemini 3.8 Flash overall 90/90 $0.01122 $0.01018 to $0.01239 100.0% to 100.0%
Gemini 3.8 Flash extraction 30/30 $0.00989 $0.00862 to $0.01130 100.0% to 100.0%
Gemini 3.8 Flash reconciliation 30/30 $0.01027 $0.00840 to $0.01299 100.0% to 100.0%
Gemini 3.8 Flash support 30/30 $0.01351 $0.01169 to $0.01569 100.0% to 100.0%
GPT-6 Astra overall 90/90 $0.03959 $0.03686 to $0.04237 100.0% to 100.0%
GPT-6 Astra extraction 30/30 $0.03858 $0.03357 to $0.04388 100.0% to 100.0%
GPT-6 Astra reconciliation 30/30 $0.04146 $0.03721 to $0.04614 100.0% to 100.0%
GPT-6 Astra support 30/30 $0.03871 $0.03505 to $0.04334 100.0% to 100.0%
GLM-5.3 Flash overall 89/90 $0.00078 $0.00073 to $0.00083 96.7% to 100.0%
GLM-5.3 Flash extraction 30/30 $0.00068 $0.00063 to $0.00074 100.0% to 100.0%
GLM-5.3 Flash reconciliation 30/30 $0.00070 $0.00063 to $0.00077 100.0% to 100.0%
GLM-5.3 Flash support 29/30 $0.00096 $0.00085 to $0.00107 90.0% to 100.0%

Runtime and token usage

Input-token fields below use Gateway native prompt usage; cache reads and cache writes are separate fields. Reasoning is separately reported by Gateway. Provider token definitions can differ. Null means unavailable, never inferred zero. Usage totals for a model with unresolved billing cover only recorded usage and are incomplete; they are not totals for its unknown request.

Model Mean seconds Tool calls Prompt tokens Output tokens Reasoning Cache read Cache write
DeepSeek V4 Flash 6.87 473 101601 69576 0 448256 0
GLM-5.3 Flash 11.58 401 323990 32854 5242 49664 0
Qwen 3.8 Flash 14.93 474 2178 57801 28317 485687 133946
GPT-5.6 Luna 8.93 378 117947 27053 8288 182047 98429
DeepSeek V4 Pro 12.78 442 247552 42702 17068 254909 0
GPT-5.6 Sol 10.3 387 122703 26317 3987 168313 101579
Gemini 3.8 Flash 17.75 480 601765 30156 118874 0 0
Qwen 3.8 Max 34.21 588 2304 71843 33912 537594 147358
Claude Sonnet 5 12.63 384 556 58041 11637 423114 252484
Claude Opus 5 12.44 435 604 62366 5308 547873 174763
GPT-6 Astra 14.15 392 137742 27685 0 210057 47279

Provider routing and billing

Provider counts refer to reconciled inference requests, including failures and retries. Gateway-internal attempts are also retained in finish metadata where observable.

Model Resolved provider Requests
Qwen 3.8 Flash alibaba 363
Qwen 3.8 Max alibaba 388
Claude Opus 5 claudeaws 302
Claude Sonnet 5 claudeaws 278
DeepSeek V4 Flash blackbox 334
DeepSeek V4 Pro fireworks 325
Gemini 3.8 Flash vertex 387
GPT-5.6 Luna openai 363
GPT-5.6 Sol bedrock 1
GPT-5.6 Sol openai 364
GPT-6 Astra azure 1
GPT-6 Astra openai 372
GLM-5.3 Flash baseten 291

Setup work excluded from scored metrics

Development/preflight: 26 trials, including earlier runtime revisions. Pilot: 72 separate runs. Combined reconciled setup spend across all 12 tested models: $0.74671. All recorded charges remain in the source journals. Before the original scored runs, a reviewed budget-only amendment preserved the completed ten-model qualification work. Astra and GLM subsequently passed their own smoke checks and completed the same six-case pilot.

Concurrency phases

In the original 900-run study, the first 134 scored attempts used two workers and the remaining 766 used eight. The subsequent Astra and GLM extension ran all 180 attempts with eight workers using the same frozen benchmark. No active trial was aborted or replayed at the handoff. This is a chronological comparison with a different task mix, not a randomized test of worker count. The per-task request/output limits and deadlines stayed fixed.

Model Worker limit Runs Passed Mean seconds Spend
DeepSeek V4 Flash 2 13 13 6.20 $0.00357
DeepSeek V4 Flash 8 77 76 6.98 $0.02107
GLM-5.3 Flash 8 90 89 11.58 $0.06914
Qwen 3.8 Flash 2 13 13 11.81 $0.00955
Qwen 3.8 Flash 8 77 77 15.46 $0.06583
GPT-5.6 Luna 2 14 13 7.61 $0.01460
GPT-5.6 Luna 8 76 75 9.18 $0.07965
DeepSeek V4 Pro 2 13 12 10.05 $0.06534
DeepSeek V4 Pro 8 77 77 13.24 $0.50934
GPT-5.6 Sol 2 13 13 11.53 $0.11854
GPT-5.6 Sol 8 77 77 10.09 $0.71756
Gemini 3.8 Flash 2 14 14 17.79 $0.14002
Gemini 3.8 Flash 8 76 76 17.74 $0.87017
Qwen 3.8 Max 2 13 12 29.54 $0.14457
Qwen 3.8 Max 8 77 76 34.99 $0.95436
Claude Sonnet 5 2 13 13 11.24 $0.19057
Claude Sonnet 5 8 77 74 12.86 $1.22316
Claude Opus 5 2 14 14 11.42 $0.44758
Claude Opus 5 8 76 75 12.63 $2.61349
GPT-6 Astra 8 90 90 14.15 $3.56271

Reproduction fingerprint

Scored report SHA-256: 541e13cc381c6524a911a4a37c49fcdbdf88b2ac20b72390f6317b46a8a5db18.

Frozen runtime/input fingerprint: d56bab72dbd583e5b9658c4beff81c0393a25709a991ffca01f31fe947b14a5f.

Extension fingerprint: 6155047e4ebea551c0b3fd0d05339f24cc5dacdd33de6b8b69478bcc90513d6a.

Full protocol and limitations. Raw JSON report. CSV summary. Readable traces.