Measured results
Every request in the 990-run comparison has a reconciled Gateway bill.
Original synthetic data. All 990 included scored runs, 11 exact model IDs. The 12-model source contains 1,080 runs; Muse Spark is excluded only from the comparison because its billing is incomplete. To reproduce the reports and figures, follow the instructions in the downloadable benchmark archive.
Actual charges include Sol's promotional Gateway pricing. See the billed-versus-market comparison before treating the ordering as a comparison at undiscounted rates.
Overall
| Model | Passed | Cost / success | Total bill |
|---|---|---|---|
| DeepSeek V4 Flash | 89/90 | $0.00028 | $0.02464 |
| GLM-5.3 Flash | 89/90 | $0.00078 | $0.06914 |
| Qwen 3.8 Flash | 90/90 | $0.00084 | $0.07538 |
| GPT-5.6 Luna | 88/90 | $0.00107 | $0.09425 |
| DeepSeek V4 Pro | 89/90 | $0.00646 | $0.57467 |
| GPT-5.6 Sol | 90/90 | $0.00929 | $0.83610 |
| Gemini 3.8 Flash | 90/90 | $0.01122 | $1.01019 |
| Qwen 3.8 Max | 88/90 | $0.01249 | $1.09892 |
| Claude Sonnet 5 | 87/90 | $0.01625 | $1.41372 |
| Claude Opus 5 | 89/90 | $0.03439 | $3.06108 |
| GPT-6 Astra | 90/90 | $0.03959 | $3.56271 |
Workflow estimates and uncertainty
| Model | Workflow | Passed | Cost / success | 95% cost interval | 95% success interval |
|---|---|---|---|---|---|
| GPT-5.6 Luna | overall | 88/90 | $0.00107 | $0.00101 to $0.00114 | 94.4% to 100.0% |
| GPT-5.6 Luna | extraction | 29/30 | $0.00102 | $0.00090 to $0.00120 | 90.0% to 100.0% |
| GPT-5.6 Luna | reconciliation | 30/30 | $0.00103 | $0.00098 to $0.00109 | 100.0% to 100.0% |
| GPT-5.6 Luna | support | 29/30 | $0.00116 | $0.00107 to $0.00128 | 90.0% to 100.0% |
| GPT-5.6 Sol | overall | 90/90 | $0.00929 | $0.00894 to $0.00969 | 100.0% to 100.0% |
| GPT-5.6 Sol | extraction | 30/30 | $0.00874 | $0.00813 to $0.00935 | 100.0% to 100.0% |
| GPT-5.6 Sol | reconciliation | 30/30 | $0.00979 | $0.00940 to $0.01020 | 100.0% to 100.0% |
| GPT-5.6 Sol | support | 30/30 | $0.00934 | $0.00857 to $0.01027 | 100.0% to 100.0% |
| Claude Opus 5 | overall | 89/90 | $0.03439 | $0.03195 to $0.03714 | 96.7% to 100.0% |
| Claude Opus 5 | extraction | 30/30 | $0.02905 | $0.02490 to $0.03281 | 100.0% to 100.0% |
| Claude Opus 5 | reconciliation | 30/30 | $0.03071 | $0.02692 to $0.03458 | 100.0% to 100.0% |
| Claude Opus 5 | support | 29/30 | $0.04373 | $0.03883 to $0.05162 | 90.0% to 100.0% |
| Claude Sonnet 5 | overall | 87/90 | $0.01625 | $0.01475 to $0.01804 | 91.1% to 100.0% |
| Claude Sonnet 5 | extraction | 30/30 | $0.01358 | $0.01209 to $0.01497 | 100.0% to 100.0% |
| Claude Sonnet 5 | reconciliation | 29/30 | $0.01488 | $0.01230 to $0.01892 | 90.0% to 100.0% |
| Claude Sonnet 5 | support | 28/30 | $0.02053 | $0.01730 to $0.02521 | 80.0% to 100.0% |
| DeepSeek V4 Flash | overall | 89/90 | $0.00028 | $0.00026 to $0.00030 | 96.7% to 100.0% |
| DeepSeek V4 Flash | extraction | 30/30 | $0.00025 | $0.00023 to $0.00027 | 100.0% to 100.0% |
| DeepSeek V4 Flash | reconciliation | 30/30 | $0.00024 | $0.00021 to $0.00027 | 100.0% to 100.0% |
| DeepSeek V4 Flash | support | 29/30 | $0.00034 | $0.00030 to $0.00039 | 90.0% to 100.0% |
| DeepSeek V4 Pro | overall | 89/90 | $0.00646 | $0.00601 to $0.00699 | 96.7% to 100.0% |
| DeepSeek V4 Pro | extraction | 30/30 | $0.00566 | $0.00515 to $0.00619 | 100.0% to 100.0% |
| DeepSeek V4 Pro | reconciliation | 29/30 | $0.00615 | $0.00543 to $0.00693 | 90.0% to 100.0% |
| DeepSeek V4 Pro | support | 30/30 | $0.00755 | $0.00657 to $0.00868 | 100.0% to 100.0% |
| Qwen 3.8 Flash | overall | 90/90 | $0.00084 | $0.00077 to $0.00091 | 100.0% to 100.0% |
| Qwen 3.8 Flash | extraction | 30/30 | $0.00071 | $0.00061 to $0.00080 | 100.0% to 100.0% |
| Qwen 3.8 Flash | reconciliation | 30/30 | $0.00074 | $0.00065 to $0.00083 | 100.0% to 100.0% |
| Qwen 3.8 Flash | support | 30/30 | $0.00107 | $0.00091 to $0.00123 | 100.0% to 100.0% |
| Qwen 3.8 Max | overall | 88/90 | $0.01249 | $0.01159 to $0.01343 | 94.4% to 100.0% |
| Qwen 3.8 Max | extraction | 28/30 | $0.01101 | $0.01012 to $0.01182 | 83.3% to 100.0% |
| Qwen 3.8 Max | reconciliation | 30/30 | $0.01144 | $0.01008 to $0.01284 | 100.0% to 100.0% |
| Qwen 3.8 Max | support | 30/30 | $0.01491 | $0.01290 to $0.01711 | 100.0% to 100.0% |
| Gemini 3.8 Flash | overall | 90/90 | $0.01122 | $0.01018 to $0.01239 | 100.0% to 100.0% |
| Gemini 3.8 Flash | extraction | 30/30 | $0.00989 | $0.00862 to $0.01130 | 100.0% to 100.0% |
| Gemini 3.8 Flash | reconciliation | 30/30 | $0.01027 | $0.00840 to $0.01299 | 100.0% to 100.0% |
| Gemini 3.8 Flash | support | 30/30 | $0.01351 | $0.01169 to $0.01569 | 100.0% to 100.0% |
| GPT-6 Astra | overall | 90/90 | $0.03959 | $0.03686 to $0.04237 | 100.0% to 100.0% |
| GPT-6 Astra | extraction | 30/30 | $0.03858 | $0.03357 to $0.04388 | 100.0% to 100.0% |
| GPT-6 Astra | reconciliation | 30/30 | $0.04146 | $0.03721 to $0.04614 | 100.0% to 100.0% |
| GPT-6 Astra | support | 30/30 | $0.03871 | $0.03505 to $0.04334 | 100.0% to 100.0% |
| GLM-5.3 Flash | overall | 89/90 | $0.00078 | $0.00073 to $0.00083 | 96.7% to 100.0% |
| GLM-5.3 Flash | extraction | 30/30 | $0.00068 | $0.00063 to $0.00074 | 100.0% to 100.0% |
| GLM-5.3 Flash | reconciliation | 30/30 | $0.00070 | $0.00063 to $0.00077 | 100.0% to 100.0% |
| GLM-5.3 Flash | support | 29/30 | $0.00096 | $0.00085 to $0.00107 | 90.0% to 100.0% |
Runtime and token usage
Input-token fields below use Gateway native prompt usage; cache reads and cache writes are separate fields. Reasoning is separately reported by Gateway. Provider token definitions can differ. Null means unavailable, never inferred zero. Usage totals for a model with unresolved billing cover only recorded usage and are incomplete; they are not totals for its unknown request.
| Model | Mean seconds | Tool calls | Prompt tokens | Output tokens | Reasoning | Cache read | Cache write |
|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash | 6.87 | 473 | 101601 | 69576 | 0 | 448256 | 0 |
| GLM-5.3 Flash | 11.58 | 401 | 323990 | 32854 | 5242 | 49664 | 0 |
| Qwen 3.8 Flash | 14.93 | 474 | 2178 | 57801 | 28317 | 485687 | 133946 |
| GPT-5.6 Luna | 8.93 | 378 | 117947 | 27053 | 8288 | 182047 | 98429 |
| DeepSeek V4 Pro | 12.78 | 442 | 247552 | 42702 | 17068 | 254909 | 0 |
| GPT-5.6 Sol | 10.3 | 387 | 122703 | 26317 | 3987 | 168313 | 101579 |
| Gemini 3.8 Flash | 17.75 | 480 | 601765 | 30156 | 118874 | 0 | 0 |
| Qwen 3.8 Max | 34.21 | 588 | 2304 | 71843 | 33912 | 537594 | 147358 |
| Claude Sonnet 5 | 12.63 | 384 | 556 | 58041 | 11637 | 423114 | 252484 |
| Claude Opus 5 | 12.44 | 435 | 604 | 62366 | 5308 | 547873 | 174763 |
| GPT-6 Astra | 14.15 | 392 | 137742 | 27685 | 0 | 210057 | 47279 |
Provider routing and billing
Provider counts refer to reconciled inference requests, including failures and retries. Gateway-internal attempts are also retained in finish metadata where observable.
| Model | Resolved provider | Requests |
|---|---|---|
| Qwen 3.8 Flash | alibaba | 363 |
| Qwen 3.8 Max | alibaba | 388 |
| Claude Opus 5 | claudeaws | 302 |
| Claude Sonnet 5 | claudeaws | 278 |
| DeepSeek V4 Flash | blackbox | 334 |
| DeepSeek V4 Pro | fireworks | 325 |
| Gemini 3.8 Flash | vertex | 387 |
| GPT-5.6 Luna | openai | 363 |
| GPT-5.6 Sol | bedrock | 1 |
| GPT-5.6 Sol | openai | 364 |
| GPT-6 Astra | azure | 1 |
| GPT-6 Astra | openai | 372 |
| GLM-5.3 Flash | baseten | 291 |
Setup work excluded from scored metrics
Development/preflight: 26 trials, including earlier runtime revisions. Pilot: 72 separate runs. Combined reconciled setup spend across all 12 tested models: $0.74671. All recorded charges remain in the source journals. Before the original scored runs, a reviewed budget-only amendment preserved the completed ten-model qualification work. Astra and GLM subsequently passed their own smoke checks and completed the same six-case pilot.
Concurrency phases
In the original 900-run study, the first 134 scored attempts used two workers and the remaining 766 used eight. The subsequent Astra and GLM extension ran all 180 attempts with eight workers using the same frozen benchmark. No active trial was aborted or replayed at the handoff. This is a chronological comparison with a different task mix, not a randomized test of worker count. The per-task request/output limits and deadlines stayed fixed.
| Model | Worker limit | Runs | Passed | Mean seconds | Spend |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | 2 | 13 | 13 | 6.20 | $0.00357 |
| DeepSeek V4 Flash | 8 | 77 | 76 | 6.98 | $0.02107 |
| GLM-5.3 Flash | 8 | 90 | 89 | 11.58 | $0.06914 |
| Qwen 3.8 Flash | 2 | 13 | 13 | 11.81 | $0.00955 |
| Qwen 3.8 Flash | 8 | 77 | 77 | 15.46 | $0.06583 |
| GPT-5.6 Luna | 2 | 14 | 13 | 7.61 | $0.01460 |
| GPT-5.6 Luna | 8 | 76 | 75 | 9.18 | $0.07965 |
| DeepSeek V4 Pro | 2 | 13 | 12 | 10.05 | $0.06534 |
| DeepSeek V4 Pro | 8 | 77 | 77 | 13.24 | $0.50934 |
| GPT-5.6 Sol | 2 | 13 | 13 | 11.53 | $0.11854 |
| GPT-5.6 Sol | 8 | 77 | 77 | 10.09 | $0.71756 |
| Gemini 3.8 Flash | 2 | 14 | 14 | 17.79 | $0.14002 |
| Gemini 3.8 Flash | 8 | 76 | 76 | 17.74 | $0.87017 |
| Qwen 3.8 Max | 2 | 13 | 12 | 29.54 | $0.14457 |
| Qwen 3.8 Max | 8 | 77 | 76 | 34.99 | $0.95436 |
| Claude Sonnet 5 | 2 | 13 | 13 | 11.24 | $0.19057 |
| Claude Sonnet 5 | 8 | 77 | 74 | 12.86 | $1.22316 |
| Claude Opus 5 | 2 | 14 | 14 | 11.42 | $0.44758 |
| Claude Opus 5 | 8 | 76 | 75 | 12.63 | $2.61349 |
| GPT-6 Astra | 8 | 90 | 90 | 14.15 | $3.56271 |
Reproduction fingerprint
Scored report SHA-256: 541e13cc381c6524a911a4a37c49fcdbdf88b2ac20b72390f6317b46a8a5db18.
Frozen runtime/input fingerprint: d56bab72dbd583e5b9658c4beff81c0393a25709a991ffca01f31fe947b14a5f.
Extension fingerprint: 6155047e4ebea551c0b3fd0d05339f24cc5dacdd33de6b8b69478bcc90513d6a.
Full protocol and limitations. Raw JSON report. CSV summary. Readable traces.