Shared-success matched costs
The shared-success comparison contains 81 matched runs per model. We requested 85, but only 81 of the 90 aligned task/trial positions passed for every included model. We kept the existing runs and added none. This post-hoc subset compares costs conditional on shared success; it cannot estimate reliability or replace the full cost-per-success results in the full study.
All dollar amounts are actual Gateway charges. Sol's charges reflect promotional rates; this is not an undiscounted market-rate comparison. Billing details.

| Model | Matched runs | Total cost | Cost / included run |
|---|---|---|---|
| DeepSeek V4 Flash | 81 | $0.02195 | $0.00027 |
| GLM-5.3 Flash | 81 | $0.06191 | $0.00076 |
| Qwen 3.8 Flash | 81 | $0.06776 | $0.00084 |
| GPT-5.6 Luna | 81 | $0.08475 | $0.00105 |
| DeepSeek V4 Pro | 81 | $0.51572 | $0.00637 |
| GPT-5.6 Sol | 81 | $0.75519 | $0.00932 |
| Gemini 3.8 Flash | 81 | $0.89666 | $0.01107 |
| Qwen 3.8 Max | 81 | $0.99063 | $0.01223 |
| Claude Sonnet 5 | 81 | $1.25760 | $0.01553 |
| Claude Opus 5 | 81 | $2.75375 | $0.03400 |
| GPT-6 Astra | 81 | $3.18065 | $0.03927 |
Selection uses the same fixture ID and trial index for every model. All selected outputs passed the existing graders. No trial was retried, removed from the source data, or changed. Selection was made after observing the outcomes and favors jointly solved cases. Whole-task bootstrap intervals describe cost variation only within this selected subset.