Shared-success matched costs

The shared-success comparison contains 81 matched runs per model. We requested 85, but only 81 of the 90 aligned task/trial positions passed for every included model. We kept the existing runs and added none. This post-hoc subset compares costs conditional on shared success; it cannot estimate reliability or replace the full cost-per-success results in the full study.

All dollar amounts are actual Gateway charges. Sol's charges reflect promotional rates; this is not an undiscounted market-rate comparison. Billing details.

Costs on the shared-success subset
Original synthetic data. Open the chart to inspect the full-size values.
Model Matched runs Total cost Cost / included run
DeepSeek V4 Flash 81 $0.02195 $0.00027
GLM-5.3 Flash 81 $0.06191 $0.00076
Qwen 3.8 Flash 81 $0.06776 $0.00084
GPT-5.6 Luna 81 $0.08475 $0.00105
DeepSeek V4 Pro 81 $0.51572 $0.00637
GPT-5.6 Sol 81 $0.75519 $0.00932
Gemini 3.8 Flash 81 $0.89666 $0.01107
Qwen 3.8 Max 81 $0.99063 $0.01223
Claude Sonnet 5 81 $1.25760 $0.01553
Claude Opus 5 81 $2.75375 $0.03400
GPT-6 Astra 81 $3.18065 $0.03927

Selection uses the same fixture ID and trial index for every model. All selected outputs passed the existing graders. No trial was retried, removed from the source data, or changed. Selection was made after observing the outcomes and favors jointly solved cases. Whole-task bootstrap intervals describe cost variation only within this selected subset.

Exact selected and excluded positions. Full comparison.