Same model. Same tasks. Cheaper tokens. 1.4× the cost.
Measured by CanaryOne · GLM 5.2 · 11 routes ·
10 coding tasks each · 14 August 2026
Advertised token price $0.750 vs $1.400 DeepInfra fp4 46 per cent cheaper
Tasks completed 10 of 10 vs
10 of 10 Same outcome Cost per successful task
DeepInfra fp4
$0.0134
1.4× higher
Z.AI fp8
$0.0096
Why the gap? On the same job, one route spent
49 per cent more than the other in reasoning tokens per task —
105 against 71.
All 11 routes measured
Sail Research fp8 · $0.0095 per successful task · 10 of 10 passed Z.AI fp8 · $0.0096 per successful task · 10 of 10 passed Alibaba fp8 · $0.0098 per successful task · 10 of 10 passed StreamLake fp8 · $0.0098 per successful task · 10 of 10 passed Decart fp4 · $0.0106 per successful task · 9 of 10 passed Novita fp8 · $0.0108 per successful task · 10 of 10 passed CoreWeave fp4 · $0.0126 per successful task · 10 of 10 passed GMICloud fp8 · $0.0132 per successful task · 10 of 10 passed DeepInfra fp4 · $0.0134 per successful task · 10 of 10 passed DigitalOcean · $0.0173 per successful task · 9 of 10 passed SiliconFlow fp8 · $0.0191 per successful task · 10 of 10 passed $0.0096 Z.AI fp8 $0.0134 DeepInfra fp4
Measured workload, run on the night of
14 August 2026. Total spend divided by the tasks that
completed, so failed work stays in the cost. These figures cost money to reproduce
and hold that date until the next run. The size of a gap repeats from night to
night; which host sits at either end does not, so read the multiple rather than the
names.
What does this look like on your workload? Benchmark your workload →
The same weights are priced very differently.
Each bar runs from the cheapest company serving that model to the dearest, with a tick
at the average of all of them. Dearest on average first. The axis is logarithmic: the
models are three hundred times apart in price.
13 models · 220 live endpoints
Cheapest company Average of all of them Dearest company $0.01 $0.10 $1 $10 Kimi K3 Kimi K3 · $3.16 per 1M averaged across 13 companies Kimi K3 · cheapest listed host $2.60 per 1M Kimi K3 · dearest listed host $6.00 per 1M 2.3× DeepSeek V4 Pro 0423 DeepSeek V4 Pro 0423 · $1.37 per 1M averaged across 18 companies DeepSeek V4 Pro 0423 · cheapest listed host $0.660 per 1M DeepSeek V4 Pro 0423 · dearest listed host $1.91 per 1M 2.9× GLM 5.2 GLM 5.2 · $1.27 per 1M averaged across 33 companies GLM 5.2 · cheapest listed host $0.500 per 1M GLM 5.2 · dearest listed host $2.31 per 1M 4.6× GLM 5.1 GLM 5.1 · $1.22 per 1M averaged across 18 companies GLM 5.1 · cheapest listed host $0.910 per 1M GLM 5.1 · dearest listed host $1.54 per 1M 1.7× Kimi K2.7 Code Kimi K2.7 Code · $0.902 per 1M averaged across 15 companies Kimi K2.7 Code · cheapest listed host $0.670 per 1M Kimi K2.7 Code · dearest listed host $1.90 per 1M 2.8× Kimi K2.6 Kimi K2.6 · $0.822 per 1M averaged across 21 companies Kimi K2.6 · cheapest listed host $0.568 per 1M Kimi K2.6 · dearest listed host $1.20 per 1M 2.1× MiniMax M3 MiniMax M3 · $0.325 per 1M averaged across 12 companies MiniMax M3 · cheapest listed host $0.230 per 1M MiniMax M3 · dearest listed host $0.750 per 1M 3.3× Qwen3 Coder Next Qwen3 Coder Next · $0.200 per 1M averaged across 4 companies Qwen3 Coder Next · cheapest listed host $0.120 per 1M Qwen3 Coder Next · dearest listed host $0.300 per 1M 2.5× Nemotron 3 Super Nemotron 3 Super · $0.183 per 1M averaged across 3 companies Nemotron 3 Super · cheapest listed host $0.085 per 1M Nemotron 3 Super · dearest listed host $0.300 per 1M 3.5× DeepSeek V4 Flash 0423 DeepSeek V4 Flash 0423 · $0.151 per 1M averaged across 18 companies DeepSeek V4 Flash 0423 · cheapest listed host $0.068 per 1M DeepSeek V4 Flash 0423 · dearest listed host $0.440 per 1M 6.5× DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 · $0.145 per 1M averaged across 28 companies DeepSeek V4 Flash 0731 · cheapest listed host $0.079 per 1M DeepSeek V4 Flash 0731 · dearest listed host $0.440 per 1M 5.6× gpt-oss-120b gpt-oss-120b · $0.116 per 1M averaged across 20 companies gpt-oss-120b · cheapest listed host $0.030 per 1M gpt-oss-120b · dearest listed host $0.350 per 1M 12× Ling-3.0-flash Ling-3.0-flash · $0.040 per 1M averaged across 2 companies Ling-3.0-flash · cheapest listed host $0.021 per 1M Ling-3.0-flash · dearest listed host $0.060 per 1M 2.9×
Listed input price per 1M tokens · right column is dearest divided by cheapest
There is no model here whose hosts agree on a price. At the narrowest, GLM 5.1,
the dearest company still charges 69 per cent above the cheapest. At the widest,
gpt-oss-120b, 20 companies serve the same weights and the dearest
charges more than eleven and a half times the cheapest. Metadata read 17 August 2026 at 12:55 UTC.
Check the widest one on
OpenRouter →
The same price doesn't buy the same speed.
Every point is one company serving the same weights. Lower and further left is cheaper
and faster, and the shaded band holds the companies charging an identical price.
Model gpt-oss-120b GLM 5.2 DeepSeek V4 Flash 0731 Kimi K2.6 DeepSeek V4 Flash 0423 DeepSeek V4 Pro 0423 GLM 5.1 Kimi K2.7 Code Kimi K3 MiniMax M3
+5 more
Show fewer
20 companies · 12× price spread
· 6.7× first-token spread
Same price: 6 hosts at $0.15 0.0s
0.9s
1.7s
2.6s
$0.02 $0.05 $0.1 $0.2 $0.5 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
coreweave/fp4 · $0.030 per 1M · 0.5s first token deepinfra/bf16 · $0.037 per 1M · 0.3s first token akashml/bf16 · $0.037 per 1M · 0.6s first token novita/fp4 · $0.050 per 1M · 0.5s first token · not serving siliconflow/fp8 · $0.050 per 1M · 1.6s first token digitalocean · $0.055 per 1M · 0.7s first token mancer/fp8 · $0.080 per 1M · 0.8s first token google-vertex/global · $0.090 per 1M · 1.3s first token baseten/fp4 · $0.100 per 1M · 0.2s first token parasail/fp4 · $0.100 per 1M · 0.4s first token sambanova · $0.140 per 1M · 1.0s first token · not serving amazon-bedrock · $0.150 per 1M · 0.4s first token deepinfra/turbo · $0.150 per 1M · 0.4s first token together · $0.150 per 1M · 0.3s first token nebius/fp4 · $0.150 per 1M · 0.4s first token phala · $0.150 per 1M · 1.0s first token groq · $0.150 per 1M · 0.2s first token mara · $0.150 per 1M · 2.4s first token · not serving cerebras/fp16 · $0.350 per 1M · 0.3s first token
20 companies serve gpt-oss-120b, and the dearest charges
more than eleven and a half times the cheapest.
Six of them charge exactly 15 cents per
million tokens, and inside that group the first token arrives after
0.2 seconds at one end and
1.0 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 16 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
33 companies · 4.6× price spread
· 6.3× first-token spread
Same price: 10 hosts at $1.40 0.0s
2.2s
4.4s
6.6s
$0.5 $1 $2 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
sail-research/fp8 · $0.500 per 1M · 1.3s first token novita/fp8 · $0.665 per 1M · 2.8s first token streamlake/fp8 · $0.666 per 1M · 2.0s first token digitalocean · $0.700 per 1M · 1.3s first token decart/fp4 · $0.720 per 1M · 1.2s first token gmicloud/fp8 · $0.742 per 1M · 1.7s first token deepinfra/fp4 · $0.750 per 1M · 1.2s first token inceptron/fp4 · $0.750 per 1M · 0.7s first token coreweave/fp4 · $0.760 per 1M · 0.7s first token akashml/fp8 · $0.770 per 1M · 1.0s first token alibaba/fp8 · $0.966 per 1M · 1.4s first token ambient/fp8 · $1.05 per 1M · 2.1s first token morph/fp4 · $1.10 per 1M · 1.6s first token phala/fp8 · $1.13 per 1M · 2.9s first token siliconflow/fp8 · $1.19 per 1M · 2.0s first token wafer · $1.26 per 1M · 4.2s first token · not serving atlas-cloud/fp8 · $1.26 per 1M · 3.0s first token z-ai/fp8 · $1.40 per 1M · 3.1s first token fireworks · $1.40 per 1M · 1.6s first token baidu/fp8 · $1.40 per 1M · 0.8s first token cloudflare · $1.40 per 1M · 2.8s first token friendli · $1.40 per 1M · 0.6s first token parasail/fp4 · $1.40 per 1M · 0.6s first token venice/fp8 · $1.40 per 1M · 2.2s first token together · $1.40 per 1M · 0.9s first token crusoe/fp8 · $1.40 per 1M · 0.5s first token baseten/fp8 · $1.40 per 1M · 1.9s first token wafer/fast · $2.10 per 1M · 2.7s first token fireworks/fast · $2.10 per 1M · 1.0s first token cloudflare/fast · $2.10 per 1M · 1.5s first token baseten/fast · $2.10 per 1M · 1.8s first token modelrun/fp4 · $2.20 per 1M · 6.0s first token · not serving alibaba/fast · $2.31 per 1M · 1.5s first token
33 companies serve GLM 5.2, and the dearest charges
more than four and a half times the cheapest.
Ten of them charge exactly $1.40 per
million tokens, and inside that group the first token arrives after
0.5 seconds at one end and
3.1 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 31 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
28 companies · 5.6× price spread
· 7.6× first-token spread
Same price: 9 hosts at $0.14 0.0s
1.2s
2.4s
3.7s
$0.05 $0.1 $0.2 $0.5 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
streamlake/fp8 · $0.079 per 1M · 2.4s first token · not serving decart/fp4 · $0.079 per 1M · 1.1s first token · not serving digitalocean · $0.080 per 1M · 0.7s first token deepinfra/fp8 · $0.080 per 1M · 1.0s first token open-inference/fp4 · $0.080 per 1M · 2.2s first token gmicloud/fp8 · $0.084 per 1M · 3.3s first token sail-research/fp4 · $0.090 per 1M · 2.9s first token relace/fp4 · $0.105 per 1M · 0.9s first token · not serving baseten/fp8 · $0.130 per 1M · 0.4s first token coreweave/fp8 · $0.130 per 1M · 0.5s first token inceptron/fp4 · $0.130 per 1M · 1.1s first token morph/bf16 · $0.139 per 1M · 1.4s first token fireworks · $0.140 per 1M · 1.0s first token akashml/fp8 · $0.140 per 1M · 1.2s first token novita/fp8 · $0.140 per 1M · 1.5s first token together · $0.140 per 1M · 0.9s first token parasail/fp8 · $0.140 per 1M · 0.8s first token atlas-cloud/fp4 · $0.140 per 1M · 1.3s first token siliconflow/fp8 · $0.140 per 1M · 1.8s first token ambient/fp4 · $0.140 per 1M · 0.7s first token baidu/fp8 · $0.140 per 1M · 0.9s first token mancer/fp8 · $0.140 per 1M · 1.2s first token · not serving io-net/fp8 · $0.149 per 1M · 1.6s first token venice · $0.175 per 1M · 1.4s first token phala · $0.200 per 1M · 1.6s first token deepseek/fp8 · $0.220 per 1M · 0.9s first token wafer/fast · $0.280 per 1M · 1.3s first token cloudflare · $0.440 per 1M · 0.7s first token
28 companies serve DeepSeek V4 Flash 0731, and the dearest charges
five and a half times the cheapest.
Nine of them charge exactly 14 cents per
million tokens, and inside that group the first token arrives after
0.7 seconds at one end and
1.8 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 24 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
21 companies · 2.1× price spread
· 5.8× first-token spread
Same price: 5 hosts at $0.95 0.0s
0.9s
1.8s
2.7s
$0.5 $0.7 $1 $1.5 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
decart/fp4 · $0.568 per 1M · 1.4s first token chutes/int4 · $0.580 per 1M · 2.4s first token streamlake/fp8 · $0.598 per 1M · 1.7s first token · not serving inceptron/int4 · $0.600 per 1M · 0.5s first token coreweave/fp4 · $0.650 per 1M · 0.5s first token crusoe/bf16 · $0.700 per 1M · 0.4s first token parasail/int4 · $0.750 per 1M · 1.0s first token venice/int4 · $0.750 per 1M · 1.5s first token · not serving deepinfra/fp4 · $0.750 per 1M · 1.2s first token digitalocean · $0.760 per 1M · 0.9s first token · not serving siliconflow/fp8 · $0.770 per 1M · 1.4s first token novita · $0.800 per 1M · 1.6s first token atlas-cloud/int4 · $0.950 per 1M · 1.4s first token moonshotai/int4 · $0.950 per 1M · 2.0s first token baidu/fp4 · $0.950 per 1M · 0.7s first token cloudflare · $0.950 per 1M · 1.0s first token sail-research/int4 · $1.00 per 1M · 1.0s first token · not serving phala · $1.09 per 1M · 1.5s first token together · $1.20 per 1M · 0.8s first token fireworks · $0.950 per 1M · 0.5s first token
21 companies serve Kimi K2.6, and the dearest charges
two times the cheapest.
Five of them charge exactly 95 cents per
million tokens, and inside that group the first token arrives after
0.5 seconds at one end and
2.0 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 16 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
18 companies · 6.5× price spread
· 7.2× first-token spread
Same price: 6 hosts at $0.14 0.0s
1.2s
2.5s
3.7s
$0.05 $0.1 $0.2 $0.5 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
digitalocean · $0.068 per 1M · 1.5s first token streamlake/fp8 · $0.083 per 1M · 1.4s first token gmicloud/fp8 · $0.084 per 1M · 2.7s first token deepinfra/fp8 · $0.090 per 1M · 0.7s first token sail-research/fp4 · $0.090 per 1M · 0.8s first token siliconflow/fp8 · $0.130 per 1M · 1.2s first token alibaba/fp8 · $0.134 per 1M · 0.9s first token venice · $0.138 per 1M · 1.4s first token parasail/fp8 · $0.140 per 1M · 0.6s first token novita/fp8 · $0.140 per 1M · 1.0s first token atlas-cloud/fp4 · $0.140 per 1M · 1.0s first token coreweave/fp8 · $0.140 per 1M · 0.5s first token baidu/fp8 · $0.140 per 1M · 0.7s first token mancer/fp8 · $0.140 per 1M · 1.0s first token phala · $0.200 per 1M · 3.4s first token azure/us · $0.210 per 1M · 2.0s first token · not serving deepseek · $0.220 per 1M · 0.9s first token cloudflare · $0.440 per 1M · 0.8s first token
18 companies serve DeepSeek V4 Flash 0423, and the dearest charges
six and a half times the cheapest.
Six of them charge exactly 14 cents per
million tokens, and inside that group the first token arrives after
0.5 seconds at one end and
1.0 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 17 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
18 companies · 2.9× price spread
· 7.4× first-token spread
Same price: 4 hosts at $1.74 0.0s
1.2s
2.3s
3.5s
$0.5 $1 $2 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
deepseek · $0.660 per 1M · 1.2s first token streamlake/fp8 · $0.694 per 1M · 2.5s first token gmicloud/fp8 · $0.696 per 1M · 3.2s first token digitalocean · $0.870 per 1M · 1.9s first token ionstream/fp4 · $1.13 per 1M · 1.8s first token coreweave/fp8 · $1.15 per 1M · 0.6s first token deepinfra/fp8 · $1.30 per 1M · 0.9s first token alibaba/fp8 · $1.42 per 1M · 1.6s first token novita/fp8 · $1.44 per 1M · 1.5s first token siliconflow/fp8 · $1.50 per 1M · 1.7s first token venice · $1.65 per 1M · 1.7s first token atlas-cloud/fp4 · $1.68 per 1M · 1.5s first token baidu/fp8 · $1.69 per 1M · 0.8s first token baseten/fp4 · $1.74 per 1M · 0.4s first token parasail/fp8 · $1.74 per 1M · 0.8s first token together · $1.74 per 1M · 0.7s first token fireworks · $1.74 per 1M · 1.6s first token azure/us · $1.91 per 1M · 1.5s first token · not serving
18 companies serve DeepSeek V4 Pro 0423, and the dearest charges
nearly three times the cheapest.
Four of them charge exactly $1.74 per
million tokens, and inside that group the first token arrives after
0.4 seconds at one end and
1.6 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 17 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
18 companies · 1.7× price spread
· 38× first-token spread
Same price: 5 hosts at $1.40 0.0s
3.4s
6.9s
10.3s
$0.7 $1 $1.5 $2 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
gmicloud/fp8 · $0.910 per 1M · 2.0s first token streamlake/fp8 · $0.966 per 1M · 3.3s first token digitalocean · $0.975 per 1M · 1.2s first token · not serving chutes/fp8 · $0.980 per 1M · 2.1s first token · not serving wafer/fp4 · $1.00 per 1M · 0.9s first token deepinfra/fp4 · $1.05 per 1M · 0.9s first token siliconflow/fp8 · $1.19 per 1M · 1.9s first token crusoe/fp8 · $1.20 per 1M · 0.6s first token phala · $1.21 per 1M · 2.0s first token · not serving atlas-cloud/fp8 · $1.26 per 1M · 1.5s first token alibaba/fp8 · $1.33 per 1M · 2.0s first token novita/fp8 · $1.38 per 1M · 2.5s first token nebius/fp8 · $1.40 per 1M · 0.7s first token parasail/fp8 · $1.40 per 1M · 0.8s first token friendli · $1.40 per 1M · 0.2s first token z-ai/fp8 · $1.40 per 1M · 9.4s first token baidu/fp8 · $1.40 per 1M · 0.8s first token venice/fp8 · $1.54 per 1M · 1.5s first token
18 companies serve GLM 5.1, and the dearest charges
69 per cent above the cheapest.
Five of them charge exactly $1.40 per
million tokens, and inside that group the first token arrives after
0.2 seconds at one end and
9.4 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 15 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
15 companies · 2.8× price spread
· 7.7× first-token spread
Same price: 5 hosts at $0.95 0.0s
1.1s
2.2s
3.3s
$0.5 $1 $2 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
inceptron/int4 · $0.670 per 1M · 0.4s first token deepinfra/fp4 · $0.680 per 1M · 0.8s first token ambient · $0.690 per 1M · 0.8s first token coreweave/int4 · $0.710 per 1M · 3.0s first token venice/int4 · $0.750 per 1M · 1.2s first token parasail/int4 · $0.760 per 1M · 0.6s first token modelrun/fp4 · $0.850 per 1M · 0.5s first token siliconflow/fp8 · $0.859 per 1M · 1.6s first token novita/int4 · $0.912 per 1M · 1.4s first token moonshotai/int4 · $0.950 per 1M · 1.8s first token cloudflare · $0.950 per 1M · 1.0s first token atlas-cloud/int4 · $0.950 per 1M · 0.8s first token together · $0.950 per 1M · 0.6s first token alibaba/fp8 · $0.950 per 1M · 2.5s first token moonshotai/highspeed · $1.90 per 1M · 1.8s first token
15 companies serve Kimi K2.7 Code, and the dearest charges
more than two and a half times the cheapest.
Five of them charge exactly 95 cents per
million tokens, and inside that group the first token arrives after
0.6 seconds at one end and
2.5 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 15 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
13 companies · 2.3× price spread
· 6.6× first-token spread
Same price: 8 hosts at $3.00 0.0s
1.8s
3.6s
5.5s
$2 $3 $5 $7 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
sail-research/fp4 · $2.60 per 1M · 3.7s first token morph/fp4 · $2.80 per 1M · 2.3s first token digitalocean · $2.85 per 1M · 3.5s first token deepinfra/bf16 · $2.85 per 1M · 0.8s first token fireworks · $3.00 per 1M · 2.0s first token chutes/mxfp4 · $3.00 per 1M · 1.7s first token together · $3.00 per 1M · 1.6s first token moonshotai/mxfp4 · $3.00 per 1M · 5.0s first token wafer · $3.00 per 1M · 2.0s first token modal/mxfp4 · $3.00 per 1M · 2.9s first token baseten/fp8 · $3.00 per 1M · 1.8s first token phala · $3.00 per 1M · 3.5s first token morph/fast · $6.00 per 1M · 1.4s first token
13 companies serve Kimi K3, and the dearest charges
nearly two and a half times the cheapest.
Eight of them charge exactly $3.00 per
million tokens, and inside that group the first token arrives after
1.6 seconds at one end and
5.0 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 13 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
12 companies · 3.3× price spread
· 4.8× first-token spread
Same price: 8 hosts at $0.30 0.0s
1.1s
2.2s
3.3s
$0.2 $0.5 $1 First token, seconds Not serving requests at this read
Listed price per 1M tokens, logarithmic
coreweave/fp4 · $0.230 per 1M · 0.6s first token gmicloud/fp8 · $0.240 per 1M · 3.0s first token deepinfra/fp8 · $0.280 per 1M · 2.2s first token novita/fp8 · $0.300 per 1M · 1.8s first token venice/fp8 · $0.300 per 1M · 1.8s first token minimax/fp8 · $0.300 per 1M · 1.1s first token atlas-cloud/fp8 · $0.300 per 1M · 1.5s first token together · $0.300 per 1M · 1.0s first token streamlake/fp8 · $0.300 per 1M · 1.2s first token parasail/fp8 · $0.300 per 1M · 0.7s first token morph/fp4 · $0.300 per 1M · 1.4s first token modelrun/fp4 · $0.750 per 1M · 1.9s first token
12 companies serve MiniMax M3, and the dearest charges
more than three times the cheapest.
Eight of them charge exactly 30 cents per
million tokens, and inside that group the first token arrives after
0.7 seconds at one end and
1.8 seconds at the other. Paying the same does
not buy the same. First-token figures cover the 12 companies serving requests at
this read, each one their own reported median over the previous half hour.
Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
Same weights. Different capabilities. 5 of 20 gpt-oss-120b hosts don't advertise tool calling.
One bar per model, longest where most companies serve it. Filled advertises tool
calling and outlined does not. An agent that needs to call a function does not degrade
on a host without it; the loop stops.
gpt-oss-120b 15 / 20 Kimi K3 10 / 13 MiniMax M3 9 / 12 Kimi K2.6 20 / 21 GLM 5.1 17 / 18 Nemotron 3 Super 2 / 3 Seven other models: every tracked host advertises tool calling.
Stated precision, across all 220 live endpoints
10 16-bit 78 8-bit 60 4-bit 72 not stated 14 of the 205 company-and-model pairs we track do not advertise tool
calling, and 72 of 220 endpoints do not say what precision they serve
at. Both are provider-advertised, not a CanaryOne functional test, and whether four bits
changes an answer is a question about output that nothing on this page measures.
Metadata read 17 August 2026 at 12:55 UTC.
Check the capabilities on
OpenRouter →
The market moves underneath you.
One box per company, showing what it charged for
GLM 5.2 across 40 reads since
4 August 2026. Identical weights in all 32. Every box
shares one vertical scale, and a line that drops away is a discount.
Changed price during the week
9 of 32 companies
Sail Research fp8 $0.500
Novita fp8 $0.665
StreamLake fp8 $0.666
DigitalOcean $0.700
Decart fp4 $0.720
GMICloud fp8 $0.742
Inceptron fp4 $0.750
Phala $1.13
Baidu fp8 $1.40
The other 23 never moved. One dot per company, stacked where several
charge the same price. Hover a dot for its precision and context window.
Held the same price all week $0.750 to $2.31 per 1M 23 of 32 companies
DeepInfra fp4 · $0.750 per 1M · fp4 precision · 1024K context CoreWeave fp4 · $0.760 per 1M · fp4 precision · 256K context AkashML fp8 · $0.770 per 1M · fp8 precision · 95K context Alibaba fp8 · $0.966 per 1M · fp8 precision · 1024K context Ambient fp8 · $1.05 per 1M · fp8 precision · 198K context Morph · $1.10 per 1M · precision not stated SiliconFlow fp8 · $1.19 per 1M · fp8 precision · 1024K context AtlasCloud fp8 · $1.26 per 1M · fp8 precision · 1024K context Wafer · $1.26 per 1M · fp4 precision · 1024K context BaseTen fp8 · $1.40 per 1M · fp8 precision · 1024K context Cloudflare · $1.40 per 1M · precision not stated · 256K context Crusoe fp8 · $1.40 per 1M · fp8 precision · 1024K context Fireworks · $1.40 per 1M · precision not stated · 1024K context Friendli · $1.40 per 1M · precision not stated · 1024K context Parasail fp4 · $1.40 per 1M · fp4 precision · 256K context Together · $1.40 per 1M · precision not stated · 500K context Venice fp8 · $1.40 per 1M · fp8 precision · 977K context Z.AI fp8 · $1.40 per 1M · fp8 precision · 1024K context BaseTen fast · $2.10 per 1M · fp8 precision · 1024K context Cloudflare fast · $2.10 per 1M · precision not stated · 256K context Fireworks fast · $2.10 per 1M · precision not stated · 1024K context Wafer fast · $2.10 per 1M · fp4 precision · 1024K context Alibaba fast · $2.31 per 1M · fp8 precision · 1024K context $1.40 9 companies
$2.10 4 companies
The 23 companies that held their price, as a table
The steepest move was a discount that cut one company's price by
93 per cent before expiring above where the week
started. Every price here is what a buyer pays after whatever discount was running at
that moment. Metadata read 17 August 2026 at 12:55 UTC.
Check these prices on
OpenRouter →
How we measure it.
Two layers, kept apart. Host metadata is public: prices, capabilities, advertised
throughput and availability, read several times a day. Measured workloads are ours:
the same task suite run against multiple hosts serving identical weights, where total
spend is divided by the tasks that actually completed, so failed work stays in the
economics instead of disappearing from it.
We compare hosts within one model and never treat different weights as equivalent
workloads. Timestamps sit beside measurements. Snapshots are not trends, and this
record is days old rather than months. Relative positions between hosts change
between runs, so we report the size of a spread rather than naming a best host.
Prices, capabilities and precision are what a company publishes about itself, and a
published capability is a claim rather than a test we have run. First-token figures
are each company's own reported median over the half hour before we read it.
What happens on your workload?
Market averages can tell you where to look. Your own workload tells you what to deploy.
Benchmark your workload → Get early access →