CANARYONE MARKET

The open-weight market, measured.

Open-weight models are portable. Their economics aren't. We track identical model weights across the companies serving them to see what actually changes: price, latency, capabilities, and the cost of completing real work.

14 models · 220 live endpoints · market read several times a day · measured workload sweeps

Same model. Same tasks. Cheaper tokens. 1.4× the cost.

Measured by CanaryOne · GLM 5.2 · 11 routes · 10 coding tasks each · 14 August 2026

Advertised token price
$0.750 vs $1.400
DeepInfra fp4 46 per cent cheaper
Tasks completed
10 of 10 vs 10 of 10
Same outcome

Cost per successful task

DeepInfra fp4

$0.0134

1.4× higher

Z.AI fp8

$0.0096

Why the gap? On the same job, one route spent 49 per cent more than the other in reasoning tokens per task — 105 against 71.

All 11 routes measured

Sail Research fp8 · $0.0095 per successful task · 10 of 10 passed Z.AI fp8 · $0.0096 per successful task · 10 of 10 passed Alibaba fp8 · $0.0098 per successful task · 10 of 10 passed StreamLake fp8 · $0.0098 per successful task · 10 of 10 passed Decart fp4 · $0.0106 per successful task · 9 of 10 passed Novita fp8 · $0.0108 per successful task · 10 of 10 passed CoreWeave fp4 · $0.0126 per successful task · 10 of 10 passed GMICloud fp8 · $0.0132 per successful task · 10 of 10 passed DeepInfra fp4 · $0.0134 per successful task · 10 of 10 passed DigitalOcean · $0.0173 per successful task · 9 of 10 passed SiliconFlow fp8 · $0.0191 per successful task · 10 of 10 passed $0.0096 Z.AI fp8 $0.0134 DeepInfra fp4

Measured workload, run on the night of 14 August 2026. Total spend divided by the tasks that completed, so failed work stays in the cost. These figures cost money to reproduce and hold that date until the next run. The size of a gap repeats from night to night; which host sits at either end does not, so read the multiple rather than the names.

What does this look like on your workload? Benchmark your workload →

The same weights are priced very differently.

Each bar runs from the cheapest company serving that model to the dearest, with a tick at the average of all of them. Dearest on average first. The axis is logarithmic: the models are three hundred times apart in price.

13 models · 220 live endpoints

Cheapest company Average of all of them Dearest company $0.01 $0.10 $1 $10 Kimi K3 Kimi K3 · $3.16 per 1M averaged across 13 companies Kimi K3 · cheapest listed host $2.60 per 1M Kimi K3 · dearest listed host $6.00 per 1M 2.3× DeepSeek V4 Pro 0423 DeepSeek V4 Pro 0423 · $1.37 per 1M averaged across 18 companies DeepSeek V4 Pro 0423 · cheapest listed host $0.660 per 1M DeepSeek V4 Pro 0423 · dearest listed host $1.91 per 1M 2.9× GLM 5.2 GLM 5.2 · $1.27 per 1M averaged across 33 companies GLM 5.2 · cheapest listed host $0.500 per 1M GLM 5.2 · dearest listed host $2.31 per 1M 4.6× GLM 5.1 GLM 5.1 · $1.22 per 1M averaged across 18 companies GLM 5.1 · cheapest listed host $0.910 per 1M GLM 5.1 · dearest listed host $1.54 per 1M 1.7× Kimi K2.7 Code Kimi K2.7 Code · $0.902 per 1M averaged across 15 companies Kimi K2.7 Code · cheapest listed host $0.670 per 1M Kimi K2.7 Code · dearest listed host $1.90 per 1M 2.8× Kimi K2.6 Kimi K2.6 · $0.822 per 1M averaged across 21 companies Kimi K2.6 · cheapest listed host $0.568 per 1M Kimi K2.6 · dearest listed host $1.20 per 1M 2.1× MiniMax M3 MiniMax M3 · $0.325 per 1M averaged across 12 companies MiniMax M3 · cheapest listed host $0.230 per 1M MiniMax M3 · dearest listed host $0.750 per 1M 3.3× Qwen3 Coder Next Qwen3 Coder Next · $0.200 per 1M averaged across 4 companies Qwen3 Coder Next · cheapest listed host $0.120 per 1M Qwen3 Coder Next · dearest listed host $0.300 per 1M 2.5× Nemotron 3 Super Nemotron 3 Super · $0.183 per 1M averaged across 3 companies Nemotron 3 Super · cheapest listed host $0.085 per 1M Nemotron 3 Super · dearest listed host $0.300 per 1M 3.5× DeepSeek V4 Flash 0423 DeepSeek V4 Flash 0423 · $0.151 per 1M averaged across 18 companies DeepSeek V4 Flash 0423 · cheapest listed host $0.068 per 1M DeepSeek V4 Flash 0423 · dearest listed host $0.440 per 1M 6.5× DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 · $0.145 per 1M averaged across 28 companies DeepSeek V4 Flash 0731 · cheapest listed host $0.079 per 1M DeepSeek V4 Flash 0731 · dearest listed host $0.440 per 1M 5.6× gpt-oss-120b gpt-oss-120b · $0.116 per 1M averaged across 20 companies gpt-oss-120b · cheapest listed host $0.030 per 1M gpt-oss-120b · dearest listed host $0.350 per 1M 12× Ling-3.0-flash Ling-3.0-flash · $0.040 per 1M averaged across 2 companies Ling-3.0-flash · cheapest listed host $0.021 per 1M Ling-3.0-flash · dearest listed host $0.060 per 1M 2.9× Listed input price per 1M tokens · right column is dearest divided by cheapest

There is no model here whose hosts agree on a price. At the narrowest, GLM 5.1, the dearest company still charges 69 per cent above the cheapest. At the widest, gpt-oss-120b, 20 companies serve the same weights and the dearest charges more than eleven and a half times the cheapest. Metadata read 17 August 2026 at 12:55 UTC. Check the widest one on OpenRouter →

The same price doesn't buy the same speed.

Every point is one company serving the same weights. Lower and further left is cheaper and faster, and the shaded band holds the companies charging an identical price.

Model

20 companies · 12× price spread · 6.7× first-token spread

Same price: 6 hosts at $0.15 0.0s 0.9s 1.7s 2.6s $0.02$0.05$0.1$0.2$0.5 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic coreweave/fp4 · $0.030 per 1M · 0.5s first token deepinfra/bf16 · $0.037 per 1M · 0.3s first token akashml/bf16 · $0.037 per 1M · 0.6s first token novita/fp4 · $0.050 per 1M · 0.5s first token · not serving siliconflow/fp8 · $0.050 per 1M · 1.6s first token digitalocean · $0.055 per 1M · 0.7s first token mancer/fp8 · $0.080 per 1M · 0.8s first token google-vertex/global · $0.090 per 1M · 1.3s first token baseten/fp4 · $0.100 per 1M · 0.2s first token parasail/fp4 · $0.100 per 1M · 0.4s first token sambanova · $0.140 per 1M · 1.0s first token · not serving amazon-bedrock · $0.150 per 1M · 0.4s first token deepinfra/turbo · $0.150 per 1M · 0.4s first token together · $0.150 per 1M · 0.3s first token nebius/fp4 · $0.150 per 1M · 0.4s first token phala · $0.150 per 1M · 1.0s first token groq · $0.150 per 1M · 0.2s first token mara · $0.150 per 1M · 2.4s first token · not serving cerebras/fp16 · $0.350 per 1M · 0.3s first token

20 companies serve gpt-oss-120b, and the dearest charges more than eleven and a half times the cheapest. Six of them charge exactly 15 cents per million tokens, and inside that group the first token arrives after 0.2 seconds at one end and 1.0 seconds at the other. Paying the same does not buy the same. First-token figures cover the 16 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

33 companies · 4.6× price spread · 6.3× first-token spread

Same price: 10 hosts at $1.40 0.0s 2.2s 4.4s 6.6s $0.5$1$2 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic sail-research/fp8 · $0.500 per 1M · 1.3s first token novita/fp8 · $0.665 per 1M · 2.8s first token streamlake/fp8 · $0.666 per 1M · 2.0s first token digitalocean · $0.700 per 1M · 1.3s first token decart/fp4 · $0.720 per 1M · 1.2s first token gmicloud/fp8 · $0.742 per 1M · 1.7s first token deepinfra/fp4 · $0.750 per 1M · 1.2s first token inceptron/fp4 · $0.750 per 1M · 0.7s first token coreweave/fp4 · $0.760 per 1M · 0.7s first token akashml/fp8 · $0.770 per 1M · 1.0s first token alibaba/fp8 · $0.966 per 1M · 1.4s first token ambient/fp8 · $1.05 per 1M · 2.1s first token morph/fp4 · $1.10 per 1M · 1.6s first token phala/fp8 · $1.13 per 1M · 2.9s first token siliconflow/fp8 · $1.19 per 1M · 2.0s first token wafer · $1.26 per 1M · 4.2s first token · not serving atlas-cloud/fp8 · $1.26 per 1M · 3.0s first token z-ai/fp8 · $1.40 per 1M · 3.1s first token fireworks · $1.40 per 1M · 1.6s first token baidu/fp8 · $1.40 per 1M · 0.8s first token cloudflare · $1.40 per 1M · 2.8s first token friendli · $1.40 per 1M · 0.6s first token parasail/fp4 · $1.40 per 1M · 0.6s first token venice/fp8 · $1.40 per 1M · 2.2s first token together · $1.40 per 1M · 0.9s first token crusoe/fp8 · $1.40 per 1M · 0.5s first token baseten/fp8 · $1.40 per 1M · 1.9s first token wafer/fast · $2.10 per 1M · 2.7s first token fireworks/fast · $2.10 per 1M · 1.0s first token cloudflare/fast · $2.10 per 1M · 1.5s first token baseten/fast · $2.10 per 1M · 1.8s first token modelrun/fp4 · $2.20 per 1M · 6.0s first token · not serving alibaba/fast · $2.31 per 1M · 1.5s first token

33 companies serve GLM 5.2, and the dearest charges more than four and a half times the cheapest. Ten of them charge exactly $1.40 per million tokens, and inside that group the first token arrives after 0.5 seconds at one end and 3.1 seconds at the other. Paying the same does not buy the same. First-token figures cover the 31 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

28 companies · 5.6× price spread · 7.6× first-token spread

Same price: 9 hosts at $0.14 0.0s 1.2s 2.4s 3.7s $0.05$0.1$0.2$0.5 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic streamlake/fp8 · $0.079 per 1M · 2.4s first token · not serving decart/fp4 · $0.079 per 1M · 1.1s first token · not serving digitalocean · $0.080 per 1M · 0.7s first token deepinfra/fp8 · $0.080 per 1M · 1.0s first token open-inference/fp4 · $0.080 per 1M · 2.2s first token gmicloud/fp8 · $0.084 per 1M · 3.3s first token sail-research/fp4 · $0.090 per 1M · 2.9s first token relace/fp4 · $0.105 per 1M · 0.9s first token · not serving baseten/fp8 · $0.130 per 1M · 0.4s first token coreweave/fp8 · $0.130 per 1M · 0.5s first token inceptron/fp4 · $0.130 per 1M · 1.1s first token morph/bf16 · $0.139 per 1M · 1.4s first token fireworks · $0.140 per 1M · 1.0s first token akashml/fp8 · $0.140 per 1M · 1.2s first token novita/fp8 · $0.140 per 1M · 1.5s first token together · $0.140 per 1M · 0.9s first token parasail/fp8 · $0.140 per 1M · 0.8s first token atlas-cloud/fp4 · $0.140 per 1M · 1.3s first token siliconflow/fp8 · $0.140 per 1M · 1.8s first token ambient/fp4 · $0.140 per 1M · 0.7s first token baidu/fp8 · $0.140 per 1M · 0.9s first token mancer/fp8 · $0.140 per 1M · 1.2s first token · not serving io-net/fp8 · $0.149 per 1M · 1.6s first token venice · $0.175 per 1M · 1.4s first token phala · $0.200 per 1M · 1.6s first token deepseek/fp8 · $0.220 per 1M · 0.9s first token wafer/fast · $0.280 per 1M · 1.3s first token cloudflare · $0.440 per 1M · 0.7s first token

28 companies serve DeepSeek V4 Flash 0731, and the dearest charges five and a half times the cheapest. Nine of them charge exactly 14 cents per million tokens, and inside that group the first token arrives after 0.7 seconds at one end and 1.8 seconds at the other. Paying the same does not buy the same. First-token figures cover the 24 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

21 companies · 2.1× price spread · 5.8× first-token spread

Same price: 5 hosts at $0.95 0.0s 0.9s 1.8s 2.7s $0.5$0.7$1$1.5 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic decart/fp4 · $0.568 per 1M · 1.4s first token chutes/int4 · $0.580 per 1M · 2.4s first token streamlake/fp8 · $0.598 per 1M · 1.7s first token · not serving inceptron/int4 · $0.600 per 1M · 0.5s first token coreweave/fp4 · $0.650 per 1M · 0.5s first token crusoe/bf16 · $0.700 per 1M · 0.4s first token parasail/int4 · $0.750 per 1M · 1.0s first token venice/int4 · $0.750 per 1M · 1.5s first token · not serving deepinfra/fp4 · $0.750 per 1M · 1.2s first token digitalocean · $0.760 per 1M · 0.9s first token · not serving siliconflow/fp8 · $0.770 per 1M · 1.4s first token novita · $0.800 per 1M · 1.6s first token atlas-cloud/int4 · $0.950 per 1M · 1.4s first token moonshotai/int4 · $0.950 per 1M · 2.0s first token baidu/fp4 · $0.950 per 1M · 0.7s first token cloudflare · $0.950 per 1M · 1.0s first token sail-research/int4 · $1.00 per 1M · 1.0s first token · not serving phala · $1.09 per 1M · 1.5s first token together · $1.20 per 1M · 0.8s first token fireworks · $0.950 per 1M · 0.5s first token

21 companies serve Kimi K2.6, and the dearest charges two times the cheapest. Five of them charge exactly 95 cents per million tokens, and inside that group the first token arrives after 0.5 seconds at one end and 2.0 seconds at the other. Paying the same does not buy the same. First-token figures cover the 16 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

18 companies · 6.5× price spread · 7.2× first-token spread

Same price: 6 hosts at $0.14 0.0s 1.2s 2.5s 3.7s $0.05$0.1$0.2$0.5 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic digitalocean · $0.068 per 1M · 1.5s first token streamlake/fp8 · $0.083 per 1M · 1.4s first token gmicloud/fp8 · $0.084 per 1M · 2.7s first token deepinfra/fp8 · $0.090 per 1M · 0.7s first token sail-research/fp4 · $0.090 per 1M · 0.8s first token siliconflow/fp8 · $0.130 per 1M · 1.2s first token alibaba/fp8 · $0.134 per 1M · 0.9s first token venice · $0.138 per 1M · 1.4s first token parasail/fp8 · $0.140 per 1M · 0.6s first token novita/fp8 · $0.140 per 1M · 1.0s first token atlas-cloud/fp4 · $0.140 per 1M · 1.0s first token coreweave/fp8 · $0.140 per 1M · 0.5s first token baidu/fp8 · $0.140 per 1M · 0.7s first token mancer/fp8 · $0.140 per 1M · 1.0s first token phala · $0.200 per 1M · 3.4s first token azure/us · $0.210 per 1M · 2.0s first token · not serving deepseek · $0.220 per 1M · 0.9s first token cloudflare · $0.440 per 1M · 0.8s first token

18 companies serve DeepSeek V4 Flash 0423, and the dearest charges six and a half times the cheapest. Six of them charge exactly 14 cents per million tokens, and inside that group the first token arrives after 0.5 seconds at one end and 1.0 seconds at the other. Paying the same does not buy the same. First-token figures cover the 17 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

18 companies · 2.9× price spread · 7.4× first-token spread

Same price: 4 hosts at $1.74 0.0s 1.2s 2.3s 3.5s $0.5$1$2 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic deepseek · $0.660 per 1M · 1.2s first token streamlake/fp8 · $0.694 per 1M · 2.5s first token gmicloud/fp8 · $0.696 per 1M · 3.2s first token digitalocean · $0.870 per 1M · 1.9s first token ionstream/fp4 · $1.13 per 1M · 1.8s first token coreweave/fp8 · $1.15 per 1M · 0.6s first token deepinfra/fp8 · $1.30 per 1M · 0.9s first token alibaba/fp8 · $1.42 per 1M · 1.6s first token novita/fp8 · $1.44 per 1M · 1.5s first token siliconflow/fp8 · $1.50 per 1M · 1.7s first token venice · $1.65 per 1M · 1.7s first token atlas-cloud/fp4 · $1.68 per 1M · 1.5s first token baidu/fp8 · $1.69 per 1M · 0.8s first token baseten/fp4 · $1.74 per 1M · 0.4s first token parasail/fp8 · $1.74 per 1M · 0.8s first token together · $1.74 per 1M · 0.7s first token fireworks · $1.74 per 1M · 1.6s first token azure/us · $1.91 per 1M · 1.5s first token · not serving

18 companies serve DeepSeek V4 Pro 0423, and the dearest charges nearly three times the cheapest. Four of them charge exactly $1.74 per million tokens, and inside that group the first token arrives after 0.4 seconds at one end and 1.6 seconds at the other. Paying the same does not buy the same. First-token figures cover the 17 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

18 companies · 1.7× price spread · 38× first-token spread

Same price: 5 hosts at $1.40 0.0s 3.4s 6.9s 10.3s $0.7$1$1.5$2 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic gmicloud/fp8 · $0.910 per 1M · 2.0s first token streamlake/fp8 · $0.966 per 1M · 3.3s first token digitalocean · $0.975 per 1M · 1.2s first token · not serving chutes/fp8 · $0.980 per 1M · 2.1s first token · not serving wafer/fp4 · $1.00 per 1M · 0.9s first token deepinfra/fp4 · $1.05 per 1M · 0.9s first token siliconflow/fp8 · $1.19 per 1M · 1.9s first token crusoe/fp8 · $1.20 per 1M · 0.6s first token phala · $1.21 per 1M · 2.0s first token · not serving atlas-cloud/fp8 · $1.26 per 1M · 1.5s first token alibaba/fp8 · $1.33 per 1M · 2.0s first token novita/fp8 · $1.38 per 1M · 2.5s first token nebius/fp8 · $1.40 per 1M · 0.7s first token parasail/fp8 · $1.40 per 1M · 0.8s first token friendli · $1.40 per 1M · 0.2s first token z-ai/fp8 · $1.40 per 1M · 9.4s first token baidu/fp8 · $1.40 per 1M · 0.8s first token venice/fp8 · $1.54 per 1M · 1.5s first token

18 companies serve GLM 5.1, and the dearest charges 69 per cent above the cheapest. Five of them charge exactly $1.40 per million tokens, and inside that group the first token arrives after 0.2 seconds at one end and 9.4 seconds at the other. Paying the same does not buy the same. First-token figures cover the 15 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

15 companies · 2.8× price spread · 7.7× first-token spread

Same price: 5 hosts at $0.95 0.0s 1.1s 2.2s 3.3s $0.5$1$2 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic inceptron/int4 · $0.670 per 1M · 0.4s first token deepinfra/fp4 · $0.680 per 1M · 0.8s first token ambient · $0.690 per 1M · 0.8s first token coreweave/int4 · $0.710 per 1M · 3.0s first token venice/int4 · $0.750 per 1M · 1.2s first token parasail/int4 · $0.760 per 1M · 0.6s first token modelrun/fp4 · $0.850 per 1M · 0.5s first token siliconflow/fp8 · $0.859 per 1M · 1.6s first token novita/int4 · $0.912 per 1M · 1.4s first token moonshotai/int4 · $0.950 per 1M · 1.8s first token cloudflare · $0.950 per 1M · 1.0s first token atlas-cloud/int4 · $0.950 per 1M · 0.8s first token together · $0.950 per 1M · 0.6s first token alibaba/fp8 · $0.950 per 1M · 2.5s first token moonshotai/highspeed · $1.90 per 1M · 1.8s first token

15 companies serve Kimi K2.7 Code, and the dearest charges more than two and a half times the cheapest. Five of them charge exactly 95 cents per million tokens, and inside that group the first token arrives after 0.6 seconds at one end and 2.5 seconds at the other. Paying the same does not buy the same. First-token figures cover the 15 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

13 companies · 2.3× price spread · 6.6× first-token spread

Same price: 8 hosts at $3.00 0.0s 1.8s 3.6s 5.5s $2$3$5$7 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic sail-research/fp4 · $2.60 per 1M · 3.7s first token morph/fp4 · $2.80 per 1M · 2.3s first token digitalocean · $2.85 per 1M · 3.5s first token deepinfra/bf16 · $2.85 per 1M · 0.8s first token fireworks · $3.00 per 1M · 2.0s first token chutes/mxfp4 · $3.00 per 1M · 1.7s first token together · $3.00 per 1M · 1.6s first token moonshotai/mxfp4 · $3.00 per 1M · 5.0s first token wafer · $3.00 per 1M · 2.0s first token modal/mxfp4 · $3.00 per 1M · 2.9s first token baseten/fp8 · $3.00 per 1M · 1.8s first token phala · $3.00 per 1M · 3.5s first token morph/fast · $6.00 per 1M · 1.4s first token

13 companies serve Kimi K3, and the dearest charges nearly two and a half times the cheapest. Eight of them charge exactly $3.00 per million tokens, and inside that group the first token arrives after 1.6 seconds at one end and 5.0 seconds at the other. Paying the same does not buy the same. First-token figures cover the 13 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

12 companies · 3.3× price spread · 4.8× first-token spread

Same price: 8 hosts at $0.30 0.0s 1.1s 2.2s 3.3s $0.2$0.5$1 First token, seconds Not serving requests at this read Listed price per 1M tokens, logarithmic coreweave/fp4 · $0.230 per 1M · 0.6s first token gmicloud/fp8 · $0.240 per 1M · 3.0s first token deepinfra/fp8 · $0.280 per 1M · 2.2s first token novita/fp8 · $0.300 per 1M · 1.8s first token venice/fp8 · $0.300 per 1M · 1.8s first token minimax/fp8 · $0.300 per 1M · 1.1s first token atlas-cloud/fp8 · $0.300 per 1M · 1.5s first token together · $0.300 per 1M · 1.0s first token streamlake/fp8 · $0.300 per 1M · 1.2s first token parasail/fp8 · $0.300 per 1M · 0.7s first token morph/fp4 · $0.300 per 1M · 1.4s first token modelrun/fp4 · $0.750 per 1M · 1.9s first token

12 companies serve MiniMax M3, and the dearest charges more than three times the cheapest. Eight of them charge exactly 30 cents per million tokens, and inside that group the first token arrives after 0.7 seconds at one end and 1.8 seconds at the other. Paying the same does not buy the same. First-token figures cover the 12 companies serving requests at this read, each one their own reported median over the previous half hour. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

Same weights. Different capabilities.

5 of 20 gpt-oss-120b hosts don't advertise tool calling.

One bar per model, longest where most companies serve it. Filled advertises tool calling and outlined does not. An agent that needs to call a function does not degrade on a host without it; the loop stops.

  • gpt-oss-120b 15 / 20
  • Kimi K3 10 / 13
  • MiniMax M3 9 / 12
  • Kimi K2.6 20 / 21
  • GLM 5.1 17 / 18
  • Nemotron 3 Super 2 / 3

Seven other models: every tracked host advertises tool calling.

Stated precision, across all 220 live endpoints

  • 10 16-bit
  • 78 8-bit
  • 60 4-bit
  • 72 not stated

14 of the 205 company-and-model pairs we track do not advertise tool calling, and 72 of 220 endpoints do not say what precision they serve at. Both are provider-advertised, not a CanaryOne functional test, and whether four bits changes an answer is a question about output that nothing on this page measures. Metadata read 17 August 2026 at 12:55 UTC. Check the capabilities on OpenRouter →

The market moves underneath you.

One box per company, showing what it charged for GLM 5.2 across 40 reads since 4 August 2026. Identical weights in all 32. Every box shares one vertical scale, and a line that drops away is a discount.

Changed price during the week 9 of 32 companies

The other 23 never moved. One dot per company, stacked where several charge the same price. Hover a dot for its precision and context window.

Held the same price all week $0.750 to $2.31 per 1M 23 of 32 companies DeepInfra fp4 · $0.750 per 1M · fp4 precision · 1024K context CoreWeave fp4 · $0.760 per 1M · fp4 precision · 256K context AkashML fp8 · $0.770 per 1M · fp8 precision · 95K context Alibaba fp8 · $0.966 per 1M · fp8 precision · 1024K context Ambient fp8 · $1.05 per 1M · fp8 precision · 198K context Morph · $1.10 per 1M · precision not stated SiliconFlow fp8 · $1.19 per 1M · fp8 precision · 1024K context AtlasCloud fp8 · $1.26 per 1M · fp8 precision · 1024K context Wafer · $1.26 per 1M · fp4 precision · 1024K context BaseTen fp8 · $1.40 per 1M · fp8 precision · 1024K context Cloudflare · $1.40 per 1M · precision not stated · 256K context Crusoe fp8 · $1.40 per 1M · fp8 precision · 1024K context Fireworks · $1.40 per 1M · precision not stated · 1024K context Friendli · $1.40 per 1M · precision not stated · 1024K context Parasail fp4 · $1.40 per 1M · fp4 precision · 256K context Together · $1.40 per 1M · precision not stated · 500K context Venice fp8 · $1.40 per 1M · fp8 precision · 977K context Z.AI fp8 · $1.40 per 1M · fp8 precision · 1024K context BaseTen fast · $2.10 per 1M · fp8 precision · 1024K context Cloudflare fast · $2.10 per 1M · precision not stated · 256K context Fireworks fast · $2.10 per 1M · precision not stated · 1024K context Wafer fast · $2.10 per 1M · fp4 precision · 1024K context Alibaba fast · $2.31 per 1M · fp8 precision · 1024K context $1.40 9 companies $2.10 4 companies
The 23 companies that held their price, as a table
Company Price per 1M Precision Context
DeepInfra fp4 $0.750 fp4 1024K
CoreWeave fp4 $0.760 fp4 256K
AkashML fp8 $0.770 fp8 95K
Alibaba fp8 $0.966 fp8 1024K
Ambient fp8 $1.050 fp8 198K
Morph $1.100 not stated
SiliconFlow fp8 $1.190 fp8 1024K
Wafer $1.260 fp4 1024K
AtlasCloud fp8 $1.260 fp8 1024K
Z.AI fp8 $1.400 fp8 1024K
Fireworks $1.400 not stated 1024K
Cloudflare $1.400 not stated 256K
Friendli $1.400 not stated 1024K
Parasail fp4 $1.400 fp4 256K
Venice fp8 $1.400 fp8 977K
Together $1.400 not stated 500K
Crusoe fp8 $1.400 fp8 1024K
BaseTen fp8 $1.400 fp8 1024K
Wafer fast $2.100 fp4 1024K
Fireworks fast $2.100 not stated 1024K
Cloudflare fast $2.100 not stated 256K
BaseTen fast $2.100 fp8 1024K
Alibaba fast $2.310 fp8 1024K

The steepest move was a discount that cut one company's price by 93 per cent before expiring above where the week started. Every price here is what a buyer pays after whatever discount was running at that moment. Metadata read 17 August 2026 at 12:55 UTC. Check these prices on OpenRouter →

How we measure it.

Two layers, kept apart. Host metadata is public: prices, capabilities, advertised throughput and availability, read several times a day. Measured workloads are ours: the same task suite run against multiple hosts serving identical weights, where total spend is divided by the tasks that actually completed, so failed work stays in the economics instead of disappearing from it.

We compare hosts within one model and never treat different weights as equivalent workloads. Timestamps sit beside measurements. Snapshots are not trends, and this record is days old rather than months. Relative positions between hosts change between runs, so we report the size of a spread rather than naming a best host. Prices, capabilities and precision are what a company publishes about itself, and a published capability is a claim rather than a test we have run. First-token figures are each company's own reported median over the half hour before we read it.

What happens on your workload?

Market averages can tell you where to look. Your own workload tells you what to deploy.

Benchmark your workload → Get early access →

EARLY ACCESS

Get early access.

We'll reach out when early access opens.

\n