Comprehensive Pricing Report • June 2026
LLM Provider Pricing Matrix
Overview
Summary: DeepSeek V4 Pro remains the cheapest high-performance option. GLM-5.2 and Kimi K2.7 Code are competitive on DeepInfra and OpenRouter. Nemotron 3 has the lowest prices on several providers but fewer options.
Full Pricing Matrix
Input / Output Pricing per 1M Tokens (June 2026)
| Model | Baseten | DeepInfra | OpenRouter | Together AI | Fireworks AI | NVIDIA NIM |
|---|---|---|---|---|---|---|
| DeepSeek V4 Pro | — | $1.74 / $3.48 $0.145 cached |
$0.435 / $0.87 | $1.74 / $3.48 $0.20 cached |
$1.74 / $3.48 | $1.39 / $2.78 |
| GLM-5.2 | $1.50 / $0.30 | $0.95 / $3.00 $0.18 cached |
$0.95 / $3.00 | $1.40 / $4.40 $0.26 cached |
$1.40 / $4.40 $0.26 cached |
— |
| Kimi K2.7 Code | $0.95 / $4.00 $0.16 cached |
$0.74 / $3.50 $0.15 cached |
— | $0.95 / $4.00 $0.19 cached |
$0.95 / $4.00 $0.19 cached |
— |
| Nemotron 3 | — | $0.10 / $0.50 | — | $0.60 / $3.60 | — | — |
Notes
DeepSeek V4 Pro: OpenRouter is dramatically cheaper. NVIDIA NIM is second best.
GLM-5.2: DeepInfra and OpenRouter are best. Baseten has unusually low output price.
Kimi K2.7 Code: DeepInfra currently offers the lowest price.
Nemotron 3: Limited public serverless options. DeepInfra is the clearest winner.
— = No public serverless pricing found (dedicated only or not yet available).