Comprehensive Pricing Report • June 2026

LLM Provider Pricing Matrix

4 models × 6 providers Per 1M tokens (input / output)
Overview

Summary: DeepSeek V4 Pro remains the cheapest high-performance option. GLM-5.2 and Kimi K2.7 Code are competitive on DeepInfra and OpenRouter. Nemotron 3 has the lowest prices on several providers but fewer options.

Full Pricing Matrix

Input / Output Pricing per 1M Tokens (June 2026)

Model Baseten DeepInfra OpenRouter Together AI Fireworks AI NVIDIA NIM
DeepSeek V4 Pro $1.74 / $3.48
$0.145 cached
$0.435 / $0.87 $1.74 / $3.48
$0.20 cached
$1.74 / $3.48 $1.39 / $2.78
GLM-5.2 $1.50 / $0.30 $0.95 / $3.00
$0.18 cached
$0.95 / $3.00 $1.40 / $4.40
$0.26 cached
$1.40 / $4.40
$0.26 cached
Kimi K2.7 Code $0.95 / $4.00
$0.16 cached
$0.74 / $3.50
$0.15 cached
$0.95 / $4.00
$0.19 cached
$0.95 / $4.00
$0.19 cached
Nemotron 3 $0.10 / $0.50 $0.60 / $3.60
Notes

DeepSeek V4 Pro: OpenRouter is dramatically cheaper. NVIDIA NIM is second best.

GLM-5.2: DeepInfra and OpenRouter are best. Baseten has unusually low output price.

Kimi K2.7 Code: DeepInfra currently offers the lowest price.

Nemotron 3: Limited public serverless options. DeepInfra is the clearest winner.

— = No public serverless pricing found (dedicated only or not yet available).