The cheapest AI APIs in 2026, ranked by intelligence per dollar

A rundown of the cheapest AI APIs in 2026 — DeepSeek V4.1 Flash, GLM-5.3 Flash, Gemini 3.8 Flash, Step 5 Preview and more — ranked by how much intelligence each dollar actually buys.

The cheapest AI APIs in 2026

Raw price per token is only half the story. A model that costs $0.08 per 1M input tokens but scores 14 on the intelligence index can be worse value than a $0.15 model scoring 42. That’s why we rank by intelligence per dollar — tokens per dollar weighted by benchmarked quality — rather than by sticker price alone.

The cheap end of the catalog, side by side

Model Input / 1M Output / 1M Intelligence index
Nemotron 3 Super $0.08 $0.45 14
Llama 4 Scout $0.10 $0.30 6
DeepSeek V4.1 Flash $0.15 $0.60 40
GLM-5.3 Flash $0.15 $0.50 42
GPT-5.6 Luna $0.20 $1.20 38
Llama 4 Maverick $0.1875 $0.6525 9
GLM-5.3 FlashX $0.37 $1.25 unbenchmarked
Mistral Large 3 $0.50 $1.50 10
DeepSeek V4 Pro $0.66 $1.98 36
Gemini 3.8 Flash $0.75 $3.75 41
Step 5 Preview $1.00 $2.70 44

Prices are per 1M tokens (standard list price; DeepSeek shows the off-peak rate — peak hours are 2x). Each model page shows its last-verified date. Indices are on Artificial Analysis’s current v4.3 scale, which is lower than the scale used before September 2026.

The standouts

DeepSeek V4.1 Flash — full pricing. A 40 index at $0.15 / $0.60 (off-peak; peak is 2x). It replaces V4 Flash — retired in September 2026 — and it tops our intelligence-per-dollar chart, which makes it the default pick for bulk generation, batch processing and anything where token volume dominates.

GLM-5.3 Flash — full pricing. The highest index under $0.75: 42 at $0.15 / $0.50, with cached input at $0.03. A native multimodal Flash tier for high-volume coding and agent workloads.

Step 5 Preview — full pricing. StepFun’s new reasoning model is the surprise of the month: an index of 44 — above Grok 4.6 and Kimi K3 — at $1.00 / $2.70, roughly a tenth of what the premium flagships charge for that class of score. It is a preview release and unusually verbose, and the price comes from Artificial Analysis rather than a StepFun price list, so treat it as a pilot candidate rather than a default.

GLM-5.3 FlashX — full pricing. Z.ai’s speed-tuned sibling of GLM-5.3 Flash: $0.37 / $1.25 for a 1M-token context at roughly 200 tokens per second. It is not on the Artificial Analysis leaderboard yet, so it has no index here — listed for completeness.

Gemini 3.8 Flash — full pricing. The best balance on this list: an index of 41 at $0.75 / $3.75 (promotional through 31 December 2026), with a 1M-token context window and around 277 tokens per second. If you want high quality without a high price, this is the one to test first.

DeepSeek V4 Pro — full pricing. Index 36 at $0.66 / $1.98 (off-peak; $1.32 / $3.96 at peak). For reasoning-heavy, cost-sensitive work it remains one of the strongest value options we track.

Nemotron 3 Super — full pricing. One of the cheapest input prices on the list at $0.08, but with an index of 14 it’s a bulk-orchestration workhorse, not a reasoning model.

The caveats

  1. The index is one benchmark. Artificial Analysis’s Intelligence Index rewards strong general capability, but your specific workload — coding, writing, long-document work — can rank these differently.
  2. Cheap tokens still add up. Intelligence per dollar favors low prices, which is why a cheap mid-tier model can out-rank an expensive frontier one. Use it as a shortlist, not a final answer.
  3. Prices move. Every figure carries a last-verified date, and we refresh the catalog weekly.

The fastest way to get a definitive answer is to set a real budget: see what $20 a month buys or open the calculator and compare every model at your exact usage.

Advertisement