The cheapest AI APIs in 2026, ranked by intelligence per dollar
A rundown of the cheapest AI APIs in 2026 — DeepSeek V4.1 Flash, GLM-5.3 Flash, Gemini 3.8 Flash, Step 5 Preview and more — ranked by how much intelligence each dollar actually buys.
The cheapest AI APIs in 2026
Raw price per token is only half the story. A model that costs $0.08 per 1M input tokens but scores 14 on the intelligence index can be worse value than a $0.15 model scoring 42. That’s why we rank by intelligence per dollar — tokens per dollar weighted by benchmarked quality — rather than by sticker price alone.
The cheap end of the catalog, side by side
| Model | Input / 1M | Output / 1M | Intelligence index |
|---|---|---|---|
| Nemotron 3 Super | $0.08 | $0.45 | 14 |
| Llama 4 Scout | $0.10 | $0.30 | 6 |
| DeepSeek V4.1 Flash | $0.15 | $0.60 | 40 |
| GLM-5.3 Flash | $0.15 | $0.50 | 42 |
| GPT-5.6 Luna | $0.20 | $1.20 | 38 |
| Llama 4 Maverick | $0.1875 | $0.6525 | 9 |
| GLM-5.3 FlashX | $0.37 | $1.25 | unbenchmarked |
| Mistral Large 3 | $0.50 | $1.50 | 10 |
| DeepSeek V4 Pro | $0.66 | $1.98 | 36 |
| Gemini 3.8 Flash | $0.75 | $3.75 | 41 |
| Step 5 Preview | $1.00 | $2.70 | 44 |
Prices are per 1M tokens (standard list price; DeepSeek shows the off-peak rate — peak hours are 2x). Each model page shows its last-verified date. Indices are on Artificial Analysis’s current v4.3 scale, which is lower than the scale used before September 2026.
The standouts
DeepSeek V4.1 Flash — full pricing. A 40 index at $0.15 / $0.60 (off-peak; peak is 2x). It replaces V4 Flash — retired in September 2026 — and it tops our intelligence-per-dollar chart, which makes it the default pick for bulk generation, batch processing and anything where token volume dominates.
GLM-5.3 Flash — full pricing. The highest index under $0.75: 42 at $0.15 / $0.50, with cached input at $0.03. A native multimodal Flash tier for high-volume coding and agent workloads.
Step 5 Preview — full pricing. StepFun’s new reasoning model is the surprise of the month: an index of 44 — above Grok 4.6 and Kimi K3 — at $1.00 / $2.70, roughly a tenth of what the premium flagships charge for that class of score. It is a preview release and unusually verbose, and the price comes from Artificial Analysis rather than a StepFun price list, so treat it as a pilot candidate rather than a default.
GLM-5.3 FlashX — full pricing. Z.ai’s speed-tuned sibling of GLM-5.3 Flash: $0.37 / $1.25 for a 1M-token context at roughly 200 tokens per second. It is not on the Artificial Analysis leaderboard yet, so it has no index here — listed for completeness.
Gemini 3.8 Flash — full pricing. The best balance on this list: an index of 41 at $0.75 / $3.75 (promotional through 31 December 2026), with a 1M-token context window and around 277 tokens per second. If you want high quality without a high price, this is the one to test first.
DeepSeek V4 Pro — full pricing. Index 36 at $0.66 / $1.98 (off-peak; $1.32 / $3.96 at peak). For reasoning-heavy, cost-sensitive work it remains one of the strongest value options we track.
Nemotron 3 Super — full pricing. One of the cheapest input prices on the list at $0.08, but with an index of 14 it’s a bulk-orchestration workhorse, not a reasoning model.
The caveats
- The index is one benchmark. Artificial Analysis’s Intelligence Index rewards strong general capability, but your specific workload — coding, writing, long-document work — can rank these differently.
- Cheap tokens still add up. Intelligence per dollar favors low prices, which is why a cheap mid-tier model can out-rank an expensive frontier one. Use it as a shortlist, not a final answer.
- Prices move. Every figure carries a last-verified date, and we refresh the catalog weekly.
The fastest way to get a definitive answer is to set a real budget: see what $20 a month buys or open the calculator and compare every model at your exact usage.