Fastest AI models in 2026
Speed is the hidden multiplier in AI work: more shots on goal per hour often beats a higher quality ceiling. Here's how the frontier ranks on raw speed.
S tier — Fastest
Over 200 tokens per second on OpenRouter — the fastest model in the world.
- Gemini 3.7 Flash — $0.75 / $3.75 per 1M tokens · Intelligence 39
A tier — Very fast
Completes tasks quickly and efficiently on real codebases.
- Grok 4.6 — $2 / $6 per 1M tokens · Intelligence 44
B tier — Fast
Luna is neck and neck with Terra in real workflows.
- GPT-5.6 Terra — $2 / $12 per 1M tokens · Intelligence 42
- GPT-5.6 Luna — $0.20 / $1.20 per 1M tokens · Intelligence 38
C tier — Average
Fable 5.1 is faster than Opus 5 but still on the slower end.
- Claude Sonnet 5 — $2 / $10 per 1M tokens · Intelligence 38
- Claude Fable 5.1 — $10 / $50 per 1M tokens · Intelligence 53
D tier — Slow
DeepSeek checks its work heavily; Kimi is limited by serving compute.
- DeepSeek V4 Pro — $0.66 / $1.98 per 1M tokens · Intelligence 36
- Qwen3.8 2.4T A95B — $2 / $6 per 1M tokens · Intelligence 40
- Kimi K3 — $3 / $15 per 1M tokens · Intelligence 44
E tier — Very slow
Thinks a lot — simple tasks can take up to 30 minutes.
- Claude Opus 5 — $5 / $25 per 1M tokens · Intelligence 51
F tier — Slowest
Z.ai doesn't yet have the compute to serve it fast.
- GLM-5.3 (max) — $1.40 / $4.40 per 1M tokens · Intelligence 45
Frequently asked questions
Which AI models are best for fastest ai models in 2026?
The top tier is currently Gemini 3.7 Flash. Over 200 tokens per second on OpenRouter — the fastest model in the world.
How are these AI model rankings made?
These rankings combine practitioner consensus from the frontier-coding community with our own live pricing data (intelligence index, tokens per dollar). They are qualitative judgments, not benchmarks — always cross-check with the calculator for your own workload.