Fastest AI models in 2026: tokens per second compared

Gemini 3.7 Flash is the fastest AI model in 2026, topping 200 tokens per second. Here's how the fastest models rank on speed — and why speed is a bigger deal than most people think.

Fastest AI models in 2026

Speed is the hidden multiplier in AI work. Because most models don’t nail a task on the first try, the number of “shots on goal” you can get per hour is often more important than the model’s ceiling quality. A fast model you can iterate on beats a slow genius you have to coax.

Here’s how the current frontier ranks on raw speed.

How this ranking was compiled: the speed tiers reflect practitioner-reported throughput, drawn from a public frontier-ranking discussion (August 2026) and cross-checked against the speed figures we verify on model pages (e.g. Gemini 3.7 Flash at 200+ tokens/sec). Real throughput varies by host and workload, so treat the tiers as a guide, not a benchmark.

The speed tier list

Tier Models
Fastest Gemini 3.7 Flash (200+ tokens/sec)
A Grok 4.6
B GPT-5.6 Terra
C Claude Sonnet 5
D DeepSeek V4 Pro, Qwen 3.8, Kimi K3
E Claude Opus 5
F GLM-5.3

The speed king: Gemini 3.7 Flash

Gemini 3.7 Flash is the fastest model in the world right now — on OpenRouter it pumps out over 200 tokens per second. That’s fast enough to do massive site refactors and iterate on designs almost in real time.

The trade-off is that it needs multiple attempts: one-shot capability is weak, so you burn that speed iterating. But if the task is “try lots of variations quickly”, nothing else comes close. It’s also one of the cheaper models we track at $0.75 / $3.75 per 1M tokens, which makes speed plus price a compelling combo for volume work.

The fast-and-strong middle

Grok 4.6 is the best speed-for-quality balance: A-tier speed at $2 / $6, and it completes tasks quickly and efficiently on real codebases. GPT-5.6 Terra lands in B tier — noticeably faster than you’d expect for its class. GPT-5.6 Sol is slow by default, but the subscription’s “fast mode” pulls it back to mid-pack.

Claude Sonnet 5 is C tier on speed, and Kimi K3 and DeepSeek V4 Pro are D tier — DeepSeek looks fast on paper (token speed), but it checks its work so much that real-world tasks run long.

The slow geniuses

Claude Opus 5 is E tier on speed: it “thinks a lot”, and practitioners report simple tasks taking up to half an hour. It’s slow — but the output quality is why it still sits near the top of the intelligence charts (index 51). Speed is the price you pay for it.

GLM-5.3 is F tier because Z.ai doesn’t have the compute to serve it quickly yet — expect that to improve as it opens up to more providers.

Why speed belongs in the value equation

Most people compare AI models only on quality and price. Speed is the third axis, and it directly affects your cost: the faster a model, the more useful work you complete per subscription, and the fewer shots you waste.

The AI intelligence calculator prices models by tokens per dollar — but when you’re comparing Gemini vs Grok or deciding between Claude and OpenAI, remember that a model that finishes in seconds changes your per-hour economics even if the per-token price is identical.

Fastest overall: Gemini 3.7 Flash · Best speed + quality: Grok 4.6 · Slowest but best quality: Claude Opus 5.

Advertisement