The open-weight frontier: DeepSeek vs GLM vs Qwen in 2026
DeepSeek V4, GLM-5.3 and Qwen 3.8 are the open-weight models to beat. Compared on price, design and serving speed.
The open-weight frontier: DeepSeek vs GLM vs Qwen
The closed-model race (Claude vs GPT vs Grok) gets the attention, but the open-weight race is where the value lives. Three families dominate in 2026: DeepSeek, Z.ai’s GLM and Alibaba’s Qwen. Here’s how they compare.
The prices
| Model | Price in / out per 1M | Intelligence index |
|---|---|---|
| DeepSeek V4 Pro | $0.66 / $1.98 | 53 |
| DeepSeek V4.1 Flash | ~$0.15 / $0.60 | 40 |
| GLM-5.3 (max) | $1.40 / $4.40 | 60 |
| Qwen 3.8 (2.4T A95B) | $2 / $6 | 58 |
DeepSeek — the value king
DeepSeek V4 Pro and the new V4.1 Flash are the cost leaders, full stop: V4.1 Flash scores 40 on the index for $0.15 / $0.60 off-peak, and V4 Pro trades speed for reasoning depth at $0.66 / $1.98. Practitioners report building entire websites for cents — DeepSeek retired the older V4 Flash in September 2026 and now serves that name from V4.1 Flash. The trade-off: DeepSeek checks its work a lot, so real-world tasks run slower than the token speed suggests.
GLM-5.3 — the design pick
GLM-5.3 is Z.ai’s flagship: index 45 (matching Kimi K3) at $1.40 / $4.40, with genuinely strong front-end design — one of the best-designing open-weight models. A cheaper GLM-5.3 Flash tier (index 42 at $0.15 / $0.50) now covers high-volume work. The weakness is serving: Z.ai doesn’t yet have the compute to run it fast, so expect slow responses and quick usage burn.
Qwen 3.8 — the strong design competitor
Alibaba’s Qwen 3.8 (2.4T A95B) is a capable open-weight model (index 40) with strong design work, and the proprietary Qwen3.8 Max sits alongside it at the same $2 / $6. Their catch is subscription usage: practitioners report burning 57% of a plan in just a handful of prompts.
Which one to pick
- Maximum value: DeepSeek V4 Pro — the cheapest frontier option we track.
- Best design: GLM-5.3 or Qwen 3.8.
- Local / offline: DeepSeek V4.1 Flash — it still runs on hardware like the NVIDIA DGX Spark.
All three are open-weight, so you can also self-host any of them. Price out your workload in the AI intelligence calculator to see which wins on tokens per dollar for you.