The open-weight frontier: DeepSeek vs GLM vs Qwen in 2026

DeepSeek V4, GLM-5.3 and Qwen 3.8 are the open-weight models to beat. Compared on price, design and serving speed.

The open-weight frontier: DeepSeek vs GLM vs Qwen

The closed-model race (Claude vs GPT vs Grok) gets the attention, but the open-weight race is where the value lives. Three families dominate in 2026: DeepSeek, Z.ai’s GLM and Alibaba’s Qwen. Here’s how they compare.

The prices

Model Price in / out per 1M Intelligence index
DeepSeek V4 Pro $0.66 / $1.98 53
DeepSeek V4.1 Flash ~$0.15 / $0.60 40
GLM-5.3 (max) $1.40 / $4.40 60
Qwen 3.8 (2.4T A95B) $2 / $6 58

DeepSeek — the value king

DeepSeek V4 Pro and the new V4.1 Flash are the cost leaders, full stop: V4.1 Flash scores 40 on the index for $0.15 / $0.60 off-peak, and V4 Pro trades speed for reasoning depth at $0.66 / $1.98. Practitioners report building entire websites for cents — DeepSeek retired the older V4 Flash in September 2026 and now serves that name from V4.1 Flash. The trade-off: DeepSeek checks its work a lot, so real-world tasks run slower than the token speed suggests.

GLM-5.3 — the design pick

GLM-5.3 is Z.ai’s flagship: index 45 (matching Kimi K3) at $1.40 / $4.40, with genuinely strong front-end design — one of the best-designing open-weight models. A cheaper GLM-5.3 Flash tier (index 42 at $0.15 / $0.50) now covers high-volume work. The weakness is serving: Z.ai doesn’t yet have the compute to run it fast, so expect slow responses and quick usage burn.

Qwen 3.8 — the strong design competitor

Alibaba’s Qwen 3.8 (2.4T A95B) is a capable open-weight model (index 40) with strong design work, and the proprietary Qwen3.8 Max sits alongside it at the same $2 / $6. Their catch is subscription usage: practitioners report burning 57% of a plan in just a handful of prompts.

Which one to pick

  • Maximum value: DeepSeek V4 Pro — the cheapest frontier option we track.
  • Best design: GLM-5.3 or Qwen 3.8.
  • Local / offline: DeepSeek V4.1 Flash — it still runs on hardware like the NVIDIA DGX Spark.

All three are open-weight, so you can also self-host any of them. Price out your workload in the AI intelligence calculator to see which wins on tokens per dollar for you.

Advertisement