GLM-5.3 FlashX pricing in 2026
What one task costs on GLM-5.3 FlashX
Per-million-token prices are hard to feel, so here are typical token counts for common jobs priced with GLM-5.3 FlashX. Figures are USD at the rates above and exclude prompt caching and batch discounts.
| Task | Tokens (in / out) | Cost |
|---|---|---|
| Code review of a 500-line pull request | 60,000 / 4,000 | $0.027 |
| Debug one failing test | 40,000 / 6,000 | $0.022 |
| Generate a landing page in one shot | 3,000 / 30,000 | $0.039 |
| Draft a 1,200-word article | 2,000 / 2,500 | $0.0039 |
| Answer a question from a 20-page document | 25,000 / 800 | $0.010 |
| Multi-step agent run (10 tool calls) | 250,000 / 25,000 | $0.124 |
| All of the above, once | $0.226 |
GLM-5.3 FlashX by Z.ai costs $0.37 per 1M input tokens and $1.25 per 1M output tokens — about 970.9k tokens per $1. Prices last verified 2026-09-19.
The verdict
The high-speed variant of GLM-5.3 Flash, at $0.37 / $1.25 per 1M tokens with a 1M-token context window and roughly 200 tokens per second. Artificial Analysis has not benchmarked it yet, so we cannot place it on the intelligence index — treat the rating here as a tier estimate, not a measurement. If it lands anywhere near GLM-5.3 Flash's index of 42, it will be a strong value endpoint.
A concrete example: sending 100,000 input tokens and getting back 20,000 output tokens with GLM-5.3 FlashX costs about $0.06 at current prices.
Who is GLM-5.3 FlashX for?
GLM-5.3 FlashX is Z.ai's speed-first variant of GLM-5.3 Flash, announced 18 September 2026. It keeps the 1M-token context window and the same hybrid sparse-and-linear attention design, but is tuned for throughput — around 200 tokens per second — at $0.37 per 1M input and $1.25 per 1M output, with cache hits at $0.075. That is roughly 2.5x GLM-5.3 Flash's per-token price for a big jump in serving speed.
The honest caveat is measurement: FlashX is not on the Artificial Analysis leaderboard yet, so it has no Intelligence Index and the intelligence-per-dollar figure on this page is based on a tier estimate rather than a benchmark. It is a reasonable pick for latency-sensitive batch and agent workloads where GLM-5.3 Flash's throughput is the bottleneck — but if you need a verified index today, GLM-5.3 Flash (42) or GLM-5.3 (45) are the safer choices.
Strengths & weaknesses
- Strength: GLM-5.3 quality at high throughput, with a first-party price
- Best for: Bulk, Agentic, Long-context, Value
- Weakness: Not yet benchmarked on Artificial Analysis, so it has no Intelligence Index (the score falls back to a tier estimate). Costs ~2.5x GLM-5.3 Flash per token.
GLM-5.3 FlashX key facts
| Provider | GLM pricing |
|---|---|
| Input price | $0.37 / 1M tokens |
| Output price | $1.25 / 1M tokens |
| Tokens / $1 | 970.9k |
| Intelligence / $1 | 776,699 |
| Context window | 1.0M tokens |
| Tier | balanced |
| Speed | Very fast (200+ tokens/sec) |
| Released | 2026-09-18 |
| Last verified | 2026-09-19 |
What GLM-5.3 FlashX delivers for your budget
For a coding workload, here is roughly what each monthly budget buys:
| Budget | Tokens / month | Months of workload |
|---|---|---|
| $20 | 34.9M | 53.7x |
| $50 | 87.2M | 134.2x |
| $100 | 174.5M | 268.5x |
Other GLM models
How GLM-5.3 FlashX compares to its closest rivals
Rather than repeating the full catalog on every page, here are the three current models nearest to GLM-5.3 FlashX on the intelligence index — and why you might still pick one instead:
- Fugu Max (Sakana) — $2 input / $6 output per 1M tokens. Why you might pick it: Multi-agent orchestration as a single endpoint.
- Fugu Ultra v2 (Sakana) — $5 input / $30 output per 1M tokens. Why you might pick it: Highest-capacity Fugu tier for hard multi-step tasks.
- GPT-6 Astra (ChatGPT) — $10 input / $50 output per 1M tokens. Why you might pick it: Frontier reasoning inside ChatGPT Plus and Pro.
GLM-5.3 FlashX — frequently asked questions
How much does GLM-5.3 FlashX cost?
GLM-5.3 FlashX by Z.ai costs $0.37 per 1M input tokens and $1.25 per 1M output tokens.
How many tokens per dollar does GLM-5.3 FlashX give you?
On a typical 1:3 input-to-output mix, GLM-5.3 FlashX delivers around 970.9k tokens per dollar — an intelligence-per-dollar score of 776699.
What is GLM-5.3 FlashX's context window?
GLM-5.3 FlashX supports a context window of 1.0M tokens.
Is GLM-5.3 FlashX worth it in 2026?
GLM-5.3 FlashX is GLM's balanced tier model. Use the AI Intelligence Calculator to see how its value compares to every other model at your budget.