GLM-5.3-Flash (NVFP4)
GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model (MIT licence, August 2026). Despite the name it is a large mixture-of-experts network — 320B parameters, 18B active per token — built for coding and long-horizon agent work over long context at low serving cost. Reasoning effort is selectable (low, high, max), and it reads images alongside text. The developer declares English and Chinese; test other languages on your own material. It is fulfilled through global providers under our European contract, and requests go only to backends that do not retain prompts or completions.
model id: glm53-flash
from openai import OpenAI client = OpenAI( api_key="hb-...", base_url="https://api.heabsy.com/v1",) r = client.chat.completions.create( model="glm53-flash", messages=[{"role": "user", "content": "Hello"}],)print(r.choices[0].message.content)
curl https://api.heabsy.com/v1/chat/completions \ -H "Authorization: Bearer $HEABSY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "glm53-flash", "messages": [{"role": "user", "content": "Hello"}] }'
import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.HEABSY_API_KEY, baseURL: "https://api.heabsy.com/v1",}); const r = await client.chat.completions.create({ model: "glm53-flash", messages: [{ role: "user", content: "Hello" }],});console.log(r.choices[0].message.content);
Price list, updated 12 September 2026. No minimum spend, no subscription.
About the model
- Developer
- Z.ai
- Released
- August 2026
- Licence
- MIT
- Architecture
- mixture of experts · 320B parameters, 18B active per token
- Input
- text, images
- Reasoning
- selectable effort
- Languages named by the developer
- English, Chinese
Source: the developer's model card, as of 17 September 2026.
Licence: what it means for your company
MIT. A permissive open-source licence: you may use, modify and sell products built on the model, including commercially; when you redistribute the weights, the licence and copyright notices go with them.
A summary of the licence, not legal advice.
Benchmarks reported by the developer
- DeepSWE v1.1
- 63.4
- AutomationBench
- 48.8
Figures from the developer's model card, not our measurement. Source: docs.z.ai/guides/vlm/glm-5.3-flash
Settings recommended by the developer
- temperature 1, top_p 0.95
Technical details
- Context window
- 1M tokens
- Max output
- 32,768 tokens
- Throughput limit
- 50M tok/min
- Tokenizer
- Other
- Streaming
- yes
- Open weights
- zai-org/GLM-5.3-Flash
- max_tokens
- 1 … 32768
- temperature
- 0 … 2
- top_p
- 0 … 1
- frequency_penalty
- -2 … 2
- presence_penalty
- -2 … 2
- stop
- array
- seed
- integer
This model is fulfilled through global providers under our European contract: one DPA with an EU company, one invoice for every model in the catalog, no US counterparty on your paperwork. Compute may run outside the EEA — if you need strict EEA-only processing, use our own-hardware model.
Questions about this model
›Where is GLM-5.3-Flash computed?
Outside the EEA, through a global provider under our European contract. You get one DPA with an EU company and one invoice, but the compute itself does not run in the EEA — we state that instead of hiding it.
›Does GLM-5.3-Flash store prompts and completions?
No. This model runs with zero data retention: request content is not written to durable storage and nothing is used for training.
›How large is the GLM-5.3-Flash context window?
1M tokens, with up to 32,768 tokens of output.
›What does GLM-5.3-Flash cost?
$0.04 per million input tokens and $0.12 per million output tokens. No minimum spend and no subscription; the same price list covers the API and the invoice.
›Is there a throughput limit on GLM-5.3-Flash?
Yes, 50M tokens per minute. Higher limits are a question of volume, not of principle.
›Which languages does GLM-5.3-Flash support?
The developer names English, Chinese. Languages not on that list may still work, but they are not claimed — test on your own material first.
›Under which licence is GLM-5.3-Flash released?
MIT. A permissive open-source licence: you may use, modify and sell products built on the model, including commercially; when you redistribute the weights, the licence and copyright notices go with them.
›Can reasoning be switched off in GLM-5.3-Flash?
The developer lets you choose how much the model reasons; the card does not document turning reasoning off completely.
Compare with similar models
| Model | Developer | Parameters | Context | Licence | Input / 1M |
|---|---|---|---|---|---|
| GLM-5.3-Flash (NVFP4) | Z.ai | 320B / 18B | 1M | MIT | $0.04 |
| GLM-4.7-Flash | Z.ai | 30B / 3B | 195K | MIT | $0.09 |
| DeepSeek-V4-Flash-0731 | DeepSeek | 284B / 13B | 1.3M | MIT | $0.06 |
| DeepSeek V4 Flash | DeepSeek | 284B / 13B | 1M | MIT | $0.25 |
30 more models behind the same key and the same invoice.
Browse the catalog