Checked October 2, 2026. GLM-5.3 costs $1.40 per 1M input tokens, $0.26 per 1M cached input tokens and $4.40 per 1M output tokens at Z.ai; on OpenRouter (z-ai/glm-5.3) 39 provider endpoints list between $0.12 and $2.80 for input. The weights have been public on Hugging Face since August 25 under a permissive license: 753 billion parameters, about 756 GB in FP8. On a Mac that leaves only a Mac Studio with 256 or 512 GB of unified memory, because the smallest Unsloth GGUF quantization is already 216.7 GB.
When ai-on-mac.com first checked GLM-5.3 on August 21, Z.ai was still holding back the weights for a safety review, and the local requirements could only be estimated. This version replaces those estimates with the published checkpoint, the license text, the real GGUF file sizes and the current API prices. We have not run GLM-5.3 on a Mac ourselves, so this article contains no Apple silicon speed figures.
GLM-5.3 at a glance
| Item | GLM-5.3, as of October 2, 2026 | Source |
|---|---|---|
| Developer | Z.ai | Z.ai docs |
| Release | API launch on August 14, 2026; official Hugging Face repository dated August 25 | Hugging Face |
| Base model | same base model as GLM-5.2, all gains from post-training | Z.ai docs |
| Parameters | 753.3 billion in the checkpoint; mixture of experts with 256 experts, 8 plus 1 shared expert active per token | Hugging Face, config.json |
| Weights | FP8, about 756 GB; separate BF16 repository | Hugging Face |
| License | GLM-5.3 License: free use, modification and commercial use; security review only for Model-as-a-Service providers above $10 billion annual revenue | LICENSE |
| Input and output | text in, text out | Z.ai docs |
| Context and output limit | 1M tokens context, up to 128K output tokens | Z.ai docs |
| Reasoning | always on; effort low, high or max, default max | Z.ai docs |
| Z.ai API price | $1.40 input, $0.26 cached input, $4.40 output per 1M tokens | Z.ai pricing |
| OpenRouter | z-ai/glm-5.3, 39 provider endpoints, $0.12 to $2.80 input | OpenRouter |
| Ollama | only glm-5.3:cloud, which runs on Ollama’s servers | Ollama |
| Local on a Mac | only a Mac Studio with 256 or 512 GB, GGUF files from 216.7 GB | Unsloth GGUF, Apple |
Can GLM-5.3 run locally on a Mac?
Yes, but only on the largest Mac Studio configurations and only in heavily compressed form. The official FP8 checkpoint needs about 756 GB, which no Mac offers. Apple currently sells the Mac Studio with M5 Ultra and 256 GB or 512 GB of unified memory; every MacBook Pro, Mac mini and the Mac Studio with M5 Max stop at 128 GB.
For llama.cpp, Unsloth publishes GGUF quantizations of GLM-5.3. The file sizes below come from the Hugging Face file listing on October 2, 2026. Apple’s 256 GB and 512 GB are binary gigabytes, about 275 GB and 550 GB in the decimal units Hugging Face uses. macOS, the runtime and the KV cache for long contexts need the remaining space.
| Unsloth GGUF | File size | Realistic Mac (our assessment) |
|---|---|---|
| UD-IQ1_S | 216.7 GB | Mac Studio with 256 GB, little room for context |
| UD-IQ2_M | 238.6 GB | Mac Studio with 256 GB, very tight |
| UD-Q2_K_XL | 253.9 GB | Mac Studio with 512 GB |
| UD-Q3_K_XL | 343.0 GB | Mac Studio with 512 GB |
| UD-IQ4_XS | 365.3 GB | Mac Studio with 512 GB |
| UD-Q4_K_XL | 467.3 GB | Mac Studio with 512 GB, little room for context |
| Q8_0 | 801.4 GB | no Mac |
How much quality the 1-bit to 3-bit files keep compared with Z.ai’s FP8 deployment has not been measured here. Before relying on a local quantization, compare its answers with the API on your own tasks. The unified-memory guide explains why weights alone do not decide whether a model fits.
Per token, GLM-5.3 activates only 8 of its 256 experts plus one shared expert. For the GLM-5 base, Z.ai gives 40 billion active parameters out of 744 billion. That keeps the computation per token far below that of a dense model of the same size, but every generated token still has to read active weights from memory, so memory bandwidth limits the speed on Apple silicon.
Three more points for Mac users:
- Ollama: The library lists only
glm-5.3:cloud.ollama run glm-5.3:cloudsends your prompts to Ollama’s servers; nothing runs on the Mac. - MLX: As of October 2 there is no mlx-community release of the full GLM-5.3, only conversions by individual users. The smaller GLM-5.3-Flash exists as an mlx-community 4-bit version, but that is 204 GB as well.
- API from Mac tools: Z.ai’s Coding Plan starts at $18 per month, uses a points-based quota and, according to Z.ai, works with Claude Code, Kilo Code, Cline and OpenCode. Off-peak calls cost half the points.
For Macs with up to 128 GB of unified memory, the list of open models for 8 to 64 GB is the better starting point for local work.
What GLM-5.3 costs on the API
Z.ai’s own pricing page now lists GLM-5.3 directly. On August 21 it still showed only GLM-5.2.
| Token type, Z.ai | Price per 1M tokens |
|---|---|
| Input | $1.40 |
| Cached input | $0.26 |
| Storage of cached input | free for a limited time |
| Output | $4.40 |
With these prices, a short session with 1M input and 0.2M output tokens costs $2.28. A medium agent run with 10M input and 2M output tokens costs $22.80. In a cache-heavy workflow with 100M input tokens, of which 80M are cached, plus 20M output tokens, the bill is $136.80 instead of $228 without caching, a saving of 40 percent.
Because the weights are open, many providers now serve GLM-5.3. OpenRouter lists 39 endpoints. A selection from October 2, 2026:
| Provider on OpenRouter | Format | Input | Output | Cached input |
|---|---|---|---|---|
| Z.AI | FP8 | $1.40 | $4.40 | $0.26 |
| Together | not stated | $1.40 | $4.40 | $0.26 |
| DeepInfra | FP4 | $0.56 | $2.50 | $0.12 |
| Novita | FP8 | $0.69 | $2.16 | $0.13 |
| Baidu | FP8 | $0.16 | $0.49 | $0.03 |
| Alibaba (fast) | not stated | $2.80 | $8.80 | $0.56 |
Source: OpenRouter endpoint list. FP4 and NVFP4 endpoints run a more compressed version than Z.ai’s FP8 deployment, so test the results before you route production traffic to the cheapest provider. OpenRouter also offers z-ai/glm-5.3:batch at $0.45 input and $2.00 output for jobs that can wait, and z-ai/glm-5.3-prime, a faster variant at $2.80 and $8.80.
Benchmarks: where GLM-5.3 leads and where it does not
All figures in the following table come from Z.ai’s model card. The comparison column shows the best value among the eight models Z.ai lists.
| Benchmark (Z.ai) | GLM-5.2 | GLM-5.3 | Best listed model |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6 | 28.3 | GPT-5.6 Sol: 34.6 |
| DeepSWE v1.1 | 46.2 | 66.9 | GPT-5.6 Sol: 72.7 |
| SWE-Marathon v1.1 | 19.4 | 42.5 | Opus 4.8: 48.8 |
| CyberGym | 77.2 | 84.5 | GLM-5.3 |
| ExploitBench | 24.4 | 54.4 | Fable 5 (with fallback): 78.0 |
| AutomationBench v1.0.6 | 26.2 | 48.2 | GLM-5.3 |
| GDPval-AA v2 | 1,508 | 1,769 | GLM-5.3 |
The pattern matches Z.ai’s training story. The largest gains appear in long terminal sessions, multi-step software work, automation and security tasks. The 50 percent improvement in Z.ai’s headline refers to Z.ai Code Bench, an internal benchmark: GLM-5.3 reaches 34.5 percent there instead of 23.4 percent for GLM-5.2, with fewer output tokens per task. Outside parties cannot reproduce that test.
Independent data is still thin. OpenRouter shows Artificial Analysis indices for GLM-5.3 as of October 2: 44.8 for intelligence, 74.8 for coding and 53.1 for agentic tasks. On launch day, Reuters noted that the cybersecurity results had not been verified independently. Z.ai’s agent results also depend on the harness: the model card names Claude Code 2.1.207, maximum reasoning effort and contexts of up to 1M tokens.
Why Z.ai held back the weights at launch
GLM-5.3 uses the same base model as GLM-5.2. Z.ai attributes every gain to further post-training with more executable environments and more reinforcement learning, built on its open-source slime framework. As post-training scaled, the cybersecurity results rose faster than Z.ai expected: CyberGym from 77.2 to 84.5, ExploitBench from 24.4 to 54.4.
Because exploitation skills serve attackers as well as defenders, Z.ai released the API on August 14 but announced the weights for about two weeks later, after additional safety evaluation. WIRED and Axios covered the delay. The official Hugging Face repository is dated August 25, and the Unsloth GGUF repository followed on August 28.
Calling GLM-5.3 through the API
Z.ai documents three protocols for GLM-5.3:
| Protocol | Base URL |
|---|---|
| OpenAI Chat Completions | https://api.z.ai/api/coding/paas/v4 |
| OpenAI Responses | https://api.z.ai/api/v1 |
| Anthropic Messages | https://api.z.ai/api/anthropic |
Reasoning cannot be switched off anymore. If an existing integration sends thinking.type: "disabled", Z.ai’s migration notice says to change it to enabled and set reasoning_effort to low before switching the model ID to glm-5.3; otherwise the request fails. For complex coding tasks Z.ai recommends max.
curl -X POST "https://api.z.ai/api/coding/paas/v4/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $ZAI_API_KEY" \
-d '{
"model": "glm-5.3",
"messages": [
{"role": "user", "content": "Inspect this repository and propose a verified bug-fix plan."}
],
"thinking": {"type": "enabled"},
"reasoning_effort": "max",
"max_tokens": 4096
}'
If you serve the weights yourself, note one detail from the model card: in GLM-5.3’s chat template clear_thinking defaults to false. For chat use, Z.ai recommends passing clear_thinking=true.
Who GLM-5.3 suits, and who should pick something else
GLM-5.3 is worth testing if you:
- run coding agents on long, multi-step tasks,
- resend large repository context and benefit from cached input at $0.26,
- already use GLM-5.2 through Z.ai, where the price stayed the same,
- own a Mac Studio with 512 GB and want to keep code on your own machine, accepting a heavily quantized model.
Pick something else if you:
- need image or video input, since GLM-5.3 reads text only (GLM-5.3-Flash is multimodal),
- need answers without reasoning tokens,
- want to run a model locally on a Mac with 128 GB or less,
- need a long independent benchmark record before choosing a model.
What GLM-5.3 means for Mac users
For almost every Mac, GLM-5.3 is an API model: from $1.40 input and $4.40 output per 1M tokens at Z.ai, and cheaper at some OpenRouter providers that run more compressed versions. The open weights matter for providers and for owners of a Mac Studio with 256 or 512 GB, who can run GGUF quantizations of 217 to 467 GB. Before you buy hardware for it, compare the quality of these quantizations with the API on your own tasks; ai-on-mac.com has not measured them on a Mac.
Frequently Asked Questions
Are the GLM-5.3 weights available?
Yes. Z.ai publishes them on Hugging Face as zai-org/GLM-5.3 in FP8, plus a separate BF16 repository. The official repository is dated August 25, 2026, eleven days after the API launch.
Is GLM-5.3 open source?
It is an open-weight model under the GLM-5.3 License. Use, modification, redistribution and commercial use are allowed; only Model-as-a-Service providers with more than $10 billion revenue in twelve months need a security review by Z.ai first.
Can GLM-5.3 run locally on a Mac?
Only on a Mac Studio with 256 GB or 512 GB of unified memory and with a heavily quantized GGUF file. The smallest Unsloth quantization is 216.7 GB, so Macs with up to 128 GB have to use the API.
Is GLM-5.3 available in Ollama?
Only as glm-5.3:cloud. That tag runs on Ollama's servers, not locally on your Mac.
How much does GLM-5.3 cost?
Z.ai charges $1.40 per 1M input tokens, $0.26 per 1M cached input tokens and $4.40 per 1M output tokens. On OpenRouter, 39 provider endpoints list between $0.12 and $2.80 for input as of October 2, 2026.
How much context and output does GLM-5.3 support?
Z.ai documents a 1M-token context window and up to 128K output tokens. OpenRouter lists 1,048,576 tokens of context.
Can reasoning be disabled?
No. GLM-5.3 always reasons. The supported effort levels are low, high and max, and max is the default.
Which model ID do I use on OpenRouter?
Use z-ai/glm-5.3. The alias ~z-ai/glm-latest currently points to the same model but can move to a newer GLM version later.
Is GLM-5.3 better than GLM-5.2?
Z.ai reports large gains on coding and agent benchmarks, for example Terminal-Bench 3.0 from 4.6 to 28.3 and SWE-Marathon from 19.4 to 42.5. These are vendor results; independent evaluations are still limited.
Transparency
Sources and review basis
These primary and reference sources form the basis of the technical assessment. Vendor claims and external benchmarks are identified as such in the article.
- z.ai blog / glm-5.3
- docs.z.ai llm / glm-5.3
- docs.z.ai overview / pricing
- docs.z.ai overview / migrate-to-glm-new
- docs.z.ai devpack / overview
- huggingface.co zai-org / GLM-5.3
- huggingface.co main / LICENSE
- huggingface.co main / config.json
- huggingface.co zai-org / GLM-5.3-BF16
- huggingface.co zai-org / GLM-5
- huggingface.co unsloth / GLM-5.3-GGUF
- unsloth.ai models / GLM-5.3
- ollama.com library / glm-5.3
- openrouter.ai z-ai / glm-5.3
- openrouter.ai v1 / models
- openrouter.ai glm-5.3-20260816 / endpoints
- apple.com mac-studio / specs
- github.com THUDM / slime
- github.com sunblaze-ucb / cybergym
- github.com exploitbench / exploitbench
- reuters.com technology / chinas-zai-says-new-model-nears-anthropics-mythos-5-cyber-defence-tests-2026-08-14
- wired.com story / zai-open-weight-ai-models-release-cybersecurity-hacking
- axios.com 14 / china-open-source-ai-glm-53