Reviewed October 8, 2026. GLM-5.2 remains available as an API and as MIT-licensed open weights, although GLM-5.3 is now the newer generation. Z.ai lists a 1M-token context, with $1.40 input and $4.40 output per million tokens for direct API use. OpenRouter features changing prices by provider and promotion; treating its cheapest visible rate as permanent would mislead U.S. buyers. Z.ai specs · official USD price list.
Choose the route before comparing prices
| Access path | Model ID | What matters |
|---|---|---|
| Z.ai direct | glm-5.2 | Z.ai USD list rates and account limits |
| OpenRouter | z-ai/glm-5.2 | selected provider, rate, quantization and context |
| Self-hosted weights | zai-org/GLM-5.2 | large server infrastructure and runtime support |
GLM-5.2 being open-weight does not establish that it runs conveniently on an ordinary MacBook with 16–64 GB unified memory. The direct cloud API is a different product from downloading the checkpoint.
Z.ai API list rates vs OpenRouter offers
The Z.ai rate table lists the following US-dollar rates per million tokens on October 8, 2026:
| Token type | Direct Z.ai |
|---|---|
| Ordinary input | $1.40 |
| Cached input | $0.26 |
| Output | $4.40 |
| Cached-input storage | Temporarily free, per Z.ai |
These are direct Z.ai prices, not a fixed OpenRouter charge. The OpenRouter GLM-5.2 page serves the model through multiple providers and sometimes displays substantial promotional discounts. Prices, available context, cache billing and provider settings can change. Confirm those values at checkout and when routing production requests; OpenRouter billing follows its own terms.
U.S. API customers should also distinguish pay-as-you-go tokens from subscriptions or coding plans. Applicable sales tax can be additional depending on the state, the seller and the invoice arrangements.
One token-cost illustration
At direct Z.ai list rates, 100,000 uncached input tokens plus 10,000 output tokens cost $0.184: $0.14 for input and $0.044 for output. This is not a predicted bill for a large coding agent. It excludes cache effects, tools, retries and additional requests, which can dominate a workflow that keeps sending back repository files and logs.
A 1M context window is a limit, not a score
Z.ai documents 1 million tokens of context and up to 128K output tokens. A gateway provider may impose its own maximum context or response length. An enormous prompt can be useful for a complex multi-file problem, but stuffing unrelated files into it does not increase correctness. Bigger requests can mean higher latency and costs.
When evaluating long-horizon coding, start with the relevant files, failing tests, a precise objective and limits on tool permissions. Keep request sizes and harness settings comparable across models. A model that can read a full codebase still needs an actual build and test suite to verify its output.
Open weights are not a tested local Mac configuration
Z.ai publishes the official GLM-5.2 weights on Hugging Face under the MIT license. Those weights can be self-hosted on suitable infrastructure. The provided deployment instructions focus on large-scale frameworks including SGLang, vLLM and Transformers, not a routine Apple Silicon install.
No independent local GLM-5.2 Mac speed or quality measurements were performed for this article. A quantized community conversion may have a substantially different footprint or output quality. Before buying a Mac for this model, establish the exact file sizes, quantization, KV-cache requirements and compatible runtime. Our unified-memory guide and smaller-model guide are more useful for normal Mac configurations.
Understand the vendor benchmark labels
The official model card lists these GLM-5.2 scores:
| Evaluation | Score | Important condition |
|---|---|---|
| Terminal-Bench 2.1 | 81.0 | Terminus-2 harness |
| SWE-bench Pro | 62.1 | vendor/model-card report |
| Humanity’s Last Exam, no tools | 40.5 | no agent tools used |
These are vendor-reported, not ai-on-mac.com results. Z.ai also reports separate best-harness Terminal-Bench figures; mixing them without labels changes the comparison. Different tests measure different tasks and can use different scaffolds. A single benchmark bar chart without a common scale and method would make the results look more comparable than they are.
Direct API request from macOS
The following command calls Z.ai directly, not OpenRouter. Store your key in ZAI_API_KEY as a secret and never commit it:
curl -fsS https://api.z.ai/api/paas/v4/chat/completions \
-H "Authorization: Bearer $ZAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.2",
"messages": [{"role":"user","content":"List three tests for this code patch."}],
"thinking": {"type":"enabled"},
"reasoning_effort": "max",
"max_tokens": 1024
}'
This is based on Z.ai’s request examples. An OpenRouter call uses a different base URL, API key and model slug. A successful response here is cloud inference, not a local Mac benchmark.
U.S. data handling requires provider checks
An editor or coding agent sends selected files to the remote API. A U.S. developer account, U.S.-language documentation or a U.S. network endpoint does not prove that every processor is in the United States. Verify the chosen provider, subprocessors, privacy terms, retention and application-level tool permissions before uploading proprietary code.
When to use GLM-5.2
Keep it if an existing evaluated workflow works well and migration offers no measured benefit. For a new cloud application, compare it against GLM-5.3 on real tasks and total billed usage. If data must stay offline on a typical Mac, start with a smaller open-weight model.
Sources verified October 8, 2026: Z.ai model, Z.ai pricing, OpenRouter, official weights.
Frequently Asked Questions
What is the direct Z.ai price?
Z.ai lists $1.40 input, $0.26 cached input and $4.40 output per million tokens; OpenRouter providers may charge differently.
Is GLM-5.2 available locally on Mac?
The weights are MIT-licensed, but the model is very large and official inference instructions focus on server frameworks; a standard Apple Silicon setup is not independently verified here.
Does every OpenRouter route support 1M context?
Not necessarily. Z.ai documents 1M context and 128K maximum output for the model; provider settings and limits must be checked separately.
Which model ID should I use?
Use glm-5.2 at Z.ai, z-ai/glm-5.2 on OpenRouter, and zai-org/GLM-5.2 for the official weights.