Reviewed October 8, 2026. GLM-5.2 remains available as an API and as MIT-licensed open weights, although GLM-5.3 is now the newer generation. Z.ai lists a 1M-token context, with $1.40 input and $4.40 output per million tokens for direct API use. OpenRouter features changing prices by provider and promotion; treating its cheapest visible rate as permanent would mislead U.S. buyers. Z.ai specs · official USD price list.

Choose the route before comparing prices

Access pathModel IDWhat matters
Z.ai directglm-5.2Z.ai USD list rates and account limits
OpenRouterz-ai/glm-5.2selected provider, rate, quantization and context
Self-hosted weightszai-org/GLM-5.2large server infrastructure and runtime support

GLM-5.2 being open-weight does not establish that it runs conveniently on an ordinary MacBook with 16–64 GB unified memory. The direct cloud API is a different product from downloading the checkpoint.

Z.ai API list rates vs OpenRouter offers

The Z.ai rate table lists the following US-dollar rates per million tokens on October 8, 2026:

Token typeDirect Z.ai
Ordinary input$1.40
Cached input$0.26
Output$4.40
Cached-input storageTemporarily free, per Z.ai
GLM-5.2 direct Z.ai price per million tokens: $1.40 input, $4.40 output and $0.26 cache read. OpenRouter pricing varies by provider and discount.

These are direct Z.ai prices, not a fixed OpenRouter charge. The OpenRouter GLM-5.2 page serves the model through multiple providers and sometimes displays substantial promotional discounts. Prices, available context, cache billing and provider settings can change. Confirm those values at checkout and when routing production requests; OpenRouter billing follows its own terms.

U.S. API customers should also distinguish pay-as-you-go tokens from subscriptions or coding plans. Applicable sales tax can be additional depending on the state, the seller and the invoice arrangements.

One token-cost illustration

At direct Z.ai list rates, 100,000 uncached input tokens plus 10,000 output tokens cost $0.184: $0.14 for input and $0.044 for output. This is not a predicted bill for a large coding agent. It excludes cache effects, tools, retries and additional requests, which can dominate a workflow that keeps sending back repository files and logs.

A 1M context window is a limit, not a score

Z.ai documents 1 million tokens of context and up to 128K output tokens. A gateway provider may impose its own maximum context or response length. An enormous prompt can be useful for a complex multi-file problem, but stuffing unrelated files into it does not increase correctness. Bigger requests can mean higher latency and costs.

When evaluating long-horizon coding, start with the relevant files, failing tests, a precise objective and limits on tool permissions. Keep request sizes and harness settings comparable across models. A model that can read a full codebase still needs an actual build and test suite to verify its output.

Open weights are not a tested local Mac configuration

Z.ai publishes the official GLM-5.2 weights on Hugging Face under the MIT license. Those weights can be self-hosted on suitable infrastructure. The provided deployment instructions focus on large-scale frameworks including SGLang, vLLM and Transformers, not a routine Apple Silicon install.

No independent local GLM-5.2 Mac speed or quality measurements were performed for this article. A quantized community conversion may have a substantially different footprint or output quality. Before buying a Mac for this model, establish the exact file sizes, quantization, KV-cache requirements and compatible runtime. Our unified-memory guide and smaller-model guide are more useful for normal Mac configurations.

Understand the vendor benchmark labels

The official model card lists these GLM-5.2 scores:

EvaluationScoreImportant condition
Terminal-Bench 2.181.0Terminus-2 harness
SWE-bench Pro62.1vendor/model-card report
Humanity’s Last Exam, no tools40.5no agent tools used

These are vendor-reported, not ai-on-mac.com results. Z.ai also reports separate best-harness Terminal-Bench figures; mixing them without labels changes the comparison. Different tests measure different tasks and can use different scaffolds. A single benchmark bar chart without a common scale and method would make the results look more comparable than they are.

Direct API request from macOS

The following command calls Z.ai directly, not OpenRouter. Store your key in ZAI_API_KEY as a secret and never commit it:

curl -fsS https://api.z.ai/api/paas/v4/chat/completions \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.2",
    "messages": [{"role":"user","content":"List three tests for this code patch."}],
    "thinking": {"type":"enabled"},
    "reasoning_effort": "max",
    "max_tokens": 1024
  }'

This is based on Z.ai’s request examples. An OpenRouter call uses a different base URL, API key and model slug. A successful response here is cloud inference, not a local Mac benchmark.

U.S. data handling requires provider checks

An editor or coding agent sends selected files to the remote API. A U.S. developer account, U.S.-language documentation or a U.S. network endpoint does not prove that every processor is in the United States. Verify the chosen provider, subprocessors, privacy terms, retention and application-level tool permissions before uploading proprietary code.

When to use GLM-5.2

Keep it if an existing evaluated workflow works well and migration offers no measured benefit. For a new cloud application, compare it against GLM-5.3 on real tasks and total billed usage. If data must stay offline on a typical Mac, start with a smaller open-weight model.

Sources verified October 8, 2026: Z.ai model, Z.ai pricing, OpenRouter, official weights.

Frequently Asked Questions

What is the direct Z.ai price?

Z.ai lists $1.40 input, $0.26 cached input and $4.40 output per million tokens; OpenRouter providers may charge differently.

Is GLM-5.2 available locally on Mac?

The weights are MIT-licensed, but the model is very large and official inference instructions focus on server frameworks; a standard Apple Silicon setup is not independently verified here.

Does every OpenRouter route support 1M context?

Not necessarily. Z.ai documents 1M context and 128K maximum output for the model; provider settings and limits must be checked separately.

Which model ID should I use?

Use glm-5.2 at Z.ai, z-ai/glm-5.2 on OpenRouter, and zai-org/GLM-5.2 for the official weights.

Researched and written with AI, sources cited, original measurements marked. AI transparency

Published: October 8, 2026 Updated: October 8, 2026

About the author