Technical research with cited sources. Original measurements are identified in the article.

This article was researched and written with AI assistance. Editorial responsibility: Julian Dominic Altmann. How this site is made

Published: October 6, 2026 Updated: October 6, 2026

About the author

Checked October 6, 2026. Mistral AI has opened Mistral Large 4 as a Public Preview. The official model card describes a granular mixture-of-experts system with 1.05 trillion total parameters, 49 billion active parameters, a 1.6B-parameter vision encoder, and a 1-million-token context window on Mistral’s model profile.[S01]

The headline number that matters most for Mac users is not 49B. A sparse model can activate only a fraction of its network per token while still requiring access to the full expert pool. At 4 bits per parameter, 1.05T parameters work out to roughly 525GB of raw weight data before quantization metadata, runtime memory, KV cache or macOS are counted.

Large 4 is usable through the preview API today. Public model weights are planned for October 27, 2026 according to Reuters, WSJ and Axios.[S09][S10][S11] There is no verified full local Mac setup to recommend yet, and a conventional Q4 representation would already exceed 512GB before overhead.

Mistral Large 4: 1.05T total parameters, 49B active and roughly 525GB at theoretical 4-bit storage

Mistral Large 4 specifications that are actually confirmed

ItemVerified status on Oct. 6, 2026
ModelMistral Large 4
Versionv26.10
LifecyclePublic Preview
Architecturegranular mixture of experts, natively multimodal
Total parameters1.05T
Active parameters49B
Vision encoder1.6B
Mistral model-card context1M tokens
OpenRouter deployment context524,288 tokens
Current Mistral endpoint input$0.68 / 1M tokens
Cached input$0.07 / 1M tokens
Output$2.09 / 1M tokens
Public weightsplanned for Oct. 27
Final weight licensenot confirmed in the checked sources

The parameter, modality, context and feature claims come from Mistral’s current model card.[S01] The 524,288-token number is specifically OpenRouter’s current mistralai/mistral-large-4-0 deployment, not a replacement for Mistral’s own 1M specification.[S16]

Why “49B active” does not make Large 4 a 49B local model

MoE architectures route each token through a subset of experts. That is valuable for inference compute: the whole 1.05T network does not have to participate in every token calculation.

Storage is a different problem. Unless a future runtime streams or offloads experts in a particularly aggressive way, the system still needs access to the much larger set of expert weights.

That creates a simple rule for local planning:

  • use active parameters as a clue about compute per token;
  • use total parameters and checkpoint format as a clue about model storage and memory pressure.
Weight-memory estimates for Mistral Large 4

The Mac memory math: 525GB at four bits is only the starting point

These are calculated weight-only estimates, not measured Large 4 RAM usage:

RepresentationCalculationRaw weights
BF161.05T × 2 bytes2.10TB
8-bit1.05T × 1 byte1.05TB
6-bit1.05T × 0.75 byte787.5GB
5-bit1.05T × 0.625 byte656.25GB
4-bit1.05T × 0.5 byte525GB
3-bit1.05T × 0.375 byte393.75GB
2-bit1.05T × 0.25 byte262.5GB

Real quantized files are not just a neat pile of parameter bits. Group scales, metadata, runtime allocations, scratch buffers, the KV cache and the operating system all consume additional memory.

Would a 512GB Mac Studio run it?

A conventional full Q4 build is already larger than 512GB in the weight-only calculation. That makes “512GB should be enough because only 49B parameters are active” a bad assumption.

A theoretical 3-bit representation falls below 512GB, but that is not yet a product recommendation. We do not know, as of October 6, which quantizations will exist, what quality loss they will have, or whether MLX, llama.cpp or another runtime will efficiently support the final architecture.

For current local models, use the existing Mac unified-memory guide and the Unified-Memory Workbench.[S19][S20]

1M context on Mistral, 512K on other deployments

There is no need to force these numbers into a false contradiction.

Mistral’s model card lists 1M tokens.[S01] Vercel’s Mistral provider also shows 1M.[S17] OpenRouter currently exposes 524,288 tokens with up to 256K output.[S16] Vals documents a 512K context for the configuration it evaluated.[S15]

Provider-specific context limits for Mistral Large 4

For production use, the endpoint limit is what matters. If your agent workflow depends on a near-million-token prompt, verify the exact provider rather than relying on the family headline.

Current API price: unusually cheap, but the discount is time-sensitive

Mistral’s Large 4 model page currently displays:[S01]

  • $0.68 / 1M input tokens
  • $0.07 / 1M cached input tokens
  • $2.09 / 1M output tokens

OpenRouter labels the Mistral endpoint 50% off and shows $1.36 / $0.14 / $4.18 alongside those discounted values.[S16] Treat that as a time-sensitive commercial detail, not a permanent property of the model.

Mistral Large 4 current API pricing and cache savings

What that means in real token volumes

A workload with 10M input tokens + 1M output tokens costs:

10 × $0.68 + 1 × $2.09 = $8.89

If 80% of the input is billed at the cached-input rate:

2 × $0.68 + 8 × $0.07 + $2.09 = $4.01

At 100M input + 10M output, the same math gives $88.90 without caching or $40.10 with an 80% cache assumption.

These are our calculations, not an estimate of a typical user’s monthly bill. Cache reuse varies dramatically by workload.

EU regional inference is a separate deployment choice

Mistral documents EU and US regional endpoints at a 1.1× price multiplier over standard list pricing.[S03] Regional endpoints also have feature differences: the current documentation says Function Calling is the supported regional tool, while stateful features including Agents, Batch and Files are not available there.[S03]

That makes “EU-hosted” a deployment decision with concrete feature and cost trade-offs, not just a privacy badge.

Large 4 vs Large 3: much more total capacity, modestly more active compute

Mistral Large 3 shipped in December 2025 with 675B total parameters, 41B active parameters, a 256K context window and Apache 2.0 weights.[S07][S08]

Large 3Large 4 Preview
Total parameters675B1,050B
Active parameters41B49B
Official Mistral context256K1M
Input price$0.50/M$0.68/M current
Output price$1.50/M$2.09/M current
LifecycleGAPublic Preview
Public weightsyesplanned Oct. 27
LicenseApache 2.0not confirmed yet

Large 4 increases total parameters by roughly 55.6%, while active parameters rise by only about 19.5%. That is the architectural story in one line: Mistral expanded the expert pool much faster than the per-token active footprint.

Do not copy Large 3’s Apache 2.0 label onto Large 4. Frandroid’s launch coverage explicitly lists the Large 4 weight license as unspecified at the time of the announcement.[S12]

Benchmark reality: strong launch claims, mixed independent data

Mistral’s launch materials highlight preliminary results on DeepSWE v1.1, Finch/FinWorkBench, Harvey Legal Agent, Dense200 and DIOR-RSVG. Several launch reports reproduce those charts.[S12][S13][S14][S18]

VentureBeat provides an important methodological warning: some competitor scores shown in Mistral’s materials could not yet be independently matched to public benchmark sources for the exact configurations.[S14]

Independent Vals results

Vals currently reports 48.05% ±1.11 on its aggregate index for the Large 4 Preview, ranking the tested configuration #32 of 44 at the time of checking.[S15]

Vals benchmarkLarge 4 Preview
Harvey Legal Agent15.83% ±2.96
Finance Agent v254.68% ±0.58
Vibe Code Bench v1.178.40% ±3.55
BioMysteryBench67.04% ±2.67
IOI45.28% ±3.78
Terminal-Bench 4.022.73% ±0.88

Vals lists temperature 1, top-p 0.95, a 256K maximum output and high reasoning effort for the profile, while noting that individual benchmarks can use different providers and parameters.[S15]

That evidence supports a more useful conclusion than “Large 4 wins” or “Large 4 disappoints”: performance is workload-dependent, and the model is still a mutable preview.

Public Preview means the model can change underneath you

Mistral’s lifecycle policy says Public Preview models can receive silent updates and have no guaranteed path to General Availability.[S02]

Axios reports that Mistral plans additional reinforcement learning and safety testing before the October 27 weight release.[S11] Any benchmark, latency observation or behavior note published now therefore needs a date and the word Preview attached to it.

Mistral Large 4 release timeline

Are the weights downloadable today?

Not as the final public checkpoint this article could verify on October 6. Reuters states that Mistral Large 4 will be publicly available on October 27; WSJ says the company plans to release the trained weights that day, and Axios describes further work before the release.[S09][S10][S11]

That is why this guide does not invent an Ollama tag, GGUF download, MLX command or final license.

The useful local questions begin after the checkpoint arrives:

  1. What are the official file sizes and formats?
  2. What license is attached to the weights?
  3. Are FP8, NVFP4 or community GGUF/MLX quantizations available?
  4. Which runtimes support the MoE routing correctly?
  5. What is actual host/unified-memory overhead?
  6. Can experts be streamed or offloaded efficiently?
  7. What does real Apple Silicon throughput look like?

Who should care about Large 4 right now?

API-heavy coding and agent workflows: worth testing. Mistral lists Structured Outputs, Function Calling, Document QnA, Agents & Conversations and built-in tools on the model card.[S01]

Long-context document work: potentially attractive, but verify whether your provider gives you 1M or ~512K context.

EU data-location workloads: Mistral offers EU regional inference, but it carries the 1.1× pricing factor and feature constraints.[S03]

Local Mac inference: wait for the weights. The model’s storage footprint is the gating factor, not the 49B active count.

Verdict: the API is ready before the Mac story is

Mistral Large 4 is a useful example of why MoE numbers need context. 49B active parameters describe compute; 1.05T total parameters define a radically different storage problem. At four bits, the weight-only estimate is roughly 525GB.

That makes Large 4 immediately interesting as a hosted model and much less straightforward as a local Mac model. The API is live, the current discounted token rates are aggressive, and Mistral exposes a modern tool/agent stack.[S01] For local inference, the meaningful date is October 27: only the released weights can tell us which formats, licenses, quantizations and runtimes are real rather than hypothetical.

Next step: use the Unified-Memory Workbench for models that can actually be downloaded today. This page should be updated as soon as the Large 4 weights land with real file sizes and a reproducible Apple Silicon test.

Frequently Asked Questions

How many parameters does Mistral Large 4 have?

Mistral reports **1.05 trillion total parameters**, with **49 billion active parameters** per token and a 1.6B vision encoder.

Is Mistral Large 4 a 49B model?

Not for storage planning. 49B is the active subset; the full expert pool contains 1.05T parameters.

How much memory would a 4-bit Mistral Large 4 need?

The raw parameter math is about **525GB**. A real runtime would need more.

Can a 512GB Mac run Mistral Large 4?

A conventional full Q4 build already exceeds 512GB before overhead. Future lower-bit quantization or expert-offloading methods may change what is possible, but those claims cannot be verified before the weights arrive.

Is the context window 1M or 512K?

Mistral lists **1M**. OpenRouter’s current deployment is **524,288**, and Vals evaluated a 512K configuration.

When will Mistral Large 4 weights be released?

The announced date is **October 27, 2026**.

Is Mistral Large 4 Apache 2.0?

That was **not confirmed for the upcoming Large 4 weights** in the sources checked on October 6. Large 3 was Apache 2.0, which is not evidence for Large 4’s final license.

Transparency

Sources and review basis

20

These primary and reference sources form the basis of the technical assessment. Vendor claims and external benchmarks are identified as such in the article.

  1. docs.mistral.ai models / mistral-large-4-0
  2. docs.mistral.ai inference / model-lifecycle
  3. docs.mistral.ai inference / regional-inference
  4. docs.mistral.ai conversations / structured-output
  5. docs.mistral.ai agents / introduction
  6. docs.mistral.ai agents / agents-api
  7. mistral.ai news / mistral-3
  8. docs.mistral.ai models / mistral-large-3-25-12
  9. reuters.com china / mistral-ceo-says-new-ai-model-beats-chinese-ones-some-areas-2026-10-06
  10. wsj.com ai / mistral-to-release-new-ai-model-to-better-compete-with-u-s-rivals-3f7c8a3d
  11. axios.com 06 / reflection-mistral-open-weight-ai-models-china
  12. frandroid.com intelligence-artificielle / 3274547_mistral-annonce-large-4-le-grand-retour-face-aux-cadors-chinois
  13. thenextweb.com news / mistral-releases-large-4-a-1-trillion-parameter-open-weight-ai-model
  14. venturebeat.com technology / mistral-debuts-large-4-le-chonk-a-1-trillion-parameter-text-output-model-with-high-benchmarks-planned-for-open-weights-release
  15. vals.ai models / mistralai_mistral-large-4
  16. openrouter.ai mistralai / mistral-large-4-0
  17. vercel.com models / mistral-large-4
  18. numerama.com tech / 2347595-mistral-relance-la-france-dans-la-course-a-lia-et-devoile-le-chonk-un-modele-de-frontiere-a-1000-milliards-de-parametres.html
  19. ai-on-mac.com artikel / unified-memory-mac
  20. ai-on-mac.com hardware / matchmaker