LongCat-2.0 is Meituan’s new open-weight language model for coding, long-running agents and complex tool workflows. It combines a Mixture-of-Experts architecture with 1.6 trillion total parameters, approximately 48 billion activated parameters per token, a context window of up to one million tokens, and a maximum API output of 128,000 tokens.12
Updated October 8, 2026. LongCat-2.0 is Meituan’s publicly released 1.6-trillion-parameter sparse MoE, with roughly 48 billion active parameters per token, a one-million-token API context, and MIT-licensed model weights. Full self-hosting still requires data-center-scale infrastructure. Meituan also offers LongCat-2.5-Preview, a distinct API model ID with image understanding; do not confuse that newer preview with the open-weight 2.0 checkpoint.123
Meituan’s official pricing remains published, but it distinguishes standard list rates from limited-time promotional rates. Use the current rate card and actual account billing rather than treating the discount as permanent.4
LongCat-2.0 versus LongCat-2.5-Preview: current platform status
Meituan’s official changelog dates the production LongCat-2.0 API to June 30, 2026 and the separate LongCat-2.5-Preview release to September 25, 2026. The current API quickstart lists both model IDs.
| Decision point | LongCat-2.0 | LongCat-2.5-Preview |
|---|---|---|
| API model ID | LongCat-2.0 | LongCat-2.5-Preview |
| Availability | active API, published 2.0 weights | newer preview API model |
| Inputs | text-oriented language model | image understanding is now documented |
| API context / max output | 1M / 128K tokens | 1M / 128K tokens |
| Published weights | MIT-licensed 2.0 checkpoint on Hugging Face | not the same as the 2.0 checkpoint; verify a separate release |
| Limited-time pay-as-you-go (USD/M) | $0.30 input / $0.006 cached / $1.20 output | $0.30 input / $0.006 cached / $1.20 output on the current preview rate card |
These are distinct model IDs, not benchmark-equivalent checkpoints. Migrate a coding agent only after verifying tool-call behavior, reasoning mode, latency, cache hits and accuracy on your own repository. Image understanding is a reason to evaluate the preview separately, not evidence that its coding benchmarks universally exceed 2.0.35
LongCat-2.0 at a glance
| Feature | Official specification | Why it matters |
|---|---|---|
| Total parameters | 1.6 trillion | Mixture-of-Experts model with 1.6T total parameters |
| Activated parameters | approximately 48 billion per token | Compute path is far smaller than the full parameter pool |
| Context window | 1,000,000 tokens | Designed for large repositories, long documents and agent histories |
| Maximum output | 128,000 tokens | Exceptionally high documented API output ceiling |
| Pretraining | more than 35 trillion tokens | Meituan says training ran on AI ASIC superpods |
| N-gram embeddings | 135 billion parameters | Additional sparse capacity outside the regular MoE path |
| License | MIT | Broad reuse under the license terms |
| API formats | OpenAI and Anthropic | Easier integration with existing software |
| Main focus | coding and agents | Tool calls, multi-step tasks and repository changes |
What is LongCat-2.0?
LongCat-2.0 is part of Meituan’s LongCat model family. Its predecessor, LongCat-Flash, also used a Mixture-of-Experts architecture, but was considerably smaller at 560 billion total parameters and an average of roughly 27 billion activated parameters per token. LongCat-2.0 expands both the parameter pool and the long-context architecture.6
A preview API was introduced in April 2026. Meituan’s changelog records the regular release and paid API service on June 30, 2026. The platform also retired several older LongCat API models to focus resources on LongCat-2.0.3
The core distinction from a dense language model is simple: the model does not compute all 1.6 trillion parameters for every token. A routing mechanism selects a subset of relevant experts. About 48 billion parameters are activated for a token. That keeps the active compute path much smaller than the full model, although the complete weight set still has to be stored or made available across a distributed inference system.
Architecture: Why 1.6 trillion total parameters do not mean 1.6 trillion active parameters
A dense Transformer uses all of its model parameters during each forward pass. A Mixture-of-Experts model distributes parts of the feed-forward computation across many expert networks and activates only a selected subset for each token.
LongCat-2.0 reports approximately 48 billion active parameters per token. That is roughly three percent of the stated 1.6 trillion total. The design separates two requirements:
- Compute: only the selected expert path is evaluated for a token.
- Storage and infrastructure: the entire expert pool still needs to be available, usually across many accelerators.
The active-parameter figure is therefore not the model’s download size. It explains how such a large model can be served with less computation than a dense 1.6T model, but it does not turn LongCat-2.0 into an ordinary local 48B model.
LongCat Sparse Attention
Standard self-attention compares very large numbers of token pairs as sequences grow. In a naive implementation, the cost increases quadratically with sequence length. At one million tokens, that approach is impractical. LongCat-2.0 therefore introduces LongCat Sparse Attention (LSA).1
LSA combines three techniques:
Streaming-aware Indexing (SI) mixes contiguous, hardware-aligned memory access with dynamically selected tokens. The aim is to replace fragmented reads with more predictable sequential HBM access.
Cross-Layer Indexing (CLI) takes advantage of the tendency for attention saliency to remain relatively stable across neighboring layers. One indexing pass can serve several consecutive layers, reducing repeated indexing work. Meituan says cross-layer distillation enables this behavior during training.
Hierarchical Indexing (HI) uses a coarse-to-fine selection process. A first stage recalls promising blocks using approximate scoring. A second stage performs fine-grained token selection only within those candidates, reducing the search space for each query.
Meituan also extends the sparse indexing system to a three-step Multi-Token Prediction module used for speculative decoding. Multiple upcoming tokens can be proposed while parts of the indexing work are shared.1
135 billion parameters of N-gram embeddings
LongCat-2.0 inherits N-gram embeddings from LongCat-Flash-Lite. Meituan says 135 billion parameters are allocated to this component. It expands capacity along a sparse dimension that is separate from the standard MoE experts.1
In simplified terms, frequently occurring token sequences can receive additional learned representations. This may model local patterns efficiently without putting every extra parameter into more experts. Meituan claims the chosen balance is stronger than an equivalently sized pure MoE model, although independent LongCat-2.0 ablation results remain limited at launch.
One million tokens of context: what it means in practice
The LongCat API documents a 1,000,000-token context window and a maximum output length of 128,000 tokens.2 This supports:
- analyzing large software repositories
- processing long terminal and log histories
- combining many documents in one workflow
- agents that retain numerous intermediate steps and tool results
- large code migrations spanning many dependent files
- long research and automation sessions
A large context limit is not the same as perfect use of every token. Retrieval accuracy, positional robustness, irrelevant information, prompt structure, latency and cost still matter. Production systems should continue to segment large inputs, retrieve the most relevant material and verify critical outputs.
How much text fits into one million tokens?
There is no exact conversion to words or pages because tokenization depends on language, code, punctuation and formatting. As a rough order of magnitude, one million English tokens may represent several hundred thousand words. Source code can tokenize very differently because of short identifiers, symbols and indentation.
For agents, the session history matters alongside raw document length: tool outputs, terminal logs, patches and intermediate state accumulate during long sessions. A larger window reduces immediate compression pressure, but it does not replace context management.
Training on AI ASIC superpods
Meituan states that both full training and large-scale deployment were built entirely on AI ASIC superpods. Pretraining reportedly covered more than 35 trillion tokens and millions of accelerator-days. The company also reports no rollbacks or irrecoverable loss spikes during the run.1
Meituan presents LongCat-2.0 as evidence that frontier-scale training can be executed on an alternative accelerator platform. However, the model card does not provide a complete hardware bill of materials, independently audited energy consumption or a full training-cost breakdown. Efficiency claims should therefore remain limited to the published evidence.
LongCat-2.0 benchmarks
Meituan reports results across coding agents, general agents, search, instruction following, writing, mathematics and scientific reasoning. Selected LongCat-2.0 scores include:
| Benchmark | LongCat-2.0 |
|---|---|
| Terminal-Bench 2.1 | 70.8 |
| SWE-bench Pro | 59.5 |
| SWE-bench Multilingual | 77.3 |
| FORTE | 73.2 |
| BrowseComp | 79.9 |
| RWSearch | 78.8 |
| IFEval | 90.0 |
| WritingBench | 83.8 |
| IMO-AnswerBench | 81.8 |
| GPQA Diamond | 88.9 |
How should the benchmark results be interpreted?
The results vary by benchmark:
- On SWE-bench Pro, LongCat-2.0 scores 59.5 in Meituan’s table, above the listed values for Gemini 3.1 Pro, GPT-5.5 and Claude Opus 4.6, but below Claude Opus 4.7 and 4.8.
- On SWE-bench Multilingual, its 77.3 is close to the listed Claude Opus 4.6 score and above Gemini, but below Opus 4.7 and 4.8.
- On Terminal-Bench 2.1, 70.8 is roughly level with the listed Gemini result and below GPT-5.5 and Opus 4.8.
- IFEval 90.0 is below several comparison models in the official table.
- GPQA Diamond 88.9 trails multiple listed competitors.
These are not fully independent comparisons. Meituan says values without an asterisk were measured in-house under a unified harness, while asterisked comparison values came from the respective vendors’ official reports. Agent frameworks, tool sets, token budgets, sampling settings and evaluation dates can all affect results.1 The table is evidence that LongCat-2.0 is competitive, not a definitive universal ranking.
Why agent benchmarks are difficult to compare
Classic question-answering benchmarks have a relatively constrained evaluation pipeline. Agent results depend much more heavily on the complete runtime:
- Which tools are available?
- How many steps and tokens can the model use?
- Can it retry tests and repair errors?
- Which file, container and network permissions are granted?
- Is the result based on one sample or multiple attempts?
- Which benchmark and harness versions were used?
Teams should therefore test LongCat-2.0 on their own repositories, tickets and automation tasks.
LongCat-2.0 API: OpenAI and Anthropic compatibility
The LongCat API Platform supports two interface styles:
- OpenAI-compatible Chat Completions
- Anthropic-compatible Messages API
That allows many existing applications to switch providers with limited changes.2
Python example using the OpenAI SDK
import os
from openai import OpenAI
api_key = os.environ.get("LONGCAT_API_KEY")
if not api_key:
raise RuntimeError("Set LONGCAT_API_KEY before running this script.")
client = OpenAI(
api_key=api_key,
base_url="https://api.longcat.chat/openai",
)
response = client.chat.completions.create(
model="LongCat-2.0",
messages=[
{
"role": "system",
"content": (
"You are a careful coding assistant. Verify assumptions "
"and list every modified file."
),
},
{
"role": "user",
"content": (
"Analyze this repository and plan a safe migration "
"to Python 3.14."
),
},
],
max_tokens=4000,
)
print(response.choices[0].message.content)
The API key should be supplied through an environment variable or secret manager. It should never be committed to a repository or exposed in client-side JavaScript.
Reliable requests for long agent sessions
For long responses, the API documentation recommends streaming. Applications should also implement sensible timeouts, exponential backoff for HTTP 429 responses and monitoring for token usage, latency and error rates.7
A production agent should additionally:
- create a plan before changing files,
- load relevant files instead of the entire project by default,
- make changes in small logical units,
- run tests after each unit,
- validate patches and tool outputs,
- stop or request human approval when confidence is low.
Using LongCat-2.0 with Claude Code
Meituan documents a direct configuration through its Anthropic-compatible base URL. The ~/.claude/settings.json file can set variables including ANTHROPIC_AUTH_TOKEN, ANTHROPIC_BASE_URL and the LongCat model name.8
This makes LongCat-2.0 an alternative model inside the Claude Code harness. The responses are generated by LongCat, not by a Claude model. API-format compatibility also does not guarantee identical model behavior or support for every provider-specific feature.
LongCat-2.0 API pricing: official list versus promotion
As checked on October 8, 2026, Meituan’s first-party rate card still separates standard list rates and limited-time discounted rates, in USD per million tokens. The checkout/account billing record takes precedence; the promotion has no guaranteed end date.4
| API token class | Standard list rate | Limited-time rate |
|---|---|---|
| Uncached input | $0.75 | $0.30 |
| Cached input | $0.015 | $0.006 |
| Output | $2.95 | $1.20 |
Reproducible cache-aware cost calculations
Scenario A: 100,000 uncached input tokens plus 10,000 output tokens cost $0.1045 at list rates or $0.042 at published promotional rates.
Scenario B – long-running coding agent: 50M input tokens, including 40M confirmed cache hits, plus 5M output tokens:
| Charge | List rates | Promotional rates |
|---|---|---|
| 10M uncached input | $7.50 | $3.00 |
| 40M cached input | $0.60 | $0.24 |
| 5M output | $14.75 | $6.00 |
| Total | $22.85 | $9.24 |
Formula: (total input − cached input) × uncached rate + cached input × cached rate + output × output rate. Cached tokens are a subset of input, not a second input surcharge. Count only cache hits confirmed by API billing records.
Meituan also sells Token Packs valid for 30 days. For verified international enterprise accounts, its FAQ lists graduated pay-as-you-go discounts up to 40% once actual monthly consumption reaches $75 / $300 / $1,200 / $6,000. Do not assume these enterprise tiers automatically stack with a temporary promotion or apply to individual U.S. accounts. U.S. sales tax, if applicable, is not incorporated into the USD list rates.7
The one-million-token cost trap
The maximum context window is a capacity limit, not a recommended default request size. Full-history agent calls amplify prefill time, cache misses and retry costs. The official 128,000-token output limit is a ceiling, not guaranteed reply length. Meituan documents exponential backoff with retry_after on HTTP 429 and smaller inputs when context_length_exceeded occurs.27
Can LongCat-2.0 run locally on a Mac?
For a normal Mac, the practical answer is not as the complete model.
The official 1.6-trillion-parameter figure implies the following theoretical raw weight sizes:
| Format | Theoretical minimum weight size |
|---|---|
| BF16 / FP16 | approximately 3.2 TB |
| 8-bit | approximately 1.6 TB |
| 4-bit | approximately 800 GB |
Quantization metadata, runtime buffers, KV cache, routing, temporary activations and framework overhead require additional memory. The roughly 48 billion active parameters reduce computation per token; they do not remove the need to store the complete expert pool.
Even a high-memory Mac Studio is therefore not a comfortable platform for full local deployment. Aggressive offloading might enable experiments, but performance would be far from the intended superpod environment. For Mac users, the LongCat API is the realistic access method.
The model card documents GPU and NPU deployment through SGLang and SGLang-FluentLLM. Hugging Face also displays integration options for vLLM, SGLang and available quantizations. Those options do not imply that an average computer can host the full checkpoint efficiently.1
Why “48B active” does not mean “fits like a 48B model”
Each token activates only part of the expert pool, but the router may choose different experts for the next token. The system therefore needs access to the complete set of experts. A local machine cannot simply discard all currently inactive weights.
LongCat-2.0 strengths
1. A genuinely large context window
One million tokens can support tasks that quickly exceed 32K or 128K limits. Large repositories and long-running agent histories are the clearest use cases.
2. Integration with real agent harnesses
Meituan names concrete integrations including Claude Code, Hermes, OpenClaw, OpenCode and Kilo Code rather than discussing tool calling only in abstract terms.
3. MIT-licensed model weights
The model card releases the weights under the MIT License. That is more permissive than many community licenses containing usage or revenue restrictions. It does not grant rights to Meituan trademarks or patents.1
4. Flexible API compatibility
OpenAI- and Anthropic-compatible endpoints reduce migration work. Many applications can be adapted by changing the base URL, API key and model name.
5. Published API prices
The published cached-input rate is substantially lower than the uncached-input rate. That can make LongCat-2.0 relevant for long reusable contexts if latency, reliability and output quality meet the application’s requirements.
Limitations and open questions
Independent evaluations are still limited
LongCat-2.0 is new. Much of the available performance evidence comes directly from Meituan. External reproductions using identical agent budgets and harness versions are needed for stronger conclusions.
A 1M limit does not guarantee 1M-token recall
Maximum input length, effective retrieval and stable reasoning over the full sequence are separate properties. Practical needle-in-a-haystack, multi-hop and repository tests remain necessary.
Self-hosting requires enormous infrastructure
The MIT license makes the weights accessible; it does not reduce their storage requirements. API access is more realistic than local hosting for most developers.
Language and safety performance can vary
Meituan notes that performance can differ across languages and downstream applications. Medical, legal, financial, employment and infrastructure use cases require dedicated evaluation and human oversight.1
Data and operational requirements
Hosted API users should review privacy terms, data location, logging, retention, contractual protections and compliance. API compatibility alone does not answer those questions.
LongCat-2.0 vs. LongCat-Flash
| Feature | LongCat-Flash | LongCat-2.0 |
|---|---|---|
| Total parameters | 560B | 1.6T |
| Activated parameters | 18.6B–31.3B, approximately 27B average | approximately 48B |
| Context | up to 128K in the Flash generation | up to 1M |
| Main focus | efficient general and agentic use | long agent tasks, coding and repository work |
| Attention | earlier LongCat design | LongCat Sparse Attention |
| N-gram embeddings | introduced with Flash-Lite | 135B parameters |
| License | MIT for released models | MIT |
LongCat-2.0 increases both total and active capacity while redesigning the long-context path. The trade-off is an even larger infrastructure requirement.
Who is LongCat-2.0 for?
LongCat-2.0 fits these scenarios:
- developers testing long coding-agent sessions
- teams working with very large repositories
- agent-platform and tool-harness developers
- researchers studying sparse attention and MoE scaling
- companies maintaining provider-neutral OpenAI or Anthropic interfaces
- Hermes and OpenClaw users evaluating an additional model provider
- organizations that need open weights but can deploy on server infrastructure
It is less suitable for users looking for a fast local LLM on a MacBook, Mac mini or ordinary desktop. Smaller dense models and compact MoE models are much more practical for that use case.
Conclusion: Open weights, but cloud API is practical for most Macs
LongCat-2.0 combines 1.6 trillion total parameters, approximately 48 billion active parameters, a one-million-token context window, an MIT license and direct optimization for coding agents. LongCat Sparse Attention, N-gram embeddings and Multi-Token Prediction extend the long-context and inference path beyond simply increasing model scale.
The benchmark results remain primarily vendor measurements, not independently replicated workload scores. Despite MIT-licensed weights, LongCat-2.0 is a cloud API model for practical Mac usage. The separate LongCat-2.5-Preview has added image understanding since September 2026; it is a separate evaluation target, not a retroactive rename of the 2.0 checkpoint. Compare real tool reliability, latency, confirmed cache hits, billed cost and security on your own tasks.
How much unified memory a model this large really needs is covered in the RAM guide, and realistic open alternatives in the open-weight roundup.
Sources
Additional official links: GitHub Repository · Technical Blog
Footnotes
Frequently Asked Questions
Is LongCat-2.0 released under an open-source license?
The model weights and repository contributions are released under the MIT License according to the official model card. Open weight may be the more precise term because the release does not automatically include every training dataset, data pipeline and internal training system.
How many parameters does LongCat-2.0 have?
Meituan reports 1.6 trillion total parameters and approximately 48 billion activated parameters per token.
What is the LongCat-2.0 context window?
The API documents up to 1,000,000 context tokens and a maximum output of 128,000 tokens.
Can LongCat-2.0 run on a Mac?
Not practically as the complete model on ordinary Mac hardware. Even theoretical four-bit weights would require about 800 GB before runtime overhead.
Does LongCat-2.0 support tool calling?
Yes. Meituan describes native tool calls, multi-step reasoning and integrations with Claude Code, Hermes, OpenClaw, OpenCode and other agent tools.
How much does the LongCat-2.0 API cost?
Checked October 8, 2026: Meituan lists $0.75 / $0.015 / $2.95 per million uncached, cached and output tokens. The limited-time rates are $0.30 / $0.006 / $1.20; your actual platform bill is authoritative.
How does LongCat-2.5-Preview differ?
LongCat-2.5-Preview is a separate API model released September 25, 2026 with image understanding. LongCat-2.0 remains listed; its MIT weights are not the same thing as a released 2.5-Preview checkpoint.
Which API formats are supported?
The platform provides OpenAI-compatible Chat Completions and an Anthropic-compatible Messages API.