Research date: September 27, 2026. MiniMax M3.1 is no longer just a rumored successor name. The official MiniMax Code repository now contains concrete MiniMax-M3.1 model fixtures and configuration tests, while users are reporting a “M3.1 Flash Preview” inside MiniMax Code today. The missing pieces matter just as much: at the time of review, MiniMax has not published a public M3.1 model card, an official M3.1 benchmark table, or a verified M3.1 API price sheet. [S7][S8][S21][S23]
M3.1 Flash Preview appears to be a real preview deployment, but there is not enough verified evidence for a clean quality ranking yet. The strongest evidence covers the model identifier, context/output fixtures and reasoning controls in MiniMax’s own code. Early performance reports are anecdotal. Space Bunny Alpha is an intriguing candidate signal, but its scores must not be relabeled as confirmed M3.1 benchmarks. [S7][S10][S24][S26]
What is confirmed about MiniMax M3.1?
MiniMax founder Yan Junjie was reported on August 26 as saying that M3.1, M3 Pro and H3.1 were close to completion. The same report characterized M3.1 as an effort to improve stability, output quality, inference efficiency and agent generalization. A follow-up report described it as a continuation of the M3 base with more post-training and a focus on coding, tool use and agent workloads. These are reports from the earnings-call context, not an M3.1 technical model card. [S15][S16]
The more concrete evidence is now in MiniMax’s own code. An official MiniMax Code test fixture names MiniMax-M3.1 and includes a 512,000-token catalog context, a 1,000,000-token selectable context option and a 128,000-token output cap. Other fixtures persist M3.1 selections with 1M context and reasoning settings. This is strong evidence of product integration, but a test fixture is not the same thing as a public API guarantee. [S7][S9][S10]
What does “Flash Preview” tell us?
Same-day reports on September 27 show the label MiniMax-M3.1-Flash-Preview in MiniMax Code. A TRAE community thread records the name while also noting that the standard MiniMax website did not yet provide useful details. A same-day Linux DO report describes the preview in active use. That is useful field evidence, not a controlled benchmark. MiniMax’s own code artifacts separately show multi-level reasoning configuration for M3.1. [S7][S21][S22]
For now, “Flash” should be treated as a preview/product label, not as proof of a smaller parameter count, a new architecture or a particular price tier.
Confirmed, observed and unknown
| Item | Status on Sep. 27, 2026 | Evidence |
|---|---|---|
MiniMax-M3.1 model identifier | Confirmed | official MiniMax Code artifacts [S7][S8] |
| “M3.1 Flash Preview” label | Observed, not yet documented in a public model card | same-day field reports [S21][S23] |
| Context | 512K/1M in official code tests | [S7][S10] |
| Max output | 128K in official code tests | [S7] |
| Reasoning effort | multi-level evidence in official code artifacts | [S7][S9] |
| Image/video input | confirmed for M3; not yet publicly specified for M3.1 Preview | M3 card [S2] |
| Parameter count | Unknown | no public M3.1 card found |
| Architecture | Unknown | no public M3.1 card found |
| API pricing | Unknown | no verified M3.1 price sheet found |
| Open weights | Not confirmed | no M3.1 weights found |
| Official M3.1 benchmark table | Not found | research through Sep. 27 |
Do not automatically copy M3’s 428B/23B specification to M3.1
The official M3 card describes roughly 428B total parameters and ~23B activated parameters, one-million-token context and native text/image/video input. NVIDIA independently lists 428B total and roughly 22B active for M3. Until MiniMax publishes M3.1 specifications, those numbers remain M3 specifications, not M3.1 facts. [S2][S34]
Are there real M3.1 benchmarks yet?
Not benchmarks that can be cleanly labeled as verified M3.1 results.
MiniMax has not published an M3.1 benchmark table at the time of review. Artificial Analysis currently tracks M3 rather than a verified M3.1 entry. Its current M3 page reports an Intelligence Index score of 29 in the current index version, around 179 output tokens/s, a 1M context window and a listed M3 API price of $0.30/$1.20 per million input/output tokens. These are useful M3 baseline measurements, not M3.1 scores. [S30]
Benchmark versioning matters. Artificial Analysis scored M3 at 55 in its June 8 launch-era index. That does not mean the model suddenly lost half its capability; the index, benchmark composition and comparison set changed. Any M3.1 claim should therefore include benchmark version, date and methodology. [S31]
Is Space Bunny Alpha actually M3.1?
This is the most interesting unresolved lead.
OpenRouter has listed Space Bunny Alpha since September 23 as an anonymous stealth model with 1M context, text/image/video input, adjustable reasoning and a free preview. OpenRouter explicitly says the provider is anonymous and that OpenRouter is not the model’s developer or owner. [S24]
Independent tokenizer probes point strongly toward the MiniMax model family. YFarmX reports an exact MiniMax-family token-count match across a broader probe set. Other independent identity reviews reach the same limitation: the tokenizer evidence supports MiniMax-family lineage, but does not reliably distinguish older and newer M generations. It therefore cannot prove “M3.1.” [S26][S27][S28]
The appearance of an actual M3.1 Flash Preview in MiniMax Code only days later makes the hypothesis more plausible. It still does not make it proven. Until MiniMax or the stealth provider confirms the identity, Space Bunny results should be labeled candidate data. [S21][S24][S26][S27][S28]
Space Bunny candidate data — not confirmed M3.1 results
An independent Space Bunny benchmark page reports:
| Evaluation | Result | Caveat |
|---|---|---|
| AI BENCHY, high | 7.0/10 | 12/22 tasks fully passed; custom suite |
| GPQA Diamond | 82.0% | 60-question subset |
| MMLU-Pro | 75% | independent run, not MiniMax official |
| Humanity’s Last Exam | 46.1% | 300-question subset; not directly comparable with full-set scores |
These numbers describe Space Bunny Alpha. They become M3.1 evidence only if the identity is independently established. [S25]
The M3 baseline is still useful
MiniMax’s M3 launch report listed 59.0% on SWE-Bench Pro, 66.0% on Terminal-Bench 2.1, 34.8% on SWE-fficiency, 28.8% on KernelBench Hard and 74.2% on MCP Atlas. These are vendor-reported M3 results under MiniMax’s documented evaluation setups. They are useful as a baseline for any future M3.1 delta, but should not be presented as M3.1 performance. [S1][S36]
M3 also introduced MiniMax Sparse Attention. The MSA paper documents block-sparse attention and large long-context compute/runtime improvements in its experimental setup, and M3 is the publicly released production model built around that attention line. [S5]
Early hands-on reports are encouraging, but weak evidence
A Linux DO report on September 27 describes ongoing hands-on use of the preview. Same-day community observations are useful signals, but they lack standardized prompts, controlled routing, repeated trials and blind scoring. They should be separated from reproducible benchmarks. [S22]
How to evaluate M3.1 Flash Preview now
For agentic coding, a small reproducible harness is more valuable than a leaderboard guess:
- 10 previously solved real bugs, each pinned to the same repository commit.
- Tool-use tasks covering files, shell, search and at least one deliberately failing tool call.
- Long-context checks at 128K, 512K and 1M, where the preview actually permits them.
- Reasoning-effort A/B runs at low, high and max, recording success rate, output tokens and wall time.
- Cost per successful task once M3.1 pricing is published, rather than price per million tokens alone.
That dataset would create a stronger reason to revisit an article than simply repeating vendor benchmark tables.
Bottom line: M3.1 Flash Preview is real, but undocumented
MiniMax M3.1 Flash Preview is real enough to investigate, but not documented enough to rank responsibly yet. The strongest public technical evidence is MiniMax’s own repository, which already contains MiniMax-M3.1, 512K/1M context options, a 128K output fixture and reasoning configuration. Early community reports indicate the preview is usable, while Space Bunny Alpha remains a plausible candidate for the same model family rather than a confirmed identity. [S7][S10][S22][S26]
The next meaningful update should happen when MiniMax publishes an official M3.1 model card, price sheet, benchmark table or public API specification.
Methodology and source note
Research carried out on September 27, 2026. The model identifier, the context and output caps and the reasoning configuration were checked against MiniMax’s official code repository. The early preview observations come from community sources and are explicitly labeled as field observations. Space Bunny Alpha candidate values are not presented as M3.1 benchmarks. Vendor-reported M3 numbers are labeled as such and used only as a baseline. The full research log with every opened source ships with the source package.
Frequently Asked Questions
What is the MiniMax M3.1 Flash Preview?
A preview build of MiniMax M3.1 that has shown up in MiniMax Code since September 27, 2026. The model itself is verifiable in MiniMax's official code repository; the Flash Preview label itself is so far only documented through community reports.
What is officially confirmed about MiniMax M3.1?
MiniMax's own code repository contains concrete MiniMax-M3.1 fixtures: 512,000 tokens of default context, a 1,000,000-token option, a 128,000-token output cap and multi-level reasoning-effort configuration. Test fixtures are not an API specification, though.
How large is the M3.1 context window?
MiniMax's code tests carry 512,000 tokens as the default and catalog context plus a 1,000,000-token option. Whether the preview route actually serves those tiers in production is not guaranteed.
Are there benchmarks for MiniMax M3.1?
No — none that can be labeled verified M3.1 results. MiniMax has not published an M3.1 benchmark table, and independent trackers currently list M3 rather than a verified M3.1.
Is Space Bunny Alpha secretly MiniMax M3.1?
Plausible but unproven. Independent tokenizer probes place Space Bunny Alpha in the MiniMax family, but they do not reliably separate older from newer M generations. Without confirmation from MiniMax or the provider, those scores stay candidate data.
How much do M3.1 tokens cost?
Unknown. MiniMax has published no M3.1 price sheet as of the review date. The M3 prices of roughly $0.30 per million input and $1.20 per million output tokens are an M3 baseline and must not be carried over to M3.1.
Can I run MiniMax M3.1 locally on a Mac?
Not officially. No M3.1 weights have been released and the preview runs inside MiniMax Code. For local M-family models, M3 and the already published weights remain the option.
Is M3.1 Flash Preview worth testing for coding agents?
Yes, as a test candidate. A fixed eval set of real bug fixes, tool-use tests including failure cases and a long-context check will tell you more than any leaderboard guess.
Transparency
Sources and review basis
These primary and reference sources form the basis of the technical assessment. Vendor claims and external benchmarks are identified as such in the article.
- minimax.io blog / minimax-m3
- huggingface.co MiniMaxAI / MiniMax-M3
- huggingface.co main / config.json
- huggingface.co main / LICENSE
- arxiv.org abs / 2606.13392
- github.com MiniMax-AI / minimax-code
- github.com catalog / agent-model-selection.test.ts
- github.com catalog / list-models.test.ts
- github.com resolution / model-ref.test.ts
- github.com integration / session-system-queue-repository.integration.test.ts
- minimax.io news / minimax-announces-first-half-2026-financial-results-1787744160
- ir.minimax.io news-events / new-releases
- ir.minimax.io corporate-filings / quarterly-results
- minimax.io news / minimax-m2
- finance.sina.com.cn 2026-08-26 / doc-inipsezp9148443.shtml
- finance.sina.com.cn 2026-08-27 / doc-inipuqhs0676243.shtml
- datalearner.com pretrained-models / minimax-m3-1
- ai.kakarot.net issue-215
- epoch0.tokyo daily / 2026-09-20
- aireiter.com blog / minimax-m3-1-release-api
- forum.trae.cn topic / 182583
- linux.do topic / 2957314
- reddit.com 1wrbvjw / minimax_m31flash_preview
- openrouter.ai stealth / space-bunny-alpha
- spacebunnyalpha.com spacebunnyalpha.com
- yfarmx.com stealth / space-bunny-alpha
- spacebunnyai.com is-space-bunny-alpha-minimax.html
- spacebunnyai.com space-bunny-alpha-vs-minimax-m3.html
- reddit.com 1wpv45g / m31_is_space_bunny_alpha
- artificialanalysis.ai models / minimax-m3
- artificialanalysis.ai articles / minimax-m3
- artificialanalysis.ai comparisons / minimax-m3-vs-gemini-3-flash-reasoning
- artificialanalysis.ai comparisons / minimax-m3-vs-gemini-3-1-flash-lite-preview
- developer.nvidia.com blog / deploy-long-context-reasoning-and-agentic-workflows-with-minimax-m3-on-nvidia-accelerated-infrastructure
- huggingface.co nvidia / MiniMax-M3-NVFP4
- marktechpost.com 01 / minimax-releases-minimax-m3-with-msa-architecture-supporting-1m-token-context-native-multimodality-and-agentic-coding