Xiaomi MiMo-V2.6-Pro ships 1T open weights under MIT
Xiaomi's MiMo-V2.6-Pro has 1.02 trillion parameters, 42 billion active, a 1M-token context and MIT-licensed weights, but needs about 680 GB of GPU memory.
3 min read

By the numbers
- total parameters in MiMo-V2.6-Pro
- 1.02T
- parameters active per token
- 42B
- token context window
- 1M
- GPU memory needed to serve it with vLLM
- 680 GB
Xiaomi has released MiMo-V2.6-Pro and MiMo-V2.6-Flash, two AI models whose weights anyone can download under the MIT license. The Pro model has 1.02 trillion parameters, the internal values a model learns during training. DataNorth dates the release to September 21, 2026, while Forkast and LLM Rumors give September 22.
Open weights mean a company can run the model on its own hardware and change it freely. The catch, as LLM Rumors puts it, is that "open weights" does not mean simple self-hosting. Running the Pro model takes about 680 GB of GPU memory.
What Xiaomi released
Both models use a mixture-of-experts design. In this design, the model holds many specialist sub-networks, called experts, and uses only a few of them for each token, a short chunk of text. That keeps the cost of each answer far below what the total size suggests.
| MiMo-V2.6-Pro | MiMo-V2.6-Flash | |
|---|---|---|
| Total parameters | 1.02 trillion | 309 billion |
| Active per token | 42 billion | 15 billion |
| API price per million input tokens | $0.435 | $0.14 |
| API price per million output tokens | $0.87 | $0.28 |
DataNorth reports both models take text, images, audio and video as input and write text out. The context window, the amount of text a model can consider at once, is 1,048,576 tokens. The weights are on Hugging Face, and the models also run on Xiaomi AI Studio, Xiaomi's MiMo API and OpenRouter.
Forkast adds detail on the design. It describes a frozen-router mixture of experts with hybrid attention, plus a five-layer speculative decoder. A speculative decoder drafts several tokens ahead, so the model can produce text faster.
How it scores
DataNorth and Forkast report these results for the Pro model:
| Benchmark | Score |
|---|---|
| Artificial Analysis Intelligence Index v4.3 | 46.32, tied with Grok 4.7 |
| Terminal Bench 2.1 | 89.9 |
| DeepSWE v1.1 | 71.9 |
| AutomationBench | 53.1 |
| CyberGym | 94.0, per Forkast |
The Artificial Analysis index combines many tests into one score. Matching Grok 4.7 there puts an MIT-licensed model level with xAI's model on that measure.
The training kit that came with it
Xiaomi released more than 7,000 reinforcement learning task environments alongside the weights, Forkast and LLM Rumors report. Reinforcement learning trains a model by rewarding good results on tasks. These environments are the tasks.
Forkast describes the training framework as fully asynchronous and based on GRPO, a reinforcement learning method. It used 1,568 training prompts and 16 attempts, or rollouts, per step. LLM Rumors adds that Xiaomi also published a mimoagent repository with building blocks for agents and evaluation tools.
What it takes to run
Self-hosting the Pro model is a data-center job. LLM Rumors reports a 566 GB checkpoint. Serving it with vLLM, a popular inference server, needs about 680 GB of GPU memory in total. That is roughly eight Nvidia H200 or four AMD MI355X accelerators, the outlet says. Setups using SGLang, another server, spread it across two machines.
DataNorth notes that Xiaomi gives no hardware guidance and no speed figures of its own.
What this means for developers
Start with the API, not your own GPUs. At $0.435 per million input tokens and $0.87 per million output tokens, the hosted Pro model is cheap to test. Run your own tasks through it on Xiaomi's API or OpenRouter before you think about hardware.
Try Flash for high-volume work. With 15 billion active parameters and prices of $0.14 and $0.28 per million tokens, Flash suits classification, extraction and other jobs where cost matters more than peak quality.
Budget honestly for self-hosting. Eight H200-class GPUs is a serious purchase or rental. Self-host only if you need data to stay on your own machines, or if your volume makes the API more expensive than the hardware.
Check the license with your lawyers, then use it. MIT is one of the most permissive licenses, which makes MiMo-V2.6 easy to fine-tune and ship inside commercial products. Confirm the license file in the release matches what you expect.
Use the training environments even if you skip the model. Teams building their own agents can reuse the 7,000-plus task environments and the published framework to test and train other models.
Sources
Related articles

Qwen3.8-Omni-Flash cuts audio input costs by 98%
Alibaba's new omnimodal model reads text, images, audio and video inside a 1M-token context, and prices audio input more than 98% below its predecessor.

AstaBrief 8B: Ai2's open model writes cited reports
Ai2 released AstaBrief 8B, an Apache 2.0 model on Qwen3-8B that writes cited research reports in 51.1 seconds, 3.5x faster than Asta's Thinking mode.

OpenAI links reasoning-extraction campaign to Moonshot AI
OpenAI says a campaign it ties to Moonshot AI, maker of Kimi, sent 16,000 requests from 4,000+ users in July to copy its models' hidden reasoning.
The daily brief
Three to five stories a day, and what each one means for the people who build software. Free, no spam.