Global technology company Xiaomi Corp. released and open-sourced its MiMo-V2.6 series generative artificial intelligence models. The release includes two natively omnimodal models that aim to balance intelligence, efficiency, and cost.
The MiMo-V2.6 series features a new flagship model named MiMo-V2.6-Pro alongside a smaller, efficient Flash variant. Xiaomi is also rolling out a slimmed-down Pro-UltraSpeed variant that delivers up to 20 times faster output than the Pro model at the same quality.
The Pro and Flash models are native omnimodal systems capable of accepting text, image, video, and audio in the same family. Both feature 1 million-token context windows. The Pro architecture lists a 681 million-parameter vision encoder, which includes a 308 million-parameter AudioTokenizer and a 127 million-parameter audio patch encoder that can distinguish speech.
This design helps larger models meet the growing demands of agentic AI use cases and task instructions, particularly for computer use. A computer-use agent can read a task instruction, analyze a screenshot of a user interface or video, reason about the content, and decide on the next step within a single model loop.
Coding and design agents can combine source code, lengthy repository context, screenshots, UI mockups, and reference images in one session. Business applications allow the model to ingest numerous documents containing multimodal components such as call audio, screenshots, logs, and notes without requiring external tools to break down transcripts.
Xiaomi demonstrated the model managing game world building, Blender-based 3D modeling, embodied simulation, and video-music workflows in an announcement blog post. The V2.6-Pro model scored 46.32 on the Artificial Analysis Intelligence Index, outperforming Kimi K3 and Qwen3.8 Max as the highest-ranked open-source model at launch. However, it still trails advanced proprietary closed-source models such as Claude Fable 5.1 and GPT-6 Astra on aggregate measures.
On agentic tasks, V2.6-Pro scored 53.1 on AutomationBench compared to Claude Opus 5 at 50.3. It tied Opus 5 with 31.6 on Agents’ Last Exam and scored 89.9 on Terminal Bench 2.1, slightly ahead of the 89.1 score for Opus 5.
Xiaomi will maintain the standard application programming interface pricing used for MiMo-V2.5. Model cards list availability through AI Studio, MiMo Desktop, and MiMo Code.
The Pro and Flash models are also available on OpenRouter with a 1.05-million-token context through the unified API. V2.6-Flash is priced at $0.14 per 1 million input tokens and $0.28 per output token, while the Pro model costs $0.435 for input and $0.87 for output. The Pro-UltraSpeed variant costs $4.35 for input and $8.70 for output.
Xiaomi claims the Pro model costs roughly one-20th to one-60th as much as overseas models at comparable intelligence when accounting for reduced cached-token costs. OpenAI Group PBC GPT-6 Astra and Anthropic PBC Claude Fable 5.1 cost $10 to $50 per million input and output tokens.


