MiMo-V2.5
Xiaomi· released 22 Apr 2026
xiaomi/mimo-v2.5- Input / 1M tokens
- $0.14
- Output / 1M tokens
- $0.28
Longest answer: 131K tokens
About
MiMo-V2.5 is Xiaomi’s native omnimodal model, with text output and text, audio, image and video inputs. It has a 1,050,000-token context window, supports tools and reasoning, and is listed with open weights. Xiaomi describes it as delivering Pro-level agentic performance at roughly half the inference cost, and says it surpasses MiMo-V2-Omni in multimodal perception across image and video understanding. Input costs $0.14 USD per 1M tokens and output costs $0.28 USD per 1M tokens. The model id is xiaomi/mimo-v2.5, and its Hugging Face id is XiaomiMiMo/MiMo-V2.5. It belongs to the xiaomi family and was released on 2026-04-22. Teams considering multimodal or agentic tasks can weigh its input breadth and open weights alongside the stated token rates.
Who it is for
Teams building workflows that take audio, images, video or text as input may find its modality support relevant. Open weights and tool support are also listed.
What is good
- Accepts text, audio, images and video
- 1,050,000-token context window
- Open weights listed
- Tools and reasoning supported
What to know first
- Output is text only
- Input costs $0.14 USD per 1M tokens
Inferse review
MiMo-V2.5: the full review
MiMo-V2.5 supports four input modalities, a million-plus-token context window, tools and reasoning. Its listed rates are $0.14 USD per 1M input tokens and $0.28 USD per 1M output tokens.
Overview
MiMo-V2.5 is Xiaomi’s native omnimodal model, built to take text, audio, images and video as input and produce text. It combines reasoning and tool support with a context window of 1,050,000 tokens, making it relevant to teams weighing multimodal inputs, long-context work and agentic workflows in one model.
Xiaomi positions MiMo-V2.5 as delivering Pro-level agentic performance at roughly half the inference cost, and says it surpasses MiMo-V2-Omni in multimodal perception across image and video understanding. Those are the maker’s claims; the available specifications do not provide benchmark results or define the comparison’s cost baseline.
The model was released on April 22, 2026. It is listed as open-weight, with the Hugging Face identifier XiaomiMiMo/MiMo-V2.5. Its model ID is xiaomi/mimo-v2.5.
Key features
Multimodal input with text output
MiMo-V2.5 accepts text, audio, image and video inputs, while its listed output is text. The broad input mix suits applications that need to bring different media types into a language-model workflow. The specifications do not describe particular audio or video tasks, so teams should assess whether the model supports their required use cases before adopting it.
Long context and reasoning
A 1,050,000-token context window is a defining specification for this model. Reasoning is listed as supported, but no further details are given about reasoning modes or task-specific performance. Maximum output is 131,072 tokens, a separate limit from the context window.
Tool support and open weights
Tools are listed as supported, alongside open weights. This makes the model a candidate for workflows that combine model responses with tool use, as well as teams that want access to model weights. The available facts do not specify supported tool protocols, deployment requirements or the terms governing use of the weights.
Pricing
Listed API pricing is USD 0.14 per 1 million input tokens and USD 0.28 per 1 million output tokens. Input and output are priced separately, so a workload’s token mix will affect its total cost. Xiaomi’s roughly half-cost positioning is not tied here to a named baseline, and should not be treated as a substitute for estimating costs against a team’s own usage.
Platforms
MiMo-V2.5 is listed as an API model with the ID xiaomi/mimo-v2.5. Its weights are marked open, and the listed Hugging Face identifier is XiaomiMiMo/MiMo-V2.5. The available details do not specify hosting options, SDKs, API compatibility, or local hardware requirements, so teams should confirm those fit their stack.
It also belongs among open-weight models, Xiaomi models, reasoning models, vision language models, longest-context models and models with tool calling.
Who it's for
MiMo-V2.5 is worth considering for teams building applications that need text generation from text, audio, image or video inputs, especially where a very large context window, reasoning and tool support are useful together. Its open-weight listing may also matter to builders who need access to weights, though deployment specifics are not provided here.
It is a less complete fit for buyers who need documented deployment options, defined tool interfaces, modality-specific performance results or a clear cost comparison before choosing a model. Those details are not included in the available specifications.
Pros and cons
Pros
- Accepts four listed input types: text, audio, image and video.
- Offers a 1,050,000-token context window and a 131,072-token maximum output.
- Reasoning and tool support are listed.
- Open weights are listed, with a Hugging Face identifier.
- Input and output token prices are stated separately and are low in absolute terms.
Cons
- Text is the only listed output modality, despite the broad input coverage.
- No benchmark figures or detailed modality-specific capabilities are provided.
- The claimed cost advantage has no named comparison baseline in the available facts.
- Deployment, hardware and tool-interface details are not specified.
Alternatives
For a different model family, Llama 4 Scout is one alternative to consider. Xiaomi’s MiMo-V2.5-Pro, MiMo-V2.6-Flash and MiMo-V2.6-Pro are other options in the maker’s model lineup.
Other listed alternatives include LongCat 2.0, DeepSeek V4.1 Flash, its batch variant, and DeepSeek V4 Flash 0423. No comparable specifications are provided here, so these options should be evaluated against the same workload and stack requirements.
Verdict
MiMo-V2.5 brings an unusually broad input list, a million-token-scale context window, reasoning, tool support and open weights together at clearly stated per-token prices. That combination makes it a credible candidate for teams exploring multimodal or long-context applications. But the specification leaves important integration questions unanswered, and Xiaomi’s performance and cost positioning is not accompanied here by benchmark detail or a defined comparison. Treat it as a promising option to qualify against the demands of your workload, not as a proven fit based on the headline claims alone.
Details
- Lab
- Xiaomiopenrouter.ai · 3 Oct 2026
- Context
- 1,050Kopenrouter.ai · 3 Oct 2026
- Input price
- $0.14 / 1Mopenrouter.ai · 3 Oct 2026
- Output price
- $0.28 / 1Mopenrouter.ai · 3 Oct 2026
- Max output
- 131,072 tokensopenrouter.ai · 3 Oct 2026
- Inputs
- text, audio, image, videoopenrouter.ai · 3 Oct 2026
- Open weights
- Yesopenrouter.ai · 3 Oct 2026
More newest models
See the listListed on Inferse
- Newest Models in 2026466 listed
- Longest-Context Models in 2026466 listed
- Cheapest Language Models in 2026437 listed
- Models with Tool Calling in 2026395 listed
- Reasoning Models in 2026332 listed
- Vision Language Models in 2026295 listed
- Open-Weight Models in 2026178 listed
- Xiaomi Models: Pricing per 1M Tokens and Context (2026)5 listed
Sources
- openrouter.ai/xiaomi/mimo-v2.5· checked 3 Oct 2026

