Mixtral 8x22B Instruct

Mistral· released 17 Apr 2024

mistralai/mixtral-8x22b-instruct
Input / 1M tokens
$2
Output / 1M tokens
$6
Context window66K tokens

Longest answer: 52K tokens

reads textreads fileopen weightstool calling

About

Mixtral 8x22B Instruct is Mistral’s instruct fine-tune of Mixtral 8x22B, with open weights and text input and output. It uses 39 billion active parameters from a total of 141 billion. The model accepts text and files, supports tools, and has a 65,536-token context window with a maximum output of 52,428 tokens. It was released on April 17, 2024, and is identified on Hugging Face as mistralai/Mixtral-8x22B-Instruct-v0.1. Listed API rates are $2 USD per 1 million input tokens and $6 USD per 1 million output tokens. The listing places it in the mistralai family and identifies Mistral as both maker and lab. Its text-only modality makes it a language model entry rather than a multimodal one; file input is also listed. These details give teams concrete limits and rates to consider when assessing it for instruction-following workloads.

Who it is for

It suits teams looking for an open-weight Mistral instruction model with tool support and file input. Its context window and separate input and output rates are relevant to teams planning token budgets.

What is good

  • Open weights are available.
  • Supports tools and file input.
  • 65,536-token context window.
  • Maximum output is 52,428 tokens.
  • Input and output rates are listed.

What to know first

  • Text is the only listed output modality.
  • Output tokens cost $6 USD per million.

Inferse review

Mixtral 8x22B Instruct: the full review

Mixtral 8x22B Instruct combines open weights, tool support, and a 65,536-token context window. Teams should account for its $2 USD input and $6 USD output rates per million tokens.

Overview

Mixtral 8x22B Instruct is Mistral’s instruction-tuned model in the Mixtral 8x22B family. Released on April 17, 2024, it has open weights and is identified on Hugging Face as mistralai/Mixtral-8x22B-Instruct-v0.1. Its architecture has 141 billion parameters, with 39 billion active for a given operation. The model accepts text and files as input and produces text output.

For teams evaluating models against an existing stack, the useful starting points are its text-focused interface, 65,536-token context window, and support for tools. Those facts describe available capabilities, not a guarantee of performance for any particular workload.

It belongs in a review of Mistral models, and its open weights also make it relevant to teams comparing open-weight models.

Key features

Active parameter count

Mixtral 8x22B Instruct uses 39 billion active parameters out of 141 billion total. The model’s stated positioning emphasizes cost efficiency relative to its size; the available facts do not provide a benchmark or workload-specific measure to quantify that claim.

Context and inputs

The context window is 65,536 tokens. Inputs can be text or files, while the output modality is text. This combination gives teams room to work with longer inputs than a short prompt, though the context limit alone does not establish how accurately the model handles a particular task.

Tool support

Tools are supported, making the model a candidate for applications that need model-assisted tool use. Teams comparing that capability across models can browse models with tool calling.

Pricing

Input costs USD 2 per 1 million tokens, and output costs USD 6 per 1 million tokens. The maximum output is 52,428 tokens. Since output is priced at a higher per-token rate than input, teams estimating usage should account for both sides of a request rather than treating a single token rate as the full cost.

For broader price-oriented comparisons, see cheapest language models. That list provides context for shopping around; these listed rates are the specific prices for Mixtral 8x22B Instruct.

Platforms

The supplied facts identify the model through the id mistralai/mixtral-8x22b-instruct and the Hugging Face id mistralai/Mixtral-8x22B-Instruct-v0.1. They do not specify deployment platforms, SDKs, hosting arrangements, or integration requirements, so teams should verify those details against their intended environment before adopting it.

Who it's for

This model is worth evaluating for teams that want an open-weight, text-generating instruct model with file and text inputs, a substantial context window, and tool support. Its active-parameter design and listed token rates may also interest teams balancing model scale with per-token cost. Whether it fits a production workload depends on requirements not established here, including task quality, latency, hosting, and operational constraints.

Teams choosing primarily by context capacity can compare it with other entries in longest-context models. Its release date may also matter to teams surveying newest models, though release recency alone says nothing about suitability.

Pros and cons

Pros

  • Open weights, with a published Hugging Face model identifier.
  • 65,536-token context window and support for text and file inputs.
  • Tool support and a maximum output of 52,428 tokens.
  • Explicit input and output prices: USD 2 and USD 6 per 1 million tokens, respectively.

Cons

  • The model’s 141 billion total parameters make its scale an important factor for teams assessing deployment fit; the available facts do not specify hosting or hardware requirements.
  • Pricing alone does not establish total operating cost, and no latency, quality, or benchmark results are provided.
  • The listed modalities are limited to text input and text output, with files also accepted as inputs.

Alternatives

Within Mistral’s catalog, Mistral Medium 3.5 and Mistral Small 4 are alternatives to consider. Batch variants are listed separately as Mistral Medium 3.5 (batch) and Mistral Small 4 (batch).

Other named options include Devstral 2 2512, Ministral 3 14B 2512, Ministral 3 3B 2512, and Ministral 3 8B 2512. The supplied facts do not include their prices or capabilities, so a direct feature or cost comparison requires checking each model’s details.

Verdict

Mixtral 8x22B Instruct has a clear set of decision-relevant attributes: open weights, tool support, text and file inputs, a 65,536-token context window, and explicit token pricing. Its 39-billion active parameter count provides useful context for understanding the model’s design, but does not by itself determine deployment cost or task performance. It is a plausible candidate for teams whose requirements align with those stated capabilities; validate quality, hosting fit, and total cost for the intended use before selecting it.

Details

Lab
Mistralopenrouter.ai · 3 Oct 2026
Context
66Kopenrouter.ai · 3 Oct 2026
Input price
$2 / 1Mopenrouter.ai · 3 Oct 2026
Output price
$6 / 1Mopenrouter.ai · 3 Oct 2026
Max output
52,428 tokensopenrouter.ai · 3 Oct 2026
Inputs
text, fileopenrouter.ai · 3 Oct 2026
Open weights
Yesopenrouter.ai · 3 Oct 2026

More longest-context models

See the list