GLM 5.3 FlashX

Z.ai· released 18 Sept 2026

z-ai/glm-5.3-flashx
Input / 1M tokens
$0.37
Output / 1M tokens
$1.25
Context window1M tokens

Longest answer: 131K tokens

reads textreads imagereads videotool callingvisionreasoning

About

GLM 5.3 FlashX is a text-output model from Z.ai in the z-ai family, designed as a high-speed variant of GLM-5.3-Flash. Its summary describes a native multimodal model with inference speeds up to 200 tokens per second and a hybrid sparse and linear attention architecture. It accepts text, images, and video, and supports vision and reasoning. The context window is 1,048,576 tokens, with a maximum output of 131,072 tokens. The model identifier is z-ai/glm-5.3-flashx. Input costs 0.37 USD per 1M tokens and output costs 1.25 USD per 1M tokens. Z.ai is the listed maker and lab. The model was released and created on 2026-09-18. Its official listing is at https://openrouter.ai/z-ai/glm-5.3-flashx.

Who it is for

It suits teams that need text generation with image or video inputs and reasoning support. Its large context window and stated inference speed may also fit workloads that need substantial input capacity or fast responses.

What is good

  • Accepts text, image, and video inputs
  • Supports reasoning and vision
  • Context window of 1,048,576 tokens
  • Maximum output of 131,072 tokens

What to know first

  • Input costs 0.37 USD per 1M tokens
  • Output costs 1.25 USD per 1M tokens

Inferse review

GLM 5.3 FlashX: the full review

GLM 5.3 FlashX pairs multimodal input and reasoning with a large context window and a stated speed of up to 200 tokens per second. Its listed token prices are 0.37 USD per 1M input tokens and 1.25 USD per 1M output tokens.

Overview

GLM 5.3 FlashX is Z.ai’s high-speed variant of GLM-5.3-Flash. It is a multimodal model with text, image and video inputs, and text output. Z.ai lists inference speeds of up to 200 tokens per second. That speed is a stated ceiling, not a guarantee for every request or deployment.

The model uses the same hybrid sparse and linear attention architecture as GLM-5.3-Flash. Its context window is 1,048,576 tokens, and its maximum output is 131,072 tokens. Those limits make context and response length important parts of the fit for teams assessing it alongside latency and token cost.

GLM 5.3 FlashX supports reasoning and tools, and includes vision. Its model ID is z-ai/glm-5.3-flashx. It was released on September 18, 2026. For teams building around an existing model stack, the practical profile is clear: text generation, image and video inputs, a very large context allowance, and tool support, with pricing that distinguishes input from output tokens.

It also fits several distinct model-selection needs: vision language models, reasoning models, longest-context models and models with tool calling. Those categories describe capabilities, not a claim about performance on a particular workload.

Key features

  • High stated speed: Z.ai describes inference speeds of up to 200 tokens per second.
  • Multimodal inputs: the listed input types are text, image and video; output is text.
  • Large context: the context window is 1,048,576 tokens.
  • Long responses: the maximum output is 131,072 tokens.
  • Reasoning and tools: both are listed as supported capabilities.
  • Hybrid attention architecture: FlashX shares the hybrid sparse and linear attention architecture of GLM-5.3-Flash.

These specifications help narrow down whether the model belongs in an evaluation shortlist. They do not establish quality for a particular task, nor do they describe availability or behavior in a specific deployment.

Pricing

Listed pricing is $0.37 USD per 1 million input tokens and $1.25 USD per 1 million output tokens. Input and output rates differ, so teams estimating usage should account for both sides of a request, including how much text they expect the model to generate. The supplied pricing does not specify other charges or terms.

Platforms

The listed model identifier is z-ai/glm-5.3-flashx, and the official model page is on OpenRouter. The available facts do not specify SDKs, deployment options, regional availability or integrations, so teams should verify those details against the platform and stack they plan to use.

Who it's for

GLM 5.3 FlashX is worth considering for teams that need a model to accept text alongside images or video, and that want reasoning and tool support in the same model profile. Its large context and output limits may matter for workflows that need to provide substantial input or request long text responses. The listed limits alone do not show how well it handles any specific workload.

It may also suit teams weighing speed against token spend. The stated ceiling of up to 200 tokens per second gives a speed reference, while the separate input and output rates make cost estimation possible from expected token volumes. Actual latency and total spend will depend on usage and deployment details not included here.

Pros and cons

Pros

  • Supports text, image and video input, with text output.
  • Lists reasoning and tool support.
  • Offers a 1,048,576-token context window and up to 131,072 output tokens.
  • Input pricing is $0.37 USD per 1 million tokens.

Cons

  • The stated speed is a maximum, not a per-request performance promise.
  • Output tokens cost more than input tokens: $1.25 versus $0.37 USD per 1 million tokens.
  • The available specifications do not establish task quality, deployment choices or integrations.

Alternatives

Within Z.ai’s lineup, GLM 5.3 Flash is the closest named comparison because FlashX is described as its high-speed variant. Other nearby options include GLM 5.3 Prime, GLM Flash Latest and GLM 5.3 Flash (batch). The supplied facts do not include their specifications or prices, so a direct cost or capability ranking would require further verification.

For broader comparisons, the lineup also includes GLM Latest, GLM 5.3, GLM 5.3 (batch) and GLM 5.2. See the full Z.ai models list or browse newest models for additional candidates.

Verdict

GLM 5.3 FlashX has a well-defined specification profile: text, image and video inputs; text output; reasoning and tool support; a million-token-scale context; and a stated speed ceiling of up to 200 tokens per second. Its listed rates are $0.37 USD per 1 million input tokens and $1.25 USD per 1 million output tokens. That makes token mix central to budget planning.

It is a credible candidate for teams whose requirements align with those documented capabilities, but the specifications do not answer questions about workload quality, real-world latency or deployment fit. Compare it with related Z.ai models against your required inputs, output length, cost and platform constraints before choosing.

Details

Lab
Z.aiopenrouter.ai · 3 Oct 2026
Context
1,049Kopenrouter.ai · 3 Oct 2026
Input price
$0.37 / 1Mopenrouter.ai · 3 Oct 2026
Output price
$1.25 / 1Mopenrouter.ai · 3 Oct 2026
Max output
131,072 tokensopenrouter.ai · 3 Oct 2026
Inputs
text, image, videoopenrouter.ai · 3 Oct 2026
Open weights
Noopenrouter.ai · 3 Oct 2026

More newest models

See the list