GLM 5.3 Prime

Z.ai · released 23 Sept 2026

Input$2.8per 1M tokens
Output$8.8per 1M tokens
Context1,000Ktokens
WeightsClosed
context on a 1K–2M scale

About

GLM 5.3 Prime is Z.ai’s high-speed variant of GLM-5.3. It uses inference acceleration to deliver 1.5–2× the output throughput, while inheriting GLM-5.3’s full capabilities. The model takes text input and produces text output, with a context window of 1,000,000 tokens and a maximum output of 131,072 tokens. It supports reasoning and tools. Pricing is USD 2.8 per 1M input tokens and USD 8.8 per 1M output tokens. GLM 5.3 Prime was released on 2026-09-23 and has the model ID z-ai/glm-5.3-prime. Its official listing is on OpenRouter. For developers and product teams evaluating it, the key listed details are its text-only input and output, throughput description, context capacity, output cap, and separate input and output costs.

Who it is for

GLM 5.3 Prime may suit teams seeking Z.ai’s faster-output variant and working with text input and output. Its million-token context window may be relevant for applications that need substantial context.

What is good

  • 1,000,000-token context window
  • Described as delivering 1.5–2× output throughput
  • Supports tools and reasoning

What to know first

  • Text is the only listed input modality
  • Output limited to 131,072 tokens
  • Output costs USD 8.8 per 1M tokens

Verdict

GLM 5.3 Prime pairs a large context window with a stated throughput increase over GLM-5.3. Weigh that against its text-only input and USD 8.8 per 1M output-token price.

Details

Lab
Z.aiopenrouter.ai · 3 Oct 2026
Context
1,000Kopenrouter.ai · 3 Oct 2026
Input price
$2.8 / 1Mopenrouter.ai · 3 Oct 2026
Output price
$8.8 / 1Mopenrouter.ai · 3 Oct 2026
Max output
131,072 tokensopenrouter.ai · 3 Oct 2026
Inputs
textopenrouter.ai · 3 Oct 2026
Open weights
Noopenrouter.ai · 3 Oct 2026

More longest-context models

See the list