Hermes 4 405B

Nous· released 26 Aug 2025

nousresearch/hermes-4-405b
Input / 1M tokens
$1
Output / 1M tokens
$3
Context window131K tokens

Longest answer: 118K tokens

reads textopen weightsreasoning

About

Hermes 4 405B is a large-scale reasoning model from Nous, built on Meta-Llama-3.1-405B. Its summary describes a hybrid reasoning mode in which the model can choose internal deliberation. It accepts text and produces text, with a 131,072-token context window and maximum output of 117,964 tokens. Reasoning is supported, while no tool support is listed. The model has open weights and is identified as nousresearch/hermes-4-405b, with the Hugging Face id NousResearch/Hermes-4-405B. It was released on 2025-08-26. Pricing is $1 USD per 1M input tokens and $3 USD per 1M output tokens. For stack selection, the practical boundaries are text-only input and output, the stated context and output limits, and its listed per-token costs.

Who it is for

Teams evaluating a text-only reasoning model with open weights may consider Hermes 4 405B. Its described hybrid reasoning mode may be relevant where internal deliberation is useful.

What is good

  • Open weights are available.
  • Built on Meta-Llama-3.1-405B.
  • Supports reasoning.
  • Context window is 131,072 tokens.

What to know first

  • Accepts text inputs only.
  • No tool support is listed.
  • Input costs $1 USD per 1M tokens.
  • Output costs $3 USD per 1M tokens.

Inferse review

Hermes 4 405B: the full review

Hermes 4 405B offers open weights and a described hybrid reasoning mode, with text-only input and output. Its 131,072-token context, output limit, and listed costs define the practical boundaries.

Overview

Hermes 4 405B is a text-only reasoning model from Nous, built on Meta-Llama-3.1-405B. Nous released it on August 26, 2025. It has open weights, giving teams a model they can examine and deploy within workflows suited to their infrastructure. The listed model identifier is nousresearch/hermes-4-405b, and the Hugging Face identifier is NousResearch/Hermes-4-405B.

The model's defining idea is hybrid reasoning: it can choose whether to deliberate internally. That makes it relevant to builders evaluating a large model for reasoning-oriented text tasks, while its substantial context and output limits leave room for extended inputs and responses. The available specifications do not detail deployment requirements or the mechanics of its reasoning mode, so those should be assessed against a team's own constraints.

Teams comparing this release with others from Nous can browse Nous models. It also belongs in comparisons of open-weight models and reasoning models.

Key features

Hybrid reasoning

Hermes 4 405B is marked as a reasoning model and supports a hybrid approach in which it can choose to deliberate internally. The listed facts establish that capability, but do not specify how users control it or which tasks benefit most.

Room for long inputs and outputs

The context window is 131,072 tokens, and the maximum output is 117,964 tokens. Those limits provide capacity for large text inputs and lengthy generated responses, subject to the needs of the application using the model.

Open weights and text modality

Open weights are listed as available. Both the input and output modality are text, so this model is positioned for text-based use rather than image, audio, or video workflows.

Pricing

The listed rates are USD $1 per 1 million input tokens and USD $3 per 1 million output tokens. Input and output are billed at different rates, so expected generation volume matters when estimating usage costs. The available pricing facts do not specify other fees or terms.

For cost-focused comparisons, see cheapest language models.

Platforms

The listed model ID is nousresearch/hermes-4-405b, and the official listing is on OpenRouter. Its Hugging Face ID is NousResearch/Hermes-4-405B. The facts identify open weights, but do not specify hardware requirements, licensing terms, or a particular self-hosting setup; teams should verify those details before planning deployment.

Who it's for

Hermes 4 405B is suited for teams considering a large, open-weight text model with reasoning capability and a long context window. It may fit evaluations involving substantial text inputs or responses, as long as the team's stack can accommodate the model and the application can work within the stated token limits.

Its listed input rate is lower than its output rate, making the input-to-generation mix an important part of cost planning. Teams needing non-text modalities or concrete deployment guidance will need to establish whether the model fits those requirements before choosing it.

Pros and cons

Pros

  • Open weights are listed.
  • Hybrid reasoning allows the model to choose whether to deliberate internally.
  • A 131,072-token context window and 117,964-token maximum output support large text exchanges.
  • Pricing is specified separately for input and output tokens.

Cons

  • Text is the only listed input and output modality.
  • The available facts do not explain how to control hybrid reasoning or provide hardware and licensing details.
  • Output tokens cost more per million than input tokens.

Alternatives

For a smaller model in the same family, consider Hermes 3 70B Instruct. Another earlier family option is Hermes 3 405B Instruct. Their specifications are not included here, so compare their current capabilities and pricing before selecting between them.

For broader comparisons, browse newest models and longest-context models.

Verdict

Hermes 4 405B combines open weights, text-only reasoning, hybrid deliberation, and a context window of 131,072 tokens. Its stated rates—USD $1 per 1 million input tokens and USD $3 per 1 million output tokens—make usage costs dependent on how much the application generates. It is a credible candidate for teams evaluating large text models, but the listed facts leave deployment requirements and reasoning controls unspecified. Confirm those fit your stack before committing.

Details

Lab
Nousopenrouter.ai · 3 Oct 2026
Context
131Kopenrouter.ai · 3 Oct 2026
Input price
$1 / 1Mopenrouter.ai · 3 Oct 2026
Output price
$3 / 1Mopenrouter.ai · 3 Oct 2026
Max output
117,964 tokensopenrouter.ai · 3 Oct 2026
Inputs
textopenrouter.ai · 3 Oct 2026
Open weights
Yesopenrouter.ai · 3 Oct 2026

More longest-context models

See the list