Hermes 4 405B
Nous· released 26 Aug 2025
nousresearch/hermes-4-405b- Input / 1M tokens
- $1
- Output / 1M tokens
- $3
Longest answer: 118K tokens
About
Hermes 4 405B is a large-scale reasoning model from Nous, built on Meta-Llama-3.1-405B. Its summary describes a hybrid reasoning mode in which the model can choose internal deliberation. It accepts text and produces text, with a 131,072-token context window and maximum output of 117,964 tokens. Reasoning is supported, while no tool support is listed. The model has open weights and is identified as nousresearch/hermes-4-405b, with the Hugging Face id NousResearch/Hermes-4-405B. It was released on 2025-08-26. Pricing is $1 USD per 1M input tokens and $3 USD per 1M output tokens. For stack selection, the practical boundaries are text-only input and output, the stated context and output limits, and its listed per-token costs.
Who it is for
Teams evaluating a text-only reasoning model with open weights may consider Hermes 4 405B. Its described hybrid reasoning mode may be relevant where internal deliberation is useful.
What is good
- Open weights are available.
- Built on Meta-Llama-3.1-405B.
- Supports reasoning.
- Context window is 131,072 tokens.
What to know first
- Accepts text inputs only.
- No tool support is listed.
- Input costs $1 USD per 1M tokens.
- Output costs $3 USD per 1M tokens.
Inferse review
Hermes 4 405B: the full review
Hermes 4 405B offers open weights and a described hybrid reasoning mode, with text-only input and output. Its 131,072-token context, output limit, and listed costs define the practical boundaries.
Overview
Hermes 4 405B is a text-only reasoning model from Nous, built on Meta-Llama-3.1-405B. Nous released it on August 26, 2025. It has open weights, giving teams a model they can examine and deploy within workflows suited to their infrastructure. The listed model identifier is nousresearch/hermes-4-405b, and the Hugging Face identifier is NousResearch/Hermes-4-405B.
The model's defining idea is hybrid reasoning: it can choose whether to deliberate internally. That makes it relevant to builders evaluating a large model for reasoning-oriented text tasks, while its substantial context and output limits leave room for extended inputs and responses. The available specifications do not detail deployment requirements or the mechanics of its reasoning mode, so those should be assessed against a team's own constraints.
Teams comparing this release with others from Nous can browse Nous models. It also belongs in comparisons of open-weight models and reasoning models.
Key features
Hybrid reasoning
Hermes 4 405B is marked as a reasoning model and supports a hybrid approach in which it can choose to deliberate internally. The listed facts establish that capability, but do not specify how users control it or which tasks benefit most.
Room for long inputs and outputs
The context window is 131,072 tokens, and the maximum output is 117,964 tokens. Those limits provide capacity for large text inputs and lengthy generated responses, subject to the needs of the application using the model.
Open weights and text modality
Open weights are listed as available. Both the input and output modality are text, so this model is positioned for text-based use rather than image, audio, or video workflows.
Pricing
The listed rates are USD $1 per 1 million input tokens and USD $3 per 1 million output tokens. Input and output are billed at different rates, so expected generation volume matters when estimating usage costs. The available pricing facts do not specify other fees or terms.
For cost-focused comparisons, see cheapest language models.
Platforms
The listed model ID is nousresearch/hermes-4-405b, and the official listing is on OpenRouter. Its Hugging Face ID is NousResearch/Hermes-4-405B. The facts identify open weights, but do not specify hardware requirements, licensing terms, or a particular self-hosting setup; teams should verify those details before planning deployment.
Who it's for
Hermes 4 405B is suited for teams considering a large, open-weight text model with reasoning capability and a long context window. It may fit evaluations involving substantial text inputs or responses, as long as the team's stack can accommodate the model and the application can work within the stated token limits.
Its listed input rate is lower than its output rate, making the input-to-generation mix an important part of cost planning. Teams needing non-text modalities or concrete deployment guidance will need to establish whether the model fits those requirements before choosing it.
Pros and cons
Pros
- Open weights are listed.
- Hybrid reasoning allows the model to choose whether to deliberate internally.
- A 131,072-token context window and 117,964-token maximum output support large text exchanges.
- Pricing is specified separately for input and output tokens.
Cons
- Text is the only listed input and output modality.
- The available facts do not explain how to control hybrid reasoning or provide hardware and licensing details.
- Output tokens cost more per million than input tokens.
Alternatives
For a smaller model in the same family, consider Hermes 3 70B Instruct. Another earlier family option is Hermes 3 405B Instruct. Their specifications are not included here, so compare their current capabilities and pricing before selecting between them.
For broader comparisons, browse newest models and longest-context models.
Verdict
Hermes 4 405B combines open weights, text-only reasoning, hybrid deliberation, and a context window of 131,072 tokens. Its stated rates—USD $1 per 1 million input tokens and USD $3 per 1 million output tokens—make usage costs dependent on how much the application generates. It is a credible candidate for teams evaluating large text models, but the listed facts leave deployment requirements and reasoning controls unspecified. Confirm those fit your stack before committing.
Details
- Lab
- Nousopenrouter.ai · 3 Oct 2026
- Context
- 131Kopenrouter.ai · 3 Oct 2026
- Input price
- $1 / 1Mopenrouter.ai · 3 Oct 2026
- Output price
- $3 / 1Mopenrouter.ai · 3 Oct 2026
- Max output
- 117,964 tokensopenrouter.ai · 3 Oct 2026
- Inputs
- textopenrouter.ai · 3 Oct 2026
- Open weights
- Yesopenrouter.ai · 3 Oct 2026
More longest-context models
See the listListed on Inferse
- Longest-Context Models in 2026466 listed
- Newest Models in 2026466 listed
- Cheapest Language Models in 2026437 listed
- Reasoning Models in 2026332 listed
- Open-Weight Models in 2026178 listed
- Nous Models: Pricing per 1M Tokens and Context (2026)3 listed
Sources
- openrouter.ai/nousresearch/hermes-4-405b· checked 3 Oct 2026


