Hermes 3 70B Instruct
Nous· released 18 Aug 2024
nousresearch/hermes-3-llama-3.1-70b- Input / 1M tokens
- $0.7
- Output / 1M tokens
- $0.7
Longest answer: 16K tokens
About
Hermes 3 70B Instruct is a text-in, text-out language model from Nous, released on 2024-08-18. It has open weights, with the Hugging Face ID NousResearch/Hermes-3-Llama-3.1-70B. The listed description identifies it as a generalist model and notes capabilities in agentic tasks, roleplay, reasoning, multi-turn conversation, and long-context coherence. Its context window is 131,072 tokens, and its maximum output is 16,384 tokens. The model ID is nousresearch/hermes-3-llama-3.1-70b. Listed input and output pricing are both USD 0.7 per 1 million tokens. No tool support is listed here, so teams considering the model can ground their fit assessment in its stated text focus, generalist scope, context and output limits, open weights, and token rates.
Who it is for
It suits teams looking for an open-weight text model for generalist, conversational, reasoning, or roleplay tasks. Its stated context length may also fit work requiring long-context coherence.
What is good
- Open weights are available.
- 131,072-token context window.
- Listed for reasoning and multi-turn conversation.
- Input and output rates are both USD 0.7 per 1 million tokens.
What to know first
- Maximum output is 16,384 tokens.
- Inputs and outputs are text only.
Inferse review
Hermes 3 70B Instruct: the full review
Hermes 3 70B Instruct is positioned as a generalist text model with open weights and a substantial context window. Its listed output ceiling and text-only modality help define the fit.
Overview
Hermes 3 70B Instruct is a text-only model from Nous in the nousresearch family. Released on 2024-08-18, it has open weights and is identified on Hugging Face as NousResearch/Hermes-3-Llama-3.1-70B. Its context window is 131,072 tokens, with a maximum output of 16,384 tokens.
Nous describes Hermes 3 as a general-purpose language model and positions it as an advance over Hermes 2. The stated improvements span agentic capabilities, roleplay, reasoning, multi-turn conversation and coherence over long contexts. Those areas make it relevant to teams evaluating a model for varied text-based tasks, though the available specifications do not quantify the improvements.
For stack decisions, the clearest fit signals are its open weights, text input and output, and long context window. The listing gives token pricing but does not establish deployment requirements or provide performance benchmarks.
Key features
- Open weights: Weights are listed as open, with the Hugging Face identifier NousResearch/Hermes-3-Llama-3.1-70B.
- Long context: The 131,072-token context window is suited to workflows that need to provide substantial text context in one interaction.
- Text in, text out: Both the listed input and output modality are text.
- Output limit: A response can contain up to 16,384 tokens according to the listing.
- Generalist focus: Nous highlights reasoning, roleplay, agentic capabilities, multi-turn conversation and long-context coherence as areas of improvement.
Pricing
Listed pricing is $0.7 USD per 1M input tokens and $0.7 USD per 1M output tokens. Input and output are priced at the same rate. The supplied pricing details do not specify other charges or hosting terms.
Platforms
The model listing identifies Hermes 3 70B Instruct by the model ID nousresearch/hermes-3-llama-3.1-70b and links it to an OpenRouter page. Its open-weight status and Hugging Face ID are useful identifiers for teams assessing integration options, but the available details do not name supported SDKs, deployment platforms or hardware requirements.
Who it's for
Hermes 3 70B Instruct is worth considering for teams looking for an open-weight, text-focused generalist with a large context window and straightforward per-token pricing. Its stated focus areas may suit applications involving extended conversation, roleplay or reasoning, as well as agentic workflows. Teams should treat those as positioning rather than quantified performance claims and verify that the model's operational requirements fit their environment.
Builders comparing Nous releases can also review Nous models. Teams weighing openness, cost or context size across a wider field can browse open-weight models, cheapest language models and longest-context models.
Pros and cons
Pros
- Open weights and a supplied Hugging Face identifier provide useful starting points for technical evaluation.
- A 131,072-token context window and 16,384-token maximum output support substantial text exchanges.
- Input and output rates are equal at $0.7 USD per 1M tokens each.
- The model is presented as a generalist rather than for a single narrow text task.
Cons
- The available facts provide no benchmark results or quantified evidence for the claimed improvements.
- Deployment requirements, supported platforms and additional hosting costs are not specified.
- Text-only input and output exclude non-text modalities from the listed capabilities.
Alternatives
Within the Hermes line, Hermes 3 405B Instruct is another listed model to compare, while Hermes 4 405B offers a different release in the family. The supplied facts do not include specifications for those alternatives, so compare their listings directly rather than assuming differences in capability, context or cost.
For a broader view of recent additions, browse newest models.
Verdict
Hermes 3 70B Instruct has a clear profile: an open-weight, text-only generalist from Nous, a 131,072-token context window, a 16,384-token output ceiling and equal input and output pricing of $0.7 USD per 1M tokens. The listed focus on reasoning, roleplay, agentic work and long conversations gives teams several plausible use cases to evaluate. The missing details on benchmarks and deployment mean it is best approached as a candidate for stack-level assessment, not as a model whose relative performance is established by these specifications alone.
Details
- Lab
- Nousopenrouter.ai · 3 Oct 2026
- Context
- 131Kopenrouter.ai · 3 Oct 2026
- Input price
- $0.7 / 1Mopenrouter.ai · 3 Oct 2026
- Output price
- $0.7 / 1Mopenrouter.ai · 3 Oct 2026
- Max output
- 16,384 tokensopenrouter.ai · 3 Oct 2026
- Inputs
- textopenrouter.ai · 3 Oct 2026
- Open weights
- Yesopenrouter.ai · 3 Oct 2026
More longest-context models
See the listListed on Inferse
- Longest-Context Models in 2026466 listed
- Newest Models in 2026466 listed
- Cheapest Language Models in 2026437 listed
- Open-Weight Models in 2026178 listed
- Nous Models: Pricing per 1M Tokens and Context (2026)3 listed
Sources
- openrouter.ai/nousresearch/hermes-3-llama-3.1-70b· checked 3 Oct 2026


