Ling 3.0 Flash
inclusionAI· released 23 Jul 2026
inclusionai/ling-3.0-flash- Input / 1M tokens
- $0.021
- Output / 1M tokens
- $0.063
Longest answer: 33K tokens
About
Ling 3.0 Flash is an inclusionAI model described as a 124B-parameter mixture-of-experts model, with about 5.1B parameters active for each token. Its design priorities are token efficiency and production-scale agentic inference. The model accepts text and returns text, with a context window of 262,144 tokens and a maximum output of 32,768 tokens. It supports tools and reasoning. Open weights are available, and its Hugging Face ID is inclusionAI/Ling-3.0-flash. Its model ID is inclusionai/ling-3.0-flash. Input costs $0.021 USD per 1M tokens and output costs $0.063 USD per 1M tokens. The maker and lab are inclusionAI. It was released on 2026-07-23. Those parameters and token rates are useful starting points for teams evaluating text-based agentic inference; the listed input modality is text only. The official listing is https://openrouter.ai/inclusionai/ling-3.0-flash.
Who it is for
It suits teams looking at text-based agentic inference, especially where token efficiency is a priority. Open weights may also matter to builders choosing a model they can access in that form.
What is good
- Open weights are available
- Tools and reasoning are supported
- Context window is 262,144 tokens
- Input costs $0.021 USD per 1M tokens
What to know first
- Accepts text input only
- Maximum output is 32,768 tokens
Verdict
Ling 3.0 Flash combines an MoE design with a stated focus on token efficiency and agentic inference. Its text-only input and 32,768-token output limit are important constraints to check against a workload.
Details
- Lab
- inclusionAIopenrouter.ai · 3 Oct 2026
- Context
- 262Kopenrouter.ai · 3 Oct 2026
- Input price
- $0.021 / 1Mopenrouter.ai · 3 Oct 2026
- Output price
- $0.063 / 1Mopenrouter.ai · 3 Oct 2026
- Max output
- 32,768 tokensopenrouter.ai · 3 Oct 2026
- Inputs
- textopenrouter.ai · 3 Oct 2026
- Open weights
- Yesopenrouter.ai · 3 Oct 2026
More newest models
See the listListed on Inferse
- Newest Models in 2026466 listed
- Longest-Context Models in 2026466 listed
- Cheapest Language Models in 2026437 listed
- Models with Tool Calling in 2026395 listed
- Reasoning Models in 2026332 listed
- Open-Weight Models in 2026178 listed
- inclusionAI Models: Pricing per 1M Tokens and Context (2026)5 listed
Sources
- openrouter.ai/inclusionai/ling-3.0-flash· checked 3 Oct 2026

