Ling 3.0 Flash VL

inclusionAI · released 10 Sept 2026

Input$0.021per 1M tokens
Output$0.062per 1M tokens
Context262Ktokens
WeightsOpen
context on a 1K–2M scale

About

Ling 3.0 Flash VL is a text-output model from inclusionAI that accepts text, images and video. It builds on Ling 3.0 Flash, described as a 124B-total, 5.5B-active mixture-of-experts model, and strengthens language capabilities while adding native visual perception. The supplied summary also refers to advanced visual capabilities but does not specify them further. Vision, tool use and reasoning are listed. Its context window is 262,144 tokens, with a maximum output of 32,768 tokens. Input costs 0.021 USD per 1M tokens and output costs 0.0616 USD per 1M tokens. It has open weights; the Hugging Face ID is inclusionAI/Ling-3.0-flash-VL. The model was released on 2026-09-10.

Who it is for

It suits teams handling text, images or video that want a model with visual perception, tool use and reasoning. Open weights may also matter to teams evaluating model access options.

What is good

  • Accepts text, image and video inputs
  • Open weights listed
  • 262,144-token context window
  • Tool use and reasoning listed

What to know first

  • Maximum output is 32,768 tokens
  • Context window is smaller than 1,048,576 tokens

Verdict

Ling 3.0 Flash VL pairs text, image and video input with visual perception and open weights. Its 262,144-token context and 32,768-token output limit define the scale of prompts and responses it can handle.

Details

Lab
inclusionAIopenrouter.ai · 3 Oct 2026
Context
262Kopenrouter.ai · 3 Oct 2026
Input price
$0.021 / 1Mopenrouter.ai · 3 Oct 2026
Output price
$0.062 / 1Mopenrouter.ai · 3 Oct 2026
Max output
32,768 tokensopenrouter.ai · 3 Oct 2026
Inputs
text, image, videoopenrouter.ai · 3 Oct 2026
Open weights
Yesopenrouter.ai · 3 Oct 2026

More newest models

See the list