Vision Language Models in 2026

295 on record. Sorted by context; every figure comes from the OpenRouter model catalogue, with the page it was read from on each entry.

295listed
#ModelContext windowInput / 1MOutput / 1M
251Sonar ProPerplexitytext, image200K$31$151
252Sonar Pro SearchPerplexitytext, image200K$31$151
253Command A+Coheretext, image192K$0.31$1.51
254Llama Guard 4 12BMetaopen weightsimage, text164K$0.181$0.181
255Gemma 3 12BGoogleopen weightstext, image131K$0.051$0.151
256Gemma 3 27BGoogleopen weightstext, image131K$0.081$0.451
257Gemma 3 4BGoogleopen weightstext, image131K$0.051$0.11
258GLM 4.6VZ.aiopen weightsimage, text, video131K$0.31$0.91
259Ministral 3 3B 2512Mistralopen weightstext, image131K$0.11$0.11
260Mistral Medium 3Mistraltext, image, file131K$0.41$21
261Mistral Medium 3.1Mistraltext, image, file131K$0.41$21
262Mistral Medium 3.1 (batch)Mistraltext, image, file131K$0.21$11
263Muse Glimmer 30BMetaopen weightstext, image131K$0.351$1.51
264Nano Banana 2 (Gemini 3.1 Flash Image)Googleimage, text131K$0.51$31
265Nano Banana Pro (Gemini 3 Pro Image)Googleimage, text131K$21$121
266Nemotron 3.5 Content SafetyNVIDIAopen weightstext, image131K$0.21$0.21
267Qwen3 VL 235B A22B ThinkingQwenopen weightstext, image131K$0.41$41
268Qwen3 VL 32B InstructQwenopen weightstext, image131K$0.104$0.416
269Qwen3 VL 8B ThinkingQwenopen weightsimage, text131K$0.181$2.11
270GPT-4oOpenAItext, image, file128K$2.51$101
271GPT-4o (2024-05-13)OpenAItext, image, file128K$51$151
272GPT-4o (2024-08-06)OpenAItext, image, file128K$2.51$101
273GPT-4o (2024-11-20)OpenAItext, image, file128K$2.51$101
274GPT-4o (batch)OpenAItext, image, file128K$1.251$51
275GPT-4o-miniOpenAItext, image, file128K$0.151$0.61

Context bars share one log scale, 1K to 2M tokens. Prices are what each provider publishes per million tokens; 'varies' is a router that bills the model it picks.

More in Models

All models lists