Vision Model
5 items tagged with "vision-model"
Models5
Skild S1
Skild S1 is a robot foundation model designed to teach robots previously unseen, long-horizon tasks from a single video. It was announced through NVIDIA's Physical AI ecosystem coverage.
DeepSeek V4 Flash Vision Exp
Experimental vision-capable variant of DeepSeek V4 Flash available on OpenRouter, adding multimodal image understanding to the Flash model line with a 1M-token context window.
Gemini 3.1 Flash-Lite Image
A Google Gemini 3.1 Flash-Lite image model added to OpenRouter, providing image-focused multimodal capabilities with a 65,536-token context window.
Gemini 3.1 Flash Image
Google Gemini image-focused model listed on OpenRouter with a 131,072-token context window. It is positioned as a Flash-tier multimodal/image model for lower-latency image-centric workloads.
Gemini 3 Pro Image
Google Gemini Pro-tier image-focused model listed on OpenRouter with a 65,536-token context window. It targets higher-capability multimodal and image-generation use cases than Flash-tier variants.