VisionSqueezer
Providers

Popular multimodal model catalog

Every model alias accepted by --model, from Claude and GPT-6 to Pixtral, Gemma, InternVL, Phi-4, and other popular vision models.

VisionSqueezer accepts aliases for the most-used current multimodal families. There is no universal “top 20” ranking, so this catalog follows current OpenRouter vision usage and Hugging Face image-text-to-text trends.

Exact profiles

gpt6, claude, gemini, llama, qwen, deepseek, and deepseek-local use provider or open-weight formulas documented in the individual provider pages. kimi supports native vision, but its estimate is advisory because Moonshot does not publish a stable image billing grid.

The following aliases are accepted and use one deterministic 28px/2048px generic profile:

glm, pixtral, mistral, gemma, internvl, minicpm, molmo, aya, phi4, granite, llava, falcon, minimax, step, ling, and voyage.

Use the full family alias when desired, for example glm-5.3-flash, pixtral-large, gemma-4, internvl3, llava-onevision, or minimax-vl. These estimates are not provider billing claims; resizing and compression remain exact.

Sources: OpenRouter vision collection, Hugging Face trending image-text-to-text models.