Popular multimodal model catalog
VisionSqueezer accepts aliases for the most-used current multimodal families. There is no universal “top 20” ranking, so this catalog follows current OpenRouter vision usage and Hugging Face image-text-to-text trends.
Exact profiles
gpt6, claude, gemini, llama, qwen, deepseek, and deepseek-local use provider or open-weight formulas documented in the individual provider pages. kimi supports native vision, but its estimate is advisory because Moonshot does not publish a stable image billing grid.
Popular aliases with advisory estimates
The following aliases are accepted and use one deterministic 28px/2048px generic profile:
glm, pixtral, mistral, gemma, internvl, minicpm, molmo, aya, phi4, granite, llava, falcon, minimax, step, ling, and voyage.
Use the full family alias when desired, for example glm-5.3-flash, pixtral-large, gemma-4, internvl3, llava-onevision, or minimax-vl. These estimates are not provider billing claims; resizing and compression remain exact.
Sources: OpenRouter vision collection, Hugging Face trending image-text-to-text models.
