VisionSqueezer
Guides

Token Budget for Vision LLM Images

Cut vision token cost with --max-tokens. Downscale images until the target model's estimate fits, with measured Claude, GPT-6, and Gemini savings.

Providers bill images by pixel dimensions, not file size. A 12 MP photo and its 1 MB JPEG cost the same tokens unless the pixels change. The token budget is the knob that changes them: --max-tokens N downscales the image until the target model's token estimate fits inside N.

Defaults

SurfaceDefaultDisable
MCP server (optimize_image, optimize_image_batch)1600max_tokens: 0
CLIno capomit --max-tokens
Rust crate (ProcessConfig::max_tokens)no capNone or 0

The Python bindings do not expose max_tokens yet.

The budget is measured with --model (target_model in MCP). With no target it uses Claude's 28px patch estimate.

vision-squeezer screenshot.png --max-tokens 1600
vision-squeezer photo.jpg --max-tokens 1000 --model gpt6

What it saves

Measured with vision-squeezer <image> --dry-run --json. Token counts are the tool's estimates from each provider's documented rules, not live API invoices.

ImageRunClaude 4.7+ tokensGPT-6 tokensGemini tokensFile size
2400×1670 photono budget4,674 → 4,070 (−13%)2,903 → 2,9423,096 → 1,548 (−50%)−29%
2400×1670 photo--max-tokens 16004,674 → 1,584 (−66%)2,903 → 1,462 (−50%)3,096 → 1,032 (−67%)−67%
2400×1670 photo--max-tokens 10004,674 → 988 (−79%)2,903 → 939 (−68%)3,096 → 516 (−83%)−79%
4096×3072 photono budget4,661 → 4,6982,942 → 2,924 (−1%)6,192 → 5,160 (−17%)−40%
4096×3072 photo--max-tokens 16004,661 → 1,564 (−66%)2,942 → 1,476 (−50%)6,192 → 1,032 (−83%)−88%
4096×3072 photo--max-tokens 10004,661 → 999 (−79%)2,942 → 951 (−68%)6,192 → 516 (−92%)−92%

Without a budget, tokens barely move: providers already downscale oversized images themselves, so a squeezed file is smaller but the bill is the same. File size is a side effect, not the goal.

Choosing a budget

  • 1600 (default). Keeps composition, colour, and large text. In testing, a 2880×1800 code-and-error screenshot stayed fully legible at 1600, including red error highlights.
  • 1000. Body text stayed readable in the same screenshot. Distant signs and small print start to blur.
  • Around 600. Small text breaks up. Use it only for images where layout matters more than text.
  • Fine detail goes first. Distant windows, crowds, and thin lines are lost before large shapes and colours.

Things to know

  • The budget follows the target model's own grid. --model gemini --max-tokens 1600 lands on Gemini's 768px tile grid (1,548 tokens), which is large for Claude. Match the budget model to the model that will read the image.
  • Colour is preserved. auto mode behaves like standard. Black-and-white output only happens with an explicit mode: ocr.
  • Aspect ratio is preserved. The image is scaled uniformly and the few leftover pixels are cropped evenly from the edges. Coarse tile grids (for example Llama's 560px) can still stretch.
  • --max-tiles is a different unit. It counts the target model's own tiles or patches. Prefer --max-tokens.