Token Budget for Vision LLM Images
Providers bill images by pixel dimensions, not file size. A 12 MP photo and its 1 MB JPEG cost the same tokens unless the pixels change. The token budget is the knob that changes them: --max-tokens N downscales the image until the target model's token estimate fits inside N.
Defaults
| Surface | Default | Disable |
|---|---|---|
MCP server (optimize_image, optimize_image_batch) | 1600 | max_tokens: 0 |
| CLI | no cap | omit --max-tokens |
Rust crate (ProcessConfig::max_tokens) | no cap | None or 0 |
The Python bindings do not expose max_tokens yet.
The budget is measured with --model (target_model in MCP). With no target it uses Claude's 28px patch estimate.
vision-squeezer screenshot.png --max-tokens 1600
vision-squeezer photo.jpg --max-tokens 1000 --model gpt6
What it saves
Measured with vision-squeezer <image> --dry-run --json. Token counts are the tool's estimates from each provider's documented rules, not live API invoices.
| Image | Run | Claude 4.7+ tokens | GPT-6 tokens | Gemini tokens | File size |
|---|---|---|---|---|---|
| 2400×1670 photo | no budget | 4,674 → 4,070 (−13%) | 2,903 → 2,942 | 3,096 → 1,548 (−50%) | −29% |
| 2400×1670 photo | --max-tokens 1600 | 4,674 → 1,584 (−66%) | 2,903 → 1,462 (−50%) | 3,096 → 1,032 (−67%) | −67% |
| 2400×1670 photo | --max-tokens 1000 | 4,674 → 988 (−79%) | 2,903 → 939 (−68%) | 3,096 → 516 (−83%) | −79% |
| 4096×3072 photo | no budget | 4,661 → 4,698 | 2,942 → 2,924 (−1%) | 6,192 → 5,160 (−17%) | −40% |
| 4096×3072 photo | --max-tokens 1600 | 4,661 → 1,564 (−66%) | 2,942 → 1,476 (−50%) | 6,192 → 1,032 (−83%) | −88% |
| 4096×3072 photo | --max-tokens 1000 | 4,661 → 999 (−79%) | 2,942 → 951 (−68%) | 6,192 → 516 (−92%) | −92% |
Without a budget, tokens barely move: providers already downscale oversized images themselves, so a squeezed file is smaller but the bill is the same. File size is a side effect, not the goal.
Choosing a budget
- 1600 (default). Keeps composition, colour, and large text. In testing, a 2880×1800 code-and-error screenshot stayed fully legible at 1600, including red error highlights.
- 1000. Body text stayed readable in the same screenshot. Distant signs and small print start to blur.
- Around 600. Small text breaks up. Use it only for images where layout matters more than text.
- Fine detail goes first. Distant windows, crowds, and thin lines are lost before large shapes and colours.
Things to know
- The budget follows the target model's own grid.
--model gemini --max-tokens 1600lands on Gemini's 768px tile grid (1,548 tokens), which is large for Claude. Match the budget model to the model that will read the image. - Colour is preserved.
automode behaves likestandard. Black-and-white output only happens with an explicitmode: ocr. - Aspect ratio is preserved. The image is scaled uniformly and the few leftover pixels are cropped evenly from the edges. Coarse tile grids (for example Llama's 560px) can still stretch.
--max-tilesis a different unit. It counts the target model's own tiles or patches. Prefer--max-tokens.
