Vision encoders and token budgets
Why this matters
Your model has never seen a single image you've sent it. Not one. What it actually gets is a compressed summary — and if you don't know that, you'll burn hours rewriting prompts to fix a bug that lives in pixels, not words.
The agent that misread a button
Last month an agent of mine was driving a UI test suite. It kept clicking 'Cancel' when the screenshot clearly showed 'Confirm'. I rewrote the prompt three times — few-shot examples, 'look carefully,' chain-of-thought. Nothing moved. Turned out reasoning had nothing to do with it: the screenshot was 1920x1080, downsampled to fit the vision encoder's token budget, and at that resolution the anti-aliased text on both buttons blurred into near-identical shapes. The model wasn't confused. It was blind. It never had the pixels to tell the difference.
Day 19: opening the box on 'multi' in multimodal
Eighteen days on text — tokens, context windows, attention, agent loops. Today we add a second sense: vision, and by extension audio. Forget 'can the model see' — the real question is what survives the trip from pixels to tokens, because that's the only thing the LLM ever reasons over.
The two-stage pipeline
A multimodal model isn't one brain processing pixels and words side by side. It's two separate stages, bolted together:
- ▹Stage 1 — the vision tower (encoder): a separate neural network, usually a ViT (Vision Transformer), that takes your image, chops it into fixed-size patches, and crushes the whole thing down to a fixed token budget — commonly a few hundred up to ~1500, depending on the model and the resolution setting.
- ▹Stage 2 — the LLM: gets those image tokens interleaved with your text tokens and reasons over them exactly the way it reasons over words. It has zero access to the original pixels — only to whatever the encoder decided was worth keeping.
- ▹The two stages are usually trained separately and glued together with a projection layer. The LLM can't 'go back and look again' at the raw image. It only ever has what made it through the bottleneck.
Where information dies
Three lossy steps happen before the LLM even starts reasoning — and none of them show up in your prompt:
- ▹Resolution downsampling — most APIs resize your image to fit a max dimension (often somewhere between ~768 and 2000px, depending on provider and mode) before encoding even starts. Fine print, small icons, dense UI elements — gone before the model gets a look.
- ▹Patch tokenization — the image gets cut into a grid of patches (say, 14x14 pixels each), and each patch turns into roughly one token. Send a dashboard screenshot crammed with 40 text labels and a photo of a cat at the same resolution, and they get the same token budget. The encoder has no idea one of them is dense with information and the other isn't.
- ▹Fixed token budget per resolution — the token count the encoder emits is tied to pixel dimensions, not to how complex the content is. There's no 'this image is dense, give it more tokens' logic anywhere in there. Some providers raise the ceiling through tiling at higher resolutions, but at a given resolution/tile setting, a text-packed screenshot and a simple photo compress into roughly the same token count — meaning the dense one loses proportionally more.
- ▹Audio works the same way: a waveform gets chunked into fixed-length windows and compressed into a bounded set of tokens (or embeddings) before the LLM ever gets to reason about 'what was said.' Accents, overlapping speech, background noise — those are lossy-encoder problems, not reasoning problems.
Diagnostic habit: ask the encoder question first
When a multimodal agent gets something visually wrong, resist the urge to rewrite the prompt first. Ask yourself: what did the encoder actually keep? Here's how to test that fast:
- ▹Zoom crop — crop down to just the region that matters and re-send only that. If accuracy jumps, you were looking at a resolution/token-budget problem, not a reasoning one.
- ▹Higher-res tiling — for dense images (screenshots, documents, diagrams), split the image into tiles and send each as its own input, or use a provider mode built for high-res/tiled encoding (the 'detail: high' style params). You're spending more of your token budget on the pixels that actually matter.
- ▹OCR fallback — for text-heavy images (receipts, dashboards, code screenshots), don't ask the vision encoder to read text at all. Run OCR separately and feed the extracted text in as text tokens. Text tokens are lossless. Image tokens of text are not.
- ▹If none of that moves the output, now you've got a genuine reasoning or prompt problem — and you've earned the right to go iterate on the prompt.
Takeaway
Knowing where the lossy boundary sits changes how you design inputs, not just how you phrase prompts. Crop before you prompt-engineer. Tile dense images instead of describing them harder. Prefer OCR-to-text over asking the model to read pixels directly. Tomorrow we follow this thread into retrieval over multimodal content — how you index and search things that were never fully 'seen' in the first place.
Extend your knowledge
- ▹Read the original Vision Transformer paper, 'An Image is Worth 16x16 Words' — patchification, straight from the source.
- ▹Check your provider's docs for image input limits — max resolution, tiling/detail modes, token cost per image. OpenAI, Anthropic, and Google each document this differently, so don't assume.
- ▹Run the test yourself: take a screenshot with small text, crop it down to just that region, and compare the model's answers on the full image versus the crop.
- ▹Building document understanding? Run a pure vision-encoder approach against an OCR-plus-text-LLM pipeline on the same dataset and watch the accuracy gap show up firsthand.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.