Skip to content

Vision Tokens: Tiles, Not Pixels Explained

Vision models do not see your screenshot — they see tiles. The image is cut into 512-pixel squares, each becoming a fixed block of tokens; a high-detail pass doubles the grid. That is why one screenshot can cost 800+ tokens while a tight crop costs a fraction.

Animated diagrams · 2 flows

Compare all 2

Sort by any column to find the right fit. Click a name to open its details.

Crop, Don't ShrinkFull screenshot vs cropped regionNoFewer tiles, faster turnsSame upload pathA crop tool awayAgents that read screens and dashboards
Tiles to TokensImage → tile grid → token blocksNoScales with tile countVision encoderOne detail settingBudgeting image inputs; choosing crops over shrinks