1024 × 1024 px · high detail · 4 tiles των 512 px · 85 base + 680 detail = 765 tokens
(Τεκμηρίωση στα αγγλικά)
What it does
The Image Token Calculator estimates how many tokens an image will cost before you attach it to a vision-model request. Vision APIs don’t charge images by file size — they downscale the image, cut it into 512 px tiles, and bill a fixed number of tokens per tile. This tool runs that exact pipeline: low detail is a flat 85 tokens whatever the size; high detail first fits the image inside a 2048 × 2048 square, then caps the shortest side at 768 px, then charges 170 tokens per 512 px tile plus an 85-token base. You get the tile grid, the scaled dimensions, and the base + detail breakdown, so you can see why an image costs what it costs. Everything runs 100% client-side: nothing is uploaded anywhere.
How to use it
- Enter the image’s pixel Width (px) and Height (px), or tap an example chip (1024 × 1024, 1920 × 1080, 800 × 600, 4096 × 4096).
- Pick a Detail mode: Low (flat 85 tokens), High (full tile math), or Auto (the model’s own heuristic — low detail while both sides are ≤ 512 px, otherwise high detail).
- Read the tokens headline, then the breakdown: tiles used (e.g. 2 × 2), base tokens, detail tokens, and the scaled pixel dimensions the tiles were counted on.
- Copy share link encodes the width, height, and detail mode in the URL (
?w=...&h=...&d=...), so a link reproduces the exact calculation.
Examples
A square screenshot at high detail
Input: 1024 × 1024, high → 765 tokens (scaled to 768 × 768, 4 tiles: 85 + 4 × 170)
A full-HD wallpaper
Input: 1920 × 1080, auto → 1105 tokens (shortest side capped at 768 → 1365 × 768, 3 × 2 = 6 tiles: 85 + 6 × 170)
A tall 4K screenshot — the classic 1105 case
Input: 2048 × 4096, high → 1105 tokens (fit inside 2048 square → 1024 × 2048, then shortest side 1024 → 768 × 1536, 6 tiles)
Same image, low detail
Input: 4096 × 8192, low → 85 tokens (low detail ignores the size entirely)
Good to know
- Low detail is always 85 tokens. If the model only needs the gist of an image (is there a dog in this photo?),
lowskips the tile math completely — a 4096 × 8192 scan and a 16 × 16 icon cost the same. - Big images collapse to the same bill. After the two downscaling steps, any square image ≥ 768 px becomes 768 × 768 → 4 tiles → 765 tokens. A 1024 × 1024 and a 4096 × 4096 cost the same in high detail.
- Auto is an estimate of the model’s choice. Real
detail: "auto"decisions are made server-side; this tool uses the conventional heuristic (low when both sides ≤ 512 px). Billing follows the detail level the model actually used. - Model coverage: the 85-base + 170-per-tile formula covers tile-based vision models (GPT-4o, GPT-4.1, GPT-4V, o-series except o4-mini). Newer patch-based models (o4-mini, gpt-4.1-mini/nano snapshots, gpt-5.x minis) tokenize images as 32 px patches instead — numbers will differ there. Claude vision models use a different (w×h)/750 estimate.
- Shareable: the dimensions and detail mode travel in the URL, so a link reproduces the exact estimate.
- Private: runs 100% client-side — no uploads, no requests.
- Related tools: Token Estimator, Context Window Planner, Secure Token Generator.