Image Token Calculator — Java source
Estimate the vision token cost of an image before sending it to an LLM - low/high/auto detail modes, the 512px tile math, the 2048/768 downscaling steps, and a full base + tiles + detail breakdown. Runs entirely in your browser.
This is the Java implementation — the same logic the interactive tool runs, in a shareable, citable form.
// Image Token Calculator — estimate the vision token cost of an image using
// OpenAI-style tile math.
//
// Language: Java (Java 17, standard library only)
// Source: CosmoDev polyglot showcase port of the Image Token Calculator
// tool, ported from src/lib/imageTokenCalculator.ts (the canonical
// TypeScript implementation).
// Live at: https://dev.cosmolabs.org/tools/image-token-calculator
// License: display source — part of CosmoDev's polyglot tool pages.
//
// Design goals:
// - Pure + deterministic; invalid input throws IllegalArgumentException
// (as the TS reference throws).
// - Functionally equivalent to the TS reference: same inputs -> same outputs.
// - Self-contained: the JDK only (no external dependencies).
//
// Rounding note: the TS reference uses Math.round (half up); shrink spells it
// as Math.floor(x + 0.5) for exact parity.
//
// The enclosing class is package-private so the file compiles standalone as
// java.java (javac requires the public class name to match the file name).
package org.cosmolabs.cosmodev.polyglot;
final class ImageTokenCalculator {
/** Fixed token cost of the low-resolution image view. */
static final int LOW_DETAIL_TOKENS = 85;
/** Token cost of one high-resolution 512 px tile. */
static final int TILE_TOKENS = 170;
/** Images are first scaled to fit inside this square. */
static final int MAX_SIDE = 2048;
/** Then the shortest side is capped at this length. */
static final int MAX_SHORT_SIDE = 768;
/** Tile edge length in pixels. */
static final int TILE_SIZE = 512;
/** Both dimensions at or under this -> AUTO stays low detail. */
static final int AUTO_LOW_MAX = 512;
/** Requested detail mode of an image (AUTO mirrors the TS default). */
enum DetailLevel { LOW, HIGH, AUTO }
/** A width x height pair returned by {@link #preprocessImage}. */
record Dimensions(int width, int height) {}
/** Mirrors the TokenBreakdown interface in the TS lib. */
record TokenBreakdown(
/** Detail level actually applied ("low" or "high"). */
String detail,
/** Dimensions after the high-detail downscaling pipeline (identity for low). */
int scaledWidth,
int scaledHeight,
/** 512 px tiles along each axis (both 1 in low detail). */
int tilesX,
int tilesY,
/** Total 512 px tiles used (tilesX * tilesY). */
int tiles,
/** Fixed base cost of the low-resolution view, in tokens. */
int base,
/** Extra tokens for the high-resolution tile views (0 in low detail). */
int detailTokens,
/** Total estimated tokens: base + detailTokens. */
int total) {}
private ImageTokenCalculator() {}
/** JS Math.round parity, floored at 1 px: half up, never zero. */
private static int shrink(int side, double scale) {
return Math.max(1, (int) Math.floor(side * scale + 0.5));
}
/** ceil(n / d) for positive integers, without floating point. */
private static int ceilDiv(int n, int d) {
return (n + d - 1) / d;
}
/**
* Scale (width, height) per the vision preprocessing pipeline:
* 1. fit inside a MAX_SIDE x MAX_SIDE square (longest side capped), then
* 2. cap the shortest side at MAX_SHORT_SIDE.
*
* Aspect ratio is preserved; each step is skipped when already satisfied.
*/
static Dimensions preprocessImage(int width, int height) {
int w = width;
int h = height;
int longest = Math.max(w, h);
if (longest > MAX_SIDE) {
double scale = (double) MAX_SIDE / longest;
w = shrink(w, scale);
h = shrink(h, scale);
}
int shortest = Math.min(w, h);
if (shortest > MAX_SHORT_SIDE) {
double scale = (double) MAX_SHORT_SIDE / shortest;
w = shrink(w, scale);
h = shrink(h, scale);
}
return new Dimensions(w, h);
}
/**
* Estimate the token cost of a width x height image at the given detail
* level.
*
* - LOW: fixed LOW_DETAIL_TOKENS, whatever the size.
* - HIGH: the image is downscaled by preprocessImage, tiled into TILE_SIZE
* squares, and each tile costs TILE_TOKENS on top of the base.
* - AUTO: low when both dimensions are <= AUTO_LOW_MAX, otherwise high.
*
* @throws IllegalArgumentException for non-positive dimensions or an
* unknown detail level
*/
static TokenBreakdown imageTokens(int width, int height, DetailLevel detail) {
if (width <= 0 || height <= 0) {
throw new IllegalArgumentException("Width and height must be greater than zero");
}
boolean resolvedHigh = switch (detail) {
case LOW -> false;
case HIGH -> true;
case AUTO -> width > AUTO_LOW_MAX || height > AUTO_LOW_MAX;
};
if (!resolvedHigh) {
return new TokenBreakdown(
"low", width, height, 1, 1, 1, LOW_DETAIL_TOKENS, 0, LOW_DETAIL_TOKENS);
}
Dimensions scaled = preprocessImage(width, height);
int tilesX = ceilDiv(scaled.width(), TILE_SIZE);
int tilesY = ceilDiv(scaled.height(), TILE_SIZE);
int tiles = tilesX * tilesY;
int detailTokens = tiles * TILE_TOKENS;
return new TokenBreakdown(
"high",
scaled.width(),
scaled.height(),
tilesX,
tilesY,
tiles,
LOW_DETAIL_TOKENS,
detailTokens,
LOW_DETAIL_TOKENS + detailTokens);
}
/**
* Parse a detail-level string ("low" | "high" | "auto") — the bridge from
* the TS string union to DetailLevel.
*
* @throws IllegalArgumentException if the string is not a known level
*/
static DetailLevel parseDetail(String detail) {
return switch (detail) {
case "low" -> DetailLevel.LOW;
case "high" -> DetailLevel.HIGH;
case "auto" -> DetailLevel.AUTO;
default -> throw new IllegalArgumentException("Unknown detail level: " + detail);
};
}
}
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →