Skip to content

Image Token Calculator — Ruby source

Estimate the vision token cost of an image before sending it to an LLM - low/high/auto detail modes, the 512px tile math, the 2048/768 downscaling steps, and a full base + tiles + detail breakdown. Runs entirely in your browser.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# frozen_string_literal: true

# Image Token Calculator — estimate the vision token cost of an image using
# OpenAI-style tile math.
#
# Language: Ruby (3.2+, standard library only)
# Source:   CosmoDev polyglot showcase port of the Image Token Calculator
#           tool, ported from src/lib/imageTokenCalculator.ts (the canonical
#           TypeScript implementation).
# Live at:  https://dev.cosmolabs.org/tools/image-token-calculator
# License:  display source — part of CosmoDev's polyglot tool pages.
#
# Design goals:
#   - Pure + deterministic; raises ArgumentError on invalid input (as the TS
#     reference throws).
#   - Functionally equivalent to the TS reference: same inputs -> same outputs.
#   - Self-contained: core + stdlib only (no gems).
#
# Rounding note: the TS reference uses Math.round (half up). This port spells
# it explicitly as (x + 0.5).floor for exact parity with the other languages.

module ImageTokenCalculator
  # Fixed token cost of the low-resolution image view.
  LOW_DETAIL_TOKENS = 85
  # Token cost of one high-resolution 512 px tile.
  TILE_TOKENS = 170
  # Images are first scaled to fit inside this square.
  MAX_SIDE = 2048
  # Then the shortest side is capped at this length.
  MAX_SHORT_SIDE = 768
  # Tile edge length in pixels.
  TILE_SIZE = 512
  # Both dimensions at or under this -> 'auto' stays low detail.
  AUTO_LOW_MAX = 512

  # Mirrors the TokenBreakdown interface in the TS lib (Ruby 3.2 Data class).
  TokenBreakdown = Data.define(
    :detail,
    :scaled_width,
    :scaled_height,
    :tiles_x,
    :tiles_y,
    :tiles,
    :base,
    :detail_tokens,
    :total
  )

  class << self
    # JS Math.round parity, floored at 1 px: half up, never zero.
    private

    def shrink(side, scale)
      [(side * scale + 0.5).floor, 1].max
    end

    public

    # Scale (width, height) per the vision preprocessing pipeline:
    #
    # 1. fit inside a MAX_SIDE x MAX_SIDE square (longest side capped), then
    # 2. cap the shortest side at MAX_SHORT_SIDE.
    #
    # Aspect ratio is preserved; each step is skipped when already satisfied.
    def preprocess_image(width, height)
      w = width
      h = height
      longest = [w, h].max
      if longest > MAX_SIDE
        scale = MAX_SIDE.to_f / longest
        w = shrink(w, scale)
        h = shrink(h, scale)
      end
      shortest = [w, h].min
      if shortest > MAX_SHORT_SIDE
        scale = MAX_SHORT_SIDE.to_f / shortest
        w = shrink(w, scale)
        h = shrink(h, scale)
      end
      [w, h]
    end

    # Estimate the token cost of a width x height image at a detail level.
    #
    # - 'low':  fixed LOW_DETAIL_TOKENS, whatever the size.
    # - 'high': the image is downscaled by preprocess_image, tiled into
    #   TILE_SIZE squares, and each tile costs TILE_TOKENS on top of the base.
    # - 'auto': low when both dimensions are <= AUTO_LOW_MAX, else high.
    #
    # Raises ArgumentError for non-positive/non-integer dimensions or an
    # unknown detail level.
    def image_tokens(width, height, detail = 'auto')
      [width, height].each do |value|
        raise ArgumentError, 'Width and height must be whole pixels' unless value.is_a?(Integer)

        raise ArgumentError, 'Width and height must be greater than zero' if value <= 0
      end

      resolved =
        case detail
        when 'low', 'high' then detail
        when 'auto'
          width <= AUTO_LOW_MAX && height <= AUTO_LOW_MAX ? 'low' : 'high'
        else
          raise ArgumentError, "Unknown detail level: #{detail.inspect}"
        end

      if resolved == 'low'
        return TokenBreakdown.new(
          detail: 'low',
          scaled_width: width,
          scaled_height: height,
          tiles_x: 1,
          tiles_y: 1,
          tiles: 1,
          base: LOW_DETAIL_TOKENS,
          detail_tokens: 0,
          total: LOW_DETAIL_TOKENS
        )
      end

      scaled_width, scaled_height = preprocess_image(width, height)
      tiles_x = (scaled_width.to_f / TILE_SIZE).ceil
      tiles_y = (scaled_height.to_f / TILE_SIZE).ceil
      tiles = tiles_x * tiles_y
      detail_tokens = tiles * TILE_TOKENS
      TokenBreakdown.new(
        detail: 'high',
        scaled_width: scaled_width,
        scaled_height: scaled_height,
        tiles_x: tiles_x,
        tiles_y: tiles_y,
        tiles: tiles,
        base: LOW_DETAIL_TOKENS,
        detail_tokens: detail_tokens,
        total: LOW_DETAIL_TOKENS + detail_tokens
      )
    end
  end
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →