Skip to content

Token Estimator — Ruby source

Estimate LLM token counts for any text or code - per-content-type heuristics (prose, code, JSON, CJK) with a ±15% range, plus chat-framing overhead. Runs entirely in your browser.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# token-estimator — Ruby port: tokenizer-free LLM token estimation.
#
# Display snippet: ports the line classifier and estimator core from the
# TypeScript lib (src/lib/tokenEstimator.ts). Each non-empty line is
# classified (prose / code / json / cjk) and divided by that type's
# chars-per-token rate; the estimate carries a ±15% band. String#length
# counts characters here (the TS reference counts UTF-16 code units); the
# full result shape (chars/words/lines/framing) lives in TS/Go, as does the
# whole-text JSON gate (parses as JSON => json throughout).

module TokenEstimator
  # Average characters per token by content type (CHARS_PER_TOKEN in TS).
  CHARS_PER_TOKEN = { prose: 4.0, code: 3.5, json: 3.0, cjk: 1.5 }.freeze
  ESTIMATE_TOLERANCE = 0.15
  CODE_SYMBOLS = '{}();=<>[]#'.chars.freeze
  CONTENT_TYPES = %w[prose code json cjk].freeze

  # CJK ideographs (U+4E00..U+9FFF), kana (U+3040..U+30FF), Hangul (U+AC00..U+D7AF).
  CJK_RE = /[\u{4E00}-\u{9FFF}\u{3040}-\u{30FF}\u{AC00}-\u{D7AF}]/

  module_function

  # Classify a line by its shape. Order: json, cjk, code, prose.
  def detect_line_type(line)
    t = line.strip
    if '{[}"'.include?(t[0]) && (line.include?(':') || line.include?(','))
      return :json
    end
    return :cjk if CJK_RE.match?(line)
    density = line.count(CODE_SYMBOLS.join).to_f / [line.length, 1].max
    return :code if density > 0.08 || %w[; { }].include?(t[-1])
    :prose
  end

  # Sum per-line estimates for every non-empty line of text.
  def estimate_tokens(text)
    breakdown = { prose: 0.0, code: 0.0, json: 0.0, cjk: 0.0 }
    tokens = 0.0
    text.split(/\r?\n/).each do |line|
      next if line.strip.empty?
      type = detect_line_type(line)
      lt = (line.length / CHARS_PER_TOKEN[type]).round
      lt = 1.0 if lt < 1.0 # max(1, round(len / rate))
      tokens += lt
      breakdown[type] += lt
    end
    # Dominant type: strictly-greater scan keeps ties on :prose, as in TS.
    dominant = CONTENT_TYPES.reduce(:prose) do |best, type|
      breakdown[type.to_sym] > breakdown[best.to_sym] ? type.to_sym : best
    end
    {
      tokens: tokens,
      low: (tokens * (1.0 - ESTIMATE_TOLERANCE)).round,
      high: (tokens * (1.0 + ESTIMATE_TOLERANCE)).round,
      dominant: dominant,
      breakdown: breakdown,
    }
  end
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →