Token Estimator — Ruby source
Estimate LLM token counts for any text or code - per-content-type heuristics (prose, code, JSON, CJK) with a ±15% range, plus chat-framing overhead. Runs entirely in your browser.
This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.
# token-estimator — Ruby port: tokenizer-free LLM token estimation.
#
# Display snippet: ports the line classifier and estimator core from the
# TypeScript lib (src/lib/tokenEstimator.ts). Each non-empty line is
# classified (prose / code / json / cjk) and divided by that type's
# chars-per-token rate; the estimate carries a ±15% band. String#length
# counts characters here (the TS reference counts UTF-16 code units); the
# full result shape (chars/words/lines/framing) lives in TS/Go, as does the
# whole-text JSON gate (parses as JSON => json throughout).
module TokenEstimator
# Average characters per token by content type (CHARS_PER_TOKEN in TS).
CHARS_PER_TOKEN = { prose: 4.0, code: 3.5, json: 3.0, cjk: 1.5 }.freeze
ESTIMATE_TOLERANCE = 0.15
CODE_SYMBOLS = '{}();=<>[]#'.chars.freeze
CONTENT_TYPES = %w[prose code json cjk].freeze
# CJK ideographs (U+4E00..U+9FFF), kana (U+3040..U+30FF), Hangul (U+AC00..U+D7AF).
CJK_RE = /[\u{4E00}-\u{9FFF}\u{3040}-\u{30FF}\u{AC00}-\u{D7AF}]/
module_function
# Classify a line by its shape. Order: json, cjk, code, prose.
def detect_line_type(line)
t = line.strip
if '{[}"'.include?(t[0]) && (line.include?(':') || line.include?(','))
return :json
end
return :cjk if CJK_RE.match?(line)
density = line.count(CODE_SYMBOLS.join).to_f / [line.length, 1].max
return :code if density > 0.08 || %w[; { }].include?(t[-1])
:prose
end
# Sum per-line estimates for every non-empty line of text.
def estimate_tokens(text)
breakdown = { prose: 0.0, code: 0.0, json: 0.0, cjk: 0.0 }
tokens = 0.0
text.split(/\r?\n/).each do |line|
next if line.strip.empty?
type = detect_line_type(line)
lt = (line.length / CHARS_PER_TOKEN[type]).round
lt = 1.0 if lt < 1.0 # max(1, round(len / rate))
tokens += lt
breakdown[type] += lt
end
# Dominant type: strictly-greater scan keeps ties on :prose, as in TS.
dominant = CONTENT_TYPES.reduce(:prose) do |best, type|
breakdown[type.to_sym] > breakdown[best.to_sym] ? type.to_sym : best
end
{
tokens: tokens,
low: (tokens * (1.0 - ESTIMATE_TOLERANCE)).round,
high: (tokens * (1.0 + ESTIMATE_TOLERANCE)).round,
dominant: dominant,
breakdown: breakdown,
}
end
end
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →