Eval Metrics — Ruby source
The standard LLM-eval numbers with exact math — unbiased pass@k, precision/recall/F1 from confusion counts, exact-match and label-set micro-F1. 100% client-side.
This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.
# Eval Metrics — the standard LLM-eval metrics with exact, testable formulas.
#
# Language: Ruby (3.x, zero dependencies)
# Port of src/lib/evalMetrics.ts (the canonical TypeScript implementation).
# Tool page: https://dev.cosmolabs.org/tools/eval-metrics
#
# pass@k is the unbiased estimator from the Codex paper (Chen et al. 2021),
# identical to the HumanEval implementation's combinatorial form. Precision,
# recall, and F1 run over confusion counts and over label sets; exact-match
# rate runs over paired strings (case-sensitive).
module EvalMetrics
# Precision / recall / F1 triple (immutable Struct).
Prf1 = Struct.new(:precision, :recall, :f1, keyword_init: true)
# Confusion counts; tn is accepted but unused by P/R/F1.
ConfusionCounts = Struct.new(:tp, :fp, :fn, :tn, keyword_init: true) do
def initialize(tp:, fp:, fn:, tn: 0)
super
end
end
module_function
# Unbiased pass@k: probability that at least one of k samples drawn without
# replacement from n (of which c are correct) passes.
#
# 1 when n - c < k (a wrong draw is impossible)
# 1 - Π_{i=0..k-1} (n - c - i) / (n - i) otherwise
#
# Raises ArgumentError on impossible inputs (the TS RangeError contract).
def pass_at_k(n, c, k)
raise ArgumentError, 'n must be > 0' if n <= 0
raise ArgumentError, 'c must be in [0, n]' if c.negative? || c > n
raise ArgumentError, 'k must be in [1, n]' if k <= 0 || k > n
return 1.0 if n - c < k
product = 1.0
k.times do |i|
product *= (n - c - i).to_f / (n - i)
end
1.0 - product
end
# Precision/recall/F1 over confusion counts. Zero denominators score 0.
# Raises ArgumentError on negative counts.
def precision_recall(counts)
raise ArgumentError, 'counts must be >= 0' if counts.tp.negative? || counts.fp.negative? || counts.fn.negative?
precision = counts.tp + counts.fp > 0 ? counts.tp.to_f / (counts.tp + counts.fp) : 0.0
recall = counts.tp + counts.fn > 0 ? counts.tp.to_f / (counts.tp + counts.fn) : 0.0
f1 = precision + recall > 0.0 ? (2.0 * precision * recall) / (precision + recall) : 0.0
Prf1.new(precision: precision, recall: recall, f1: f1)
end
# Micro-averaged P/R/F1 across per-class confusion counts.
def micro_average(per_class)
sums = per_class.each_with_object(ConfusionCounts.new(tp: 0, fp: 0, fn: 0)) do |c, acc|
acc.tp += c.tp
acc.fp += c.fp
acc.fn += c.fn
end
precision_recall(sums)
end
# Exact-match rate over paired predictions/references (case-sensitive).
# Empty input scores 0; length mismatch raises ArgumentError.
def exact_match_rate(predictions, references)
if predictions.length != references.length
raise ArgumentError, 'predictions and references must have the same length'
end
return 0.0 if predictions.empty?
hits = predictions.zip(references).count { |p, r| p == r }
hits.to_f / predictions.length
end
# P/R/F1 over label SETS — the standard multi-label / extraction metric.
def set_match(prediction, reference)
p = prediction.uniq
r = reference.uniq
tp = r.count { |label| p.include?(label) }
fp = p.count { |label| !r.include?(label) }
fn = r.count { |label| !p.include?(label) }
precision_recall(ConfusionCounts.new(tp: tp, fp: fp, fn: fn))
end
end
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →