Skip to content

Eval Metrics — Ruby source

The standard LLM-eval numbers with exact math — unbiased pass@k, precision/recall/F1 from confusion counts, exact-match and label-set micro-F1. 100% client-side.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# Eval Metrics — the standard LLM-eval metrics with exact, testable formulas.
#
# Language: Ruby (3.x, zero dependencies)
# Port of src/lib/evalMetrics.ts (the canonical TypeScript implementation).
# Tool page: https://dev.cosmolabs.org/tools/eval-metrics
#
# pass@k is the unbiased estimator from the Codex paper (Chen et al. 2021),
# identical to the HumanEval implementation's combinatorial form. Precision,
# recall, and F1 run over confusion counts and over label sets; exact-match
# rate runs over paired strings (case-sensitive).

module EvalMetrics
  # Precision / recall / F1 triple (immutable Struct).
  Prf1 = Struct.new(:precision, :recall, :f1, keyword_init: true)

  # Confusion counts; tn is accepted but unused by P/R/F1.
  ConfusionCounts = Struct.new(:tp, :fp, :fn, :tn, keyword_init: true) do
    def initialize(tp:, fp:, fn:, tn: 0)
      super
    end
  end

  module_function

  # Unbiased pass@k: probability that at least one of k samples drawn without
  # replacement from n (of which c are correct) passes.
  #
  #   1                                       when n - c < k  (a wrong draw is impossible)
  #   1 - Π_{i=0..k-1} (n - c - i) / (n - i)  otherwise
  #
  # Raises ArgumentError on impossible inputs (the TS RangeError contract).
  def pass_at_k(n, c, k)
    raise ArgumentError, 'n must be > 0' if n <= 0
    raise ArgumentError, 'c must be in [0, n]' if c.negative? || c > n
    raise ArgumentError, 'k must be in [1, n]' if k <= 0 || k > n
    return 1.0 if n - c < k

    product = 1.0
    k.times do |i|
      product *= (n - c - i).to_f / (n - i)
    end
    1.0 - product
  end

  # Precision/recall/F1 over confusion counts. Zero denominators score 0.
  # Raises ArgumentError on negative counts.
  def precision_recall(counts)
    raise ArgumentError, 'counts must be >= 0' if counts.tp.negative? || counts.fp.negative? || counts.fn.negative?

    precision = counts.tp + counts.fp > 0 ? counts.tp.to_f / (counts.tp + counts.fp) : 0.0
    recall = counts.tp + counts.fn > 0 ? counts.tp.to_f / (counts.tp + counts.fn) : 0.0
    f1 = precision + recall > 0.0 ? (2.0 * precision * recall) / (precision + recall) : 0.0
    Prf1.new(precision: precision, recall: recall, f1: f1)
  end

  # Micro-averaged P/R/F1 across per-class confusion counts.
  def micro_average(per_class)
    sums = per_class.each_with_object(ConfusionCounts.new(tp: 0, fp: 0, fn: 0)) do |c, acc|
      acc.tp += c.tp
      acc.fp += c.fp
      acc.fn += c.fn
    end
    precision_recall(sums)
  end

  # Exact-match rate over paired predictions/references (case-sensitive).
  # Empty input scores 0; length mismatch raises ArgumentError.
  def exact_match_rate(predictions, references)
    if predictions.length != references.length
      raise ArgumentError, 'predictions and references must have the same length'
    end
    return 0.0 if predictions.empty?

    hits = predictions.zip(references).count { |p, r| p == r }
    hits.to_f / predictions.length
  end

  # P/R/F1 over label SETS — the standard multi-label / extraction metric.
  def set_match(prediction, reference)
    p = prediction.uniq
    r = reference.uniq
    tp = r.count { |label| p.include?(label) }
    fp = p.count { |label| !r.include?(label) }
    fn = r.count { |label| !p.include?(label) }
    precision_recall(ConfusionCounts.new(tp: tp, fp: fp, fn: fn))
  end
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →