Skip to content

Eval Metrics — JavaScript source

The standard LLM-eval numbers with exact math — unbiased pass@k, precision/recall/F1 from confusion counts, exact-match and label-set micro-F1. 100% client-side.

This is the JavaScript implementation — the same logic the interactive tool runs, in a shareable, citable form.

/**
 * Eval Metrics — the standard LLM-eval metrics with exact, testable formulas.
 *
 * Language: JavaScript (ES2022+, ES module; runs unmodified in Node 18+
 *           and modern browsers)
 * Port of src/lib/evalMetrics.ts (the canonical TypeScript implementation).
 * Tool page: https://dev.cosmolabs.org/tools/eval-metrics
 *
 * pass@k is the unbiased estimator from the Codex paper (Chen et al. 2021),
 * identical to the HumanEval implementation's combinatorial form. Precision,
 * recall, and F1 run over confusion counts and over label sets; exact-match
 * rate runs over paired strings (case-sensitive).
 */

/**
 * Unbiased pass@k: probability that at least one of k samples drawn without
 * replacement from n (of which c are correct) passes.
 *
 *   1                                       when n - c < k  (a wrong draw is impossible)
 *   1 - Π_{i=0..k-1} (n - c - i) / (n - i)  otherwise
 *
 * Throws a RangeError on impossible inputs.
 */
export function passAtK(n, c, k) {
  if (n <= 0) throw new RangeError('n must be > 0');
  if (c < 0 || c > n) throw new RangeError('c must be in [0, n]');
  if (k <= 0 || k > n) throw new RangeError('k must be in [1, n]');
  if (n - c < k) return 1;
  let product = 1;
  for (let i = 0; i < k; i++) {
    product *= (n - c - i) / (n - i);
  }
  return 1 - product;
}

/**
 * Precision/recall/F1 over confusion counts. Zero denominators score 0.
 * Throws a RangeError on negative counts.
 */
export function precisionRecall(counts) {
  const { tp, fp, fn } = counts;
  if ([tp, fp, fn].some((v) => v < 0)) throw new RangeError('counts must be >= 0');
  const precision = tp + fp > 0 ? tp / (tp + fp) : 0;
  const recall = tp + fn > 0 ? tp / (tp + fn) : 0;
  const f1 = precision + recall > 0 ? (2 * precision * recall) / (precision + recall) : 0;
  return { precision, recall, f1 };
}

/** Micro-averaged P/R/F1 across per-class confusion counts. */
export function microAverage(perClass) {
  const sums = perClass.reduce(
    (acc, c) => ({ tp: acc.tp + c.tp, fp: acc.fp + c.fp, fn: acc.fn + c.fn }),
    { tp: 0, fp: 0, fn: 0 },
  );
  return precisionRecall(sums);
}

/**
 * Exact-match rate over paired predictions/references (case-sensitive).
 * Empty input scores 0; length mismatch throws a RangeError.
 */
export function exactMatchRate(predictions, references) {
  if (predictions.length !== references.length) {
    throw new RangeError('predictions and references must have the same length');
  }
  if (predictions.length === 0) return 0;
  let hits = 0;
  for (let i = 0; i < predictions.length; i++) {
    if (predictions[i] === references[i]) hits++;
  }
  return hits / predictions.length;
}

/** P/R/F1 over label SETS — the standard multi-label / extraction metric. */
export function setMatch(prediction, reference) {
  const p = new Set(prediction);
  const r = new Set(reference);
  let tp = 0;
  for (const label of r) if (p.has(label)) tp++;
  const fp = [...p].filter((l) => !r.has(l)).length;
  const fn = [...r].filter((l) => !p.has(l)).length;
  return precisionRecall({ tp, fp, fn });
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →