Unbiased estimator: draw k of n samples, c of which pass.
What it does
The Eval Metrics tool computes the numbers every LLM evaluation reports — with the exact formulas, not approximations. pass@k uses the unbiased estimator from the Codex paper (identical to HumanEval’s implementation); precision / recall / F1 come from your confusion counts; Matches scores prediction⇥reference pairs with exact-match rate and label-set micro-F1. Runs 100% client-side.
How to use it
- Pick a mode: pass@k, P/R/F1, or Matches.
- pass@k: enter
n(samples generated),c(how many passed),k(draws per attempt). - P/R/F1: enter TP/FP/FN from your grader.
- Matches: one
prediction<TAB>referencepair per line; comma-separated values on either side are scored as label sets.
Examples
20 samples, 4 correct, k=5:
pass@5 = 1 − (16·15·14·13·12)/(20·19·18·17·16) ≈ 0.6317
A grader with TP=30, FP=10, FN=20:
precision 75.0% · recall 60.0% · F1 66.7%
Good to know
- The naive pass@k estimate (average of 1 if any of k pre-drawn samples passed) is biased on small n; the combinatorial form used here is not.
- When fewer wrong samples exist than k, pass@k is exactly 1 — the tool says so instead of making you wonder.
- Zero denominators (no predictions, no positives) score 0 rather than NaN.
- Runs 100% client-side. Your eval data never leaves the browser.
- Related: Token Estimator, LLM Cost Calculator, Model Picker.