Skip to content

Eval Metrics

The standard LLM-eval numbers with exact math — unbiased pass@k, precision/recall/F1 from confusion counts, exact-match and label-set micro-F1. 100% client-side.

AI & Agents
cosmodev ~/tools/eval-metrics-

Unbiased estimator: draw k of n samples, c of which pass.