Statistics Calculator — PHP source
Compute descriptive statistics - count, sum, mean, median, mode, min/max, range, variance, standard deviation, and quartiles (Q1/Q3/IQR) - from any list of numbers. Tolerates mixed separators and flags unparseable tokens. Choose sample (n−1) or population (n) variance. Everything runs 100% client-side.
This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.
<?php
/**
* statistics — polyglot showcase port (PHP).
*
* Pure descriptive-statistics logic for the Statistics Calculator tool on
* CosmoDev (dev.cosmolabs.org). Ported from the canonical TypeScript source at
* src/lib/statistics.ts so the tool page can display the same logic across six
* languages.
*
* Self-contained: standard library only, no external dependencies (no Composer
* packages). Requires PHP 8.1+ for readonly properties.
*
* License/usage: display source — part of CosmoDev's polyglot tool pages.
*/
namespace CosmoDev\Statistics;
/**
* Outcome of splitting a free-form number list into valid + invalid tokens.
*/
final class ParseResult
{
public function __construct(
/** Finite numbers, in the order they appeared. */
public readonly array $values = [],
/** Tokens that could not be parsed as finite numbers, in order. */
public readonly array $invalid = [],
) {}
}
/**
* Descriptive statistics over a sample of numbers. Numeric fields are `NAN`
* when `count === 0`; `mode` is `[]` when there is no mode.
*/
final class Stats
{
public function __construct(
public readonly int $count = 0,
public readonly float $sum = NAN,
public readonly float $mean = NAN,
public readonly float $median = NAN,
/** Most frequent value(s), ascending. Empty when uniform / no mode. */
public readonly array $mode = [],
public readonly float $min = NAN,
public readonly float $max = NAN,
public readonly float $range = NAN,
public readonly float $variance = NAN,
public readonly float $stddev = NAN,
public readonly float $q1 = NAN,
public readonly float $q3 = NAN,
public readonly float $iqr = NAN,
) {}
}
/**
* Split a free-form number list into finite values and unparseable tokens.
*
* Separators are any run of whitespace and/or commas. Tokens like `INF` and
* `NAN` are not finite, so they land in `$invalid`. Empty input yields no
* values and no invalid tokens.
*/
function parse_numbers(string $input): ParseResult
{
if (trim($input) === '') {
return new ParseResult();
}
// PREG_SPLIT_NO_EMPTY drops the empties that leading/trailing/doubled
// separators would otherwise produce — matching the TS `.filter(len > 0)`.
$tokens = preg_split('/[\s,]+/', $input, -1, PREG_SPLIT_NO_EMPTY);
$values = [];
$invalid = [];
foreach ($tokens as $tok) {
// is_numeric accepts decimals, exponents, signs, and leading dots.
// The float cast then resolves the value; is_finite rejects overflow
// (1e400 → INF) and the literal INF/NAN — mirroring JS Number().
if (is_numeric($tok)) {
$v = (float) $tok;
if (is_finite($v)) {
$values[] = $v;
continue;
}
}
$invalid[] = $tok;
}
return new ParseResult($values, $invalid);
}
/**
* Linear-interpolation quantile — the R-7 convention used by NumPy, Excel's
* PERCENTILE, and most stats textbooks. `$sorted` must be ascending and
* non-empty; `$p` is in [0, 1].
*/
function quantile(array $sorted, float $p): float
{
$n = count($sorted);
// Position along the n-1 gaps between the n sorted samples.
$h = ($n - 1) * $p;
$lower = (int) floor($h);
$upper = (int) ceil($h);
if ($lower === $upper) {
return $sorted[$lower];
}
// Interpolate between the two bracketing samples.
return $sorted[$lower] + ($h - $lower) * ($sorted[$upper] - $sorted[$lower]);
}
/**
* Most frequent value(s), ascending. Returns `[]` when there is no mode — i.e.
* when every value is distinct, or when all distinct values share the same
* frequency (a flat / uniform distribution with ≥2 distinct values). A single
* repeated value (e.g. [5,5,5]) does have a mode: [5].
*/
function compute_mode(array $values): array
{
$freq = [];
foreach ($values as $v) {
// Cast to a string key so floats hash like JS Map keys; PHP normalizes
// numeric string keys, so equal values collide on the same key.
$key = (string) $v;
$freq[$key] = ($freq[$key] ?? 0) + 1;
}
// A single distinct value is always the mode (covers [7] and [5,5,5]).
if (count($freq) === 1) {
return [(float) array_key_first($freq)];
}
$max = max($freq);
$modes = [];
foreach ($freq as $key => $count) {
if ($count === $max) {
$modes[] = (float) $key;
}
}
// All distinct values share the max frequency → uniform → no mode.
if (count($modes) === count($freq)) {
return [];
}
sort($modes, SORT_NUMERIC);
return $modes;
}
/**
* Compute descriptive statistics over `$values`. With `$sample = true`
* (default) variance/stddev use the sample estimator (divide by n−1); with
* `$sample = false` they use the population estimator (divide by n). Empty
* input returns count 0 with every numeric field `NAN` and `mode: []` — never
* throws.
*/
function summarize(array $values, bool $sample = true): Stats
{
$n = count($values);
if ($n === 0) {
return new Stats();
}
// Sort a copy (not the caller's array) so min/max/quantiles are cheap.
$sorted = $values;
sort($sorted, SORT_NUMERIC);
// Sum is accumulated over the ORIGINAL order to match the TS reference and
// keep floating-point summation deterministic across languages.
$sum = (float) array_sum($values);
$mean = $sum / $n;
$min = $sorted[0];
$max = $sorted[$n - 1];
$median = quantile($sorted, 0.5);
$q1 = quantile($sorted, 0.25);
$q3 = quantile($sorted, 0.75);
// Sum of squared deviations from the mean (original order).
$ss = 0.0;
foreach ($values as $x) {
$d = $x - $mean;
$ss += $d * $d;
}
// Sample variance is undefined for n < 2; population variance is always ss/n.
if ($sample) {
$variance = ($n >= 2) ? $ss / ($n - 1) : NAN;
} else {
$variance = $ss / $n;
}
$stddev = is_finite($variance) ? sqrt($variance) : NAN;
return new Stats(
count: $n,
sum: $sum,
mean: $mean,
median: $median,
mode: compute_mode($values),
min: $min,
max: $max,
range: $max - $min,
variance: $variance,
stddev: $stddev,
q1: $q1,
q3: $q3,
iqr: $q3 - $q1,
);
}
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →