Base32 / Base58 / Base62 / Base85 Encoder — PHP source
Encode text to Base32, Base58, Base62, or Ascii85 - or decode it back. UTF-8 safe, runs entirely in your browser, with a shareable link to your exact input.
This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.
<?php
/**
* base-encoder — Base32 (RFC 4648), Base58 (Bitcoin), Base62, and Base85
* (Ascii85) byte-array encoders, operating on the UTF-8 bytes of the input.
*
* Language: PHP (8.1+, standard library only — 64-bit build assumed for the
* Base85 32-bit group value; Base58/62 use a manual base-256 big-int
* that fits even on a 32-bit build)
* Source: CosmoDev polyglot showcase port of the Base Encoder tool, ported
* from cli/base-encoder/base-encoder.go (the authoritative Go twin).
* License: display source — part of CosmoDev's polyglot tool pages.
*
* Design goals:
* - Pure + deterministic; never throws (decode returns null for invalid or
* malformed input, mirroring the TS lib's null and the Go twin's errInvalid).
* - Functionally equivalent to the Go twin: same inputs -> same outputs.
* - Self-contained: stdlib only (no Composer packages, no bcmath/gmp).
*
* Arbitrary-precision note: Base58 and Base62 base-convert the whole byte
* array, which overflows PHP_INT_MAX for inputs longer than a few bytes. The
* Go twin leans on math/big and the JS/Python ports on native BigInt; PHP has
* no big-integer type without an extension, so we implement the same idea with
* a little-endian base-256 byte string and two primitives — divmod_small (peel
* a base-N digit off the little end) and muladd_small (reassemble a number from
* its base-N digits). Intermediate products stay under 2^17, so this is safe
* on any PHP build.
*
* String model: PHP strings are byte arrays, so an input string already IS its
* UTF-8 byte sequence and a decoded byte string is already the text result —
* no transcoding is needed (mirrors Go's string([]byte), which never fails).
*/
declare(strict_types=1);
/**
* One of the four supported byte-array base encodings. Mirrors the Go twin's
* `Scheme` type and the TS `Scheme` union.
*/
const BASE_ENCODER_BASE32 = 'base32';
const BASE_ENCODER_BASE58 = 'base58';
const BASE_ENCODER_BASE62 = 'base62';
const BASE_ENCODER_BASE85 = 'base85';
const BASE_ENCODER_B32_ALPHABET = 'ABCDEFGHIJKLMNOPQRSTUVWXYZ234567';
const BASE_ENCODER_B58_ALPHABET = '123456789ABCDEFGHJKLMNPQRSTUVWXYZabcdefghijkmnopqrstuvwxyz';
const BASE_ENCODER_B62_ALPHABET = '0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz';
// Data characters emitted by a final (partial) 5-byte chunk before '=' padding,
// per RFC 4648. Index = byte count (0..4). Matches the TS `outLen` table.
const BASE_ENCODER_OUT_LEN_32 = [0, 2, 4, 5, 7];
// ---------------------------------------------------------------------------
// Arbitrary-precision primitives (base-256, little-endian). Used by Base58 and
// Base62 so the port stays extension-free. The "number" is a PHP byte string
// with the LEAST-significant byte first.
// ---------------------------------------------------------------------------
/**
* Divide a little-endian base-256 unsigned integer (a byte string) by a small
* `$base` (<= 256), storing the quotient back into `$digits` (with high zero
* limbs stripped) and returning the remainder. The long-division step used to
* peel base-N digits off the little end during encoding.
*
* @param string $digits Modified in place (passed by reference).
* @param int $base
* @return int
*/
function base_encoder_divmod_small(string &$digits, int $base): int
{
$rem = 0;
$len = strlen($digits);
for ($i = $len - 1; $i >= 0; $i--) {
$cur = $rem * 256 + ord($digits[$i]);
$digits[$i] = chr((int) ($cur / $base));
$rem = $cur % $base;
}
// Strip high (trailing in LE) zero limbs — keeps the representation minimal.
$digits = rtrim($digits, "\0");
return $rem;
}
/**
* Multiply a little-endian base-256 unsigned integer by `$base` and add
* `$digit`, in place. The inverse of base_encoder_divmod_small: used to
* reassemble a number from its base-N digits (processed MSB first).
*
* @param string $digits Modified in place (passed by reference).
* @param int $base
* @param int $digit
*/
function base_encoder_muladd_small(string &$digits, int $base, int $digit): void
{
$carry = $digit;
$len = strlen($digits);
for ($i = 0; $i < $len; $i++) {
$cur = ord($digits[$i]) * $base + $carry;
$digits[$i] = chr($cur & 0xff);
$carry = $cur >> 8;
}
while ($carry > 0) {
$digits .= chr($carry & 0xff);
$carry >>= 8;
}
}
/**
* Little-endian base-256 byte string -> minimal big-endian bytes (the form the
* encoders emit and the decoders reconstruct). Reversing yields a minimal
* representation that matches Go's big.Int.Bytes().
*
* @param string $le
* @return string
*/
function base_encoder_to_be_bytes(string $le): string
{
$out = strrev($le);
// Defensive: strip any accidental leading zero (the Go twin guarantees
// minimal output, so we match that contract exactly).
$out = ltrim($out, "\0");
return $out;
}
// ---------------------------------------------------------------------------
// Base32 — RFC 4648 alphabet, padded to a multiple of 8 chars with '='.
// ---------------------------------------------------------------------------
function base_encoder_encode32(string $data): string
{
$out = '';
$len = strlen($data);
for ($i = 0; $i < $len; $i += 5) {
$chunk = substr($data, $i, 5);
$clen = strlen($chunk);
$b = [0, 0, 0, 0, 0];
for ($j = 0; $j < $clen; $j++) {
$b[$j] = ord($chunk[$j]);
}
// Pack 5 bytes (40 bits) into 8 base32 digits (5 bits each, big-endian).
$digits = [
($b[0] >> 3) & 0x1f,
(($b[0] << 2) | ($b[1] >> 6)) & 0x1f,
($b[1] >> 1) & 0x1f,
(($b[1] << 4) | ($b[2] >> 4)) & 0x1f,
(($b[2] << 1) | ($b[3] >> 7)) & 0x1f,
($b[3] >> 2) & 0x1f,
(($b[3] << 3) | ($b[4] >> 5)) & 0x1f,
$b[4] & 0x1f,
];
$outLen = ($clen === 5) ? 8 : BASE_ENCODER_OUT_LEN_32[$clen];
for ($k = 0; $k < $outLen; $k++) {
$out .= BASE_ENCODER_B32_ALPHABET[$digits[$k]];
}
for ($k = $outLen; $k < 8; $k++) {
$out .= '=';
}
}
return $out;
}
function base_encoder_decode32(string $s): ?string
{
$out = '';
$buffer = 0;
$bits = 0;
$len = strlen($s);
for ($i = 0; $i < $len; $i++) {
$c = $s[$i];
if ($c === '=') {
break; // padding marks the end
}
$idx = strpos(BASE_ENCODER_B32_ALPHABET, $c);
if ($idx === false) {
return null;
}
$buffer = ($buffer << 5) | $idx;
$bits += 5;
if ($bits >= 8) {
$bits -= 8;
$out .= chr(($buffer >> $bits) & 0xff);
$buffer &= (1 << $bits) - 1; // keep only the leftover bits
}
}
return $out;
}
// ---------------------------------------------------------------------------
// Base58 — Bitcoin alphabet. Leading 0x00 bytes -> leading '1' (count preserved).
// ---------------------------------------------------------------------------
function base_encoder_encode58(string $data): string
{
$len = strlen($data);
// Count leading zero bytes — each maps to a leading '1'.
$zeros = 0;
while ($zeros < $len && ord($data[$zeros]) === 0) {
$zeros++;
}
// Big-endian byte array (skipping the leading zeros) -> LE base-256.
$le = '';
for ($i = $zeros; $i < $len; $i++) {
base_encoder_muladd_small($le, 256, ord($data[$i]));
}
// Base-convert to 58 digits (collected least-significant first).
$digits = [];
while ($le !== '') {
$digits[] = base_encoder_divmod_small($le, 58);
}
$out = str_repeat('1', $zeros);
for ($k = count($digits) - 1; $k >= 0; $k--) {
$out .= BASE_ENCODER_B58_ALPHABET[$digits[$k]];
}
return $out;
}
function base_encoder_decode58(string $s): ?string
{
$len = strlen($s);
// Count leading '1's — each maps to a 0x00 byte.
$zeros = 0;
while ($zeros < $len && $s[$zeros] === '1') {
$zeros++;
}
$le = '';
for ($i = $zeros; $i < $len; $i++) {
$idx = strpos(BASE_ENCODER_B58_ALPHABET, $s[$i]);
if ($idx === false) {
return null;
}
base_encoder_muladd_small($le, 58, $idx);
}
// LE -> minimal big-endian bytes.
return str_repeat("\0", $zeros) . base_encoder_to_be_bytes($le);
}
// ---------------------------------------------------------------------------
// Base62 — standard base-conversion of the byte array (no leading-zero
// special-casing beyond the standard big-int).
// ---------------------------------------------------------------------------
function base_encoder_encode62(string $data): string
{
if ($data === '') {
return '';
}
$le = '';
$len = strlen($data);
for ($i = 0; $i < $len; $i++) {
base_encoder_muladd_small($le, 256, ord($data[$i]));
}
if ($le === '') {
return '0'; // value zero
}
$digits = [];
while ($le !== '') {
$digits[] = base_encoder_divmod_small($le, 62);
}
$out = '';
for ($k = count($digits) - 1; $k >= 0; $k--) {
$out .= BASE_ENCODER_B62_ALPHABET[$digits[$k]];
}
return $out;
}
function base_encoder_decode62(string $s): ?string
{
if ($s === '') {
return '';
}
$le = '';
$len = strlen($s);
for ($i = 0; $i < $len; $i++) {
$idx = strpos(BASE_ENCODER_B62_ALPHABET, $s[$i]);
if ($idx === false) {
return null;
}
base_encoder_muladd_small($le, 62, $idx);
}
return base_encoder_to_be_bytes($le);
}
// ---------------------------------------------------------------------------
// Base85 — Ascii85. 4 bytes -> 5 chars in '!'(33)..'u'(117); a full 4-zero
// group is shortened to 'z'. No <~ ~> delimiters. Partial final groups emit
// one fewer char than (bytes+1) would suggest; decode reverses, padding with
// 'u' (value 84).
// ---------------------------------------------------------------------------
function base_encoder_encode85(string $data): string
{
$out = '';
$len = strlen($data);
for ($i = 0; $i < $len; $i += 4) {
$chunk = substr($data, $i, 4);
$clen = strlen($chunk);
$isFull = ($clen === 4);
$b = [0, 0, 0, 0];
for ($j = 0; $j < $clen; $j++) {
$b[$j] = ord($chunk[$j]);
}
$u = $b[0] * 16777216 + $b[1] * 65536 + $b[2] * 256 + $b[3];
if ($isFull && $u === 0) {
$out .= 'z'; // zero-group shorthand
continue;
}
$digits = [0, 0, 0, 0, 0];
$v = $u;
for ($k = 4; $k >= 0; $k--) {
$digits[$k] = $v % 85;
$v = intdiv($v, 85);
}
$emit = $isFull ? 5 : $clen + 1; // n bytes -> n+1 chars
for ($k = 0; $k < $emit; $k++) {
$out .= chr($digits[$k] + 33);
}
}
return $out;
}
function base_encoder_decode85(string $s): ?string
{
$out = '';
$group = [];
$len = strlen($s);
for ($i = 0; $i < $len; $i++) {
$c = $s[$i];
if ($c === 'z') {
// 'z' is only valid at a group boundary (an empty accumulator).
if ($group !== []) {
return null;
}
$out .= "\x00\x00\x00\x00";
continue;
}
$code = ord($c);
if ($code < 33 || $code > 117) {
return null;
}
$group[] = $code - 33;
if (count($group) === 5) {
$v = 0;
foreach ($group as $d) {
$v = $v * 85 + $d;
}
if ($v > 0xffffffff) {
return null; // a 5-char group must fit in 32 bits
}
$out .= chr(($v >> 24) & 0xff) . chr(($v >> 16) & 0xff)
. chr(($v >> 8) & 0xff) . chr($v & 0xff);
$group = [];
}
}
// Handle a partial final group (2-4 chars -> 1-3 bytes).
if ($group !== []) {
$m = count($group);
if ($m < 2) {
return null; // a lone trailing char is malformed
}
while (count($group) < 5) {
$group[] = 84; // pad with 'u'
}
$v = 0;
foreach ($group as $d) {
$v = $v * 85 + $d;
}
if ($v > 0xffffffff) {
return null;
}
$all = [($v >> 24) & 0xff, ($v >> 16) & 0xff, ($v >> 8) & 0xff, $v & 0xff];
for ($k = 0; $k < $m - 1; $k++) {
$out .= chr($all[$k]);
}
}
return $out;
}
// ---------------------------------------------------------------------------
// Public API
// ---------------------------------------------------------------------------
/**
* Dispatch raw bytes to the chosen scheme's encoder. Mirrors the Go twin's
* private `encodeBytes`.
*/
function base_encoder_encode_bytes(string $data, string $scheme): string
{
switch ($scheme) {
case BASE_ENCODER_BASE32:
return base_encoder_encode32($data);
case BASE_ENCODER_BASE58:
return base_encoder_encode58($data);
case BASE_ENCODER_BASE62:
return base_encoder_encode62($data);
case BASE_ENCODER_BASE85:
return base_encoder_encode85($data);
default:
return '';
}
}
/**
* Dispatch an encoded string to the chosen scheme's decoder. An invalid or
* malformed input yields null (mirroring the TS lib). Mirrors the Go twin's
* private `decodeBytes`.
*/
function base_encoder_decode_bytes(string $encoded, string $scheme): ?string
{
switch ($scheme) {
case BASE_ENCODER_BASE32:
return base_encoder_decode32($encoded);
case BASE_ENCODER_BASE58:
return base_encoder_decode58($encoded);
case BASE_ENCODER_BASE62:
return base_encoder_decode62($encoded);
case BASE_ENCODER_BASE85:
return base_encoder_decode85($encoded);
default:
return null;
}
}
/**
* Returns the chosen-scheme encoding of the UTF-8 bytes of `$text`. Empty text
* encodes to "". Mirrors `Encode` in cli/base-encoder/base-encoder.go.
*
* (PHP strings are byte arrays, so `$text` already IS its UTF-8 byte sequence.)
*/
function base_encoder_encode(string $text, string $scheme): string
{
return base_encoder_encode_bytes($text, $scheme);
}
/**
* Reverses an encoded string back to UTF-8 text. Invalid characters or a
* malformed structure yield null — mirroring the Go twin's errInvalid and the
* TS lib's null. Mirrors `Decode` in cli/base-encoder/base-encoder.go.
*
* (PHP strings are byte arrays, so the decoded bytes already ARE the text —
* no transcoding, mirroring Go's string([]byte), which never fails.)
*/
function base_encoder_decode(string $encoded, string $scheme): ?string
{
return base_encoder_decode_bytes($encoded, $scheme);
}
// ---------------------------------------------------------------------------
// Showcase self-test — mirrors cli/base-encoder/base-encoder_test.go vectors.
// Run directly: `php php.php` (the backtrace check skips this when included).
// ---------------------------------------------------------------------------
if (PHP_SAPI === 'cli' && debug_backtrace(DEBUG_BACKTRACE_IGNORE_ARGS) === []) {
// Base32 — known values + RFC 4648 padding + case sensitivity.
assert(base_encoder_encode('hello', BASE_ENCODER_BASE32) === 'NBSWY3DP');
assert(base_encoder_encode('foo', BASE_ENCODER_BASE32) === 'MZXW6==='); // 3 bytes -> 5 chars + 3 '='
assert(base_encoder_decode('NBSWY3DP', BASE_ENCODER_BASE32) === 'hello');
assert(base_encoder_decode('nbswy3dp', BASE_ENCODER_BASE32) === null); // lowercase not in RFC 4648
// Base58 — each leading 0x00 byte -> a leading '1'.
assert(base_encoder_encode("\x00", BASE_ENCODER_BASE58) === '1');
assert(str_starts_with(base_encoder_encode("\x00\x00A", BASE_ENCODER_BASE58), '11'));
assert(base_encoder_decode('1', BASE_ENCODER_BASE58) === "\x00");
assert(base_encoder_decode(base_encoder_encode("\x00\x00A", BASE_ENCODER_BASE58), BASE_ENCODER_BASE58) === "\x00\x00A");
// Base62 — plain big-int base conversion (no leading-zero preservation).
assert(base_encoder_encode('A', BASE_ENCODER_BASE62) === '13'); // 1*62 + 3
assert(base_encoder_decode('13', BASE_ENCODER_BASE62) === 'A');
assert(base_encoder_encode("\x00", BASE_ENCODER_BASE62) === '0');
assert(base_encoder_decode('0', BASE_ENCODER_BASE62) === ''); // minimal rep of 0 is empty
// Base85 — Ascii85 'z' shorthand + 32-bit overflow rejection.
assert(base_encoder_encode('hello', BASE_ENCODER_BASE85) === 'BOu!rDZ');
assert(base_encoder_encode("\x00\x00\x00\x00", BASE_ENCODER_BASE85) === 'z');
assert(base_encoder_encode(str_repeat("\x00", 8), BASE_ENCODER_BASE85) === 'zz');
assert(base_encoder_decode('uuuuu', BASE_ENCODER_BASE85) === null); // 5-char group overflows 32 bits
assert(base_encoder_decode('B', BASE_ENCODER_BASE85) === null); // lone trailing char is malformed
// Cross-scheme — empty, multibyte round-trip, and invalid rejection.
foreach ([BASE_ENCODER_BASE32, BASE_ENCODER_BASE58, BASE_ENCODER_BASE62, BASE_ENCODER_BASE85] as $scheme) {
assert(base_encoder_encode('', $scheme) === '');
assert(base_encoder_decode('', $scheme) === '');
assert(base_encoder_decode(base_encoder_encode('CosmoDev 🚀', $scheme), $scheme) === 'CosmoDev 🚀');
assert(base_encoder_decode('~!not-valid!~', $scheme) === null); // '~' outside every alphabet
}
echo "ok\n";
}
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →