Skip to content

Base32 / Base58 / Base62 / Base85 Encoder — PHP source

Encode text to Base32, Base58, Base62, or Ascii85 - or decode it back. UTF-8 safe, runs entirely in your browser, with a shareable link to your exact input.

This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.

<?php
/**
 * base-encoder — Base32 (RFC 4648), Base58 (Bitcoin), Base62, and Base85
 * (Ascii85) byte-array encoders, operating on the UTF-8 bytes of the input.
 *
 * Language: PHP (8.1+, standard library only — 64-bit build assumed for the
 *           Base85 32-bit group value; Base58/62 use a manual base-256 big-int
 *           that fits even on a 32-bit build)
 * Source:   CosmoDev polyglot showcase port of the Base Encoder tool, ported
 *           from cli/base-encoder/base-encoder.go (the authoritative Go twin).
 * License:  display source — part of CosmoDev's polyglot tool pages.
 *
 * Design goals:
 *   - Pure + deterministic; never throws (decode returns null for invalid or
 *     malformed input, mirroring the TS lib's null and the Go twin's errInvalid).
 *   - Functionally equivalent to the Go twin: same inputs -> same outputs.
 *   - Self-contained: stdlib only (no Composer packages, no bcmath/gmp).
 *
 * Arbitrary-precision note: Base58 and Base62 base-convert the whole byte
 * array, which overflows PHP_INT_MAX for inputs longer than a few bytes. The
 * Go twin leans on math/big and the JS/Python ports on native BigInt; PHP has
 * no big-integer type without an extension, so we implement the same idea with
 * a little-endian base-256 byte string and two primitives — divmod_small (peel
 * a base-N digit off the little end) and muladd_small (reassemble a number from
 * its base-N digits). Intermediate products stay under 2^17, so this is safe
 * on any PHP build.
 *
 * String model: PHP strings are byte arrays, so an input string already IS its
 * UTF-8 byte sequence and a decoded byte string is already the text result —
 * no transcoding is needed (mirrors Go's string([]byte), which never fails).
 */

declare(strict_types=1);

/**
 * One of the four supported byte-array base encodings. Mirrors the Go twin's
 * `Scheme` type and the TS `Scheme` union.
 */
const BASE_ENCODER_BASE32 = 'base32';
const BASE_ENCODER_BASE58 = 'base58';
const BASE_ENCODER_BASE62 = 'base62';
const BASE_ENCODER_BASE85 = 'base85';

const BASE_ENCODER_B32_ALPHABET = 'ABCDEFGHIJKLMNOPQRSTUVWXYZ234567';
const BASE_ENCODER_B58_ALPHABET = '123456789ABCDEFGHJKLMNPQRSTUVWXYZabcdefghijkmnopqrstuvwxyz';
const BASE_ENCODER_B62_ALPHABET = '0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz';

// Data characters emitted by a final (partial) 5-byte chunk before '=' padding,
// per RFC 4648. Index = byte count (0..4). Matches the TS `outLen` table.
const BASE_ENCODER_OUT_LEN_32 = [0, 2, 4, 5, 7];

// ---------------------------------------------------------------------------
// Arbitrary-precision primitives (base-256, little-endian). Used by Base58 and
// Base62 so the port stays extension-free. The "number" is a PHP byte string
// with the LEAST-significant byte first.
// ---------------------------------------------------------------------------

/**
 * Divide a little-endian base-256 unsigned integer (a byte string) by a small
 * `$base` (<= 256), storing the quotient back into `$digits` (with high zero
 * limbs stripped) and returning the remainder. The long-division step used to
 * peel base-N digits off the little end during encoding.
 *
 * @param string $digits Modified in place (passed by reference).
 * @param int    $base
 * @return int
 */
function base_encoder_divmod_small(string &$digits, int $base): int
{
    $rem = 0;
    $len = strlen($digits);
    for ($i = $len - 1; $i >= 0; $i--) {
        $cur = $rem * 256 + ord($digits[$i]);
        $digits[$i] = chr((int) ($cur / $base));
        $rem = $cur % $base;
    }
    // Strip high (trailing in LE) zero limbs — keeps the representation minimal.
    $digits = rtrim($digits, "\0");
    return $rem;
}

/**
 * Multiply a little-endian base-256 unsigned integer by `$base` and add
 * `$digit`, in place. The inverse of base_encoder_divmod_small: used to
 * reassemble a number from its base-N digits (processed MSB first).
 *
 * @param string $digits Modified in place (passed by reference).
 * @param int    $base
 * @param int    $digit
 */
function base_encoder_muladd_small(string &$digits, int $base, int $digit): void
{
    $carry = $digit;
    $len = strlen($digits);
    for ($i = 0; $i < $len; $i++) {
        $cur = ord($digits[$i]) * $base + $carry;
        $digits[$i] = chr($cur & 0xff);
        $carry = $cur >> 8;
    }
    while ($carry > 0) {
        $digits .= chr($carry & 0xff);
        $carry >>= 8;
    }
}

/**
 * Little-endian base-256 byte string -> minimal big-endian bytes (the form the
 * encoders emit and the decoders reconstruct). Reversing yields a minimal
 * representation that matches Go's big.Int.Bytes().
 *
 * @param string $le
 * @return string
 */
function base_encoder_to_be_bytes(string $le): string
{
    $out = strrev($le);
    // Defensive: strip any accidental leading zero (the Go twin guarantees
    // minimal output, so we match that contract exactly).
    $out = ltrim($out, "\0");
    return $out;
}

// ---------------------------------------------------------------------------
// Base32 — RFC 4648 alphabet, padded to a multiple of 8 chars with '='.
// ---------------------------------------------------------------------------

function base_encoder_encode32(string $data): string
{
    $out = '';
    $len = strlen($data);
    for ($i = 0; $i < $len; $i += 5) {
        $chunk = substr($data, $i, 5);
        $clen = strlen($chunk);
        $b = [0, 0, 0, 0, 0];
        for ($j = 0; $j < $clen; $j++) {
            $b[$j] = ord($chunk[$j]);
        }
        // Pack 5 bytes (40 bits) into 8 base32 digits (5 bits each, big-endian).
        $digits = [
            ($b[0] >> 3) & 0x1f,
            (($b[0] << 2) | ($b[1] >> 6)) & 0x1f,
            ($b[1] >> 1) & 0x1f,
            (($b[1] << 4) | ($b[2] >> 4)) & 0x1f,
            (($b[2] << 1) | ($b[3] >> 7)) & 0x1f,
            ($b[3] >> 2) & 0x1f,
            (($b[3] << 3) | ($b[4] >> 5)) & 0x1f,
            $b[4] & 0x1f,
        ];
        $outLen = ($clen === 5) ? 8 : BASE_ENCODER_OUT_LEN_32[$clen];
        for ($k = 0; $k < $outLen; $k++) {
            $out .= BASE_ENCODER_B32_ALPHABET[$digits[$k]];
        }
        for ($k = $outLen; $k < 8; $k++) {
            $out .= '=';
        }
    }
    return $out;
}

function base_encoder_decode32(string $s): ?string
{
    $out = '';
    $buffer = 0;
    $bits = 0;
    $len = strlen($s);
    for ($i = 0; $i < $len; $i++) {
        $c = $s[$i];
        if ($c === '=') {
            break; // padding marks the end
        }
        $idx = strpos(BASE_ENCODER_B32_ALPHABET, $c);
        if ($idx === false) {
            return null;
        }
        $buffer = ($buffer << 5) | $idx;
        $bits += 5;
        if ($bits >= 8) {
            $bits -= 8;
            $out .= chr(($buffer >> $bits) & 0xff);
            $buffer &= (1 << $bits) - 1; // keep only the leftover bits
        }
    }
    return $out;
}

// ---------------------------------------------------------------------------
// Base58 — Bitcoin alphabet. Leading 0x00 bytes -> leading '1' (count preserved).
// ---------------------------------------------------------------------------

function base_encoder_encode58(string $data): string
{
    $len = strlen($data);
    // Count leading zero bytes — each maps to a leading '1'.
    $zeros = 0;
    while ($zeros < $len && ord($data[$zeros]) === 0) {
        $zeros++;
    }
    // Big-endian byte array (skipping the leading zeros) -> LE base-256.
    $le = '';
    for ($i = $zeros; $i < $len; $i++) {
        base_encoder_muladd_small($le, 256, ord($data[$i]));
    }
    // Base-convert to 58 digits (collected least-significant first).
    $digits = [];
    while ($le !== '') {
        $digits[] = base_encoder_divmod_small($le, 58);
    }
    $out = str_repeat('1', $zeros);
    for ($k = count($digits) - 1; $k >= 0; $k--) {
        $out .= BASE_ENCODER_B58_ALPHABET[$digits[$k]];
    }
    return $out;
}

function base_encoder_decode58(string $s): ?string
{
    $len = strlen($s);
    // Count leading '1's — each maps to a 0x00 byte.
    $zeros = 0;
    while ($zeros < $len && $s[$zeros] === '1') {
        $zeros++;
    }
    $le = '';
    for ($i = $zeros; $i < $len; $i++) {
        $idx = strpos(BASE_ENCODER_B58_ALPHABET, $s[$i]);
        if ($idx === false) {
            return null;
        }
        base_encoder_muladd_small($le, 58, $idx);
    }
    // LE -> minimal big-endian bytes.
    return str_repeat("\0", $zeros) . base_encoder_to_be_bytes($le);
}

// ---------------------------------------------------------------------------
// Base62 — standard base-conversion of the byte array (no leading-zero
// special-casing beyond the standard big-int).
// ---------------------------------------------------------------------------

function base_encoder_encode62(string $data): string
{
    if ($data === '') {
        return '';
    }
    $le = '';
    $len = strlen($data);
    for ($i = 0; $i < $len; $i++) {
        base_encoder_muladd_small($le, 256, ord($data[$i]));
    }
    if ($le === '') {
        return '0'; // value zero
    }
    $digits = [];
    while ($le !== '') {
        $digits[] = base_encoder_divmod_small($le, 62);
    }
    $out = '';
    for ($k = count($digits) - 1; $k >= 0; $k--) {
        $out .= BASE_ENCODER_B62_ALPHABET[$digits[$k]];
    }
    return $out;
}

function base_encoder_decode62(string $s): ?string
{
    if ($s === '') {
        return '';
    }
    $le = '';
    $len = strlen($s);
    for ($i = 0; $i < $len; $i++) {
        $idx = strpos(BASE_ENCODER_B62_ALPHABET, $s[$i]);
        if ($idx === false) {
            return null;
        }
        base_encoder_muladd_small($le, 62, $idx);
    }
    return base_encoder_to_be_bytes($le);
}

// ---------------------------------------------------------------------------
// Base85 — Ascii85. 4 bytes -> 5 chars in '!'(33)..'u'(117); a full 4-zero
// group is shortened to 'z'. No <~ ~> delimiters. Partial final groups emit
// one fewer char than (bytes+1) would suggest; decode reverses, padding with
// 'u' (value 84).
// ---------------------------------------------------------------------------

function base_encoder_encode85(string $data): string
{
    $out = '';
    $len = strlen($data);
    for ($i = 0; $i < $len; $i += 4) {
        $chunk = substr($data, $i, 4);
        $clen = strlen($chunk);
        $isFull = ($clen === 4);
        $b = [0, 0, 0, 0];
        for ($j = 0; $j < $clen; $j++) {
            $b[$j] = ord($chunk[$j]);
        }
        $u = $b[0] * 16777216 + $b[1] * 65536 + $b[2] * 256 + $b[3];
        if ($isFull && $u === 0) {
            $out .= 'z'; // zero-group shorthand
            continue;
        }
        $digits = [0, 0, 0, 0, 0];
        $v = $u;
        for ($k = 4; $k >= 0; $k--) {
            $digits[$k] = $v % 85;
            $v = intdiv($v, 85);
        }
        $emit = $isFull ? 5 : $clen + 1; // n bytes -> n+1 chars
        for ($k = 0; $k < $emit; $k++) {
            $out .= chr($digits[$k] + 33);
        }
    }
    return $out;
}

function base_encoder_decode85(string $s): ?string
{
    $out = '';
    $group = [];
    $len = strlen($s);
    for ($i = 0; $i < $len; $i++) {
        $c = $s[$i];
        if ($c === 'z') {
            // 'z' is only valid at a group boundary (an empty accumulator).
            if ($group !== []) {
                return null;
            }
            $out .= "\x00\x00\x00\x00";
            continue;
        }
        $code = ord($c);
        if ($code < 33 || $code > 117) {
            return null;
        }
        $group[] = $code - 33;
        if (count($group) === 5) {
            $v = 0;
            foreach ($group as $d) {
                $v = $v * 85 + $d;
            }
            if ($v > 0xffffffff) {
                return null; // a 5-char group must fit in 32 bits
            }
            $out .= chr(($v >> 24) & 0xff) . chr(($v >> 16) & 0xff)
                  . chr(($v >> 8) & 0xff) . chr($v & 0xff);
            $group = [];
        }
    }
    // Handle a partial final group (2-4 chars -> 1-3 bytes).
    if ($group !== []) {
        $m = count($group);
        if ($m < 2) {
            return null; // a lone trailing char is malformed
        }
        while (count($group) < 5) {
            $group[] = 84; // pad with 'u'
        }
        $v = 0;
        foreach ($group as $d) {
            $v = $v * 85 + $d;
        }
        if ($v > 0xffffffff) {
            return null;
        }
        $all = [($v >> 24) & 0xff, ($v >> 16) & 0xff, ($v >> 8) & 0xff, $v & 0xff];
        for ($k = 0; $k < $m - 1; $k++) {
            $out .= chr($all[$k]);
        }
    }
    return $out;
}

// ---------------------------------------------------------------------------
// Public API
// ---------------------------------------------------------------------------

/**
 * Dispatch raw bytes to the chosen scheme's encoder. Mirrors the Go twin's
 * private `encodeBytes`.
 */
function base_encoder_encode_bytes(string $data, string $scheme): string
{
    switch ($scheme) {
        case BASE_ENCODER_BASE32:
            return base_encoder_encode32($data);
        case BASE_ENCODER_BASE58:
            return base_encoder_encode58($data);
        case BASE_ENCODER_BASE62:
            return base_encoder_encode62($data);
        case BASE_ENCODER_BASE85:
            return base_encoder_encode85($data);
        default:
            return '';
    }
}

/**
 * Dispatch an encoded string to the chosen scheme's decoder. An invalid or
 * malformed input yields null (mirroring the TS lib). Mirrors the Go twin's
 * private `decodeBytes`.
 */
function base_encoder_decode_bytes(string $encoded, string $scheme): ?string
{
    switch ($scheme) {
        case BASE_ENCODER_BASE32:
            return base_encoder_decode32($encoded);
        case BASE_ENCODER_BASE58:
            return base_encoder_decode58($encoded);
        case BASE_ENCODER_BASE62:
            return base_encoder_decode62($encoded);
        case BASE_ENCODER_BASE85:
            return base_encoder_decode85($encoded);
        default:
            return null;
    }
}

/**
 * Returns the chosen-scheme encoding of the UTF-8 bytes of `$text`. Empty text
 * encodes to "". Mirrors `Encode` in cli/base-encoder/base-encoder.go.
 *
 * (PHP strings are byte arrays, so `$text` already IS its UTF-8 byte sequence.)
 */
function base_encoder_encode(string $text, string $scheme): string
{
    return base_encoder_encode_bytes($text, $scheme);
}

/**
 * Reverses an encoded string back to UTF-8 text. Invalid characters or a
 * malformed structure yield null — mirroring the Go twin's errInvalid and the
 * TS lib's null. Mirrors `Decode` in cli/base-encoder/base-encoder.go.
 *
 * (PHP strings are byte arrays, so the decoded bytes already ARE the text —
 * no transcoding, mirroring Go's string([]byte), which never fails.)
 */
function base_encoder_decode(string $encoded, string $scheme): ?string
{
    return base_encoder_decode_bytes($encoded, $scheme);
}

// ---------------------------------------------------------------------------
// Showcase self-test — mirrors cli/base-encoder/base-encoder_test.go vectors.
// Run directly: `php php.php` (the backtrace check skips this when included).
// ---------------------------------------------------------------------------
if (PHP_SAPI === 'cli' && debug_backtrace(DEBUG_BACKTRACE_IGNORE_ARGS) === []) {
    // Base32 — known values + RFC 4648 padding + case sensitivity.
    assert(base_encoder_encode('hello', BASE_ENCODER_BASE32) === 'NBSWY3DP');
    assert(base_encoder_encode('foo', BASE_ENCODER_BASE32) === 'MZXW6==='); // 3 bytes -> 5 chars + 3 '='
    assert(base_encoder_decode('NBSWY3DP', BASE_ENCODER_BASE32) === 'hello');
    assert(base_encoder_decode('nbswy3dp', BASE_ENCODER_BASE32) === null); // lowercase not in RFC 4648

    // Base58 — each leading 0x00 byte -> a leading '1'.
    assert(base_encoder_encode("\x00", BASE_ENCODER_BASE58) === '1');
    assert(str_starts_with(base_encoder_encode("\x00\x00A", BASE_ENCODER_BASE58), '11'));
    assert(base_encoder_decode('1', BASE_ENCODER_BASE58) === "\x00");
    assert(base_encoder_decode(base_encoder_encode("\x00\x00A", BASE_ENCODER_BASE58), BASE_ENCODER_BASE58) === "\x00\x00A");

    // Base62 — plain big-int base conversion (no leading-zero preservation).
    assert(base_encoder_encode('A', BASE_ENCODER_BASE62) === '13'); // 1*62 + 3
    assert(base_encoder_decode('13', BASE_ENCODER_BASE62) === 'A');
    assert(base_encoder_encode("\x00", BASE_ENCODER_BASE62) === '0');
    assert(base_encoder_decode('0', BASE_ENCODER_BASE62) === ''); // minimal rep of 0 is empty

    // Base85 — Ascii85 'z' shorthand + 32-bit overflow rejection.
    assert(base_encoder_encode('hello', BASE_ENCODER_BASE85) === 'BOu!rDZ');
    assert(base_encoder_encode("\x00\x00\x00\x00", BASE_ENCODER_BASE85) === 'z');
    assert(base_encoder_encode(str_repeat("\x00", 8), BASE_ENCODER_BASE85) === 'zz');
    assert(base_encoder_decode('uuuuu', BASE_ENCODER_BASE85) === null); // 5-char group overflows 32 bits
    assert(base_encoder_decode('B', BASE_ENCODER_BASE85) === null); // lone trailing char is malformed

    // Cross-scheme — empty, multibyte round-trip, and invalid rejection.
    foreach ([BASE_ENCODER_BASE32, BASE_ENCODER_BASE58, BASE_ENCODER_BASE62, BASE_ENCODER_BASE85] as $scheme) {
        assert(base_encoder_encode('', $scheme) === '');
        assert(base_encoder_decode('', $scheme) === '');
        assert(base_encoder_decode(base_encoder_encode('CosmoDev 🚀', $scheme), $scheme) === 'CosmoDev 🚀');
        assert(base_encoder_decode('~!not-valid!~', $scheme) === null); // '~' outside every alphabet
    }

    echo "ok\n";
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →