Skip to content

ULID Generator — PHP source

Generate Universally Unique Lexicographically Sortable Identifiers (ULID) - 26-character Crockford-base32 strings that sort by millisecond timestamp. Paste any ULID to decode its timestamp and randomness. Runs entirely in your browser.

This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.

<?php
/* ulid-generator — Universally Unique Lexicographically Sortable IDentifier.
 *
 * Language: PHP (8.0+, standard library only)
 * Source:   CosmoDev polyglot showcase port of the ULID Generator tool, ported
 *           from cli/ulid-generator/ulid-generator.go (the hand-rolled Go twin
 *           of src/lib/ulid.ts — the canonical TypeScript implementation).
 * License:  display source — part of CosmoDev's polyglot tool pages.
 *
 * Design goals:
 *   - Pure + deterministic; the generator never throws (decode throws
 *     ValueError on malformed input, matching the TS `throw` and Go error).
 *   - Functionally equivalent to the Go twin: same inputs -> same outputs for
 *     the timestamp prefix, which is the cross-language lock-step anchor.
 *   - Self-container: stdlib only (no Composer packages — no `ramsey/uuid`).
 *
 * A ULID is 26 Crockford-base32 chars: the first 10 encode a 48-bit millisecond
 * timestamp (MSB-first) and the last 16 encode 80 bits of randomness. Crockford
 * alphabet: "0123456789ABCDEFGHJKMNPQRSTVWXYZ" (excludes I, L, O, U).
 *
 * Time-encoding is the lock-step ANCHOR. Because 32^16 == 2^80 exactly,
 * packing [6 time bytes | 10 random bytes] into a 16-byte big-endian integer
 * and base32-encoding the whole 128-bit value reproduces the Go twin's separate
 * time encoding byte-for-byte in the first 10 chars — so
 * ulid_decode_time(ulid_generate(ms)) == ms holds identically. The random tail
 * is drawn from `random_bytes` (a stdlib CSPRNG, the PHP analogue of Go's
 * crypto/rand) and varies call to call, exactly as in the Go default-RNG path.
 */

declare(strict_types=1);

// The Crockford base32 alphabet (excludes I, L, O, U).
const CROCKFORD = '0123456789ABCDEFGHJKMNPQRSTVWXYZ';

/**
 * Build (once) a map from each Crockford symbol to its value, also accepting
 * lowercase a-z where the uppercase equivalent is valid. i/l/o/u are absent
 * because I/L/O/U are not in the alphabet — matching the `ulid` npm package's
 * DECODING table, so a lowercase TS-produced ULID decodes the same way here.
 */
function ulid_decode_map(): array
{
    static $map = null;
    if ($map === null) {
        $map = [];
        for ($i = 0, $n = strlen(CROCKFORD); $i < $n; $i++) {
            $map[CROCKFORD[$i]] = $i;
        }
        for ($c = ord('a'); $c <= ord('z'); $c++) {
            $up = chr($c - ord('a') + ord('A'));
            if (isset($map[$up])) {
                $map[chr($c)] = $map[$up];
            }
        }
    }
    return $map;
}

/**
 * Divide the big-endian integer stored in $b (mutated in place) by 32 and
 * return the remainder (0-31). Long division in base 256: each byte is
 * combined with the carry from the previous (more-significant) byte, the high
 * 8 bits become the new byte and the low 5 bits become the next carry.
 *
 * Mirrors `divBy32` in the Go twin so the base32 digits come out identically.
 */
function ulid_div_by_32(array &$b): int
{
    $rem = 0;
    $n = count($b);
    for ($i = 0; $i < $n; $i++) {
        $cur = ($rem << 8) | $b[$i];
        $b[$i] = $cur >> 5;
        $rem = $cur & 0x1F;
    }
    return $rem;
}

/**
 * Encode a 48-bit millisecond timestamp into 6 big-endian bytes (MSB first).
 * Exposed so showcase tests can build a deterministic ulid_encode input
 * without touching the RNG.
 *
 * @return list<int>
 */
function ulid_time_bytes(int $ms): array
{
    $u = $ms & 0xFFFFFFFFFFFF; // low 48 bits (the timestamp field width)
    $bytes = [];
    for ($j = 0; $j < 6; $j++) {
        $bytes[] = ($u >> (40 - 8 * $j)) & 0xFF;
    }
    return $bytes;
}

/**
 * Pure, deterministic ULID encoder: base32-encode the 128-bit big-endian value
 * [$timeBytes | $randomBytes] into the canonical 26 chars. No RNG, no time
 * source — the showcase tests assert exact strings through this function.
 *
 * @param list<int> $timeBytes    exactly 6 bytes
 * @param list<int> $randomBytes  exactly 10 bytes
 */
function ulid_encode(array $timeBytes, array $randomBytes): string
{
    $b = array_merge($timeBytes, $randomBytes); // 16 bytes
    $out = array_fill(0, 26, '');
    // Repeatedly divide the 128-bit value by 32, collecting remainders
    // LSB-first into out[25] down to out[0]. 26 base32 digits cover 130 bits,
    // so the most significant digit captures the leftover top 3 bits (0-7).
    for ($i = 25; $i >= 0; $i--) {
        $out[$i] = CROCKFORD[ulid_div_by_32($b)];
    }
    return implode('', $out);
}

/**
 * Generate a new ULID for the given millisecond timestamp. The Go twin of
 * generateUlid() in src/lib/ulid.ts. Never throws: $ms is truncated to its low
 * 48 bits; negative values encode the low 48 bits of the two's-complement
 * representation (defined but not meaningful).
 */
function ulid_generate(int $ms): string
{
    // random_bytes is a stdlib CSPRNG. unpack('C*', ...) is 1-indexed, so
    // array_values reindexes to 0-based to match the byte-buffer convention.
    $randomBytes = array_values(unpack('C*', random_bytes(10)));
    return ulid_encode(ulid_time_bytes($ms), $randomBytes);
}

/**
 * Extract the 48-bit millisecond timestamp encoded in the first 10 chars of
 * $id. The Go twin of decodeUlidTime() in src/lib/ulid.ts. Throws ValueError
 * if $id is not exactly 26 chars or contains a character outside the Crockford
 * base32 alphabet (I, L, O, U are invalid) — erroring on the same inputs as
 * the `ulid` npm package's decodeTime.
 */
function ulid_decode_time(string $id): int
{
    if (strlen($id) !== 26) {
        throw new ValueError('ulid: malformed id: length ' . strlen($id) . ', want 26');
    }
    $map = ulid_decode_map();
    $ts = 0;
    for ($i = 0; $i < 10; $i++) {
        $c = $id[$i];
        if (!isset($map[$c])) {
            throw new ValueError("ulid: invalid character '{$c}' at position {$i}");
        }
        $ts = $ts * 32 + $map[$c];
    }
    return $ts;
}

// ---------- showcase tests (the canonical suite lives in src/lib) ----------
// Run only when this file is executed directly from the CLI, not when imported.
if (isset($argv[0]) && realpath($argv[0]) === __FILE__) {
    $zeros = array_fill(0, 10, 0);
    $ulidRe = '/^[0-9A-HJKMNP-TV-Z]{26}$/';

    // All-zero input -> all-zero output; the canonical base32 of 0.
    assert(ulid_encode(ulid_time_bytes(0), $zeros) === str_repeat('0', 26));

    // The lock-step anchor: decode(encode(ms)) == ms, regardless of the random
    // tail. Verified across several byte boundaries in the 48-bit time field.
    foreach ([0, 1, 42, 150000, 1000000, 2000000, (1 << 48) - 1] as $ms) {
        assert(ulid_decode_time(ulid_encode(ulid_time_bytes($ms), $zeros)) === $ms);
    }

    // 150000 base32 in Crockford is "4JFG" -> zero-padded to 10 chars.
    assert(substr(ulid_encode(ulid_time_bytes(150000), $zeros), 0, 10) === '0000004JFG');

    // ulid_generate produces a well-formed 26-char Crockford ULID that round-trips.
    $gid = ulid_generate(150000);
    assert(preg_match($ulidRe, $gid) === 1);
    assert(ulid_decode_time($gid) === 150000);

    // The time prefix is the most-significant part of the string, so an older
    // timestamp sorts before a newer one regardless of the tail.
    assert(ulid_generate(1000000) < ulid_generate(2000000));

    // Malformed input throws (the canonical TS vector + an excluded letter).
    $threw = false;
    try {
        ulid_decode_time('not-a-ulid');
    } catch (ValueError $e) {
        $threw = true;
    }
    assert($threw);

    echo "ulid-generator showcase: all assertions passed\n";
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →