Skip to content

ULID Generator — Python source

Generate Universally Unique Lexicographically Sortable Identifiers (ULID) - 26-character Crockford-base32 strings that sort by millisecond timestamp. Paste any ULID to decode its timestamp and randomness. Runs entirely in your browser.

This is the Python implementation — the same logic the interactive tool runs, in a shareable, citable form.

"""ulid-generator — Universally Unique Lexicographically Sortable IDentifier.

Language: Python (3.9+, standard library only)
Source:   CosmoDev polyglot showcase port of the ULID Generator tool, ported
          from cli/ulid-generator/ulid-generator.go (the hand-rolled Go twin
          of src/lib/ulid.ts — the canonical TypeScript implementation).
License:  display source — part of CosmoDev's polyglot tool pages.

Design goals:
  - Pure + deterministic; never raises from the generator (decode raises
    ValueError on malformed input, matching the TS `throw` and Go error).
  - Functionally equivalent to the Go twin: same inputs -> same outputs for
    the timestamp prefix, which is the cross-language lock-step anchor.
  - Self-contained: stdlib only (no pip packages — no `python-ulid`).

A ULID is 26 Crockford-base32 chars: the first 10 encode a 48-bit millisecond
timestamp (MSB-first) and the last 16 encode 80 bits of randomness. Crockford
alphabet: "0123456789ABCDEFGHJKMNPQRSTVWXYZ" (excludes I, L, O, U).

Time-encoding is the lock-step ANCHOR. Because 32**16 == 2**80 exactly,
packing [6 time bytes | 10 random bytes] into a 16-byte big-endian integer
and base32-encoding the whole 128-bit value reproduces the Go twin's separate
time encoding byte-for-byte in the first 10 chars — so
decode_ulid_time(generate_ulid(ms)) == ms holds identically. The random tail
is drawn from ``secrets.token_bytes`` (a stdlib CSPRNG, the Python analogue of
Go's crypto/rand) and varies call to call, exactly as in the Go default-RNG
path.
"""
from __future__ import annotations

import secrets
from typing import List

# The Crockford base32 alphabet (excludes I, L, O, U).
CROCKFORD = "0123456789ABCDEFGHJKMNPQRSTVWXYZ"


def _build_decode_map() -> dict:
    """Map each Crockford symbol to its value, also accepting lowercase a-z
    where the uppercase equivalent is valid. ``i``/``l``/``o``/``u`` are absent
    because I/L/O/U are not in the alphabet — matching the ``ulid`` npm
    package's DECODING table, so a lowercase TS-produced ULID decodes the same
    way here."""
    table = {c: i for i, c in enumerate(CROCKFORD)}
    for c in "abcdefghijklmnopqrstuvwxyz":
        up = c.upper()
        if up in table:
            table[c] = table[up]
    return table


_DECODE = _build_decode_map()


def _div_by_32(b: List[int]) -> int:
    """Divide the big-endian integer stored in ``b`` (mutated in place) by 32
    and return the remainder (0-31). Long division in base 256: each byte is
    combined with the carry from the previous (more-significant) byte, the
    high 8 bits become the new byte and the low 5 bits become the next carry.

    Mirrors ``divBy32`` in the Go twin so the base32 digits come out
    identically."""
    rem = 0
    for i in range(len(b)):
        cur = (rem << 8) | b[i]
        b[i] = cur >> 5
        rem = cur & 0x1F
    return rem


def time_bytes(ms: int) -> List[int]:
    """Encode a 48-bit millisecond timestamp into 6 big-endian bytes (MSB
    first). Exposed so showcase tests can build a deterministic
    ``encode_ulid`` input without touching the RNG."""
    u = ms & 0xFFFFFFFFFFFF  # low 48 bits (the timestamp field width)
    return [(u >> (40 - 8 * j)) & 0xFF for j in range(6)]


def encode_ulid(time_bytes: List[int], random_bytes: List[int]) -> str:
    """Pure, deterministic ULID encoder: base32-encode the 128-bit big-endian
    value ``[time_bytes | random_bytes]`` into the canonical 26 chars. No RNG,
    no time source — the showcase tests assert exact strings through this
    function."""
    b = list(time_bytes) + list(random_bytes)  # 16 bytes
    out = [""] * 26
    # Repeatedly divide the 128-bit value by 32, collecting remainders
    # LSB-first into out[25] down to out[0]. 26 base32 digits cover 130 bits,
    # so the most significant digit captures the leftover top 3 bits (0-7).
    for i in range(25, -1, -1):
        out[i] = CROCKFORD[_div_by_32(b)]
    return "".join(out)


def generate_ulid(ms: int) -> str:
    """Generate a new ULID for the given millisecond timestamp. The Go twin of
    ``generateUlid()`` in src/lib/ulid.ts. Never raises: ``ms`` is truncated to
    its low 48 bits; negative values encode the low 48 bits of the two's-
    complement representation (defined but not meaningful)."""
    random_bytes = list(secrets.token_bytes(10))  # stdlib CSPRNG
    return encode_ulid(time_bytes(ms), random_bytes)


def decode_ulid_time(id_: str) -> int:
    """Extract the 48-bit millisecond timestamp encoded in the first 10 chars
    of ``id_``. The Go twin of ``decodeUlidTime()`` in src/lib/ulid.ts. Raises
    ``ValueError`` if ``id_`` is not exactly 26 chars or contains a character
    outside the Crockford base32 alphabet (I, L, O, U are invalid) — erroring
    on the same inputs as the ``ulid`` npm package's ``decodeTime``."""
    if len(id_) != 26:
        raise ValueError(f"ulid: malformed id: length {len(id_)}, want 26")
    ts = 0
    for i in range(10):
        c = id_[i]
        v = _DECODE.get(c)
        if v is None:
            raise ValueError(f"ulid: invalid character {c!r} at position {i}")
        ts = ts * 32 + v
    return ts


if __name__ == "__main__":
    # ---------- showcase tests (the canonical suite lives in src/lib) ----------
    ZEROS = [0] * 10

    # All-zero input -> all-zero output; the canonical base32 of 0.
    assert encode_ulid(time_bytes(0), ZEROS) == "0" * 26

    # The lock-step anchor: decode(encode(ms)) == ms, regardless of the random
    # tail. Verified across several byte boundaries in the 48-bit time field.
    for ms in (0, 1, 42, 150_000, 1_000_000, 2_000_000, (1 << 48) - 1):
        assert decode_ulid_time(encode_ulid(time_bytes(ms), ZEROS)) == ms

    # 150000 base32 in Crockford is "4JFG" -> zero-padded to 10 chars.
    assert encode_ulid(time_bytes(150_000), ZEROS)[:10] == "0000004JFG"

    # generate_ulid produces a well-formed 26-char Crockford ULID that round-trips.
    import re

    ulid_re = re.compile(r"^[0-9A-HJKMNP-TV-Z]{26}$")
    gid = generate_ulid(150_000)
    assert ulid_re.match(gid)
    assert decode_ulid_time(gid) == 150_000

    # The time prefix is the most-significant part of the string, so an older
    # timestamp sorts before a newer one regardless of the tail.
    assert generate_ulid(1_000_000) < generate_ulid(2_000_000)

    # Malformed input raises (the canonical TS vector + an excluded letter).
    try:
        decode_ulid_time("not-a-ulid")
        raise AssertionError("expected ValueError")
    except ValueError:
        pass

    print("ulid-generator showcase: all assertions passed")

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →