ULID Generator — Python source
Generate Universally Unique Lexicographically Sortable Identifiers (ULID) - 26-character Crockford-base32 strings that sort by millisecond timestamp. Paste any ULID to decode its timestamp and randomness. Runs entirely in your browser.
This is the Python implementation — the same logic the interactive tool runs, in a shareable, citable form.
"""ulid-generator — Universally Unique Lexicographically Sortable IDentifier.
Language: Python (3.9+, standard library only)
Source: CosmoDev polyglot showcase port of the ULID Generator tool, ported
from cli/ulid-generator/ulid-generator.go (the hand-rolled Go twin
of src/lib/ulid.ts — the canonical TypeScript implementation).
License: display source — part of CosmoDev's polyglot tool pages.
Design goals:
- Pure + deterministic; never raises from the generator (decode raises
ValueError on malformed input, matching the TS `throw` and Go error).
- Functionally equivalent to the Go twin: same inputs -> same outputs for
the timestamp prefix, which is the cross-language lock-step anchor.
- Self-contained: stdlib only (no pip packages — no `python-ulid`).
A ULID is 26 Crockford-base32 chars: the first 10 encode a 48-bit millisecond
timestamp (MSB-first) and the last 16 encode 80 bits of randomness. Crockford
alphabet: "0123456789ABCDEFGHJKMNPQRSTVWXYZ" (excludes I, L, O, U).
Time-encoding is the lock-step ANCHOR. Because 32**16 == 2**80 exactly,
packing [6 time bytes | 10 random bytes] into a 16-byte big-endian integer
and base32-encoding the whole 128-bit value reproduces the Go twin's separate
time encoding byte-for-byte in the first 10 chars — so
decode_ulid_time(generate_ulid(ms)) == ms holds identically. The random tail
is drawn from ``secrets.token_bytes`` (a stdlib CSPRNG, the Python analogue of
Go's crypto/rand) and varies call to call, exactly as in the Go default-RNG
path.
"""
from __future__ import annotations
import secrets
from typing import List
# The Crockford base32 alphabet (excludes I, L, O, U).
CROCKFORD = "0123456789ABCDEFGHJKMNPQRSTVWXYZ"
def _build_decode_map() -> dict:
"""Map each Crockford symbol to its value, also accepting lowercase a-z
where the uppercase equivalent is valid. ``i``/``l``/``o``/``u`` are absent
because I/L/O/U are not in the alphabet — matching the ``ulid`` npm
package's DECODING table, so a lowercase TS-produced ULID decodes the same
way here."""
table = {c: i for i, c in enumerate(CROCKFORD)}
for c in "abcdefghijklmnopqrstuvwxyz":
up = c.upper()
if up in table:
table[c] = table[up]
return table
_DECODE = _build_decode_map()
def _div_by_32(b: List[int]) -> int:
"""Divide the big-endian integer stored in ``b`` (mutated in place) by 32
and return the remainder (0-31). Long division in base 256: each byte is
combined with the carry from the previous (more-significant) byte, the
high 8 bits become the new byte and the low 5 bits become the next carry.
Mirrors ``divBy32`` in the Go twin so the base32 digits come out
identically."""
rem = 0
for i in range(len(b)):
cur = (rem << 8) | b[i]
b[i] = cur >> 5
rem = cur & 0x1F
return rem
def time_bytes(ms: int) -> List[int]:
"""Encode a 48-bit millisecond timestamp into 6 big-endian bytes (MSB
first). Exposed so showcase tests can build a deterministic
``encode_ulid`` input without touching the RNG."""
u = ms & 0xFFFFFFFFFFFF # low 48 bits (the timestamp field width)
return [(u >> (40 - 8 * j)) & 0xFF for j in range(6)]
def encode_ulid(time_bytes: List[int], random_bytes: List[int]) -> str:
"""Pure, deterministic ULID encoder: base32-encode the 128-bit big-endian
value ``[time_bytes | random_bytes]`` into the canonical 26 chars. No RNG,
no time source — the showcase tests assert exact strings through this
function."""
b = list(time_bytes) + list(random_bytes) # 16 bytes
out = [""] * 26
# Repeatedly divide the 128-bit value by 32, collecting remainders
# LSB-first into out[25] down to out[0]. 26 base32 digits cover 130 bits,
# so the most significant digit captures the leftover top 3 bits (0-7).
for i in range(25, -1, -1):
out[i] = CROCKFORD[_div_by_32(b)]
return "".join(out)
def generate_ulid(ms: int) -> str:
"""Generate a new ULID for the given millisecond timestamp. The Go twin of
``generateUlid()`` in src/lib/ulid.ts. Never raises: ``ms`` is truncated to
its low 48 bits; negative values encode the low 48 bits of the two's-
complement representation (defined but not meaningful)."""
random_bytes = list(secrets.token_bytes(10)) # stdlib CSPRNG
return encode_ulid(time_bytes(ms), random_bytes)
def decode_ulid_time(id_: str) -> int:
"""Extract the 48-bit millisecond timestamp encoded in the first 10 chars
of ``id_``. The Go twin of ``decodeUlidTime()`` in src/lib/ulid.ts. Raises
``ValueError`` if ``id_`` is not exactly 26 chars or contains a character
outside the Crockford base32 alphabet (I, L, O, U are invalid) — erroring
on the same inputs as the ``ulid`` npm package's ``decodeTime``."""
if len(id_) != 26:
raise ValueError(f"ulid: malformed id: length {len(id_)}, want 26")
ts = 0
for i in range(10):
c = id_[i]
v = _DECODE.get(c)
if v is None:
raise ValueError(f"ulid: invalid character {c!r} at position {i}")
ts = ts * 32 + v
return ts
if __name__ == "__main__":
# ---------- showcase tests (the canonical suite lives in src/lib) ----------
ZEROS = [0] * 10
# All-zero input -> all-zero output; the canonical base32 of 0.
assert encode_ulid(time_bytes(0), ZEROS) == "0" * 26
# The lock-step anchor: decode(encode(ms)) == ms, regardless of the random
# tail. Verified across several byte boundaries in the 48-bit time field.
for ms in (0, 1, 42, 150_000, 1_000_000, 2_000_000, (1 << 48) - 1):
assert decode_ulid_time(encode_ulid(time_bytes(ms), ZEROS)) == ms
# 150000 base32 in Crockford is "4JFG" -> zero-padded to 10 chars.
assert encode_ulid(time_bytes(150_000), ZEROS)[:10] == "0000004JFG"
# generate_ulid produces a well-formed 26-char Crockford ULID that round-trips.
import re
ulid_re = re.compile(r"^[0-9A-HJKMNP-TV-Z]{26}$")
gid = generate_ulid(150_000)
assert ulid_re.match(gid)
assert decode_ulid_time(gid) == 150_000
# The time prefix is the most-significant part of the string, so an older
# timestamp sorts before a newer one regardless of the tail.
assert generate_ulid(1_000_000) < generate_ulid(2_000_000)
# Malformed input raises (the canonical TS vector + an excluded letter).
try:
decode_ulid_time("not-a-ulid")
raise AssertionError("expected ValueError")
except ValueError:
pass
print("ulid-generator showcase: all assertions passed")
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →