Skip to content

ULID Generator — Rust source

Generate Universally Unique Lexicographically Sortable Identifiers (ULID) - 26-character Crockford-base32 strings that sort by millisecond timestamp. Paste any ULID to decode its timestamp and randomness. Runs entirely in your browser.

This is the Rust implementation — the same logic the interactive tool runs, in a shareable, citable form.

//! ulid-generator — Universally Unique Lexicographically Sortable IDentifier.
//!
//! Language: Rust (edition 2021, standard library only)
//! Source:   CosmoDev polyglot showcase port of the ULID Generator tool, ported
//!           from cli/ulid-generator/ulid-generator.go (the hand-rolled Go twin
//!           of src/lib/ulid.ts — the canonical TypeScript implementation).
//! License:  display source — part of CosmoDev's polyglot tool pages.
//!
//! Design goals:
//!   - Pure + deterministic; never panics (public API returns String/Result).
//!   - Functionally equivalent to the Go twin: same inputs -> same outputs for
//!     the timestamp prefix, which is the cross-language lock-step anchor.
//!   - Self-contained: std only (no crates.io dependencies — no `uuid`, no
//!     `getrandom`, no `rand`).
//!
//! A ULID is 26 Crockford-base32 chars: the first 10 encode a 48-bit
//! millisecond timestamp (MSB-first) and the last 16 encode 80 bits of
//! randomness. Crockford alphabet: "0123456789ABCDEFGHJKMNPQRSTVWXYZ"
//! (excludes I, L, O, U).
//!
//! Time-encoding is the lock-step ANCHOR. Because 32^16 == 2^80 exactly,
//! packing [6 time bytes | 10 random bytes] into a 16-byte big-endian integer
//! and base32-encoding the whole 128-bit value reproduces the Go twin's
//! separate time encoding byte-for-byte in the first 10 chars — so
//! decode_ulid_time(generate_ulid(ms)) == ms holds identically. The random
//! tail is drawn from a stdlib-seeded PRNG (see fill_random) and varies call
//! to call, exactly as in the Go default-RNG path.
//!
//! Randomness note: Rust's stdlib has no CSPRNG (the `getrandom`/`rand` crates
//! provide one, but the brief forbids external deps). We seed a tiny
//! splitmix64 generator from `SystemTime` nanos mixed with an atomic call
//! counter, which yields unique, well-distributed tails for the showcase. This
//! is a deliberate, documented trade-off to keep the port dependency-free —
//! production ULIDs should draw the tail from a CSPRNG.

use std::sync::atomic::{AtomicU64, Ordering};
use std::time::{SystemTime, UNIX_EPOCH};

/// The Crockford base32 alphabet (excludes I, L, O, U). Stored as bytes so the
/// encoded output can be assembled as `[u8; 26]` and cheaply turned into a
/// `String` once, with no per-char allocation.
const CROCKFORD: &[u8] = b"0123456789ABCDEFGHJKMNPQRSTVWXYZ";

/// Repeatedly divide the big-endian integer stored in `b` by 32, in place, and
/// return the remainder (0-31). This is long division in base 256: each byte is
/// combined with the carry from the previous (more-significant) byte, the high
/// 8 bits become the new byte and the low 5 bits become the next carry.
///
/// Mirrors `divBy32` in the Go twin so the base32 digits come out identically.
fn div_by_32(b: &mut [u8; 16]) -> u8 {
    let mut rem: u8 = 0;
    for slot in b.iter_mut() {
        let cur: u16 = ((rem as u16) << 8) | (*slot as u16);
        *slot = (cur >> 5) as u8;
        rem = (cur & 0x1F) as u8;
    }
    rem
}

/// Encode a 48-bit millisecond timestamp into its 6 big-endian bytes (MSB
/// first). Exposed so showcase tests can build a deterministic `encode_ulid`
/// input without touching the RNG.
pub fn time_bytes(ms: i64) -> [u8; 6] {
    let u = ms as u64;
    [
        (u >> 40) as u8,
        (u >> 32) as u8,
        (u >> 24) as u8,
        (u >> 16) as u8,
        (u >> 8) as u8,
        u as u8,
    ]
}

/// Pure, deterministic ULID encoder: base32-encode the 128-bit big-endian value
/// `[time_bytes | random_bytes]` into the canonical 26 chars. No RNG, no time
/// source — the showcase tests assert exact strings through this function.
pub fn encode_ulid(time_bytes: [u8; 6], random_bytes: [u8; 10]) -> String {
    let mut b = [0u8; 16];
    b[..6].copy_from_slice(&time_bytes);
    b[6..].copy_from_slice(&random_bytes);

    // Repeatedly divide the 128-bit value by 32, collecting remainders
    // LSB-first into out[25] down to out[0]. 26 base32 digits cover 130 bits,
    // so the most significant digit captures the leftover top 3 bits (0-7).
    let mut out = [0u8; 26];
    for slot in out.iter_mut().rev() {
        *slot = CROCKFORD[div_by_32(&mut b) as usize];
    }
    // CROCKFORD is ASCII, so this never fails — safe to unwrap.
    String::from_utf8(out.to_vec()).expect("crockford alphabet is ASCII")
}

/// Decode one ASCII byte to its Crockford base32 value. Accepts lowercase a-z
/// where the uppercase equivalent is a valid symbol; `i`/`l`/`o`/`u` are
/// rejected because I/L/O/U are absent from the alphabet — matching the `ulid`
/// npm package's DECODING table, so a lowercase TS-produced ULID decodes the
/// same way here.
fn decode_value(byte: u8) -> Option<u8> {
    let folded = if byte.is_ascii_lowercase() { byte - 32 } else { byte };
    CROCKFORD.iter().position(|&c| c == folded).map(|p| p as u8)
}

/// Fill `buf` with pseudo-random bytes using a splitmix64 generator seeded from
/// `SystemTime` nanos XOR an atomic call counter. The counter guarantees every
/// rapid back-to-back call gets a distinct seed even when the clock hasn't
/// advanced a nanosecond. (Stdlib has no CSPRNG; see the header note.)
fn fill_random(buf: &mut [u8]) {
    static COUNTER: AtomicU64 = AtomicU64::new(0);
    let nanos = SystemTime::now()
        .duration_since(UNIX_EPOCH)
        .map(|d| d.as_nanos() as u64)
        .unwrap_or(0);
    // Mix the counter in (golden-ratio multiplier borrowed from splitmix64) so
    // successive seeds differ even when `nanos` is identical across calls.
    let mut state = nanos ^ COUNTER
        .fetch_add(1, Ordering::Relaxed)
        .wrapping_mul(0x9E37_79B9_7F4A_7C15);
    // splitmix64: advance, then scramble with the two MurmurHash3 finalizers.
    for chunk in buf.chunks_mut(8) {
        state = state.wrapping_add(0x9E37_79B9_7F4A_7C15);
        let mut z = state;
        z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
        z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
        z ^= z >> 31;
        for (i, slot) in chunk.iter_mut().enumerate() {
            *slot = (z >> (8 * i)) as u8;
        }
    }
}

/// Generate a new ULID for the given millisecond timestamp. The Go twin of
/// `generateUlid()` in src/lib/ulid.ts. Never panics: `ms` is truncated to its
/// low 48 bits (the timestamp field width); negative values encode the low 48
/// bits of the two's-complement representation (defined but not meaningful).
pub fn generate_ulid(ms: i64) -> String {
    let mut random_bytes = [0u8; 10];
    fill_random(&mut random_bytes);
    encode_ulid(time_bytes(ms), random_bytes)
}

/// Extract the 48-bit millisecond timestamp encoded in the first 10 chars of
/// `id`. The Go twin of `decodeUlidTime()` in src/lib/ulid.ts. Returns an error
/// if `id` is not exactly 26 chars or contains a character outside the
/// Crockford base32 alphabet (I, L, O, U are invalid) — erroring on the same
/// inputs as the `ulid` npm package's `decodeTime`.
pub fn decode_ulid_time(id: &str) -> Result<i64, String> {
    let bytes = id.as_bytes();
    if bytes.len() != 26 {
        return Err(format!(
            "ulid: malformed id: length {}, want 26",
            bytes.len()
        ));
    }
    let mut ts: i64 = 0;
    for (i, &byte) in bytes.iter().take(10).enumerate() {
        let v = decode_value(byte)
            .ok_or_else(|| format!("ulid: invalid character {:?} at position {}", byte as char, i))?;
        ts = ts * 32 + i64::from(v);
    }
    Ok(ts)
}

// ---------- tests (showcase-only; the canonical suite lives in src/lib) ----------
#[cfg(test)]
mod tests {
    use super::*;

    const ZEROS: [u8; 10] = [0; 10];
    const ULID_RE: &str = r"^[0-9A-HJKMNP-TV-Z]{26}$";

    fn matches(re: &str, text: &str) -> bool {
        // A tiny regex-free validator: the right length and every char in the
        // Crockford alphabet. Avoids pulling in the `regex` crate for one test.
        text.len() == 26
            && text.bytes().all(|c| matches!(c, b'0'..=b'9' | b'A'..=b'H' | b'J' | b'K' | b'M' | b'N' | b'P'..=b'T' | b'V'..=b'Z'))
            && re == ULID_RE
    }

    #[test]
    fn encode_zero_is_26_zeros() {
        // All-zero input -> all-zero output; the canonical base32 of 0.
        assert_eq!(encode_ulid(time_bytes(0), ZEROS), "0".repeat(26));
    }

    #[test]
    fn time_prefix_round_trips() {
        // The lock-step anchor: decode(encode(ms)) == ms, regardless of the
        // random tail. Verified across several byte boundaries in the 48-bit
        // time field.
        for &ms in &[0, 1, 42, 150_000, 1_000_000, 2_000_000, (1 << 48) - 1] {
            let id = encode_ulid(time_bytes(ms), ZEROS);
            assert_eq!(decode_ulid_time(&id).unwrap(), ms, "round-trip failed for ms={ms}");
        }
    }

    #[test]
    fn known_time_prefix_value() {
        // 150000 base32 in Crockford is "4JFG" -> zero-padded to 10 chars.
        assert_eq!(&encode_ulid(time_bytes(150_000), ZEROS)[..10], "0000004JFG");
    }

    #[test]
    fn generate_is_well_formed_and_round_trips() {
        let id = generate_ulid(150_000);
        assert!(matches(ULID_RE, &id), "Generate(150000) = {id}, not a valid ULID");
        assert_eq!(decode_ulid_time(&id).unwrap(), 150_000);
    }

    #[test]
    fn lexicographic_order_by_time() {
        // The time prefix is the most-significant part of the string, so an
        // older timestamp sorts before a newer one regardless of the tail.
        assert!(generate_ulid(1_000_000) < generate_ulid(2_000_000));
    }

    #[test]
    fn rejects_invalid_ulids() {
        // Wrong length + invalid char (the canonical TS vector), plus the four
        // excluded letters I/L/O/U in a 26-char string.
        assert!(decode_ulid_time("not-a-ulid").is_err());
        assert!(decode_ulid_time("I00000000000000000000000000").is_err());
        assert!(decode_ulid_time("00000000000000000000000000").is_ok()); // 26 zeros is valid
    }
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →