Skip to content

Punycode Converter — PHP source

Convert internationalized domain names (IDN) between Unicode and Punycode (xn--) ACE form. RFC 3492 compliant, runs entirely in your browser, with a shareable link to your exact input.

This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.

<?php
/**
 * punycode — RFC 3492 Punycode encode/decode + IDNA2003 toASCII/toUnicode.
 *
 * Language: PHP (8.1+, standard library only — mbstring for safe multibyte
 *           handling, which is effectively universal in modern PHP)
 * Source:   CosmoDev polyglot showcase port of the Punycode tool, ported from
 *           src/lib/punycode.ts (canonical TypeScript) and cli/punycode/punycode.go
 *           (the live Go CLI twin — the two are kept in lock-step).
 * License:  display source — part of CosmoDev's polyglot tool pages.
 *
 * Design goals:
 *   - Pure + deterministic; never throws (decode returns ?string, null on
 *     malformed input).
 *   - Functionally equivalent to the Go/TS reference: same inputs -> same outputs.
 *   - Self-contained: stdlib only (no Composer packages).
 *
 * Implements RFC 3492 (Punycode) plus the IDNA2003 toASCII/toUnicode label
 * helpers. punycode_encode_label/punycode_decode_label operate on a single label
 * (no ACE prefix); punycode_encode/punycode_decode wrap them with the "xn--"
 * prefixing and "." splitting of a full domain. PHP strings are byte buffers,
 * so we iterate code points via the mb_ functions (mb_str_split / mb_ord /
 * mb_chr) — equivalent to Go runes and the TS code-point iteration. intdiv() is
 * used for every RFC division so it matches Go's integer division exactly (all
 * operands are non-negative).
 */

declare(strict_types=1);

// RFC 3492 parameters (section 5), matching the TS/Go constants. Prefixed
// PUNY_ to avoid colliding with generic names in the global namespace.
const PUNY_BASE = 36;
const PUNY_TMIN = 1;
const PUNY_TMAX = 26;
const PUNY_SKEW = 38;
const PUNY_DAMP = 700;
const PUNY_INITIAL_BIAS = 72;
const PUNY_INITIAL_N = 128;
const PUNY_ACE_PREFIX = 'xn--';
// Mirrors Number.MAX_SAFE_INTEGER (2^53-1): the overflow guard for malformed
// decode input. Makes pathological generalized numbers (e.g. 40 nines) return
// null instead of running away.
const PUNY_MAX_INT = (1 << 53) - 1;

/**
 * Bias adaptation (RFC 3492 section 6.1). delta/numpoints are always
 * non-negative, so intdiv() matches the reference integer division exactly.
 */
function punycode_adapt(int $delta, int $numpoints, bool $firsttime): int
{
    $d = $firsttime ? intdiv($delta, PUNY_DAMP) : intdiv($delta, 2);
    $d += intdiv($d, $numpoints);
    $k = 0;
    while ($d > intdiv((PUNY_BASE - PUNY_TMIN) * PUNY_TMAX, 2)) {
        $d = intdiv($d, PUNY_BASE - PUNY_TMIN);
        $k += PUNY_BASE;
    }
    return $k + intdiv((PUNY_BASE - PUNY_TMIN + 1) * $d, $d + PUNY_SKEW);
}

/**
 * Map a digit value (0–35) to its RFC 3492 base-36 character (lowercase a–z for
 * 0–25, 0–9 for 26–35). $d is always in [0,35] on the encode path.
 */
function punycode_digit_to_char(int $d): string
{
    if ($d < 26) {
        return chr(ord('a') + $d); // a–z
    }
    return chr(ord('0') + ($d - 26)); // 0–9
}

/**
 * Map a code point to its digit value (0–35), case-insensitive, or -1 if it is
 * not a valid base-36 digit. Any non-ASCII code point yields -1, mirroring the
 * reference behavior of rejecting non-ASCII in the extension portion.
 */
function punycode_char_to_digit(int $cp): int
{
    if ($cp >= 0x61 && $cp <= 0x7A) { // a–z
        return $cp - 0x61;
    }
    if ($cp >= 0x41 && $cp <= 0x5A) { // A–Z
        return $cp - 0x41;
    }
    if ($cp >= 0x30 && $cp <= 0x39) { // 0–9
        return $cp - 0x30 + 26;
    }
    return -1;
}

/**
 * True if the string contains any non-ASCII code point (>= 128).
 */
function punycode_has_non_ascii(string $s): bool
{
    foreach (mb_str_split($s, 1, 'UTF-8') ?: [] as $ch) {
        if (mb_ord($ch, 'UTF-8') >= 128) {
            return true;
        }
    }
    return false;
}

/**
 * Safely turn a decoded code point into a character. Mirrors Go's rune(n)
 * (which emits U+FFFD for an out-of-range or surrogate value) rather than
 * returning false.
 */
function punycode_cp_to_char(int $n): string
{
    return mb_chr($n, 'UTF-8') ?: "\u{FFFD}";
}

/**
 * Punycode-encode a single label (RFC 3492) and return the encoded label with
 * no ACE prefix. Basic (ASCII) code points are emitted first, followed by a '-'
 * delimiter (only if there was at least one), then the generalized base-36
 * deltas for the non-basic code points. Twin of encodeLabel() in the TS/Go.
 */
function punycode_encode_label(string $input): string
{
    // Iterate by code point so astral characters are single elements.
    $code_points = mb_str_split($input, 1, 'UTF-8') ?: [];
    $length = count($code_points);

    $output = [];
    foreach ($code_points as $ch) {
        if (mb_ord($ch, 'UTF-8') < 128) {
            $output[] = $ch;
        }
    }
    $b = count($output);
    if ($b > 0) {
        $output[] = '-';
    }

    $n = PUNY_INITIAL_N;
    $delta = 0;
    $bias = PUNY_INITIAL_BIAS;
    $h = $b;

    while ($h < $length) {
        // Smallest code point in the input that is >= n. The while guard ensures
        // at least one such code point remains, so $m is always set.
        $m = -1;
        foreach ($code_points as $ch) {
            $c = mb_ord($ch, 'UTF-8');
            if ($c >= $n && ($m === -1 || $c < $m)) {
                $m = $c;
            }
        }
        $delta += ($m - $n) * ($h + 1);
        $n = $m;
        foreach ($code_points as $ch) {
            $c = mb_ord($ch, 'UTF-8');
            if ($c < $n) {
                $delta += 1;
            } elseif ($c === $n) {
                $q = $delta;
                $k = PUNY_BASE;
                while (true) {
                    $t = max(PUNY_TMIN, min(PUNY_TMAX, $k - $bias));
                    if ($q < $t) {
                        break;
                    }
                    $output[] = punycode_digit_to_char($t + ($q - $t) % (PUNY_BASE - $t));
                    $q = intdiv($q - $t, PUNY_BASE - $t);
                    $k += PUNY_BASE;
                }
                $output[] = punycode_digit_to_char($q);
                $bias = punycode_adapt($delta, $h + 1, $h === $b);
                $delta = 0;
                $h += 1;
            }
        }
        $delta += 1;
        $n += 1;
    }

    return implode('', $output);
}

/**
 * Punycode-decode a single label (RFC 3492). Returns the decoded label, or null
 * if the input is malformed (invalid digit, truncated generalized number,
 * non-ASCII in the basic portion, or arithmetic overflow). Twin of
 * decodeLabel() in the TS (string | null) and Go (bool form).
 */
function punycode_decode_label(string $input): ?string
{
    // mb_strrpos on the ASCII '-' gives the code-point index of the last dash;
    // since '-' is ASCII it coincides with the byte index the Go twin uses.
    $last_dash = mb_strrpos($input, '-', 0, 'UTF-8');
    $output = [];
    if ($last_dash !== false) {
        $basic = mb_substr($input, 0, $last_dash, 'UTF-8');
        foreach (mb_str_split($basic, 1, 'UTF-8') ?: [] as $ch) {
            if (mb_ord($ch, 'UTF-8') >= 128) {
                return null; // basic portion must be ASCII
            }
            $output[] = $ch;
        }
    }
    $ext = ($last_dash !== false)
        ? mb_substr($input, $last_dash + 1, null, 'UTF-8')
        : $input;
    $ext_chars = mb_str_split($ext, 1, 'UTF-8') ?: [];

    $n = PUNY_INITIAL_N;
    $i = 0;
    $bias = PUNY_INITIAL_BIAS;
    $pos = 0;
    $ext_len = count($ext_chars);

    while ($pos < $ext_len) {
        $oldi = $i;
        $w = 1;
        $k = PUNY_BASE;
        while (true) {
            if ($pos >= $ext_len) {
                return null; // truncated generalized number
            }
            $digit = punycode_char_to_digit(mb_ord($ext_chars[$pos], 'UTF-8'));
            if ($digit < 0) {
                return null; // invalid digit
            }
            $pos += 1;
            if ($digit >= intdiv(PUNY_MAX_INT, $w)) {
                return null; // overflow guard
            }
            $i += $digit * $w;
            $t = max(PUNY_TMIN, min(PUNY_TMAX, $k - $bias));
            if ($digit < $t) {
                break;
            }
            $w *= PUNY_BASE - $t;
            $k += PUNY_BASE;
        }
        $bias = punycode_adapt($i - $oldi, count($output) + 1, $oldi === 0);
        $out_len = count($output) + 1;
        $n += intdiv($i, $out_len);
        $i = $i % $out_len;
        // Insert the decoded code point n at index i (≡ output.splice(i, 0, …)).
        array_splice($output, $i, 0, [punycode_cp_to_char($n)]);
        $i += 1;
    }

    return implode('', $output);
}

/**
 * IDNA toASCII: encode a domain to Punycode ("xn--") form. Lowercases the whole
 * domain, splits on ".", ACE-encodes any label containing a non-ASCII code
 * point, leaves ASCII-only labels untouched, and rejoins with ".". Empty input
 * returns empty. Twin of encode() in the TS/Go.
 */
function punycode_encode(string $domain): string
{
    if ($domain === '') {
        return '';
    }
    $lower = mb_strtolower($domain, 'UTF-8');
    $labels = explode('.', $lower);
    foreach ($labels as $idx => $label) {
        if (punycode_has_non_ascii($label)) {
            $labels[$idx] = PUNY_ACE_PREFIX . punycode_encode_label($label);
        }
    }
    return implode('.', $labels);
}

/**
 * IDNA toUnicode: decode a Punycode ("xn--") domain back to Unicode. Splits on
 * ".", decodes any label beginning with "xn--" (case-insensitive, detected on
 * the lowercased label), leaves every other label untouched, and rejoins with
 * ".". Returns null if any "xn--" label is invalid — the whole domain is
 * rejected, matching IDNA semantics. Empty input returns "". Twin of decode()
 * in the TS/Go.
 */
function punycode_decode(string $domain): ?string
{
    if ($domain === '') {
        return '';
    }
    $out = [];
    foreach (explode('.', $domain) as $label) {
        $lower = mb_strtolower($label, 'UTF-8');
        if (str_starts_with($lower, PUNY_ACE_PREFIX)
            && mb_strlen($label, 'UTF-8') > mb_strlen(PUNY_ACE_PREFIX, 'UTF-8')) {
            $decoded = punycode_decode_label(
                mb_substr($label, mb_strlen(PUNY_ACE_PREFIX, 'UTF-8'), null, 'UTF-8')
            );
            if ($decoded === null) {
                return null;
            }
            $out[] = $decoded;
        } else {
            $out[] = $label;
        }
    }
    return implode('.', $out);
}

// Showcase vectors — run only when this file is executed directly (not when
// included as a library). The debug_backtrace() check is the idiomatic PHP
// "am I the entry script?" guard. Requires zend.assertions >= 0 (the default).
if (!debug_backtrace(DEBUG_BACKTRACE_IGNORE_ARGS)) {
    assert(punycode_encode('münchen.de') === 'xn--mnchen-3ya.de');
    assert(punycode_decode('xn--mnchen-3ya.de') === 'münchen.de');
    assert(punycode_encode('Bücher.DE') === 'xn--bcher-kva.de'); // lowercased first
    assert(punycode_encode_label('café') === 'caf-dma');
    assert(punycode_decode_label('caf-dma') === 'café');
    assert(punycode_decode('xn--!') === null);                       // invalid digit
    assert(punycode_decode_label(str_repeat('9', 40)) === null);     // overflow guard
    assert(punycode_decode(punycode_encode('café.fr')) === 'café.fr');
    echo "punycode: all showcase vectors passed\n";
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →