Punycode Converter — PHP source
Convert internationalized domain names (IDN) between Unicode and Punycode (xn--) ACE form. RFC 3492 compliant, runs entirely in your browser, with a shareable link to your exact input.
This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.
<?php
/**
* punycode — RFC 3492 Punycode encode/decode + IDNA2003 toASCII/toUnicode.
*
* Language: PHP (8.1+, standard library only — mbstring for safe multibyte
* handling, which is effectively universal in modern PHP)
* Source: CosmoDev polyglot showcase port of the Punycode tool, ported from
* src/lib/punycode.ts (canonical TypeScript) and cli/punycode/punycode.go
* (the live Go CLI twin — the two are kept in lock-step).
* License: display source — part of CosmoDev's polyglot tool pages.
*
* Design goals:
* - Pure + deterministic; never throws (decode returns ?string, null on
* malformed input).
* - Functionally equivalent to the Go/TS reference: same inputs -> same outputs.
* - Self-contained: stdlib only (no Composer packages).
*
* Implements RFC 3492 (Punycode) plus the IDNA2003 toASCII/toUnicode label
* helpers. punycode_encode_label/punycode_decode_label operate on a single label
* (no ACE prefix); punycode_encode/punycode_decode wrap them with the "xn--"
* prefixing and "." splitting of a full domain. PHP strings are byte buffers,
* so we iterate code points via the mb_ functions (mb_str_split / mb_ord /
* mb_chr) — equivalent to Go runes and the TS code-point iteration. intdiv() is
* used for every RFC division so it matches Go's integer division exactly (all
* operands are non-negative).
*/
declare(strict_types=1);
// RFC 3492 parameters (section 5), matching the TS/Go constants. Prefixed
// PUNY_ to avoid colliding with generic names in the global namespace.
const PUNY_BASE = 36;
const PUNY_TMIN = 1;
const PUNY_TMAX = 26;
const PUNY_SKEW = 38;
const PUNY_DAMP = 700;
const PUNY_INITIAL_BIAS = 72;
const PUNY_INITIAL_N = 128;
const PUNY_ACE_PREFIX = 'xn--';
// Mirrors Number.MAX_SAFE_INTEGER (2^53-1): the overflow guard for malformed
// decode input. Makes pathological generalized numbers (e.g. 40 nines) return
// null instead of running away.
const PUNY_MAX_INT = (1 << 53) - 1;
/**
* Bias adaptation (RFC 3492 section 6.1). delta/numpoints are always
* non-negative, so intdiv() matches the reference integer division exactly.
*/
function punycode_adapt(int $delta, int $numpoints, bool $firsttime): int
{
$d = $firsttime ? intdiv($delta, PUNY_DAMP) : intdiv($delta, 2);
$d += intdiv($d, $numpoints);
$k = 0;
while ($d > intdiv((PUNY_BASE - PUNY_TMIN) * PUNY_TMAX, 2)) {
$d = intdiv($d, PUNY_BASE - PUNY_TMIN);
$k += PUNY_BASE;
}
return $k + intdiv((PUNY_BASE - PUNY_TMIN + 1) * $d, $d + PUNY_SKEW);
}
/**
* Map a digit value (0–35) to its RFC 3492 base-36 character (lowercase a–z for
* 0–25, 0–9 for 26–35). $d is always in [0,35] on the encode path.
*/
function punycode_digit_to_char(int $d): string
{
if ($d < 26) {
return chr(ord('a') + $d); // a–z
}
return chr(ord('0') + ($d - 26)); // 0–9
}
/**
* Map a code point to its digit value (0–35), case-insensitive, or -1 if it is
* not a valid base-36 digit. Any non-ASCII code point yields -1, mirroring the
* reference behavior of rejecting non-ASCII in the extension portion.
*/
function punycode_char_to_digit(int $cp): int
{
if ($cp >= 0x61 && $cp <= 0x7A) { // a–z
return $cp - 0x61;
}
if ($cp >= 0x41 && $cp <= 0x5A) { // A–Z
return $cp - 0x41;
}
if ($cp >= 0x30 && $cp <= 0x39) { // 0–9
return $cp - 0x30 + 26;
}
return -1;
}
/**
* True if the string contains any non-ASCII code point (>= 128).
*/
function punycode_has_non_ascii(string $s): bool
{
foreach (mb_str_split($s, 1, 'UTF-8') ?: [] as $ch) {
if (mb_ord($ch, 'UTF-8') >= 128) {
return true;
}
}
return false;
}
/**
* Safely turn a decoded code point into a character. Mirrors Go's rune(n)
* (which emits U+FFFD for an out-of-range or surrogate value) rather than
* returning false.
*/
function punycode_cp_to_char(int $n): string
{
return mb_chr($n, 'UTF-8') ?: "\u{FFFD}";
}
/**
* Punycode-encode a single label (RFC 3492) and return the encoded label with
* no ACE prefix. Basic (ASCII) code points are emitted first, followed by a '-'
* delimiter (only if there was at least one), then the generalized base-36
* deltas for the non-basic code points. Twin of encodeLabel() in the TS/Go.
*/
function punycode_encode_label(string $input): string
{
// Iterate by code point so astral characters are single elements.
$code_points = mb_str_split($input, 1, 'UTF-8') ?: [];
$length = count($code_points);
$output = [];
foreach ($code_points as $ch) {
if (mb_ord($ch, 'UTF-8') < 128) {
$output[] = $ch;
}
}
$b = count($output);
if ($b > 0) {
$output[] = '-';
}
$n = PUNY_INITIAL_N;
$delta = 0;
$bias = PUNY_INITIAL_BIAS;
$h = $b;
while ($h < $length) {
// Smallest code point in the input that is >= n. The while guard ensures
// at least one such code point remains, so $m is always set.
$m = -1;
foreach ($code_points as $ch) {
$c = mb_ord($ch, 'UTF-8');
if ($c >= $n && ($m === -1 || $c < $m)) {
$m = $c;
}
}
$delta += ($m - $n) * ($h + 1);
$n = $m;
foreach ($code_points as $ch) {
$c = mb_ord($ch, 'UTF-8');
if ($c < $n) {
$delta += 1;
} elseif ($c === $n) {
$q = $delta;
$k = PUNY_BASE;
while (true) {
$t = max(PUNY_TMIN, min(PUNY_TMAX, $k - $bias));
if ($q < $t) {
break;
}
$output[] = punycode_digit_to_char($t + ($q - $t) % (PUNY_BASE - $t));
$q = intdiv($q - $t, PUNY_BASE - $t);
$k += PUNY_BASE;
}
$output[] = punycode_digit_to_char($q);
$bias = punycode_adapt($delta, $h + 1, $h === $b);
$delta = 0;
$h += 1;
}
}
$delta += 1;
$n += 1;
}
return implode('', $output);
}
/**
* Punycode-decode a single label (RFC 3492). Returns the decoded label, or null
* if the input is malformed (invalid digit, truncated generalized number,
* non-ASCII in the basic portion, or arithmetic overflow). Twin of
* decodeLabel() in the TS (string | null) and Go (bool form).
*/
function punycode_decode_label(string $input): ?string
{
// mb_strrpos on the ASCII '-' gives the code-point index of the last dash;
// since '-' is ASCII it coincides with the byte index the Go twin uses.
$last_dash = mb_strrpos($input, '-', 0, 'UTF-8');
$output = [];
if ($last_dash !== false) {
$basic = mb_substr($input, 0, $last_dash, 'UTF-8');
foreach (mb_str_split($basic, 1, 'UTF-8') ?: [] as $ch) {
if (mb_ord($ch, 'UTF-8') >= 128) {
return null; // basic portion must be ASCII
}
$output[] = $ch;
}
}
$ext = ($last_dash !== false)
? mb_substr($input, $last_dash + 1, null, 'UTF-8')
: $input;
$ext_chars = mb_str_split($ext, 1, 'UTF-8') ?: [];
$n = PUNY_INITIAL_N;
$i = 0;
$bias = PUNY_INITIAL_BIAS;
$pos = 0;
$ext_len = count($ext_chars);
while ($pos < $ext_len) {
$oldi = $i;
$w = 1;
$k = PUNY_BASE;
while (true) {
if ($pos >= $ext_len) {
return null; // truncated generalized number
}
$digit = punycode_char_to_digit(mb_ord($ext_chars[$pos], 'UTF-8'));
if ($digit < 0) {
return null; // invalid digit
}
$pos += 1;
if ($digit >= intdiv(PUNY_MAX_INT, $w)) {
return null; // overflow guard
}
$i += $digit * $w;
$t = max(PUNY_TMIN, min(PUNY_TMAX, $k - $bias));
if ($digit < $t) {
break;
}
$w *= PUNY_BASE - $t;
$k += PUNY_BASE;
}
$bias = punycode_adapt($i - $oldi, count($output) + 1, $oldi === 0);
$out_len = count($output) + 1;
$n += intdiv($i, $out_len);
$i = $i % $out_len;
// Insert the decoded code point n at index i (≡ output.splice(i, 0, …)).
array_splice($output, $i, 0, [punycode_cp_to_char($n)]);
$i += 1;
}
return implode('', $output);
}
/**
* IDNA toASCII: encode a domain to Punycode ("xn--") form. Lowercases the whole
* domain, splits on ".", ACE-encodes any label containing a non-ASCII code
* point, leaves ASCII-only labels untouched, and rejoins with ".". Empty input
* returns empty. Twin of encode() in the TS/Go.
*/
function punycode_encode(string $domain): string
{
if ($domain === '') {
return '';
}
$lower = mb_strtolower($domain, 'UTF-8');
$labels = explode('.', $lower);
foreach ($labels as $idx => $label) {
if (punycode_has_non_ascii($label)) {
$labels[$idx] = PUNY_ACE_PREFIX . punycode_encode_label($label);
}
}
return implode('.', $labels);
}
/**
* IDNA toUnicode: decode a Punycode ("xn--") domain back to Unicode. Splits on
* ".", decodes any label beginning with "xn--" (case-insensitive, detected on
* the lowercased label), leaves every other label untouched, and rejoins with
* ".". Returns null if any "xn--" label is invalid — the whole domain is
* rejected, matching IDNA semantics. Empty input returns "". Twin of decode()
* in the TS/Go.
*/
function punycode_decode(string $domain): ?string
{
if ($domain === '') {
return '';
}
$out = [];
foreach (explode('.', $domain) as $label) {
$lower = mb_strtolower($label, 'UTF-8');
if (str_starts_with($lower, PUNY_ACE_PREFIX)
&& mb_strlen($label, 'UTF-8') > mb_strlen(PUNY_ACE_PREFIX, 'UTF-8')) {
$decoded = punycode_decode_label(
mb_substr($label, mb_strlen(PUNY_ACE_PREFIX, 'UTF-8'), null, 'UTF-8')
);
if ($decoded === null) {
return null;
}
$out[] = $decoded;
} else {
$out[] = $label;
}
}
return implode('.', $out);
}
// Showcase vectors — run only when this file is executed directly (not when
// included as a library). The debug_backtrace() check is the idiomatic PHP
// "am I the entry script?" guard. Requires zend.assertions >= 0 (the default).
if (!debug_backtrace(DEBUG_BACKTRACE_IGNORE_ARGS)) {
assert(punycode_encode('münchen.de') === 'xn--mnchen-3ya.de');
assert(punycode_decode('xn--mnchen-3ya.de') === 'münchen.de');
assert(punycode_encode('Bücher.DE') === 'xn--bcher-kva.de'); // lowercased first
assert(punycode_encode_label('café') === 'caf-dma');
assert(punycode_decode_label('caf-dma') === 'café');
assert(punycode_decode('xn--!') === null); // invalid digit
assert(punycode_decode_label(str_repeat('9', 40)) === null); // overflow guard
assert(punycode_decode(punycode_encode('café.fr')) === 'café.fr');
echo "punycode: all showcase vectors passed\n";
}
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →