Punycode Converter — Ruby source
Convert internationalized domain names (IDN) between Unicode and Punycode (xn--) ACE form. RFC 3492 compliant, runs entirely in your browser, with a shareable link to your exact input.
This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.
# punycode - RFC 3492 Punycode encode/decode + IDNA2003 toASCII/toUnicode.
#
# Language: Ruby (2.7+, standard library only)
# Source: CosmoDev polyglot showcase port of the Punycode tool, ported from
# src/lib/punycode.ts (the canonical TypeScript implementation) and
# cli/punycode/punycode.go (the live Go CLI twin).
# License: display source - part of CosmoDev's polyglot tool pages.
#
# Design goals:
# - Pure + deterministic; encode never raises, decode returns nil on
# malformed input.
# - Functionally equivalent to the TS/Go reference: same inputs -> same
# outputs.
# - Self-contained: stdlib only. String#each_char yields full characters
# (Ruby strings are encoding-aware), so astral characters (emoji, CJK
# extensions) are single elements, matching Go runes and the TS
# code-point iteration; Integer carries the RFC 3492 arithmetic (unbounded
# in Ruby, so the 2^53-1 / Number.MAX_SAFE_INTEGER guard is kept as an
# explicit rejection of absurd inputs); code points beyond U+10FFFF reject
# the label.
module Punycode
module_function
BASE = 36
TMIN = 1
TMAX = 26
SKEW = 38
DAMP = 700
INITIAL_BIAS = 72
INITIAL_N = 128
ACE_PREFIX = 'xn--'
MAX_INT = 9_007_199_254_740_991 # 2^53-1 overflow guard (Number.MAX_SAFE_INTEGER)
MAX_CODEPOINT = 0x10FFFF
# Bias adaptation (RFC 3492 section 6.1).
def adapt(delta, numpoints, firsttime)
d = firsttime ? delta / DAMP : delta / 2 # Integer#/ floors
d += d / numpoints
k = 0
while d > (BASE - TMIN) * TMAX / 2
d /= BASE - TMIN
k += BASE
end
k + (BASE - TMIN + 1) * d / (d + SKEW)
end
# Map a digit value (0-35) to its base-36 character (lowercase).
def digit_to_char(d)
d < 26 ? ('a'.ord + d).chr : ('0'.ord + (d - 26)).chr
end
# Map a character to its digit value (0-35), case-insensitive, or -1 if invalid.
def char_to_digit(c)
o = c.ord
return o - 97 if o.between?(97, 122) # a-z
return o - 65 if o.between?(65, 90) # A-Z
return o - 48 + 26 if o.between?(48, 57) # 0-9
-1
end
# True if the string contains any non-ASCII code point (>= 128).
def has_non_ascii(s)
s.each_char.any? { |c| c.ord >= 128 }
end
# Punycode-encode a single label (RFC 3492), no ACE prefix. Basic code
# points are emitted first, then a '-' delimiter (only if there was at least
# one), then the generalized-base-36 deltas.
def encode_label(input)
cps = input.each_char.map(&:ord) # astral chars are one element each
length = cps.length
output = +''
cps.each { |cp| output << cp.chr(Encoding::UTF_8) if cp < 128 }
b = output.length
output << '-' if b.positive?
n = INITIAL_N
delta = 0
bias = INITIAL_BIAS
h = b
while h < length
m = cps.select { |cp| cp >= n }.min # smallest code point >= n
delta += (m - n) * (h + 1)
n = m
cps.each do |cp|
if cp < n
delta += 1
elsif cp == n
q = delta
k = BASE
loop do
t = [TMIN, [TMAX, k - bias].min].max
break if q < t
output << digit_to_char(t + (q - t) % (BASE - t))
q = (q - t) / (BASE - t)
k += BASE
end
output << digit_to_char(q)
bias = adapt(delta, h + 1, h == b)
delta = 0
h += 1
end
end
delta += 1
n += 1
end
output
end
# Punycode-decode a single label (RFC 3492). Returns nil when the input is
# malformed (invalid digit, truncated generalized number, non-ASCII in the
# basic portion, code point beyond U+10FFFF, or overflow).
def decode_label(input)
last_dash = input.rindex('-')
output = []
if last_dash
input[0...last_dash].each_char do |c|
return nil if c.ord >= 128 # basic portion must be ASCII
output << c
end
end
ext = last_dash ? input[(last_dash + 1)..] : input
n = INITIAL_N
i = 0
bias = INITIAL_BIAS
pos = 0
while pos < ext.length
oldi = i
w = 1
k = BASE
loop do
return nil if pos >= ext.length # truncated generalized number
digit = char_to_digit(ext[pos])
return nil if digit.negative? # invalid digit
pos += 1
return nil if digit >= MAX_INT / w # overflow guard
i += digit * w
t = [TMIN, [TMAX, k - bias].min].max
break if digit < t
w *= BASE - t
k += BASE
end
bias = adapt(i - oldi, output.length + 1, oldi.zero?)
out_len = output.length + 1
n += i / out_len
i %= out_len
return nil if n > MAX_CODEPOINT
output.insert(i, [n].pack('U'))
i += 1
end
output.join
end
# IDNA toASCII: lowercase the domain, ACE-encode ("xn--" + Punycode) any
# label containing a non-ASCII code point, leave ASCII-only labels
# untouched. Empty input returns empty. Split with limit -1 keeps empty
# labels, like the TS split('.').
def encode(domain)
return '' if domain.empty?
domain.downcase.split('.', -1).map do |label|
has_non_ascii(label) ? ACE_PREFIX + encode_label(label) : label
end.join('.')
end
# IDNA toUnicode: decode any "xn--" label (case-insensitive, prefix detected
# on the lowercased label), leave every other label untouched. Returns nil
# when any "xn--" label is invalid - the whole domain is rejected, matching
# IDNA semantics. Empty input returns empty.
def decode(domain)
return '' if domain.empty?
domain.split('.', -1).map do |label|
if label.downcase.start_with?(ACE_PREFIX) && label.length > ACE_PREFIX.length
decoded = decode_label(label[ACE_PREFIX.length..])
return nil if decoded.nil? # reject the whole domain on any invalid label
decoded
else
label
end
end.join('.')
end
end
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →