Skip to content

Punycode Converter — Ruby source

Convert internationalized domain names (IDN) between Unicode and Punycode (xn--) ACE form. RFC 3492 compliant, runs entirely in your browser, with a shareable link to your exact input.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# punycode - RFC 3492 Punycode encode/decode + IDNA2003 toASCII/toUnicode.
#
# Language:   Ruby (2.7+, standard library only)
# Source:     CosmoDev polyglot showcase port of the Punycode tool, ported from
#             src/lib/punycode.ts (the canonical TypeScript implementation) and
#             cli/punycode/punycode.go (the live Go CLI twin).
# License:    display source - part of CosmoDev's polyglot tool pages.
#
# Design goals:
#   - Pure + deterministic; encode never raises, decode returns nil on
#     malformed input.
#   - Functionally equivalent to the TS/Go reference: same inputs -> same
#     outputs.
#   - Self-contained: stdlib only. String#each_char yields full characters
#     (Ruby strings are encoding-aware), so astral characters (emoji, CJK
#     extensions) are single elements, matching Go runes and the TS
#     code-point iteration; Integer carries the RFC 3492 arithmetic (unbounded
#     in Ruby, so the 2^53-1 / Number.MAX_SAFE_INTEGER guard is kept as an
#     explicit rejection of absurd inputs); code points beyond U+10FFFF reject
#     the label.

module Punycode
  module_function

  BASE = 36
  TMIN = 1
  TMAX = 26
  SKEW = 38
  DAMP = 700
  INITIAL_BIAS = 72
  INITIAL_N = 128
  ACE_PREFIX = 'xn--'
  MAX_INT = 9_007_199_254_740_991 # 2^53-1 overflow guard (Number.MAX_SAFE_INTEGER)
  MAX_CODEPOINT = 0x10FFFF

  # Bias adaptation (RFC 3492 section 6.1).
  def adapt(delta, numpoints, firsttime)
    d = firsttime ? delta / DAMP : delta / 2 # Integer#/ floors
    d += d / numpoints
    k = 0
    while d > (BASE - TMIN) * TMAX / 2
      d /= BASE - TMIN
      k += BASE
    end
    k + (BASE - TMIN + 1) * d / (d + SKEW)
  end

  # Map a digit value (0-35) to its base-36 character (lowercase).
  def digit_to_char(d)
    d < 26 ? ('a'.ord + d).chr : ('0'.ord + (d - 26)).chr
  end

  # Map a character to its digit value (0-35), case-insensitive, or -1 if invalid.
  def char_to_digit(c)
    o = c.ord
    return o - 97   if o.between?(97, 122)  # a-z
    return o - 65   if o.between?(65, 90)   # A-Z
    return o - 48 + 26 if o.between?(48, 57) # 0-9

    -1
  end

  # True if the string contains any non-ASCII code point (>= 128).
  def has_non_ascii(s)
    s.each_char.any? { |c| c.ord >= 128 }
  end

  # Punycode-encode a single label (RFC 3492), no ACE prefix. Basic code
  # points are emitted first, then a '-' delimiter (only if there was at least
  # one), then the generalized-base-36 deltas.
  def encode_label(input)
    cps = input.each_char.map(&:ord) # astral chars are one element each
    length = cps.length

    output = +''
    cps.each { |cp| output << cp.chr(Encoding::UTF_8) if cp < 128 }
    b = output.length
    output << '-' if b.positive?

    n = INITIAL_N
    delta = 0
    bias = INITIAL_BIAS
    h = b
    while h < length
      m = cps.select { |cp| cp >= n }.min # smallest code point >= n
      delta += (m - n) * (h + 1)
      n = m
      cps.each do |cp|
        if cp < n
          delta += 1
        elsif cp == n
          q = delta
          k = BASE
          loop do
            t = [TMIN, [TMAX, k - bias].min].max
            break if q < t

            output << digit_to_char(t + (q - t) % (BASE - t))
            q = (q - t) / (BASE - t)
            k += BASE
          end
          output << digit_to_char(q)
          bias = adapt(delta, h + 1, h == b)
          delta = 0
          h += 1
        end
      end
      delta += 1
      n += 1
    end

    output
  end

  # Punycode-decode a single label (RFC 3492). Returns nil when the input is
  # malformed (invalid digit, truncated generalized number, non-ASCII in the
  # basic portion, code point beyond U+10FFFF, or overflow).
  def decode_label(input)
    last_dash = input.rindex('-')
    output = []
    if last_dash
      input[0...last_dash].each_char do |c|
        return nil if c.ord >= 128 # basic portion must be ASCII

        output << c
      end
    end
    ext = last_dash ? input[(last_dash + 1)..] : input

    n = INITIAL_N
    i = 0
    bias = INITIAL_BIAS
    pos = 0
    while pos < ext.length
      oldi = i
      w = 1
      k = BASE
      loop do
        return nil if pos >= ext.length # truncated generalized number
        digit = char_to_digit(ext[pos])
        return nil if digit.negative? # invalid digit

        pos += 1
        return nil if digit >= MAX_INT / w # overflow guard

        i += digit * w
        t = [TMIN, [TMAX, k - bias].min].max
        break if digit < t

        w *= BASE - t
        k += BASE
      end
      bias = adapt(i - oldi, output.length + 1, oldi.zero?)
      out_len = output.length + 1
      n += i / out_len
      i %= out_len
      return nil if n > MAX_CODEPOINT

      output.insert(i, [n].pack('U'))
      i += 1
    end

    output.join
  end

  # IDNA toASCII: lowercase the domain, ACE-encode ("xn--" + Punycode) any
  # label containing a non-ASCII code point, leave ASCII-only labels
  # untouched. Empty input returns empty. Split with limit -1 keeps empty
  # labels, like the TS split('.').
  def encode(domain)
    return '' if domain.empty?

    domain.downcase.split('.', -1).map do |label|
      has_non_ascii(label) ? ACE_PREFIX + encode_label(label) : label
    end.join('.')
  end

  # IDNA toUnicode: decode any "xn--" label (case-insensitive, prefix detected
  # on the lowercased label), leave every other label untouched. Returns nil
  # when any "xn--" label is invalid - the whole domain is rejected, matching
  # IDNA semantics. Empty input returns empty.
  def decode(domain)
    return '' if domain.empty?

    domain.split('.', -1).map do |label|
      if label.downcase.start_with?(ACE_PREFIX) && label.length > ACE_PREFIX.length
        decoded = decode_label(label[ACE_PREFIX.length..])
        return nil if decoded.nil? # reject the whole domain on any invalid label

        decoded
      else
        label
      end
    end.join('.')
  end
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →