Skip to content

Hex ↔ Text Converter — Ruby source

Convert text to hexadecimal and hex back to text, with delimiter options (none, spaces, 0x, backslash-x) and full UTF-8 support. 100% client-side.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# hex-converter — pure hex ↔ text conversion.
#
# Language: Ruby (3.2+, standard library only)
# Source:   CosmoDev polyglot showcase port of the hex-converter tool,
#           ported from src/lib/hexText.ts (the canonical TypeScript
#           implementation).
# License:  display source — part of CosmoDev's polyglot tool pages
#           (dev.cosmolabs.org). Deterministic, side-effect free; invalid
#           byte sequences decode to U+FFFD, matching the canonical logic.

module HexConverter
  # How encoded bytes are joined when rendered as a hex string.
  DELIM_NONE = "none"
  DELIM_SPACE = "space"
  DELIM_0X = "0x"
  DELIM_BACKSLASH_X = "backslash-x"

  # U+FFFD, substituted for malformed UTF-8 on decode.
  REPLACEMENT_CHAR = "�"

  # Outcome of decoding hex back to text. Mirrors the canonical TS surface
  # (ok / text / error) so the shape is identical across every language in
  # the polyglot showcase.
  DecodeResult = Struct.new(:ok, :text, :error, keyword_init: true) do
    def self.ok_text(text)
      new(ok: true, text: text, error: nil)
    end

    def self.fail(message)
      new(ok: false, text: "", error: message)
    end
  end

  module_function

  # UTF-8 encode a Ruby string into an array of byte values (0..255).
  #
  # Hand-rolled for byte-exact parity across every showcase language. Ruby
  # strings iterate by character natively, so astral characters encode as
  # 4-byte sequences.
  def utf8_encode(text)
    bytes = []
    text.each_char do |ch|
      cp = ch.ord
      if cp <= 0x7F
        bytes << cp
      elsif cp <= 0x7FF
        bytes << (0xC0 | (cp >> 6))
        bytes << (0x80 | (cp & 0x3F))
      elsif cp <= 0xFFFF
        bytes << (0xE0 | (cp >> 12))
        bytes << (0x80 | ((cp >> 6) & 0x3F))
        bytes << (0x80 | (cp & 0x3F))
      else
        bytes << (0xF0 | (cp >> 18))
        bytes << (0x80 | ((cp >> 12) & 0x3F))
        bytes << (0x80 | ((cp >> 6) & 0x3F))
        bytes << (0x80 | (cp & 0x3F))
      end
    end
    bytes
  end

  # Render a single code point as a String, substituting U+FFFD for any value
  # that is not a valid Unicode scalar (surrogates or out of range). Array#pack
  # with "U" raises on such values; substituting instead keeps the decoder
  # total, consistent with its stated "invalid → U+FFFD" contract.
  def char_from_code_point(cp)
    if cp.between?(0, 0x10FFFF) && !cp.between?(0xD800, 0xDFFF)
      [cp].pack("U")
    else
      REPLACEMENT_CHAR
    end
  end

  # UTF-8 decode an array of bytes into a string.
  #
  # Truncated or invalid sequences yield U+FFFD; missing continuation bytes
  # default to 0, mirroring the canonical decoder's lenient reads.
  def utf8_decode(bytes_in)
    out = +""
    i = 0
    n = bytes_in.length

    next_byte = lambda do
      return 0 if i >= n

      b = bytes_in[i]
      i += 1
      b
    end

    while i < n
      b = bytes_in[i]
      i += 1
      cp =
        if b <= 0x7F
          b
        elsif (b >> 5) == 0b110
          ((b & 0x1F) << 6) | (next_byte.call & 0x3F)
        elsif (b >> 4) == 0b1110
          b1 = next_byte.call
          b2 = next_byte.call
          ((b & 0x0F) << 12) | ((b1 & 0x3F) << 6) | (b2 & 0x3F)
        elsif (b >> 3) == 0b11110
          b1 = next_byte.call
          b2 = next_byte.call
          b3 = next_byte.call
          ((b & 0x07) << 18) | ((b1 & 0x3F) << 12) | ((b2 & 0x3F) << 6) | (b3 & 0x3F)
        else
          0xFFFD
        end
      out << char_from_code_point(cp)
    end
    out
  end

  # Render text as a hex string.
  #
  # +delimiter+ controls how per-byte hex pairs are joined:
  #   - 'none'        -> "48656c6c6f"
  #   - 'space'       -> "48 65 6c 6c 6f"
  #   - '0x'          -> "0x48 0x65 ..."
  #   - 'backslash-x' -> "\x48\x65..." (no separators, C-style)
  # Unknown delimiters fall back to no delimiter.
  def text_to_hex(text, delimiter: DELIM_NONE, uppercase: false)
    hexes = utf8_encode(text).map { |b| format("%02x", b) }
    hexes = hexes.map(&:upcase) if uppercase
    case delimiter
    when DELIM_NONE        then hexes.join
    when DELIM_SPACE       then hexes.join(" ")
    when DELIM_0X          then hexes.map { |h| "0x#{h}" }.join(" ")
    when DELIM_BACKSLASH_X then hexes.map { |h| "\\x#{h}" }.join
    else hexes.join
    end
  end

  # Strip common affixes users paste alongside hex, then lowercase.
  #
  # Removes "0x" and "\x" literals (case-insensitive, anywhere), whitespace,
  # commas, and colons (MAC-style "aa:bb:cc"). Ruby's [[:space:]] POSIX class
  # is Unicode-aware, matching the canonical TS \s.
  def sanitize_hex(text)
    no_markers = text.to_s.gsub(/0x/i, "").gsub(/\\x/i, "")
    no_markers.gsub(/[[:space:],:]/, "").downcase
  end

  # Decode a (possibly decorated) hex string back to text.
  #
  # Invalid characters and odd lengths are reported via +error+; valid input
  # that contains malformed UTF-8 still decodes with U+FFFD substitution.
  def hex_to_text(hex_str, _delimiter = DELIM_NONE)
    cleaned = sanitize_hex(hex_str)
    return DecodeResult.ok_text("") if cleaned.empty?
    unless cleaned.match?(/\A[0-9a-f]+\z/)
      return DecodeResult.fail("Hex strings may only contain 0-9 and a-f.")
    end
    unless cleaned.length.even?
      return DecodeResult.fail("Hex must have an even number of digits.")
    end

    bytes_out = cleaned.scan(/../).map { |pair| pair.to_i(16) }
    DecodeResult.ok_text(utf8_decode(bytes_out))
  end
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →