Skip to content

Base32 / Base58 / Base62 / Base85 Encoder — Ruby source

Encode text to Base32, Base58, Base62, or Ascii85 - or decode it back. UTF-8 safe, runs entirely in your browser, with a shareable link to your exact input.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# base-encoder — Base32 (RFC 4648), Base58 (Bitcoin), Base62, and Base85
# (Ascii85) byte-array encoders, operating on the UTF-8 bytes of the input
# text.
#
# Language: Ruby (3.2+, standard library only)
# Source:   CosmoDev polyglot showcase port of the Base Encoder tool, ported
#           from cli/base-encoder/base-encoder.go (the authoritative Go twin).
# License:  display source — part of CosmoDev's polyglot tool pages.
#
# Design goals:
#   - Pure + deterministic; never raises (decode returns nil for invalid or
#     malformed input, mirroring the TS lib's `null` and the Go twin's
#     `errInvalid`).
#   - Functionally equivalent to the Go twin: same inputs -> same outputs.
#   - Self-contained: stdlib only (no gems).
#
# Arbitrary-precision note: Base58 and Base62 base-convert the whole byte
# array. Ruby's Integer is arbitrary-precision (like Python's int and Go's
# math/big), so we get the exact same semantics for free — no manual bignum
# code (unlike the dependency-free Rust/C/C++ siblings).

module BaseEncoder
  B32_ALPHABET = "ABCDEFGHIJKLMNOPQRSTUVWXYZ234567".freeze
  B58_ALPHABET =
    "123456789ABCDEFGHJKLMNPQRSTUVWXYZabcdefghijkmnopqrstuvwxyz".freeze
  B62_ALPHABET =
    "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz".freeze

  # Data characters emitted by a final (partial) 5-byte chunk before '='
  # padding, per RFC 4648. Index = byte count (0..4). Matches the TS `outLen`
  # table.
  OUT_LEN_32 = [0, 2, 4, 5, 7].freeze

  # byte value -> alphabet index, for decode lookups (a missing byte is an
  # invalid character -> nil, mirroring the TS/Python `.indexOf` / `.find`).
  B32_INDEX = B32_ALPHABET.bytes.each_with_index.to_h.freeze
  B58_INDEX = B58_ALPHABET.bytes.each_with_index.to_h.freeze
  B62_INDEX = B62_ALPHABET.bytes.each_with_index.to_h.freeze

  module_function

  # Integer -> minimal big-endian bytes (matches Go's big.Int.Bytes() and
  # Python's int.to_bytes). Ruby has no direct int-to-bytes, so peel bytes off
  # the little end and reverse.
  def int_to_be_bytes(num)
    return +"".b if num.zero?

    out = +"".b
    while num.positive?
      out << (num & 0xFF)
      num >>= 8
    end
    out.reverse
  end

  # -------------------------------------------------------------------------
  # Base32 — RFC 4648 alphabet, padded to a multiple of 8 chars with '='.
  # -------------------------------------------------------------------------

  def encode32(data)
    out = +""
    i = 0
    while i < data.bytesize
      chunk = data.byteslice(i, 5)
      n = chunk.bytesize
      b = chunk.bytes
      b << 0 while b.size < 5 # zero-pad the final partial chunk
      # Pack 5 bytes (40 bits) into 8 base32 digits (5 bits each, big-endian).
      digits = [
        (b[0] >> 3) & 0x1F,
        ((b[0] << 2) | (b[1] >> 6)) & 0x1F,
        (b[1] >> 1) & 0x1F,
        ((b[1] << 4) | (b[2] >> 4)) & 0x1F,
        ((b[2] << 1) | (b[3] >> 7)) & 0x1F,
        (b[3] >> 2) & 0x1F,
        ((b[3] << 3) | (b[4] >> 5)) & 0x1F,
        b[4] & 0x1F,
      ]
      if n == 5
        out << digits.map { |d| B32_ALPHABET[d] }.join
      else
        out_len = OUT_LEN_32[n]
        # Truncate to the data chars and pad with '=' to 8 total.
        out << digits[0, out_len].map { |d| B32_ALPHABET[d] }.join << ("=" * (8 - out_len))
      end
      i += 5
    end
    out
  end

  def decode32(s)
    out = +"".b
    buffer = 0
    bits = 0
    s.each_byte do |c|
      break if c == 61 # '=' — padding marks the end

      idx = B32_INDEX[c]
      return nil if idx.nil?

      buffer = (buffer << 5) | idx
      bits += 5
      if bits >= 8
        bits -= 8
        out << ((buffer >> bits) & 0xFF)
        buffer &= (1 << bits) - 1 # keep only the leftover bits
      end
    end
    out
  end

  # -------------------------------------------------------------------------
  # Base58 — Bitcoin alphabet. Leading 0x00 bytes -> leading '1' (count
  # preserved).
  # -------------------------------------------------------------------------

  def encode58(data)
    # Count leading zero bytes — each maps to a leading '1'.
    zeros = 0
    zeros += 1 while zeros < data.bytesize && data.getbyte(zeros).zero?
    # Big-endian byte array (skipping the leading zeros) -> Integer.
    num = 0
    (zeros...data.bytesize).each { |i| num = (num << 8) | data.getbyte(i) }
    # Base-convert to 58 digits (collected least-significant first).
    digits = []
    while num.positive?
      num, rem = num.divmod(58)
      digits << rem
    end
    ("1" * zeros) + digits.reverse.map { |d| B58_ALPHABET[d] }.join
  end

  def decode58(s)
    # Count leading '1's — each maps to a 0x00 byte.
    zeros = 0
    zeros += 1 while zeros < s.length && s[zeros] == "1"
    num = 0
    s[zeros..].each_char do |c|
      idx = B58_INDEX[c.ord]
      return nil if idx.nil?

      num = num * 58 + idx
    end
    out = +"".b
    zeros.times { out << 0 }
    out << int_to_be_bytes(num)
    out
  end

  # -------------------------------------------------------------------------
  # Base62 — standard base-conversion of the byte array (no leading-zero
  # special-casing beyond the standard big-int).
  # -------------------------------------------------------------------------

  def encode62(data)
    return +"" if data.bytesize.zero?

    num = 0
    data.each_byte { |b| num = (num << 8) | b }
    return +"0" if num.zero?

    digits = []
    while num.positive?
      num, rem = num.divmod(62)
      digits << rem
    end
    digits.reverse.map { |d| B62_ALPHABET[d] }.join
  end

  def decode62(s)
    return +"".b if s.empty?

    num = 0
    s.each_char do |c|
      idx = B62_INDEX[c.ord]
      return nil if idx.nil?

      num = num * 62 + idx
    end
    int_to_be_bytes(num)
  end

  # -------------------------------------------------------------------------
  # Base85 — Ascii85. 4 bytes -> 5 chars in '!'(33)..'u'(117); a full 4-zero
  # group is shortened to 'z'. No <~ ~> delimiters. Partial final groups emit
  # one fewer char than (bytes+1) would suggest; decode reverses, padding
  # with 'u' (value 84).
  # -------------------------------------------------------------------------

  def encode85(data)
    out = +""
    i = 0
    while i < data.bytesize
      chunk = data.byteslice(i, 4)
      n = chunk.bytesize
      is_full = n == 4
      b = chunk.bytes
      b << 0 while b.size < 4 # zero-pad the final partial group
      u = b[0] * 16_777_216 + b[1] * 65_536 + b[2] * 256 + b[3]
      i += 4
      if is_full && u.zero?
        out << "z" # zero-group shorthand
        next
      end
      digits = [0, 0, 0, 0, 0]
      v = u
      4.downto(0) do |k|
        digits[k] = v % 85
        v /= 85
      end
      emit = is_full ? 5 : n + 1 # n bytes -> n+1 chars
      out << digits[0, emit].map { |d| (d + 33).chr }.join
    end
    out
  end

  def decode85(s)
    out = +"".b
    group = [] # accumulated digit values (0..84)
    s.each_byte do |c|
      if c == 122 # 'z' — only valid at a group boundary (empty accumulator)
        return nil unless group.empty?

        out << 0 << 0 << 0 << 0
        next
      end
      return nil if c < 33 || c > 117

      group << (c - 33)
      next unless group.size == 5

      v = 0
      group.each { |d| v = v * 85 + d }
      return nil if v > 0xFFFFFFFF # a 5-char group must fit in 32 bits

      out << ((v >> 24) & 0xFF) << ((v >> 16) & 0xFF) << ((v >> 8) & 0xFF) << (v & 0xFF)
      group.clear
    end
    # Handle a partial final group (2-4 chars -> 1-3 bytes).
    unless group.empty?
      m = group.size
      return nil if m < 2 # a lone trailing char is malformed

      group << 84 while group.size < 5 # pad with 'u'
      v = 0
      group.each { |d| v = v * 85 + d }
      return nil if v > 0xFFFFFFFF

      all = [(v >> 24) & 0xFF, (v >> 16) & 0xFF, (v >> 8) & 0xFF, v & 0xFF]
      out << all[0, m - 1].pack("C*")
    end
    out
  end

  # -------------------------------------------------------------------------
  # Public API
  # -------------------------------------------------------------------------

  # Dispatch raw bytes to the chosen scheme's encoder. Mirrors the Go twin's
  # private `encodeBytes`.
  def encode_bytes(data, scheme)
    case scheme
    when :base32 then encode32(data)
    when :base58 then encode58(data)
    when :base62 then encode62(data)
    when :base85 then encode85(data)
    end
  end

  # Dispatch an encoded string to the chosen scheme's decoder. An invalid or
  # malformed input yields nil (mirroring the TS `null`). Mirrors the Go
  # twin's private `decodeBytes`.
  def decode_bytes(s, scheme)
    case scheme
    when :base32 then decode32(s)
    when :base58 then decode58(s)
    when :base62 then decode62(s)
    when :base85 then decode85(s)
    end
  end

  # Encode the UTF-8 bytes of `text` per `scheme`. Empty text -> "".
  # Mirrors `Encode` in cli/base-encoder/base-encoder.go.
  def encode(text, scheme)
    encode_bytes(text.b, scheme)
  end

  # Decode `encoded` back to UTF-8 text. Invalid chars / malformed -> nil
  # (mirrors the Go twin's `errInvalid` and the TS lib's `null`).
  # Mirrors `Decode` in cli/base-encoder/base-encoder.go.
  def decode(s, scheme)
    data = decode_bytes(s.b, scheme)
    return nil if data.nil?

    # Lossy so a structurally-valid-but-non-UTF-8 payload never raises a
    # second error — scrub replaces invalid bytes with U+FFD, mirroring Go's
    # string(data) (which never fails) and Python's errors="replace".
    data.force_encoding("UTF-8").scrub
  end
end

# ---------------------------------------------------------------------------
# Showcase self-test — mirrors cli/base-encoder/base-encoder_test.go vectors.
# Run directly: `ruby ruby.rb`
# ---------------------------------------------------------------------------
if __FILE__ == $PROGRAM_NAME
  nul = "\u0000"

  # Base32 — known values + RFC 4648 padding + case sensitivity.
  raise "b32 hello" unless BaseEncoder.encode("hello", :base32) == "NBSWY3DP"
  # 3 bytes -> 5 chars + 3 '='
  raise "b32 foo" unless BaseEncoder.encode("foo", :base32) == "MZXW6==="
  raise "b32 decode" unless BaseEncoder.decode("NBSWY3DP", :base32) == "hello"
  # lowercase not in RFC 4648
  raise "b32 lowercase" unless BaseEncoder.decode("nbswy3dp", :base32).nil?

  # Base58 — each leading 0x00 byte -> a leading '1'.
  raise "b58 zero byte" unless BaseEncoder.encode(nul, :base58) == "1"
  raise "b58 two zeros" unless BaseEncoder.encode(nul + nul + "A", :base58).start_with?("11")
  raise "b58 decode 1" unless BaseEncoder.decode("1", :base58) == nul
  # round-trip preserves the leading zero bytes exactly
  raise "b58 round-trip" unless BaseEncoder.decode(BaseEncoder.encode(nul + nul + "A", :base58), :base58) == nul + nul + "A"

  # Base62 — plain big-int base conversion (no leading-zero preservation).
  raise "b62 A" unless BaseEncoder.encode("A", :base62) == "13" # 1*62 + 3
  raise "b62 decode" unless BaseEncoder.decode("13", :base62) == "A"
  raise "b62 zero" unless BaseEncoder.encode(nul, :base62) == "0"
  # minimal rep of 0 is empty
  raise "b62 minimal zero" unless BaseEncoder.decode("0", :base62) == ""

  # Base85 — Ascii85 'z' shorthand + 32-bit overflow rejection.
  raise "b85 hello" unless BaseEncoder.encode("hello", :base85) == "BOu!rDZ"
  raise "b85 z" unless BaseEncoder.encode(nul * 4, :base85) == "z"
  raise "b85 zz" unless BaseEncoder.encode(nul * 8, :base85) == "zz"
  # 5-char group overflows 32 bits
  raise "b85 overflow" unless BaseEncoder.decode("uuuuu", :base85).nil?
  # lone trailing char is malformed
  raise "b85 lone char" unless BaseEncoder.decode("B", :base85).nil?

  # Cross-scheme — empty, multibyte round-trip, and invalid rejection.
  %i[base32 base58 base62 base85].each do |scheme|
    raise "#{scheme} empty" unless BaseEncoder.encode("", scheme) == ""
    raise "#{scheme} empty decode" unless BaseEncoder.decode("", scheme) == ""
    # multibyte UTF-8 round-trips through every scheme
    raise "#{scheme} multibyte" unless BaseEncoder.decode(BaseEncoder.encode("CosmoDev 🚀", scheme), scheme) == "CosmoDev 🚀"
    # '~' outside every alphabet
    raise "#{scheme} invalid" unless BaseEncoder.decode("~!not-valid!~", scheme).nil?
  end

  puts "ok"
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →