Skip to content

Text Extractor — Ruby source

Pull URLs, emails, IPv4/IPv6 addresses, hashes (MD5/SHA-1/SHA-256/SHA-512), and domains out of logs, headers, or any pasted text.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# extract — pull URLs, emails, IPv4/IPv6 addresses, hashes, and domains
# out of arbitrary text (logs, headers, config).
#
# Language: Ruby (3.2+, standard library only)
# Source:   CosmoDev polyglot showcase port of the Extract tool, ported from
#           src/lib/extract.ts (the canonical TypeScript implementation) and
#           held in lock-step with its Go twin cli/extract/extract.go.
# License:  display source — part of CosmoDev's polyglot tool pages.
#
# Design goals:
#   - Pure + deterministic; never raises.
#   - Functionally equivalent to the TS/Go reference: same inputs -> same outputs.
#   - Self-contained: stdlib only (Regexp ships in the language; no gems).
#
# Behavior: pull URLs, emails, IPv4/IPv6 addresses, hashes (md5/sha1/sha256/
# sha512 by length), and domains out of arbitrary text. Matches are deduped
# per type preserving first-occurrence order; an email also contributes its
# domain to the domain list when both email and domain are selected. A nil or
# empty types list defaults to all six kinds.

module Extract
  # Canonical extraction kinds, in display order. Mirrors EXTRACT_TYPES in TS.
  TYPES = %w[url email ipv4 ipv6 hash domain].freeze

  # Per-kind patterns, scanned with String#scan. These mirror the RE record
  # in src/lib/extract.ts (and the Go twin) verbatim: same anchors, classes,
  # and counted repetition. Every group is non-capturing (?:...) so scan
  # returns the full match string, not group arrays. The IPv6 pattern is
  # intentionally permissive — a hex/colon run — and is post-filtered by
  # ipv6? so bare hex words, times, and MACs are rejected, exactly as in the
  # TS lib.
  RE = {
    url:    /https?:\/\/[^\s]+/,
    email:  /[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}/,
    ipv4:   /\b(?:\d{1,3}\.){3}\d{1,3}\b/,
    ipv6:   /[0-9a-fA-F:]+/,
    # md5 (32) / sha1 (40) / sha256 (64) / sha512 (128). `\b` keeps each
    # length honest, so a 64-char run does not also match as a leading
    # 32-char hash.
    hash:   /\b[a-fA-F0-9]{32}\b|\b[a-fA-F0-9]{40}\b|
             \b[a-fA-F0-9]{64}\b|\b[a-fA-F0-9]{128}\b/x,
    domain: /\b[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z]{2,})+\b/
  }.freeze

  # Validates a single 1-4 hex-digit IPv6 hextet (full-string match).
  IPV6_GROUP = /\A[0-9a-fA-F]{1,4}\z/

  # One field per kind — always all six, populated only for the selected
  # types (unselected kinds stay empty arrays). Mirrors the TS
  # `Record<ExtractType, string[]>` and the Go twin's Result struct.
  Result = Struct.new(:url, :email, :ipv4, :ipv6, :hash, :domain) do
    def self.empty
      new([], [], [], [], [], [])
    end
  end

  class << self
    # Extract every occurrence of the given `types` (default: all six) from
    # `input`. Returns an Extract::Result with one field per kind — always
    # all six, populated only for the selected types. Matches are deduped per
    # kind, preserving first-occurrence order (Array#uniq keeps first
    # occurrences — the Go twin of the uniq() helper in src/lib/extract.ts).
    # An email also contributes its domain to the domain list when both email
    # and domain are selected. Never raises; nil input reads as empty text.
    def extract(input, types = nil)
      text = input || ''
      selected = types.nil? || types.empty? ? TYPES : types

      out = Result.empty
      out.url   = text.scan(RE[:url]).uniq if selected.include?('url')
      out.email = text.scan(RE[:email]).uniq if selected.include?('email')
      out.ipv4  = text.scan(RE[:ipv4]).uniq if selected.include?('ipv4')
      if selected.include?('ipv6')
        out.ipv6 = text.scan(RE[:ipv6]).select { |run| ipv6?(run) }.uniq
      end
      out.hash = text.scan(RE[:hash]).uniq if selected.include?('hash')
      if selected.include?('domain')
        combined = text.scan(RE[:domain])
        # Cross-rule: an email also yields its domain in the domain list.
        if selected.include?('email')
          combined += text.scan(RE[:email]).map { |e| domain_of(e) }
        end
        out.domain = combined.uniq
      end
      out
    end

    # A hex/colon run is a plausible IPv6: it has a colon AND either contains
    # `::` (a compressed zero-run) or is exactly eight groups of 1-4 hex
    # digits. Mirrors isIpv6() in src/lib/extract.ts.
    def ipv6?(run)
      return false unless run.include?(':')
      return true if run.include?('::')

      groups = run.split(':')
      groups.length == 8 && groups.all? { |g| IPV6_GROUP.match?(g) }
    end

    # Domain part (after the last `@`) of a matched email. rpartition splits
    # from the right with the last separator, mirroring lastIndexOf('@')
    # exactly.
    def domain_of(email)
      email.rpartition('@').last
    end
  end
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →