Text Extractor — Ruby source
Pull URLs, emails, IPv4/IPv6 addresses, hashes (MD5/SHA-1/SHA-256/SHA-512), and domains out of logs, headers, or any pasted text.
This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.
# extract — pull URLs, emails, IPv4/IPv6 addresses, hashes, and domains
# out of arbitrary text (logs, headers, config).
#
# Language: Ruby (3.2+, standard library only)
# Source: CosmoDev polyglot showcase port of the Extract tool, ported from
# src/lib/extract.ts (the canonical TypeScript implementation) and
# held in lock-step with its Go twin cli/extract/extract.go.
# License: display source — part of CosmoDev's polyglot tool pages.
#
# Design goals:
# - Pure + deterministic; never raises.
# - Functionally equivalent to the TS/Go reference: same inputs -> same outputs.
# - Self-contained: stdlib only (Regexp ships in the language; no gems).
#
# Behavior: pull URLs, emails, IPv4/IPv6 addresses, hashes (md5/sha1/sha256/
# sha512 by length), and domains out of arbitrary text. Matches are deduped
# per type preserving first-occurrence order; an email also contributes its
# domain to the domain list when both email and domain are selected. A nil or
# empty types list defaults to all six kinds.
module Extract
# Canonical extraction kinds, in display order. Mirrors EXTRACT_TYPES in TS.
TYPES = %w[url email ipv4 ipv6 hash domain].freeze
# Per-kind patterns, scanned with String#scan. These mirror the RE record
# in src/lib/extract.ts (and the Go twin) verbatim: same anchors, classes,
# and counted repetition. Every group is non-capturing (?:...) so scan
# returns the full match string, not group arrays. The IPv6 pattern is
# intentionally permissive — a hex/colon run — and is post-filtered by
# ipv6? so bare hex words, times, and MACs are rejected, exactly as in the
# TS lib.
RE = {
url: /https?:\/\/[^\s]+/,
email: /[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}/,
ipv4: /\b(?:\d{1,3}\.){3}\d{1,3}\b/,
ipv6: /[0-9a-fA-F:]+/,
# md5 (32) / sha1 (40) / sha256 (64) / sha512 (128). `\b` keeps each
# length honest, so a 64-char run does not also match as a leading
# 32-char hash.
hash: /\b[a-fA-F0-9]{32}\b|\b[a-fA-F0-9]{40}\b|
\b[a-fA-F0-9]{64}\b|\b[a-fA-F0-9]{128}\b/x,
domain: /\b[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z]{2,})+\b/
}.freeze
# Validates a single 1-4 hex-digit IPv6 hextet (full-string match).
IPV6_GROUP = /\A[0-9a-fA-F]{1,4}\z/
# One field per kind — always all six, populated only for the selected
# types (unselected kinds stay empty arrays). Mirrors the TS
# `Record<ExtractType, string[]>` and the Go twin's Result struct.
Result = Struct.new(:url, :email, :ipv4, :ipv6, :hash, :domain) do
def self.empty
new([], [], [], [], [], [])
end
end
class << self
# Extract every occurrence of the given `types` (default: all six) from
# `input`. Returns an Extract::Result with one field per kind — always
# all six, populated only for the selected types. Matches are deduped per
# kind, preserving first-occurrence order (Array#uniq keeps first
# occurrences — the Go twin of the uniq() helper in src/lib/extract.ts).
# An email also contributes its domain to the domain list when both email
# and domain are selected. Never raises; nil input reads as empty text.
def extract(input, types = nil)
text = input || ''
selected = types.nil? || types.empty? ? TYPES : types
out = Result.empty
out.url = text.scan(RE[:url]).uniq if selected.include?('url')
out.email = text.scan(RE[:email]).uniq if selected.include?('email')
out.ipv4 = text.scan(RE[:ipv4]).uniq if selected.include?('ipv4')
if selected.include?('ipv6')
out.ipv6 = text.scan(RE[:ipv6]).select { |run| ipv6?(run) }.uniq
end
out.hash = text.scan(RE[:hash]).uniq if selected.include?('hash')
if selected.include?('domain')
combined = text.scan(RE[:domain])
# Cross-rule: an email also yields its domain in the domain list.
if selected.include?('email')
combined += text.scan(RE[:email]).map { |e| domain_of(e) }
end
out.domain = combined.uniq
end
out
end
# A hex/colon run is a plausible IPv6: it has a colon AND either contains
# `::` (a compressed zero-run) or is exactly eight groups of 1-4 hex
# digits. Mirrors isIpv6() in src/lib/extract.ts.
def ipv6?(run)
return false unless run.include?(':')
return true if run.include?('::')
groups = run.split(':')
groups.length == 8 && groups.all? { |g| IPV6_GROUP.match?(g) }
end
# Domain part (after the last `@`) of a matched email. rpartition splits
# from the right with the last separator, mirroring lastIndexOf('@')
# exactly.
def domain_of(email)
email.rpartition('@').last
end
end
end
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →