Skip to content

Text Extractor — PHP source

Pull URLs, emails, IPv4/IPv6 addresses, hashes (MD5/SHA-1/SHA-256/SHA-512), and domains out of logs, headers, or any pasted text.

This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.

<?php
/**
 * extract — pull URLs, emails, IPv4/IPv6 addresses, hashes, and domains
 * out of arbitrary text (logs, headers, config).
 *
 * Language: PHP (8.1+, standard library only — PCRE ships in the language)
 * Source:   CosmoDev polyglot showcase port of the Extract tool, ported from
 *           src/lib/extract.ts (the canonical TypeScript implementation) and
 *           held in lock-step with its Go twin cli/extract/extract.go.
 * License:  display source — part of CosmoDev's polyglot tool pages.
 *
 * Design goals:
 *   - Pure + deterministic; never throws.
 *   - Functionally equivalent to the TS/Go reference: same inputs -> same outputs.
 *   - Self-contained: stdlib only (no Composer packages).
 *
 * Behavior: pull URLs, emails, IPv4/IPv6 addresses, hashes
 * (md5/sha1/sha256/sha512 by length), and domains out of arbitrary text.
 * Matches are deduped per type preserving first-occurrence order; an email
 * also contributes its domain to the domain list when both email and domain
 * are selected. A zero-length types list defaults to all six kinds.
 *
 * PCRE note: PHP's regex engine is PCRE2, not RE2 — but every construct used
 * here (word boundaries `\b`, non-capturing `(?:...)`, counted repetition,
 * alternation) is plain vanilla and behaves identically to Go's `regexp` /
 * JS for these patterns. No backreferences, no lookahead, no lookbehind.
 */

declare(strict_types=1);

/**
 * The six canonical extraction kinds, as const strings matching the TS
 * 'url' | 'email' | 'ipv4' | 'ipv6' | 'hash' | 'domain' union and Go's Type.
 */
const EXTRACT_TYPE_URL    = 'url';
const EXTRACT_TYPE_EMAIL  = 'email';
const EXTRACT_TYPE_IPV4   = 'ipv4';
const EXTRACT_TYPE_IPV6   = 'ipv6';
const EXTRACT_TYPE_HASH   = 'hash';
const EXTRACT_TYPE_DOMAIN = 'domain';

/**
 * Canonical extraction kinds in display order. Mirrors EXTRACT_TYPES in TS.
 */
const EXTRACT_TYPES = [
    EXTRACT_TYPE_URL, EXTRACT_TYPE_EMAIL, EXTRACT_TYPE_IPV4,
    EXTRACT_TYPE_IPV6, EXTRACT_TYPE_HASH, EXTRACT_TYPE_DOMAIN,
];

/**
 * Per-kind patterns. These mirror the RE record in src/lib/extract.ts (and
 * the Go twin) verbatim, wrapped in the PCRE `/.../ ` delimiter. The IPv6
 * pattern is intentionally permissive — a hex/colon run — and is post-filtered
 * by extract_is_ipv6() so bare hex words, times, and MACs are rejected.
 */
const RE_EXTRACT = [
    'url'    => '/https?:\/\/[^\s]+/',
    'email'  => '/[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}/',
    'ipv4'   => '/\b(?:\d{1,3}\.){3}\d{1,3}\b/',
    'ipv6'   => '/[0-9a-fA-F:]+/',
    // md5 (32) / sha1 (40) / sha256 (64) / sha512 (128). `\b` keeps each length
    // honest, so a 64-char run does not also match as a leading 32-char hash.
    'hash'   => '/\b[a-fA-F0-9]{32}\b|\b[a-fA-F0-9]{40}\b|\b[a-fA-F0-9]{64}\b|\b[a-fA-F0-9]{128}\b/',
    'domain' => '/\b[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z]{2,})+\b/',
];

/** Validates a single 1–4 hex-digit IPv6 hextet (anchored full match). */
const RE_IPV6_GROUP = '/^[0-9a-fA-F]{1,4}$/';

/**
 * Deduplicate values, preserving first-occurrence order. Uses array keys for
 * O(1) de-duplication; PHP associative arrays keep insertion order, so
 * array_keys() returns first-occurrence order. Mirrors uniq() in the TS lib.
 *
 * @param list<string> $values
 * @return list<string>
 */
function extract_uniq(array $values): array
{
    $seen = [];
    foreach ($values as $v) {
        $seen[$v] = true; // repeats overwrite at the same key, preserving first position
    }
    return array_keys($seen);
}

/**
 * Every non-overlapping match of $pattern in $text. preg_match_all fills
 * $matches[0] with the full matches (PREG_PATTERN_ORDER is the default order).
 * A false return means an invalid pattern — impossible for these static
 * literals, but we degrade to an empty list rather than throwing.
 *
 * @return list<string>
 */
function extract_all_matches(string $pattern, string $text): array
{
    $matches = [];
    if (@preg_match_all($pattern, $text, $matches) !== false) {
        return $matches[0] ?? [];
    }
    return [];
}

/**
 * A hex/colon run is a plausible IPv6: it has a colon AND either contains `::`
 * (a compressed zero-run) or is exactly eight groups of 1–4 hex digits.
 * Mirrors isIpv6() in src/lib/extract.ts.
 */
function extract_is_ipv6(string $run): bool
{
    if (!str_contains($run, ':')) {
        return false;
    }
    if (str_contains($run, '::')) {
        return true;
    }
    $groups = explode(':', $run);
    if (count($groups) !== 8) {
        return false;
    }
    foreach ($groups as $g) {
        if (preg_match(RE_IPV6_GROUP, $g) !== 1) {
            return false;
        }
    }
    return true;
}

/**
 * Domain part (after the last `@`) of a matched email. strrpos returns the
 * byte offset of the last `@` (ASCII, so byte-safe on email addresses).
 * Mirrors domainOf() in src/lib/extract.ts.
 */
function extract_domain_of(string $email): string
{
    $i = strrpos($email, '@');
    return $i === false ? $email : substr($email, $i + 1);
}

/**
 * Extract every occurrence of the given $types (default: all six) from $input.
 * Returns an array with a key per type — always all six, populated only for
 * the selected types (unselected types are empty arrays). Matches are deduped
 * per type, preserving first-occurrence order. An email also contributes its
 * domain to the 'domain' list when both 'email' and 'domain' are selected.
 * Never throws; null input is treated as empty text.
 *
 * @param string|null $input
 * @param list<string> $types Extraction kinds to pull; empty = all six.
 * @return array<string, list<string>>
 */
function extract(?string $input, array $types = []): array
{
    // TS does `const text = input ?? ''`; PHP mirrors it with the nullable param.
    $text = $input ?? '';
    // TS defaults the param AND re-defaults an empty array; `$types !== []`
    // covers both the default and an explicitly empty list.
    $selected = $types !== [] ? $types : EXTRACT_TYPES;
    $want = static fn (string $t): bool => in_array($t, $selected, true);

    $out = [
        'url' => [], 'email' => [], 'ipv4' => [],
        'ipv6' => [], 'hash' => [], 'domain' => [],
    ];

    if ($want(EXTRACT_TYPE_URL)) {
        $out['url'] = extract_uniq(extract_all_matches(RE_EXTRACT['url'], $text));
    }
    if ($want(EXTRACT_TYPE_EMAIL)) {
        $out['email'] = extract_uniq(extract_all_matches(RE_EXTRACT['email'], $text));
    }
    if ($want(EXTRACT_TYPE_IPV4)) {
        $out['ipv4'] = extract_uniq(extract_all_matches(RE_EXTRACT['ipv4'], $text));
    }
    if ($want(EXTRACT_TYPE_IPV6)) {
        $runs = extract_all_matches(RE_EXTRACT['ipv6'], $text);
        $filtered = array_values(array_filter($runs, 'extract_is_ipv6'));
        $out['ipv6'] = extract_uniq($filtered);
    }
    if ($want(EXTRACT_TYPE_HASH)) {
        $out['hash'] = extract_uniq(extract_all_matches(RE_EXTRACT['hash'], $text));
    }
    if ($want(EXTRACT_TYPE_DOMAIN)) {
        $combined = extract_all_matches(RE_EXTRACT['domain'], $text);
        // Cross-rule: an email also yields its domain in the domain list.
        if ($want(EXTRACT_TYPE_EMAIL)) {
            foreach (extract_all_matches(RE_EXTRACT['email'], $text) as $e) {
                $combined[] = extract_domain_of($e);
            }
        }
        $out['domain'] = extract_uniq($combined);
    }

    return $out;
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →