Skip to content

Slugify — PHP source

Generate clean, URL-safe slugs from any text with locale-aware Unicode transliteration. Accents, emoji, and punctuation are handled automatically - runs entirely in your browser.

This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.

<?php
/**
 * slugify — URL-safe slug generator with locale-aware Unicode transliteration.
 *
 * Language: PHP (8.1+, standard library only — mbstring used for safe
 *           multibyte handling, which is effectively universal in modern PHP)
 * Source:   CosmoDev polyglot showcase port of the Slugify tool, ported from
 *           src/lib/slugify.ts (the canonical TypeScript implementation).
 * License:  display source — part of CosmoDev's polyglot tool pages.
 *
 * Design goals:
 *   - Pure + deterministic; never throws.
 *   - Functionally equivalent to the TS reference: same inputs -> same outputs.
 *   - Self-contained: stdlib only (no Composer packages).
 *
 * Pipeline: transliterate ligatures -> NFKD decompose -> strip combining
 * diacritics -> collapse non-alphanumeric runs -> split into words -> apply
 * casing -> (strip stopwords) -> join with the separator -> (truncate at a
 * word boundary). Anything that can't be transliterated to ASCII (emoji, CJK,
 * ...) collapses to a word separator.
 *
 * Unicode note: PHP's Normalizer (ext/intl) provides NFKD, but intl is not
 * universally enabled and the brief favors stdlib. We reproduce the SAME
 * observable effect the TS code relies on:
 *   (a) map the fixed set of non-decomposing ligatures (TRANSLIT) to ASCII,
 *   (b) strip every code point in the combining-diacritic block U+0300..U+036F.
 * In the TS source NFKD exists solely to split a precomposed accented letter
 * into base + U+03xx combining mark so step (b) can drop the mark. For the
 * Latin range — the script this tool targets — that decomposition always lands
 * in U+0300..U+036F, so removing the range reproduces the TS output for every
 * input the tool is designed for. This is a deliberate, documented trade-off
 * to keep the port dependency-free.
 */

declare(strict_types=1);

/**
 * Letter casing for the produced slug.
 */
const SLUG_CASE_LOWER = 'lower';
const SLUG_CASE_PRESERVE = 'preserve';
const SLUG_CASE_UPPER = 'upper';

/**
 * Letters / ligatures that NFKD does NOT decompose into an ASCII base +
 * combining mark. Mapping them up front turns "Straße" -> "strasse",
 * "Æsir" -> "aesir", "Søren" -> "soren". Accented Latin letters (á é ñ ü ...)
 * need no entry — their diacritic lands in U+0300..U+036F and is stripped.
 *
 * Stored as code-point => ASCII replacement, where the key is the integer
 * Unicode code point (so mb_chr/mb_ord drive the lookup and we avoid source
 * files that aren't valid in every editor encoding).
 */
const TRANSLIT = [
    // Germanic
    0x00DF => 'ss', // ß
    // Latin ligatures
    0x00E6 => 'ae', 0x00C6 => 'ae', // æ, Æ
    0x0153 => 'oe', 0x0152 => 'oe', // œ, Œ
    0xFB00 => 'ff', 0xFB01 => 'fi', 0xFB02 => 'fl',
    0xFB03 => 'ffi', 0xFB04 => 'ffl', 0xFB05 => 'st', 0xFB06 => 'st',
    // Nordic / insular
    0x00F0 => 'd', 0x00D0 => 'd', // ð, Ð
    0x00FE => 'th', 0x00DE => 'th', // þ, Þ
    0x00F8 => 'o', 0x00D8 => 'o', // ø, Ø
    // Eastern European / strokes
    0x0142 => 'l', 0x0141 => 'l', // ł, Ł
    0x0111 => 'd', 0x0110 => 'd', // đ, Đ
    0x0127 => 'h', 0x0126 => 'h', // ħ, Ħ
];

/**
 * Common English stopwords, lowercased and stored as array keys for O(1)
 * isset() lookup. Compared case-insensitively so preserve/upper modes still
 * drop them.
 */
const STOPWORDS = [
    'the' => true, 'a' => true, 'an' => true, 'and' => true, 'or' => true,
    'but' => true, 'of' => true, 'to' => true, 'in' => true, 'on' => true,
    'at' => true, 'for' => true, 'with' => true, 'by' => true, 'from' => true,
];

/**
 * Options for slugify(). Each key is optional; defaults match the TS source.
 *
 *   - separator:      string  default '-'      joining chars; '' concatenates
 *   - maxLength:      int     default 0        0/negative = unlimited
 *   - case:           string  default 'lower'  one of the SLUG_CASE_* consts
 *   - stripStopwords: bool    default false    drop common English stopwords
 */
function slugify_default_options(): array
{
    return [
        'separator'      => '-',
        'maxLength'      => 0,
        'case'           => SLUG_CASE_LOWER,
        'stripStopwords' => false,
    ];
}

/**
 * Reports whether the given code point is a combining diacritical mark in the
 * U+0300..U+036F block — the range the TS source strips after NFKD.
 */
function slugify_is_combining_mark(int $cp): bool
{
    return $cp >= 0x0300 && $cp <= 0x036F;
}

/**
 * Reports whether a code point is an ASCII alphanumeric ([a-zA-Z0-9]) — the
 * only characters that survive into words. Restricting to the ASCII range
 * ensures non-ASCII letters (Greek, Cyrillic, ...) and emoji act as word
 * separators, matching TS's [^a-zA-Z0-9] rule.
 */
function slugify_is_ascii_alnum(int $cp): bool
{
    if ($cp > 0x7F) {
        return false;
    }
    $c = chr($cp);
    return ($c >= 'a' && $c <= 'z')
        || ($c >= 'A' && $c <= 'Z')
        || ($c >= '0' && $c <= '9');
}

/**
 * Break text into a list of clean ASCII words: transliterated, diacritics
 * stripped, cased per options.
 *
 * Single pass over code points (mb_chr splits the string into full Unicode
 * code points regardless of byte width): transliterate non-ASCII ligatures,
 * drop combining marks, push ASCII alphanumerics into the current word, and
 * treat every other code point as a word boundary.
 *
 * @return list<string>
 */
function slugify_tokenize(string $text, array $options): array
{
    $mode = $options['case'] ?? SLUG_CASE_LOWER;

    $words = [];
    $current = '';

    // preg_split(//u) splits into a UTF-8 code point array. The mb_ functions
    // below then give us integer code points for range checks.
    $codepoints = preg_split('//u', $text, -1, PREG_SPLIT_NO_EMPTY) ?: [];

    foreach ($codepoints as $char) {
        $cp = mb_ord($char, 'UTF-8');

        if (slugify_is_combining_mark($cp)) {
            continue; // strip diacritic
        }
        if (isset(TRANSLIT[$cp])) {
            $current .= TRANSLIT[$cp];
            continue;
        }
        if (slugify_is_ascii_alnum($cp)) {
            $current .= $char;
            continue;
        }
        // Separator code point — flush the in-progress word if any.
        if ($current !== '') {
            $words[] = $current;
            $current = '';
        }
    }
    if ($current !== '') {
        $words[] = $current;
    }

    // Apply casing. Preserve is a no-op.
    if ($mode === SLUG_CASE_UPPER) {
        $words = array_map('mb_strtoupper', $words);
    } elseif ($mode === SLUG_CASE_LOWER) {
        $words = array_map('mb_strtolower', $words);
    }

    // Optionally drop English stopwords (case-insensitive).
    if (!empty($options['stripStopwords'])) {
        $words = array_values(array_filter(
            $words,
            fn ($w) => !isset(STOPWORDS[mb_strtolower($w, 'UTF-8')])
        ));
    }

    return $words;
}

/**
 * Truncate a slug to $max chars at the last whole-word boundary.
 *
 * Operates on code-point count (not bytes) to match TS's String.slice
 * semantics; the slug is ASCII-only by construction so byte/char counts
 * coincide, but we use the multibyte view for correctness.
 */
function slugify_truncate_at_word(string $slug, string $separator, int $max): string
{
    if (mb_strlen($slug, 'UTF-8') <= $max) {
        return $slug;
    }
    $cut = mb_substr($slug, 0, $max, 'UTF-8');
    if ($separator === '') {
        return $cut; // nothing to break on — hard cut
    }
    $pos = mb_strrpos($cut, $separator, 0, 'UTF-8');
    if ($pos !== false && $pos > 0) {
        return mb_substr($cut, 0, $pos, 'UTF-8');
    }
    return $cut; // no usable separator found -> hard cut
}

/**
 * Convert arbitrary text into a URL-safe slug. Never throws; an empty or
 * all-symbol input simply yields an empty string.
 */
function slugify(string $text, array $options = []): string
{
    $options = array_merge(slugify_default_options(), $options);

    $separator = $options['separator'];
    $slug = implode($separator, slugify_tokenize($text, $options));

    if (!empty($options['maxLength']) && $options['maxLength'] > 0) {
        return slugify_truncate_at_word($slug, $separator, (int) $options['maxLength']);
    }
    return $slug;
}

/**
 * Slugify each line independently (batch mode). Returns exactly one slug per
 * input line, matching the TS /\r?\n/ split.
 *
 * @return list<string>
 */
function slugify_lines(string $text, array $options = []): array
{
    // preg_split on /\r?\n/ reproduces the TS regex split exactly, handling
    // both \n and \r\n line endings (and emitting empty entries for blank
    // lines, which slugify then turns into empty slugs).
    $lines = preg_split('/\r?\n/', $text) ?: [];
    return array_map(fn ($line) => slugify($line, $options), $lines);
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →