Slugify — PHP source
Generate clean, URL-safe slugs from any text with locale-aware Unicode transliteration. Accents, emoji, and punctuation are handled automatically - runs entirely in your browser.
This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.
<?php
/**
* slugify — URL-safe slug generator with locale-aware Unicode transliteration.
*
* Language: PHP (8.1+, standard library only — mbstring used for safe
* multibyte handling, which is effectively universal in modern PHP)
* Source: CosmoDev polyglot showcase port of the Slugify tool, ported from
* src/lib/slugify.ts (the canonical TypeScript implementation).
* License: display source — part of CosmoDev's polyglot tool pages.
*
* Design goals:
* - Pure + deterministic; never throws.
* - Functionally equivalent to the TS reference: same inputs -> same outputs.
* - Self-contained: stdlib only (no Composer packages).
*
* Pipeline: transliterate ligatures -> NFKD decompose -> strip combining
* diacritics -> collapse non-alphanumeric runs -> split into words -> apply
* casing -> (strip stopwords) -> join with the separator -> (truncate at a
* word boundary). Anything that can't be transliterated to ASCII (emoji, CJK,
* ...) collapses to a word separator.
*
* Unicode note: PHP's Normalizer (ext/intl) provides NFKD, but intl is not
* universally enabled and the brief favors stdlib. We reproduce the SAME
* observable effect the TS code relies on:
* (a) map the fixed set of non-decomposing ligatures (TRANSLIT) to ASCII,
* (b) strip every code point in the combining-diacritic block U+0300..U+036F.
* In the TS source NFKD exists solely to split a precomposed accented letter
* into base + U+03xx combining mark so step (b) can drop the mark. For the
* Latin range — the script this tool targets — that decomposition always lands
* in U+0300..U+036F, so removing the range reproduces the TS output for every
* input the tool is designed for. This is a deliberate, documented trade-off
* to keep the port dependency-free.
*/
declare(strict_types=1);
/**
* Letter casing for the produced slug.
*/
const SLUG_CASE_LOWER = 'lower';
const SLUG_CASE_PRESERVE = 'preserve';
const SLUG_CASE_UPPER = 'upper';
/**
* Letters / ligatures that NFKD does NOT decompose into an ASCII base +
* combining mark. Mapping them up front turns "Straße" -> "strasse",
* "Æsir" -> "aesir", "Søren" -> "soren". Accented Latin letters (á é ñ ü ...)
* need no entry — their diacritic lands in U+0300..U+036F and is stripped.
*
* Stored as code-point => ASCII replacement, where the key is the integer
* Unicode code point (so mb_chr/mb_ord drive the lookup and we avoid source
* files that aren't valid in every editor encoding).
*/
const TRANSLIT = [
// Germanic
0x00DF => 'ss', // ß
// Latin ligatures
0x00E6 => 'ae', 0x00C6 => 'ae', // æ, Æ
0x0153 => 'oe', 0x0152 => 'oe', // œ, Œ
0xFB00 => 'ff', 0xFB01 => 'fi', 0xFB02 => 'fl',
0xFB03 => 'ffi', 0xFB04 => 'ffl', 0xFB05 => 'st', 0xFB06 => 'st',
// Nordic / insular
0x00F0 => 'd', 0x00D0 => 'd', // ð, Ð
0x00FE => 'th', 0x00DE => 'th', // þ, Þ
0x00F8 => 'o', 0x00D8 => 'o', // ø, Ø
// Eastern European / strokes
0x0142 => 'l', 0x0141 => 'l', // ł, Ł
0x0111 => 'd', 0x0110 => 'd', // đ, Đ
0x0127 => 'h', 0x0126 => 'h', // ħ, Ħ
];
/**
* Common English stopwords, lowercased and stored as array keys for O(1)
* isset() lookup. Compared case-insensitively so preserve/upper modes still
* drop them.
*/
const STOPWORDS = [
'the' => true, 'a' => true, 'an' => true, 'and' => true, 'or' => true,
'but' => true, 'of' => true, 'to' => true, 'in' => true, 'on' => true,
'at' => true, 'for' => true, 'with' => true, 'by' => true, 'from' => true,
];
/**
* Options for slugify(). Each key is optional; defaults match the TS source.
*
* - separator: string default '-' joining chars; '' concatenates
* - maxLength: int default 0 0/negative = unlimited
* - case: string default 'lower' one of the SLUG_CASE_* consts
* - stripStopwords: bool default false drop common English stopwords
*/
function slugify_default_options(): array
{
return [
'separator' => '-',
'maxLength' => 0,
'case' => SLUG_CASE_LOWER,
'stripStopwords' => false,
];
}
/**
* Reports whether the given code point is a combining diacritical mark in the
* U+0300..U+036F block — the range the TS source strips after NFKD.
*/
function slugify_is_combining_mark(int $cp): bool
{
return $cp >= 0x0300 && $cp <= 0x036F;
}
/**
* Reports whether a code point is an ASCII alphanumeric ([a-zA-Z0-9]) — the
* only characters that survive into words. Restricting to the ASCII range
* ensures non-ASCII letters (Greek, Cyrillic, ...) and emoji act as word
* separators, matching TS's [^a-zA-Z0-9] rule.
*/
function slugify_is_ascii_alnum(int $cp): bool
{
if ($cp > 0x7F) {
return false;
}
$c = chr($cp);
return ($c >= 'a' && $c <= 'z')
|| ($c >= 'A' && $c <= 'Z')
|| ($c >= '0' && $c <= '9');
}
/**
* Break text into a list of clean ASCII words: transliterated, diacritics
* stripped, cased per options.
*
* Single pass over code points (mb_chr splits the string into full Unicode
* code points regardless of byte width): transliterate non-ASCII ligatures,
* drop combining marks, push ASCII alphanumerics into the current word, and
* treat every other code point as a word boundary.
*
* @return list<string>
*/
function slugify_tokenize(string $text, array $options): array
{
$mode = $options['case'] ?? SLUG_CASE_LOWER;
$words = [];
$current = '';
// preg_split(//u) splits into a UTF-8 code point array. The mb_ functions
// below then give us integer code points for range checks.
$codepoints = preg_split('//u', $text, -1, PREG_SPLIT_NO_EMPTY) ?: [];
foreach ($codepoints as $char) {
$cp = mb_ord($char, 'UTF-8');
if (slugify_is_combining_mark($cp)) {
continue; // strip diacritic
}
if (isset(TRANSLIT[$cp])) {
$current .= TRANSLIT[$cp];
continue;
}
if (slugify_is_ascii_alnum($cp)) {
$current .= $char;
continue;
}
// Separator code point — flush the in-progress word if any.
if ($current !== '') {
$words[] = $current;
$current = '';
}
}
if ($current !== '') {
$words[] = $current;
}
// Apply casing. Preserve is a no-op.
if ($mode === SLUG_CASE_UPPER) {
$words = array_map('mb_strtoupper', $words);
} elseif ($mode === SLUG_CASE_LOWER) {
$words = array_map('mb_strtolower', $words);
}
// Optionally drop English stopwords (case-insensitive).
if (!empty($options['stripStopwords'])) {
$words = array_values(array_filter(
$words,
fn ($w) => !isset(STOPWORDS[mb_strtolower($w, 'UTF-8')])
));
}
return $words;
}
/**
* Truncate a slug to $max chars at the last whole-word boundary.
*
* Operates on code-point count (not bytes) to match TS's String.slice
* semantics; the slug is ASCII-only by construction so byte/char counts
* coincide, but we use the multibyte view for correctness.
*/
function slugify_truncate_at_word(string $slug, string $separator, int $max): string
{
if (mb_strlen($slug, 'UTF-8') <= $max) {
return $slug;
}
$cut = mb_substr($slug, 0, $max, 'UTF-8');
if ($separator === '') {
return $cut; // nothing to break on — hard cut
}
$pos = mb_strrpos($cut, $separator, 0, 'UTF-8');
if ($pos !== false && $pos > 0) {
return mb_substr($cut, 0, $pos, 'UTF-8');
}
return $cut; // no usable separator found -> hard cut
}
/**
* Convert arbitrary text into a URL-safe slug. Never throws; an empty or
* all-symbol input simply yields an empty string.
*/
function slugify(string $text, array $options = []): string
{
$options = array_merge(slugify_default_options(), $options);
$separator = $options['separator'];
$slug = implode($separator, slugify_tokenize($text, $options));
if (!empty($options['maxLength']) && $options['maxLength'] > 0) {
return slugify_truncate_at_word($slug, $separator, (int) $options['maxLength']);
}
return $slug;
}
/**
* Slugify each line independently (batch mode). Returns exactly one slug per
* input line, matching the TS /\r?\n/ split.
*
* @return list<string>
*/
function slugify_lines(string $text, array $options = []): array
{
// preg_split on /\r?\n/ reproduces the TS regex split exactly, handling
// both \n and \r\n line endings (and emitting empty entries for blank
// lines, which slugify then turns into empty slugs).
$lines = preg_split('/\r?\n/', $text) ?: [];
return array_map(fn ($line) => slugify($line, $options), $lines);
}
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →