Skip to content

文字列をスラグ化する(URLセーフ) snippet

どんなタイトルもURLセーフなスラグに変えます。小文字化、アクセントの畳み込み、英数字以外のすべてをダッシュへ、ダッシュの連結を圧縮してトリム。順序の落とし穴はどこでも効いてきます。アクセントは取り除く前に畳み込んでください。さもないと Crème は creme ではなく cr-me になります。どの言語でもパイプラインは同じです。Unicode の NFKD/NFD 分解 → 結合文字の除去 → 小文字化(ロケールに依存しない形で)→ 正規表現でダッシュへ → トリム。このサイトの対話型スラグ化ツールがフル仕様の実装で、ここにあるのは各言語のstdlibによるコア部分です。Zigは省略されています(stdlibにUnicodeテーブルがないためです)。

どんなタイトルもURLセーフなスラグに変えます。小文字化、アクセントの畳み込み、英数字以外のすべてをダッシュへ、ダッシュの連結を圧縮してトリム。順序の落とし穴はどこでも効いてきます。アクセントは取り除く前に畳み込んでください。さもないと Crème は creme ではなく cr-me になります。どの言語でもパイプラインは同じです。Unicode の NFKD/NFD 分解 → 結合文字の除去 → 小文字化(ロケールに依存しない形で)→ 正規表現でダッシュへ → トリム。このサイトの対話型スラグ化ツールがフル仕様の実装で、ここにあるのは各言語のstdlibによるコア部分です。Zigは省略されています(stdlibにUnicodeテーブルがないためです)。

Runnable recipe · 12 languagesslugify ツールを開く →
Text & Parsingslugurlunicodenormalizeseo

Every language

12 languages, copy-ready. One at a time with syntax highlighting, or all inline.

SQLSQLrunnable
SELECT trim(both '-' from
         lower(
           regexp_replace(
             regexp_replace(
               'Crème Brûlée!', '[^a-zA-Z0-9]+', '-', 'g'),
             '^-*|-*$', '', 'g')))
AS slug;

Honest limit: Postgres has no built-in unaccent — that is an extension, unavailable in the playground — so accented letters become dashes here (cr-me-br-l-e). Do the accent fold in the application layer. Run the regex chain itself in the playground.

Run in the SQL playground →
JSJavaScript
function slugify(s) {
  return s
    .normalize('NFKD')                 // split accented chars
    .replace(/[\u0300-\u036f]/g, '')   // drop the combining marks
    .toLowerCase()
    .replace(/[^a-z0-9]+/g, '-')        // everything else → dash
    .replace(/^-+|-+$/g, '');           // trim edge dashes
}

slugify('Crème Brûlée! Örebro'); // 'creme-brulee-orebro'

The \u0300-\u036f range covers Latin combining marks — enough for Western text; full Unicode needs a proper transliteration table.

TSTypeScript
export function slugify(input: string): string {
  return input
    .normalize('NFKD')
    .replace(/[\u0300-\u036f]/g, '')
    .toLowerCase()
    .replace(/[^a-z0-9]+/g, '-')
    .replace(/^-+|-+$/g, '');
}

Pure string in, string out — trivially testable. This is the same contract the site's slugify tool and its Go twin share.

GoGo
import (
	"regexp"
	"strings"
	"unicode"

	"golang.org/x/text/unicode/norm"
)

var nonAlnum = regexp.MustCompile(`[^a-z0-9]+`)

func slugify(s string) string {
	var b strings.Builder
	for _, r := range norm.NFD.String(s) { // decompose accents
		if !unicode.Is(unicode.Mn, r) { // skip combining marks
			b.WriteRune(r)
		}
	}
	s = strings.ToLower(b.String())
	s = nonAlnum.ReplaceAllString(s, "-")
	return strings.Trim(s, "-")
}

unicode.Mn is the combining-mark class. Hoist the regexp to package level — MustCompile in the function recompiles per call.

RsRust
use deunicode::deunicode;

fn slugify(s: &str) -> String {
    let ascii = deunicode(s).to_lowercase(); // café → cafe, 東 → tou
    let mut out = String::new();
    let mut prev_dash = false;
    for ch in ascii.chars() {
        let isalnum = ch.is_ascii_alphanumeric();
        if !isalnum {
            prev_dash = true;
        } else {
            if prev_dash && !out.is_empty() {
                out.push('-');
            }
            out.push(ch);
            prev_dash = false;
        }
    }
    out
}

deunicode transliterates ANY script to ASCII in one call — the Rust answer to accent folding. The `slug` crate packages this exact pipeline.

PHPPHP
function slugify(string $s): string {
    $s = \Transliterator::createFromRules(
        ':: Any-Latin; :: Latin-ASCII; :: Lower();'
    )->transliterate($s);

    $s = strtolower(preg_replace('/[^a-z0-9]+/', '-', $s) ?? $s);
    return trim($s, '-');
}

Needs ext-intl (Transliterator) — it folds any script, not just Latin accents. The iconv ASCII//TRANSLIT trick is famously locale-brittle; do not use it.

PyPython
import re
import unicodedata

def slugify(s: str) -> str:
    s = unicodedata.normalize('NFKD', s)
    s = ''.join(c for c in s if not unicodedata.combining(c))
    s = s.lower()
    s = re.sub(r'[^a-z0-9]+', '-', s)
    return s.strip('-')

slugify('Crème Brûlée! Örebro')  # 'creme-brulee-orebro'

unicodedata.combining(c) is the per-char test for combining marks. python-slugify is the batteries-included version (more scripts, custom maps).

C#C#
using System.Globalization;
using System.Text;
using System.Text.RegularExpressions;

static string Slugify(string s)
{
    var formD = s.Normalize(NormalizationForm.FormD);

    var stripped = new string(formD
        .Where(ch => CharUnicodeInfo.GetUnicodeCategory(ch)
                     != UnicodeCategory.NonSpacingMark)
        .ToArray());

    return Regex.Replace(
            stripped.ToLowerInvariant(),
            @"[^a-z0-9]+", "-")
        .Trim('-');
}

ToLowerInvariant — never ToLower() with the current culture. For hot paths, mark the regex [GeneratedRegex(...)] to source-generate it.

JvJava
import java.text.Normalizer;
import java.util.Locale;

static String slugify(String s) {
    String norm = Normalizer.normalize(s, Normalizer.Form.NFD)
            .replaceAll("\\p{M}+", ""); // marks

    return norm.toLowerCase(Locale.ROOT)
            .replaceAll("[^a-z0-9]+", "-")
            .replaceAll("^-+|-+$", "");
}

Locale.ROOT, never the default locale — toLowerCase under a Turkish locale turns I into ı̇ and silently changes the slug.

SwSwift
import Foundation

func slugify(_ s: String) -> String {
    let folded = s.folding(
        options: [.diacriticInsensitive, .caseInsensitive],
        locale: Locale(identifier: "en_US")
    ).lowercased()

    let dashed = folded.replacingOccurrences(
        of: "[^a-z0-9]+",
        with: "-",
        options: .regularExpression)

    return dashed.trimmingCharacters(in: CharacterSet(charactersIn: "-"))
}

String.folding(options: .diacriticInsensitive) is the one-call accent fold — the Swift-native alternative to NFD surgery.

KtKotlin
import java.text.Normalizer
import java.util.Locale

fun slugify(s: String): String =
    Normalizer.normalize(s, Normalizer.Form.NFD)
        .replace(Regex("\\p{M}+"), "")
        .lowercase(Locale.ROOT)
        .replace(Regex("[^a-z0-9]+"), "-")
        .trim('-')

Same JVM pipeline as Java with Kotlin's chainable spelling. String.lowercase(Locale.ROOT) — the locale argument is what keeps slugs stable across servers.

RbRuby
def slugify(s)
  s.unicode_normalize(:nfd)
   .gsub(/\p{Mn}+/, '')      # combining marks
   .downcase
   .gsub(/[^a-z0-9]+/, '-')
   .gsub(/^-+|-+$/, '')
end

\p{Mn} is the Unicode combining-mark property — regex-level Unicode classes are Ruby's cleanest route. Rails has parameterize built in.

Keep going

Try the interactive slugify tool →Read the regex-tokens cheatsheet →