Skip to content

Slugify — Ruby source

Generate clean, URL-safe slugs from any text with locale-aware Unicode transliteration. Accents, emoji, and punctuation are handled automatically - runs entirely in your browser.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# slugify — Ruby port: URL-safe slugs with locale-aware Unicode transliteration.

# Letters/ligatures NFKD does not decompose into an ASCII base + combining
# mark. Accented Latin (á é ñ …) needs no entry: NFKD splits it and the
# U+0300..U+036F strip below drops the diacritic.
TRANSLIT = {
  'ß' => 'ss',
  'æ' => 'ae', 'Æ' => 'ae', 'œ' => 'oe', 'Œ' => 'oe',
  'ff' => 'ff', 'fi' => 'fi', 'fl' => 'fl', 'ffi' => 'ffi', 'ffl' => 'ffl', 'ſt' => 'st', 'st' => 'st',
  'ð' => 'd', 'Ð' => 'd', 'þ' => 'th', 'Þ' => 'th', 'ø' => 'o', 'Ø' => 'o',
  'ł' => 'l', 'Ł' => 'l', 'đ' => 'd', 'Đ' => 'd', 'ħ' => 'h', 'Ħ' => 'h',
}.freeze

STOPWORDS = %w[the a an and or but of to in on at for with by from].freeze

# Break +text+ into clean ASCII words (transliterated, diacritics stripped,
# cased). Mirrors the TS tokenize().
def slugify_words(text, case_mode: :lower, strip_stopwords: false)
  ascii = text.gsub(/[^\x00-\x7F]/) { |ch| TRANSLIT[ch] || ch }  # 1. transliterate
              .unicode_normalize(:nfkd)                           # 2. decompose
              .gsub(/[\u{0300}-\u{036F}]/, '')                    # 3. drop marks
              .gsub(/[^a-zA-Z0-9]+/, ' ')                         # 4. collapse runs
              .strip
  words = ascii.empty? ? [] : ascii.split(' ')
  words.map!(&:upcase) if case_mode == :upper
  words.map!(&:downcase) if case_mode == :lower
  words.reject! { |w| STOPWORDS.include?(w.downcase) } if strip_stopwords
  words
end

# Truncate to max chars at the last whole-word boundary (hard cut when the
# separator is empty or absent from the head).
def truncate_at_word(slug, separator, max)
  return slug if slug.length <= max
  cut = slug[0, max]
  return cut if separator.empty?
  last = cut.rindex(separator)
  last&.positive? ? cut[0, last] : cut
end

# Convert arbitrary text into a URL-safe slug. Options mirror the TS
# SlugifyOptions; `case_mode:` spells the TS `case:` (a Ruby keyword).
def slugify(text, separator: '-', max_length: nil, case_mode: :lower, strip_stopwords: false)
  slug = slugify_words(text, case_mode:, strip_stopwords:).join(separator)
  return slug if max_length.nil? || max_length <= 0
  truncate_at_word(slug, separator, max_length)
end

# Slugify each line independently (batch mode), matching the TS /\r?\n/ split
# (the -1 limit keeps trailing empty lines; Ruby splits '' to [] where JS
# yields one empty line, so that case is special-cased).
def slugify_lines(text, **opts)
  return [''] if text.empty?

  text.split(/\r?\n/, -1).map { |line| slugify(line, **opts) }
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →