Skip to content

Text Statistics & Readability — Ruby source

Count words, sentences, paragraphs, characters, lines, and reading time, plus Flesch Reading Ease and Flesch-Kincaid grade-level readability scores.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# text-stats — text statistics & readability: chars, words, sentences, syllables, Flesch. Language: Ruby (3.1+, stdlib only). Port of src/lib/textStats.ts — String#length counts characters (code points) & \s is ASCII-only in Ruby, the TS reference counts UTF-16 units & Unicode spaces.

module TextStats
  WORD_RE = /[A-Za-z0-9'’-]+/.freeze          # U+2019 keeps contractions whole
  SENTENCE_RE = /[.!?]+(?:\s|\z)/.freeze       # \z = absolute end (JS '$' without /m)
  PARA_SEP_RE = /\n{2,}/.freeze

  # Full analysis — readability stays nil when words or sentences are too few.
  Result = Struct.new(:characters, :characters_no_spaces, :words, :sentences, :paragraphs, :lines,
                      :syllables, :reading_time_ms, :speaking_time_ms,
                      :flesch_reading_ease, :flesch_kincaid_grade, :readability_label, keyword_init: true)

  module_function

  # JS Math.round: ties toward +infinity.
  def js_round(x) = (x + 0.5).floor

  # Syllables in one word via the vowel-group heuristic.
  def count_syllables(word)
    w = word.downcase.gsub(/[^a-z]/, '')
    return 0 if w.empty?
    return 1 if w.length <= 3
    s = w.gsub(/(?:[^laeiouy]es|ed|[^laeiouy]e)$/, '') # silent trailing e ('~le' keeps its syllable)
    s = s.sub(/^y/, '')
    [1, s.scan(/[aeiouy]+/).length].max
  end

  def label_for(f)
    f >= 80 ? 'Very Easy' : f >= 70 ? 'Easy' : f >= 60 ? 'Standard'
            : f >= 50 ? 'Fairly Hard' : f >= 30 ? 'Hard' : 'Very Hard'
  end

  def analyze(input)
    text = input.to_s
    words = text.scan(WORD_RE)
    sentences = words.empty? ? 0 : [1, text.scan(SENTENCE_RE).length].max
    paragraphs = text.strip.empty? ? 0 : text.split(PARA_SEP_RE).count { |p| !p.strip.empty? }
    lines = text.empty? ? 0 : text.split("\n", -1).length
    syllables = words.sum { |w| count_syllables(w) }
    reading_ms = js_round(words.length / 200.0 * 60_000)  # 200 wpm
    speaking_ms = js_round(words.length / 130.0 * 60_000) # 130 wpm
    fre = fkg = label = nil
    if !words.empty? && sentences.positive?
      wps = words.length.to_f / sentences
      spw = syllables.to_f / words.length
      fre = js_round((206.835 - 1.015 * wps - 84.6 * spw) * 10) / 10.0
      fkg = js_round((0.39 * wps + 11.8 * spw - 15.59) * 10) / 10.0
      label = label_for(fre)
    end
    Result.new(characters: text.length, characters_no_spaces: text.gsub(/\s/, '').length, words: words.length,
               sentences:, paragraphs:, lines:, syllables:, reading_time_ms: reading_ms,
               speaking_time_ms: speaking_ms, flesch_reading_ease: fre,
               flesch_kincaid_grade: fkg, readability_label: label)
  end
end

['The quick brown fox jumps over the lazy dog.',
 "Hi.\n\nMy name is Inigo Montoya. You killed my father; prepare to die!"].each do |text|
  s = TextStats.analyze(text)
  flesch = s.readability_label ? format('%.1f (%s)', s.flesch_reading_ease, s.readability_label) : 'n/a'
  grade = s.flesch_kincaid_grade ? format('%.1f', s.flesch_kincaid_grade) : 'n/a'
  puts "chars=#{s.characters} nospace=#{s.characters_no_spaces} words=#{s.words} " \
       "sentences=#{s.sentences} paragraphs=#{s.paragraphs} lines=#{s.lines} " \
       "syllables=#{s.syllables} reading=#{s.reading_time_ms}ms speaking=#{s.speaking_time_ms}ms " \
       "flesch=#{flesch} grade=#{grade}"
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →