Skip to content

Word & Character Counter — Rust source

Live word, character, sentence, and paragraph counts plus reading-time estimate as you type.

This is the Rust implementation — the same logic the interactive tool runs, in a shareable, citable form.

// word-counter — Rust
// ===================
// Polyglot showcase port of CosmoDev's "word-counter" tool, ported from the
// canonical TypeScript at src/lib/wordCount.ts (the live web library and
// unit-test surface).
//
// Display source — part of CosmoDev's polyglot tool pages
// (dev.cosmolabs.org), where each tool's pure logic is shown in six
// languages side by side. Pure string logic; no cryptographic surface.
//
// Stdlib-only (no external crates such as `regex`), so the two regex-based
// helpers from the TypeScript are expressed as small hand-rolled scanners.
// Exported names use snake_case per Rust convention.

/// Count whitespace-separated words.
///
/// `split_whitespace` splits on runs of Unicode whitespace and trims
/// leading/trailing whitespace, yielding an empty iterator for empty or
/// whitespace-only input — so such input counts as 0 words.
pub fn count_words(text: &str) -> usize {
    text.split_whitespace().count()
}

/// Total character count, whitespace included.
///
/// Counts Unicode scalar values via `chars()`. The canonical TypeScript uses
/// `.length` (UTF-16 code units); the two agree for all BMP text and diverge
/// only for astral symbols (most emoji), which JavaScript counts as 2.
pub fn count_chars(text: &str) -> usize {
    text.chars().count()
}

/// Character count with all whitespace removed.
pub fn count_chars_no_spaces(text: &str) -> usize {
    text.chars().filter(|c| !c.is_whitespace()).count()
}

/// Number of sentences: maximal runs of non-terminal characters followed by
/// one or more sentence terminators (. ! ?).
///
/// Equivalent to the regex `/[^.!?]+[.!?]+/g`. Rust's stdlib ships no regex
/// engine, so we scan with a tiny state machine instead:
///   - accumulate non-terminal content;
///   - when content is followed by a terminator run, count one sentence;
///   - leading punctuation (no prior content) is ignored, and trailing
///     content without a terminator does not count.
pub fn count_sentences(text: &str) -> usize {
    let mut count = 0usize;
    let mut in_content = false; // have we seen non-terminal text yet?
    let mut in_terminators = false; // are we inside the trailing .!? run?

    for ch in text.chars() {
        let is_terminator = matches!(ch, '.' | '!' | '?');
        if is_terminator {
            if in_content && !in_terminators {
                // content → terminator: one sentence completed
                count += 1;
                in_terminators = true;
            }
            // extra terminators (or leading punctuation): nothing to count
        } else {
            in_content = true;
            in_terminators = false;
        }
    }
    count
}

/// Number of paragraphs separated by one or more blank lines.
///
/// Splitting on the literal two-newline string `"\n\n"` and discarding
/// segments that are blank after trimming is equivalent to splitting on
/// `/\n{2,}/`: any extra newlines in a longer run attach to a neighbouring
/// segment as trimmable whitespace and are removed.
pub fn count_paragraphs(text: &str) -> usize {
    text.split("\n\n")
        .filter(|segment| !segment.trim().is_empty())
        .count()
}

/// Estimated reading time in whole minutes at 200 words per minute.
///
/// 0 for empty input, otherwise at least 1. Integer arithmetic —
/// `(words + 100) / 200` — reproduces `round(words / 200)` exactly for
/// non-negative word counts and avoids cross-language float differences.
pub fn reading_time_min(words: usize) -> usize {
    if words == 0 {
        return 0;
    }
    let minutes = (words + 100) / 200;
    minutes.max(1)
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →