Skip to content

Découper une chaîne sur n'importe quel espace snippet

Tokeniser une entrée sur n'importe quelle suite d'espaces — espaces, tabulations, sauts de ligne — plutôt qu'un seul caractère littéral.

Tokeniser une entrée sur n'importe quelle suite d'espaces — espaces, tabulations, sauts de ligne — plutôt qu'un seul caractère littéral. Le piège, c'est de découper sur ' ' : les suites d'espaces et les blancs en tête ou en queue produisent des tokens chaîne-vide qui corrompent discrètement comptages, lookups et jointures. Environ la moitié du registre a un mode dédié façon awk (strings.Fields en Go, split_whitespace en Rust, split sans argument en Python, split en Ruby) ; le reste s'appuie sur la regex \s+ et doit d'abord trimmer, sinon un '' de tête se faufile. Méfie-toi aussi des écarts de dialecte : \s de Java est ASCII-only sans (?U), et strtok_r en C ne connaît que les octets que tu lui listes.

Recette exécutable · 15 langages
Text & Parsingstringswhitespacesplittokensregex

Every language

15 langages, copy-ready. One at a time with syntax highlighting, or all inline.

SQLSQLrunnable
-- PostgreSQL
SELECT array_remove(regexp_split_to_array('  hello  world   again ', '\s+'), '') AS words;
-- words: {hello,world,again}

Postgres dialect — the SQL standard has no string split. regexp_split_to_array emits empty leading/trailing elements when the input starts or ends with whitespace; array_remove scrubs them. regexp_split_to_table() gives one word per row instead.

Run in the SQL playground →
JSJavaScript
const words = '  hello\t world \n again '.trim().split(/\s+/);
console.log(words); // [ 'hello', 'world', 'again' ]

trim() is load-bearing: without it a leading run of whitespace yields a leading '' element. JS \s already covers NBSP and other Unicode spaces — no u flag needed.

TSTypeScript
const words: string[] = '  hello\t world \n again '.trim().split(/\s+/);
console.log(words.length); // 3

Same engine as JavaScript — split(/\s+/) returns string[], never null or undefined elements, so no null guards needed downstream.

GoGo
package main

import (
	"fmt"
	"strings"
)

func main() {
	words := strings.Fields("  hello\t world \n again ")
	fmt.Println(words) // [hello world again]
}

strings.Fields is purpose-built: splits on any run of unicode.IsSpace whitespace and never yields empty fields. strings.Split(s, " ") is the trap — it only knows single spaces and emits empties.

RsRust
fn main() {
    let words: Vec<&str> = "  hello\t world \n again ".split_whitespace().collect();
    println!("{:?}", words); // ["hello", "world", "again"]
}

split_whitespace uses char::is_whitespace (Unicode-aware) and skips runs, yielding &str borrows with zero allocation. split(' ') is the trap: it yields an empty piece per extra space.

PHPPHP
<?php
$words = preg_split('/\s+/', trim("  hello\t world   again "));
print_r($words); // ['hello', 'world', 'again']

preg_split does not skip a leading match, so trim() first or you get a leading ''. explode(' ', $s) is the trap: single-space only, empty fields for runs. \s is ASCII in PCRE without (*UCP).

PyPython
words = "  hello\t world \n again ".split()
print(words)  # ['hello', 'world', 'again']

The no-arg split is the whole trick: it collapses runs, strips both ends, and handles any Unicode whitespace. 'a b'.split(' ') gives ['a', '', 'b'] — never pass an explicit separator for this job.

CC
#include <stdio.h>
#include <string.h>

int main(void) {
    char text[] = "  hello\t world \n again ";
    char *save = NULL;
    const char *seps = " \t\n\r\f\v";
    for (char *tok = strtok_r(text, seps, &save); tok; tok = strtok_r(NULL, seps, &save)) {
        printf("%s\n", tok);
    }
    return 0;
}

POSIX strtok_r compiles with default gnu17 gcc; it skips delimiter runs (no empty tokens) but mutates the buffer in place and knows only the six ASCII whitespace bytes you list — keep the string writable, no string literals.

C++C++
#include <iostream>
#include <sstream>
#include <string>
#include <vector>

int main() {
    std::string text = "  hello\t world \n again ";
    std::istringstream in(text);
    std::vector<std::string> words;
    std::string word;
    while (in >> word) {
        words.push_back(word);
    }
    std::cout << words.size() << " words\n"; // 3 words
}

operator>> treats any run of std::isspace characters (locale-dependent) as one separator — the idiomatic pre-regex answer. std::regex works too but is ~100x slower for this.

C#C#
string text = "  hello\t world \n again ";
string[] words = text.Split((char[]?)null, StringSplitOptions.RemoveEmptyEntries);
Console.WriteLine(string.Join("|", words)); // hello|world|again

Passing a null separator is the documented whitespace mode: it splits on runs of whitespace and drops empties — no regex, no trim. Regex.Split(text.Trim(), @"\s+") is the equivalent when you need custom classes.

JvJava
public class SplitWhitespace {
    public static void main(String[] args) {
        String[] words = "  hello\t world \n again ".trim().split("\\s+");
        System.out.println(String.join("|", words)); // hello|world|again
    }
}

trim() is load-bearing: since Java 8, split() drops only trailing empties, so leading whitespace yields an empty first element. "\\s" is ASCII-only — prepend (?U) for Unicode whitespace.

SwSwift
let words = "  hello\t world \n again ".split(whereSeparator: \.isWhitespace)
print(words) // ["hello", "world", "again"]

split(whereSeparator:) omits empty subsequences by default, so runs and leading/trailing whitespace vanish. Elements are Substring (cheap views) — wrap String($0) when you need owned copies.

KtKotlin
fun main() {
    val words = "  hello\t world \n again ".trim().split(Regex("\\s+"))
    println(words) // [hello, world, again]
}

split(Regex(...)) keeps an empty first element when the string starts with whitespace — trim() first. Also avoid split(" "): like everywhere else it yields empties between doubled spaces.

RbRuby
words = "  hello\t world \n again ".split
puts words.inspect # ["hello", "world", "again"]

No-arg split is awk-mode: any whitespace run, both ends stripped, no empties. The surprise is that split(' ') behaves the same (special-cased), while split(/ /) splits on exactly one space and yields nils for runs.

ZigZig
const std = @import("std");

pub fn main() void {
    const text = "  hello\t world \n again ";
    var it = std.mem.tokenizeAny(u8, text, " \t\n\r\x0b\x0c");
    while (it.next()) |tok| {
        std.debug.print("{s}\n", .{tok});
    }
}

tokenizeAny never yields empty tokens — it advances past runs of the delimiter set. Delimiters are ASCII bytes only; list all six whitespace bytes explicitly, and prefer tokenizeScalar with a predicate when you go Unicode.

Keep going

Read the regex-tokens cheatsheet →