Skip to content

文字列を分割しつつ区切り文字を保持する snippet

文字列を分割しても区切り文字を保持する — コードのトークン化、CSV風のストリーム、句読点でテキストを失わずに区切る処理です。コツは、分割パターンが「消えるべきもの」だけを消費することです。JavaScript の split はキャプチャグループのないマッチを捨てますが、キャプチャグループは出力に結合されます。区切り文字をグループに入れると保持できます。Python では lookahead か、より新しいキャプチャ付き分割の挙動が必要です。Go の regexp.Split には保持モードがなく、同じ正規表現での find-all が定石です。繰り返し現れる驚きは、キャプチャグループ付きの split が JS/.NET/Python で [テキスト, 区切り, テキスト, 区切り...] のようにインターリーブされた形を返すことです。呼び出し側がまず予期しない形です。

文字列を分割しても区切り文字を保持する — コードのトークン化、CSV風のストリーム、句読点でテキストを失わずに区切る処理です。コツは、分割パターンが「消えるべきもの」だけを消費することです。JavaScript の split はキャプチャグループのないマッチを捨てますが、キャプチャグループは出力に結合されます。区切り文字をグループに入れると保持できます。Python では lookahead か、より新しいキャプチャ付き分割の挙動が必要です。Go の regexp.Split には保持モードがなく、同じ正規表現での find-all が定石です。繰り返し現れる驚きは、キャプチャグループ付きの split が JS/.NET/Python で [テキスト, 区切り, テキスト, 区切り...] のようにインターリーブされた形を返すことです。呼び出し側がまず予期しない形です。

Runnable recipe · 11 languages
Regex & Text Processingsplitregextokenizestringcapture-groups

Every language

11 languages, copy-ready. One at a time with syntax highlighting, or all inline.

JSJavaScript
// lookbehind: the boundary sits AFTER each separator — nothing is consumed,
// so nothing is lost:
'a=1;b=2'.split(/(?<=[;=])/);  // ['a=', '1;', 'b=', '2']

// capture form — JS JOINS capture groups into the output (interleaved):
'a=1;b=2'.split(/([;=])/);      // ['a', '=', '1', ';', 'b', '=', '2']
// glue back: for (let i = 0; i < out.length; i += 2) out[i] + (out[i + 1] ?? '')

Lookbehind is ES2018 — Safari only shipped it in 16.4 (2023), so older targets need the capture form. The capture form runs everywhere but returns the interleaved [text, sep, ...] shape with a separator at every ODD index — and a bare split drops the separator entirely, which is the data loss this recipe exists to prevent.

TSTypeScript
export function tokenize(s: string): string[] {
  // split AFTER each ; or = — the separator stays glued to its token
  return s.split(/(?<=[;=])/).filter((t) => t.length > 0);
}

tokenize('a=1;b=2'); // ['a=', '1;', 'b=', '2']

The type is the whole TS addition — the same JS engine behavior under a string[] contract. Keep the interleave shape OUT of the signature: s.split(/([;=])/)) also typechecks as string[] while returning [text, sep, ...] pairs, a shape callers of tokenize() never expect.

GoGo
import "regexp"

var sep = regexp.MustCompile(`[;=]`)

// SplitKeep slices s after each separator, keeping it glued to the token
// before it (the lookbehind split other languages get for free).
func SplitKeep(s string) []string {
	var out []string
	prev := 0
	for _, loc := range sep.FindAllStringIndex(s, -1) {
		end := loc[1] // index just past the separator
		out = append(out, s[prev:end])
		prev = end
	}
	if prev < len(s) {
		out = append(out, s[prev:])
	}
	return out
}

// SplitKeep("a=1;b=2") → ["a=" "1;" "b=" "2"]

regexp.Split(s, -1) has NO keep-separators mode, and RE2 has no lookbehind — the (?<=...) idiom cannot translate. FindAllStringIndex + slicing between match ends is the route. Hoist MustCompile to package level; compiling per call recompiles the pattern every split.

RsRust
use regex::Regex;

fn split_keep(s: &str, re: &Regex) -> Vec<String> {
    let mut out = Vec::new();
    let mut prev = 0;
    for m in re.find_iter(s) {
        out.push(s[prev..m.end()].to_owned()); // text + separator
        prev = m.end();
    }
    if prev < s.len() {
        out.push(s[prev..].to_owned()); // trailing text after the last sep
    }
    out
}

// let re = Regex::new("[;=").unwrap();
// split_keep("a=1;b=2", &re) == ["a=", "1;", "b=", "2"]

The JS/Python capture-group trick does not transfer: regex::Regex::split yields text-only pieces and IGNORES capture groups — no interleave, just silent loss. find_iter + slicing is the stable route (split_delimiter lives in nightly/external crates). For plain char sets, std's str::match_indices runs the same loop with no crate.

PHPPHP
// PREG_SPLIT_DELIM_CAPTURE is the flag — without it the separators vanish:
$parts = preg_split('/([;=])/', 'a=1;b=2', -1, PREG_SPLIT_DELIM_CAPTURE);
// ['a', '=', '1', ';', 'b', '=', '2'] — interleaved, separators at odd indexes

// glue text + separator when you want them on one token:
$pairs = [];
for ($i = 0; $i < count($parts); $i += 2) {
    $pairs[] = $parts[$i] . ($parts[$i + 1] ?? '');
}
// ['a=', '1;', 'b=', '2']

The flag only fires for CAPTURED groups — ([;=]) keeps, [;=] drops. The result has the same interleaved [text, sep, ...] shape as JS and .NET, so pair up by odd/even index; the ?? '' absorbs a dangling trailing token with no separator after it.

PyPython
import re
from itertools import zip_longest

parts = re.split(r'([;=])', 'a=1;b=2')
# ['a', '=', '1', ';', 'b', '=', '2'] — a capture group keeps the delimiters

# pair up text + separator (the zip(it, it) idiom, tail-safe):
it = iter(parts)
tokens = [a + b for a, b in zip_longest(it, it, fillvalue='')]
# ['a=', '1;', 'b=', '2']

A capture group in the pattern is what makes re.split keep the delimiters — the same regex without parens drops them. zip(it, it) is the pair-up idiom, but it silently discards a dangling odd tail; zip_longest(fillvalue='') keeps it. An uncaptured group (?:...) or a lookahead boundary drops the text again.

C#C#
using System.Text.RegularExpressions;

var parts = Regex.Split("a=1;b=2", "([;=])");
// ["a", "=", "1", ";", "b", "=", "2"] — .NET KEEPS the groups, interleaved

// glue pairs when token and separator belong together:
var tokens = new List<string>();
for (var i = 0; i < parts.Length; i += 2)
    tokens.Add(parts[i] + (i + 1 < parts.Length ? parts[i + 1] : ""));

.NET sits on the JS side of the split: Regex.Split KEEPS capture groups in the output array — String.Split(';', '=') (the char[] overload) has no keep mode at all and silently eats them. Same interleaved [text, sep, ...] shape as JS and PHP, with separators at the odd indexes.

JvJava
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

static List<String> splitKeep(String s, Pattern sep) {
    List<String> out = new ArrayList<>();
    Matcher m = sep.matcher(s);
    int prev = 0;
    while (m.find()) {
        out.add(s.substring(prev, m.end())); // text + separator, glued
        prev = m.end();
    }
    if (prev < s.length()) {
        out.add(s.substring(prev));
    }
    return out;
}

// splitKeep("a=1;b=2", Pattern.compile("[;=]")) → ["a=", "1;", "b=", "2"]

Java's split DISCARDS captured groups — Pattern.compile("([;=])").split(s) returns text-only pieces, the OPPOSITE of JS/.NET with the identical regex. Two more silent mutations: split(s) removes trailing empty strings (split(s, 0)); split(s, -1) keeps them. The find() loop is the only keep route.

SwSwift
import Foundation

func splitKeep(_ s: String) -> [String] {
    let re = /[;=]/                    // Swift 5.7 Regex literal
    var out: [String] = []
    var prev = s.startIndex
    for m in re.matches(in: s) {
        out.append(String(s[prev..<m.range.upperBound])) // text + separator
        prev = m.range.upperBound
    }
    if prev < s.endIndex {
        out.append(String(s[prev...]))
    }
    return out
}

splitKeep("a=1;b=2") // ["a=", "1;", "b=", "2"]

String.split(separator:) drops what it splits on — there is no keep flag. matches(of:)/matches(in:) on a Swift 5.7 Regex gives you each match's range as String.Index bounds, and slicing up to each upperBound is the keep route. NSRegularExpression works too, but its NSRange math (NSString bridging, UTF-16 offsets) is the trap the Regex API exists to remove.

KtKotlin
fun splitKeep(s: String, sep: Regex): List<String> =
    buildList {
        var prev = 0
        for (m in sep.findAll(s)) {
            add(s.substring(prev, m.range.last + 1)) // text + separator
            prev = m.range.last + 1
        }
        if (prev < s.length) add(s.substring(prev))
    }

// splitKeep("a=1;b=2", Regex("[;=]")) // [a=, 1;, b=, 2]

"a=1;b=2".split(Regex("([;=])")) drops the group contents — Kotlin's regex split rides on Java's Pattern.split, so the JS capture-group trick does not transfer to the JVM. findAll + buildList is the keep route; m.range is an inclusive IntRange, hence last + 1 for the end bound.

RbRuby
s = 'a=1;b=2'

# capture groups KEEP separators (limit -1 = keep trailing empty strings):
parts = s.split(/([;=])/, -1)
# ['a', '=', '1', ';', 'b', '=', '2']

# or scan with a combined pattern — separator glued to the token before it:
tokens = s.scan(/[^;=]*[;=]|[^;=]+/)
# ['a=', '1;', 'b=', '2']

Ruby's split keeps captures like JS's — but the DEFAULT limit drops trailing empty strings; pass -1 when empties matter. The scan form sidesteps both traps: one alternation matches token-plus-its-separator or a trailing bare token, so the pieces come out pre-glued with no pairing pass.

Keep going

Read the regex-tokens cheatsheet →