Skip to content

Extraire une sous-chaîne sans risque (sémantique de slice) snippet

Chaque langage tranche différemment dès que les bornes dérapent : Python et JavaScript bornent les indices hors plage et ne lèvent jamais, alors que Go panique, Rust panique deux fois (hors plage ET hors frontière de caractère), et Java, Kotlin et C# lèvent une exception quand le départ dépasse la fin.

Chaque langage tranche différemment dès que les bornes dérapent : Python et JavaScript bornent les indices hors plage et ne lèvent jamais, alors que Go panique, Rust panique deux fois (hors plage ET hors frontière de caractère), et Java, Kotlin et C# lèvent une exception quand le départ dépasse la fin. Les unités mordent aussi — Go, substr de PHP, C, C++ et Zig tranchent des octets et couperont une rune UTF-8 multi-octets en deux, tandis que JavaScript, Java et C# comptent des unités de code UTF-16 donc les caractères astraux comptent double. Ces impls donnent à chaque langage le slice borné et conscient des indices négatifs qui se dégrade en chaîne vide au lieu de planter, et signalent où tu dois passer à une API consciente des caractères.

Recette exécutable · 15 langages
Text & Parsingstringsubstringslicingbounds-checkunicode

Every language

15 langages, copy-ready. One at a time with syntax highlighting, or all inline.

SQLSQLrunnable
-- SQL-standard form (PostgreSQL, MySQL, SQLite): 1-based, bounds CLAMP.
-- A start past the end yields '', an over-running length stops at the end.
SELECT SUBSTRING('hello world' FROM 7 FOR 5)  AS mid,      -- 'world'
       SUBSTRING('hello world' FROM 7)        AS to_end,   -- 'world'
       SUBSTRING('hello world' FROM 99 FOR 5) AS past_end, -- ''
       SUBSTRING('hello world' FROM 2 FOR 99) AS over_len; -- 'ello world'

Subtle divergence: MySQL also accepts SUBSTRING(s, -5) where the start counts from the end, while PostgreSQL treats a negative start as 'before position 1' — don't port that trick blindly. SQL Server uses the comma form SUBSTRING(s, start, len). Length is in characters for text types, so no UTF-8 splitting at the cut.

Run in the SQL playground →
JSJavaScript
function slice(s, start, end = s.length) {
  return s.slice(start, end); // clamps out-of-range bounds, accepts negatives
}

console.log(slice('hello world', 6, 11)); // 'world'
console.log(slice('hello world', -5));    // 'world' — negative counts from the end
console.log(slice('hello world', 42));    // '' — no throw, ever
console.log('ab'.substring(1, 0));        // 'a' — substring() swaps bounds to (0,1); slice returns ''

String.prototype.slice is the safe primitive: clamped bounds, negative offsets. substring() swaps reversed bounds and rejects negatives; both index UTF-16 code units, so an astral character (emoji) counts as two — Array.from(s) first when that matters.

TSTypeScript
function slice(s: string, start: number, end: number = s.length): string {
  return s.slice(start, end);
}

const s: string = 'hello world';
console.log(slice(s, 6, 11)); // world
console.log(slice(s, -5));    // world — negative counts from the end
console.log(slice(s, 42));    // '' — clamped, never throws

Bounds are validated at runtime, so the signature stays string with no undefined sneaking in — unlike plain s[i], which is string | undefined under noUncheckedIndexedAccess.

GoGo
package main

import "fmt"

// slice returns s[start:end] in RUNE indices, clamped to [0, n].
// Negative indices count back from the end, Python/JS-style.
func slice(s string, start int, end ...int) string {
	r := []rune(s)
	n := len(r)
	e := n
	if len(end) > 0 {
		e = end[0]
	}
	if start < 0 {
		start += n
	}
	if e < 0 {
		e += n
	}
	start = min(max(start, 0), n) // min/max builtins: Go 1.21+
	e = min(max(e, 0), n)
	if start >= e {
		return ""
	}
	return string(r[start:e])
}

func main() {
	s := "héllo 世界"
	fmt.Println(slice(s, 1, 4)) // éll — rune indices, not bytes
	fmt.Println(slice(s, -2))   // 世界 — negative counts from the end
	fmt.Println(slice(s, 99))   // "" — clamped; s[99:] would panic
}

Native s[a:b] is byte-based, panics when out of range, and happily splits a multi-byte rune in half. Going through []rune fixes both; convert back with string(...).

RsRust
fn main() {
    let s = "héllo 世界";

    // Byte-index slicing PANICS on a non-char-boundary: &s[1..4] aborts
    // because byte 1 is the middle of 'é'. get() returns Option instead.
    match s.get(0..6) {
        Some(sub) => println!("{sub}"), // héllo
        None => println!("out of range or not a char boundary"),
    }

    // Char-index slicing: collect runes, then skip/take — clamps by design.
    let chars: Vec<char> = s.chars().collect();
    let sub: String = chars.into_iter().skip(6).take(2).collect();
    println!("{sub}"); // 世界
}

str indices are byte offsets and slicing is panic territory: out of range or mid-character aborts the process. get() opts into Option; chars().skip().take() opts into character semantics.

PHPPHP
<?php
$s = 'héllo 世界';

// mb_substr is character-aware; negative start counts from the end;
// null length means "to the end" (PHP 8). Out-of-range clamps to ''.
echo mb_substr($s, 1, 3, 'UTF-8'), PHP_EOL;     // éll
echo mb_substr($s, -2, null, 'UTF-8'), PHP_EOL; // 世界
echo mb_substr($s, 99, 5, 'UTF-8'), PHP_EOL;    // (empty line)

// substr is BYTE-based — byte 2 is mid-'é', so this prints mojibake:
echo substr($s, 0, 2), PHP_EOL;

mb_* functions come from the mbstring extension (bundled with most PHP builds). Plain substr() counts bytes, so an offset landing inside a multi-byte character yields invalid UTF-8.

PyPython
s = "héllo 世界"

print(s[1:4])  # éll — out-of-range bounds clamp, slicing never raises
print(s[-2:])  # 世界
print(repr(s[99:]))  # '' — forgiving...
# ...but single-index access is NOT:
try:
    s[99]
except IndexError as err:
    print("IndexError:", err)

s[i:j] clamps silently and returns "" when start > end; s[i] raises IndexError instead. Indices are code points, so combining marks and emoji sequences can still be split — the grapheme package handles those.

CC
#include <stdio.h>
#include <string.h>

/* Copy s[start .. start+len) into out (size bytes), clamping both bounds.
   Never reads past the end of s, always NUL-terminates. Returns out. */
char *slice(char *out, size_t size, const char *s, size_t start, size_t len) {
    size_t slen = strlen(s);
    if (start > slen) {
        start = slen;
    }
    size_t avail = slen - start;
    if (len > avail) {
        len = avail;
    }
    if (len >= size) {
        len = size - 1; /* leave room for the NUL */
    }
    memcpy(out, s + start, len);
    out[len] = '\0';
    return out;
}

int main(void) {
    char buf[16];
    printf("%s\n", slice(buf, sizeof buf, "hello world", 6, 5));  /* world */
    printf("[%s]\n", slice(buf, sizeof buf, "hello world", 42, 5)); /* [] */
    return 0;
}

No stdlib slice helper exists — the idiom is a clamping memcpy with an explicit output size; that size argument is the only thing standing between you and a buffer overflow. C strings are byte arrays, so a UTF-8 rune can straddle the cut.

C++C++
#include <iostream>
#include <string>

// Clamped substr: plain s.substr(42) throws std::out_of_range —
// this returns "" instead. An over-running len is clamped by substr itself.
std::string slice(const std::string& s, std::size_t start,
                  std::size_t len = std::string::npos) {
    if (start >= s.size()) return "";
    return s.substr(start, len);
}

int main() {
    std::string s = "hello world";
    std::cout << slice(s, 6, 5) << '\n';   // world
    std::cout << '[' << slice(s, 42) << "]\n"; // [] — no throw
}

substr(pos, len) throws std::out_of_range only when pos > size(); len is clamped for you. std::string_view slices without copying (still byte-based).

C#C#
using System;

static string Slice(string s, int start, int? end = null)
{
    int n = s.Length;
    int from = start < 0 ? n + start : start;
    int to = end is int raw ? (raw < 0 ? n + raw : raw) : n;
    from = Math.Clamp(from, 0, n);
    to = Math.Clamp(to, 0, n);
    return from < to ? s[from..to] : "";
}

Console.WriteLine(Slice("hello world", -5));    // world
Console.WriteLine(Slice("hello world", 6, 11)); // world
Console.WriteLine($"[{Slice("hello world", 42)}]"); // [] — Substring(42) would throw

Substring throws ArgumentOutOfRangeException on any out-of-range pos/len, but the range operator s[a..b] clamps like JS slice — the helper just adds Python-style negatives. Indices are UTF-16 code units; StringInfo for surrogate pairs. C# 9+ for top-level statements.

JvJava
public class SafeSlice {
    // Python-style slice: negative counts from the end; both bounds clamp.
    static String slice(String s, int start, Integer end) {
        int n = s.length();
        int from = Math.min(Math.max(start < 0 ? n + start : start, 0), n);
        int e = Math.min(Math.max(end == null ? n : (end < 0 ? n + end : end), 0), n);
        return from < e ? s.substring(from, e) : "";
    }

    public static void main(String[] args) {
        System.out.println(slice("hello world", -5, null)); // world
        System.out.println("[" + slice("hello world", 42, 50) + "]"); // [] — substring would throw
    }
}

substring throws StringIndexOutOfBoundsException when from > to or to > length — normalize bounds first. Indices are UTF-16 code units, so an astral character occupies two.

SwSwift
let s = "héllo 世界"

// String indices are opaque — never integers. offsetBy TRAPS if it runs
// off the end; use the limitedBy: variant when bounds are untrusted.
let start = s.index(s.startIndex, offsetBy: 1)
let end = s.index(s.startIndex, offsetBy: 4)
print(String(s[start..<end])) // éll — grapheme-cluster safe

// prefix/suffix are the clamping API — impossible to cut a Character:
print(String(s.prefix(3))) // hél
print(String(s.suffix(2))) // 世界

if let i = s.index(s.startIndex, offsetBy: 99, limitedBy: s.endIndex) {
    print(String(s[i...]))
} else {
    print("") // out of range — degraded, not crashed
}

String.Index exists precisely so you can't slice by code unit and split a grapheme. Need random access by integer? Array(s) turns characters into a normal array.

KtKotlin
fun slice(s: String, start: Int, end: Int = s.length): String {
    val n = s.length
    val from = (if (start < 0) n + start else start).coerceIn(0, n)
    val to = (if (end < 0) n + end else end).coerceIn(0, n)
    return if (from < to) s.substring(from, to) else ""
}

fun main() {
    println(slice("hello world", -5)) // world
    println("[" + slice("hello world", 42) + "]") // [] — substring() would throw
    println("hello world".take(5)) // hello — take/drop clamp for free
}

substring(startIndex, endIndex) throws StringIndexOutOfBoundsException the moment a bound is off — coerceIn clamps first. take(n) and drop(n) are the built-in clamped prefix/suffix.

RbRuby
s = 'héllo 世界'

puts s[1, 3]            # éll — character-based for UTF-8 strings
puts s[-2, 2]           # 世界 — negative start counts from the end
puts s[99, 5].inspect   # nil — NOT "": a start past the end yields nil
puts s[99].inspect      # nil — single-index too

String#[](start, length) never raises: a too-far start yields nil (silent — puts nil looks like a success) and an over-running length clamps to the end. Normalize with .to_s or || "" when nil is unacceptable.

ZigZig
const std = @import("std");

/// s[start..end] clamped to [0, s.len]; a >= b yields an empty slice.
/// Byte-based: a multi-byte UTF-8 rune can straddle the cut.
fn slice(s: []const u8, start: usize, end: usize) []const u8 {
    const a = @min(start, s.len);
    const b = @min(end, s.len);
    if (a >= b) return s[0..0];
    return s[a..b];
}

pub fn main() void {
    const s = "hello world";
    std.debug.print("{s}\n", .{slice(s, 6, 11)}); // world
    std.debug.print("[{s}]\n", .{slice(s, 42, 99)}); // [] — clamped
}

Slicing is bounds-checked in safe builds (panic on overflow), so @min clamping is what turns a bad bound into "". There is no string type distinct from bytes — decode with std.unicode for rune-safe cuts.