Skip to content

Extraer una subcadena de forma segura (semántica de slicing) snippet

Cada lenguaje corta de manera distinta en cuanto los límites se salen de rango: Python y JavaScript ajustan los índices fuera de rango y nunca lanzan, mientras que Go entra en pánico, Rust entra en pánico por partida doble (fuera de rango Y en un límite que no es de carácter), y Java, Kotlin y C# lanzan cuando el inicio pasa del final.

Cada lenguaje corta de manera distinta en cuanto los límites se salen de rango: Python y JavaScript ajustan los índices fuera de rango y nunca lanzan, mientras que Go entra en pánico, Rust entra en pánico por partida doble (fuera de rango Y en un límite que no es de carácter), y Java, Kotlin y C# lanzan cuando el inicio pasa del final. Las unidades también muerden — Go, el substr de PHP, C, C++ y Zig cortan bytes y partirán por la mitad un rune UTF-8 multibyte, mientras que JavaScript, Java y C# cuentan unidades de código UTF-16 así que los caracteres astrales cuentan por dos. Estas impls le dan a cada lenguaje el slice con clamping y soporte de índices negativos que degrada a cadena vacía en lugar de romperse, y señalan dónde debes cambiar a una API consciente de caracteres.

Receta ejecutable · 15 lenguajes
Text & Parsingstringsubstringslicingbounds-checkunicode

Every language

15 lenguajes, copy-ready. One at a time with syntax highlighting, or all inline.

SQLSQLrunnable
-- SQL-standard form (PostgreSQL, MySQL, SQLite): 1-based, bounds CLAMP.
-- A start past the end yields '', an over-running length stops at the end.
SELECT SUBSTRING('hello world' FROM 7 FOR 5)  AS mid,      -- 'world'
       SUBSTRING('hello world' FROM 7)        AS to_end,   -- 'world'
       SUBSTRING('hello world' FROM 99 FOR 5) AS past_end, -- ''
       SUBSTRING('hello world' FROM 2 FOR 99) AS over_len; -- 'ello world'

Subtle divergence: MySQL also accepts SUBSTRING(s, -5) where the start counts from the end, while PostgreSQL treats a negative start as 'before position 1' — don't port that trick blindly. SQL Server uses the comma form SUBSTRING(s, start, len). Length is in characters for text types, so no UTF-8 splitting at the cut.

Run in the SQL playground →
JSJavaScript
function slice(s, start, end = s.length) {
  return s.slice(start, end); // clamps out-of-range bounds, accepts negatives
}

console.log(slice('hello world', 6, 11)); // 'world'
console.log(slice('hello world', -5));    // 'world' — negative counts from the end
console.log(slice('hello world', 42));    // '' — no throw, ever
console.log('ab'.substring(1, 0));        // 'a' — substring() swaps bounds to (0,1); slice returns ''

String.prototype.slice is the safe primitive: clamped bounds, negative offsets. substring() swaps reversed bounds and rejects negatives; both index UTF-16 code units, so an astral character (emoji) counts as two — Array.from(s) first when that matters.

TSTypeScript
function slice(s: string, start: number, end: number = s.length): string {
  return s.slice(start, end);
}

const s: string = 'hello world';
console.log(slice(s, 6, 11)); // world
console.log(slice(s, -5));    // world — negative counts from the end
console.log(slice(s, 42));    // '' — clamped, never throws

Bounds are validated at runtime, so the signature stays string with no undefined sneaking in — unlike plain s[i], which is string | undefined under noUncheckedIndexedAccess.

GoGo
package main

import "fmt"

// slice returns s[start:end] in RUNE indices, clamped to [0, n].
// Negative indices count back from the end, Python/JS-style.
func slice(s string, start int, end ...int) string {
	r := []rune(s)
	n := len(r)
	e := n
	if len(end) > 0 {
		e = end[0]
	}
	if start < 0 {
		start += n
	}
	if e < 0 {
		e += n
	}
	start = min(max(start, 0), n) // min/max builtins: Go 1.21+
	e = min(max(e, 0), n)
	if start >= e {
		return ""
	}
	return string(r[start:e])
}

func main() {
	s := "héllo 世界"
	fmt.Println(slice(s, 1, 4)) // éll — rune indices, not bytes
	fmt.Println(slice(s, -2))   // 世界 — negative counts from the end
	fmt.Println(slice(s, 99))   // "" — clamped; s[99:] would panic
}

Native s[a:b] is byte-based, panics when out of range, and happily splits a multi-byte rune in half. Going through []rune fixes both; convert back with string(...).

RsRust
fn main() {
    let s = "héllo 世界";

    // Byte-index slicing PANICS on a non-char-boundary: &s[1..4] aborts
    // because byte 1 is the middle of 'é'. get() returns Option instead.
    match s.get(0..6) {
        Some(sub) => println!("{sub}"), // héllo
        None => println!("out of range or not a char boundary"),
    }

    // Char-index slicing: collect runes, then skip/take — clamps by design.
    let chars: Vec<char> = s.chars().collect();
    let sub: String = chars.into_iter().skip(6).take(2).collect();
    println!("{sub}"); // 世界
}

str indices are byte offsets and slicing is panic territory: out of range or mid-character aborts the process. get() opts into Option; chars().skip().take() opts into character semantics.

PHPPHP
<?php
$s = 'héllo 世界';

// mb_substr is character-aware; negative start counts from the end;
// null length means "to the end" (PHP 8). Out-of-range clamps to ''.
echo mb_substr($s, 1, 3, 'UTF-8'), PHP_EOL;     // éll
echo mb_substr($s, -2, null, 'UTF-8'), PHP_EOL; // 世界
echo mb_substr($s, 99, 5, 'UTF-8'), PHP_EOL;    // (empty line)

// substr is BYTE-based — byte 2 is mid-'é', so this prints mojibake:
echo substr($s, 0, 2), PHP_EOL;

mb_* functions come from the mbstring extension (bundled with most PHP builds). Plain substr() counts bytes, so an offset landing inside a multi-byte character yields invalid UTF-8.

PyPython
s = "héllo 世界"

print(s[1:4])  # éll — out-of-range bounds clamp, slicing never raises
print(s[-2:])  # 世界
print(repr(s[99:]))  # '' — forgiving...
# ...but single-index access is NOT:
try:
    s[99]
except IndexError as err:
    print("IndexError:", err)

s[i:j] clamps silently and returns "" when start > end; s[i] raises IndexError instead. Indices are code points, so combining marks and emoji sequences can still be split — the grapheme package handles those.

CC
#include <stdio.h>
#include <string.h>

/* Copy s[start .. start+len) into out (size bytes), clamping both bounds.
   Never reads past the end of s, always NUL-terminates. Returns out. */
char *slice(char *out, size_t size, const char *s, size_t start, size_t len) {
    size_t slen = strlen(s);
    if (start > slen) {
        start = slen;
    }
    size_t avail = slen - start;
    if (len > avail) {
        len = avail;
    }
    if (len >= size) {
        len = size - 1; /* leave room for the NUL */
    }
    memcpy(out, s + start, len);
    out[len] = '\0';
    return out;
}

int main(void) {
    char buf[16];
    printf("%s\n", slice(buf, sizeof buf, "hello world", 6, 5));  /* world */
    printf("[%s]\n", slice(buf, sizeof buf, "hello world", 42, 5)); /* [] */
    return 0;
}

No stdlib slice helper exists — the idiom is a clamping memcpy with an explicit output size; that size argument is the only thing standing between you and a buffer overflow. C strings are byte arrays, so a UTF-8 rune can straddle the cut.

C++C++
#include <iostream>
#include <string>

// Clamped substr: plain s.substr(42) throws std::out_of_range —
// this returns "" instead. An over-running len is clamped by substr itself.
std::string slice(const std::string& s, std::size_t start,
                  std::size_t len = std::string::npos) {
    if (start >= s.size()) return "";
    return s.substr(start, len);
}

int main() {
    std::string s = "hello world";
    std::cout << slice(s, 6, 5) << '\n';   // world
    std::cout << '[' << slice(s, 42) << "]\n"; // [] — no throw
}

substr(pos, len) throws std::out_of_range only when pos > size(); len is clamped for you. std::string_view slices without copying (still byte-based).

C#C#
using System;

static string Slice(string s, int start, int? end = null)
{
    int n = s.Length;
    int from = start < 0 ? n + start : start;
    int to = end is int raw ? (raw < 0 ? n + raw : raw) : n;
    from = Math.Clamp(from, 0, n);
    to = Math.Clamp(to, 0, n);
    return from < to ? s[from..to] : "";
}

Console.WriteLine(Slice("hello world", -5));    // world
Console.WriteLine(Slice("hello world", 6, 11)); // world
Console.WriteLine($"[{Slice("hello world", 42)}]"); // [] — Substring(42) would throw

Substring throws ArgumentOutOfRangeException on any out-of-range pos/len, but the range operator s[a..b] clamps like JS slice — the helper just adds Python-style negatives. Indices are UTF-16 code units; StringInfo for surrogate pairs. C# 9+ for top-level statements.

JvJava
public class SafeSlice {
    // Python-style slice: negative counts from the end; both bounds clamp.
    static String slice(String s, int start, Integer end) {
        int n = s.length();
        int from = Math.min(Math.max(start < 0 ? n + start : start, 0), n);
        int e = Math.min(Math.max(end == null ? n : (end < 0 ? n + end : end), 0), n);
        return from < e ? s.substring(from, e) : "";
    }

    public static void main(String[] args) {
        System.out.println(slice("hello world", -5, null)); // world
        System.out.println("[" + slice("hello world", 42, 50) + "]"); // [] — substring would throw
    }
}

substring throws StringIndexOutOfBoundsException when from > to or to > length — normalize bounds first. Indices are UTF-16 code units, so an astral character occupies two.

SwSwift
let s = "héllo 世界"

// String indices are opaque — never integers. offsetBy TRAPS if it runs
// off the end; use the limitedBy: variant when bounds are untrusted.
let start = s.index(s.startIndex, offsetBy: 1)
let end = s.index(s.startIndex, offsetBy: 4)
print(String(s[start..<end])) // éll — grapheme-cluster safe

// prefix/suffix are the clamping API — impossible to cut a Character:
print(String(s.prefix(3))) // hél
print(String(s.suffix(2))) // 世界

if let i = s.index(s.startIndex, offsetBy: 99, limitedBy: s.endIndex) {
    print(String(s[i...]))
} else {
    print("") // out of range — degraded, not crashed
}

String.Index exists precisely so you can't slice by code unit and split a grapheme. Need random access by integer? Array(s) turns characters into a normal array.

KtKotlin
fun slice(s: String, start: Int, end: Int = s.length): String {
    val n = s.length
    val from = (if (start < 0) n + start else start).coerceIn(0, n)
    val to = (if (end < 0) n + end else end).coerceIn(0, n)
    return if (from < to) s.substring(from, to) else ""
}

fun main() {
    println(slice("hello world", -5)) // world
    println("[" + slice("hello world", 42) + "]") // [] — substring() would throw
    println("hello world".take(5)) // hello — take/drop clamp for free
}

substring(startIndex, endIndex) throws StringIndexOutOfBoundsException the moment a bound is off — coerceIn clamps first. take(n) and drop(n) are the built-in clamped prefix/suffix.

RbRuby
s = 'héllo 世界'

puts s[1, 3]            # éll — character-based for UTF-8 strings
puts s[-2, 2]           # 世界 — negative start counts from the end
puts s[99, 5].inspect   # nil — NOT "": a start past the end yields nil
puts s[99].inspect      # nil — single-index too

String#[](start, length) never raises: a too-far start yields nil (silent — puts nil looks like a success) and an over-running length clamps to the end. Normalize with .to_s or || "" when nil is unacceptable.

ZigZig
const std = @import("std");

/// s[start..end] clamped to [0, s.len]; a >= b yields an empty slice.
/// Byte-based: a multi-byte UTF-8 rune can straddle the cut.
fn slice(s: []const u8, start: usize, end: usize) []const u8 {
    const a = @min(start, s.len);
    const b = @min(end, s.len);
    if (a >= b) return s[0..0];
    return s[a..b];
}

pub fn main() void {
    const s = "hello world";
    std.debug.print("{s}\n", .{slice(s, 6, 11)}); // world
    std.debug.print("[{s}]\n", .{slice(s, 42, 99)}); // [] — clamped
}

Slicing is bounds-checked in safe builds (panic on overflow), so @min clamping is what turns a bad bound into "". There is no string type distinct from bytes — decode with std.unicode for rune-safe cuts.