Skip to content

文字列を大文字小文字を無視して比較する snippet

2つの文字列がケースを無視して「同じ」かを判定する — ユーザー入力、ヘッダー名、enum 値の照合です。Unicode のせいでこれは3つの実際の判断になります: 単純 fold か完全 fold か (ドイツ語の ß と SS が等しくなるのは完全 fold のときだけ — Python の casefold は yes、JS の toLowerCase は no)、ロケール (トルコ語のドットなし ı のため 'I'.toLowerCase() はロケール依存)、そしてバイトかコードポイントか (C の strcasecmp は ASCII で止まる)。デフォルトは言語が持つロケール非依存の Unicode fold (EqualFold、OrdinalIgnoreCase、casefold、equalsIgnoreCase) を使い、テキストがユーザーに見えるものなら、そのときだけロケール対応の機構 (Intl.Collator、ICU collation、明示的な Locale) に手を伸ばします。

2つの文字列がケースを無視して「同じ」かを判定する — ユーザー入力、ヘッダー名、enum 値の照合です。Unicode のせいでこれは3つの実際の判断になります: 単純 fold か完全 fold か (ドイツ語の ß と SS が等しくなるのは完全 fold のときだけ — Python の casefold は yes、JS の toLowerCase は no)、ロケール (トルコ語のドットなし ı のため 'I'.toLowerCase() はロケール依存)、そしてバイトかコードポイントか (C の strcasecmp は ASCII で止まる)。デフォルトは言語が持つロケール非依存の Unicode fold (EqualFold、OrdinalIgnoreCase、casefold、equalsIgnoreCase) を使い、テキストがユーザーに見えるものなら、そのときだけロケール対応の機構 (Intl.Collator、ICU collation、明示的な Locale) に手を伸ばします。

Runnable recipe · 15 languages
Text & Parsingstringscase-insensitivecomparisonunicode

Every language

15 languages, copy-ready. One at a time with syntax highlighting, or all inline.

SQLSQLrunnable
-- PostgreSQL is case-SENSITIVE by default — fold explicitly:
SELECT LOWER('Hello') = LOWER('HELLO') AS plain_ci; -- true

-- ICU collation at primary strength folds ß=ss for real:
CREATE COLLATION IF NOT EXISTS ci_primary (
  provider = icu, locale = 'und-u-ks-level1', deterministic = false
);
SELECT 'Straße' = 'STRASSE' COLLATE ci_primary AS esszett_ci; -- true

Dialects diverge hard: MySQL's default utf8mb4_0900_ai_ci is ALREADY case-insensitive (=, LIKE), SQL Server depends on the column collation (_CI_/_CS_ suffix), Postgres needs LOWER() or a nondeterministic ICU collation (which also makes LIKE behave case-insensitively). Indexes on LOWER(col) are what keep this fast.

Run in the SQL playground →
JSJavaScript
const a = 'Straße';
const b = 'STRASSE';

// toLowerCase is Unicode-aware but does NOT full-fold ß to ss:
console.log(a.toLowerCase() === b.toLowerCase()); // false

// Full folding lives in Intl.Collator, not in string methods:
const de = new Intl.Collator('de', { sensitivity: 'base' });
console.log(de.compare(a, b) === 0); // true

// Locale bites on the Turkish dotless i:
console.log('I'.toLowerCase() === 'i');           // true — locale-free
console.log('I'.toLocaleLowerCase('tr') === 'i'); // false — it is 'ı'

toLowerCase()/toUpperCase() are locale-independent full mappings; toLocaleLowerCase(locale) is the culture-aware variant. For matching, prefer collator.compare(x, y) === 0 — sensitivity: 'base' ignores case AND accents, sensitivity: 'case' ignores case only.

TSTypeScript
function equalsIgnoreCase(a: string, b: string, locale?: string): boolean {
  if (locale) {
    return a.toLocaleLowerCase(locale) === b.toLocaleLowerCase(locale);
  }
  return a.toLowerCase() === b.toLowerCase();
}

console.log(equalsIgnoreCase('Hello', 'HELLO'));          // true
console.log(equalsIgnoreCase('TITLE', 'title', 'tr'));   // false — İ/ı rules

Both sides must go through the SAME fold — mixing toLowerCase() with toLocaleLowerCase(locale) compares two different mappings. Intl.Collator is the full-folding alternative when ß/ss or other expansions must match.

GoGo
package main

import (
	"fmt"
	"strings"
)

func main() {
	fmt.Println(strings.EqualFold("Hello", "HELLO"))              // true
	fmt.Println(strings.EqualFold("Straße", "STRASSE"))           // false — simple fold keeps ß
	fmt.Println(strings.EqualFold("İstanbul", "istanbul"))       // false — no locale in simple folding
}

strings.EqualFold is Unicode SIMPLE case folding, locale-free, with an ASCII fast path — the right default for identifiers and protocol strings. It is not the Turkish locale and it does not expand ß to ss; there is no stdlib full-fold compare.

RsRust
fn main() {
    let a = "Hello";
    let b = "HELLO";

    println!("{}", a.eq_ignore_ascii_case(b)); // true — ASCII only, zero-cost

    // Unicode: lowercase both (full, locale-free mapping) and compare:
    println!("{}", a.to_lowercase() == b.to_lowercase());                  // true
    println!("{}", "Straße".to_lowercase() == "STRASSE".to_lowercase());  // false — ß stays ß
}

eq_ignore_ascii_case is a memcmp-grade fast path that ignores every non-ASCII byte pair. to_lowercase() is the Unicode default case mapping; for FULL case folding (ß = ss) use the unicase crate's UniCase::new.

PHPPHP
<?php
// strcasecmp: POSIX, ASCII, byte-safe — fine for headers and enums:
var_dump(strcasecmp('Hello', 'HELLO') === 0); // true

// Unicode text: mbstring, and pass the encoding explicitly:
var_dump(mb_strtolower('HÉLLO', 'UTF-8') === mb_strtolower('héllo', 'UTF-8')); // true
var_dump(mb_stripos('HÉLLO', 'héllo', 0, 'UTF-8') === 0);                      // true

// Folding asymmetry — lowercasing never turns ß into ss:
var_dump(mb_strtolower('Straße', 'UTF-8') === mb_strtolower('STRASSE', 'UTF-8')); // false

strcasecmp/stristr on UTF-8 work only while the case difference lives in ASCII bytes — é vs É compares unequal. The mb_* family needs the encoding parameter (or a set mb_internal_encoding) or it silently truncates at the first non-ASCII byte.

PyPython
print("Hello".lower() == "HELLO".lower())             # True
print("Straße".casefold() == "STRASSE".casefold())   # True — casefold, not lower
print("I".casefold() == "i")  # casefold is locale-free — no Turkish variant exists

casefold() is the aggressive, locale-free FULL fold designed exactly for caseless matching: ß→ss, fi→fi. lower() is display casing and misses those expansions. Python has no per-locale case mapping at all — locale.setlocale() does not affect str.lower().

CC
#include <stdio.h>
#include <strings.h> /* strcasecmp lives here, NOT string.h, on POSIX */

int main(void) {
    printf("%d\n", strcasecmp("Hello", "HELLO") == 0); /* 1 — ASCII fold */

    /* Only ASCII folds — the UTF-8 bytes of É vs é differ, so two strings
       that differ ONLY by the case of an accented letter compare UNEQUAL.
       This prints 0: */
    const char *u = "caf\xC3\x89"; /* café with capital É */
    const char *l = "caf\xC3\xA9"; /* café with lowercase é */
    printf("%d\n", strcasecmp(u, l) == 0); /* 0 */
    return 0;
}

strcasecmp is POSIX (strings.h; use _stricmp/<string.h> on Windows). It folds ASCII only — for real Unicode fold to a canonical form with a library (utf8proc, ICU) or decode UTF-8 and fold code points yourself. And never feed char straight to tolower(): non-ASCII bytes are negative, which is UB.

C++C++
#include <algorithm>
#include <cctype>
#include <iostream>
#include <string>

// ASCII fold both, comparing equal lengths first. The unsigned char casts
// are load-bearing: a signed char with the high bit set is UB for tolower.
bool equals_ignore_case_ascii(const std::string& a, const std::string& b) {
    return a.size() == b.size() &&
           std::equal(a.begin(), a.end(), b.begin(), b.end(),
                      [](unsigned char x, unsigned char y) {
                          return std::tolower(x) == std::tolower(y);
                      });
}

int main() {
    std::cout << std::boolalpha
              << equals_ignore_case_ascii("Hello", "HELLO") << '\n'; // true
}

std::equal with the two-range overload is C++14. This is ASCII-only; boost::algorithm::iequals adds locales but still works on bytes. Real Unicode caseless compare is ICU territory: icu::UnicodeString::caseCompare.

C#C#
using System;

string a = "Hello", b = "HELLO";

// OrdinalIgnoreCase is the correct default — fast, stable, culture-free:
Console.WriteLine(string.Equals(a, b, StringComparison.OrdinalIgnoreCase)); // True

// Culture variants exist and diverge on exactly the characters you fear:
Console.WriteLine("ß".ToUpperInvariant() == "SS"); // True — fold asymmetries are real
Console.WriteLine(string.Equals("i", "I", StringComparison.InvariantCultureIgnoreCase)); // True — but 'ı' is not 'i' in tr-TR

Rule of thumb from the docs themselves: use OrdinalIgnoreCase for identifiers, paths, protocol strings; culture-aware ignore-case only for DISPLAYED-to-human comparisons. == on strings is ordinal case-SENSITIVE and won't compile against a StringComparison — the overload lives on string.Equals.

JvJava
import java.util.Locale;

public class IgnoreCase {
    public static void main(String[] args) {
        System.out.println("Hello".equalsIgnoreCase("HELLO")); // true

        // equalsIgnoreCase is locale-free, but (to|from)String case methods
        // are not — the classic Turkish-I production bug:
        Locale tr = Locale.forLanguageTag("tr");
        System.out.println("title".toUpperCase(tr));       // TİTLE — dotted capital İ
        System.out.println("title".toUpperCase(Locale.ROOT)); // TITLE
    }
}

equalsIgnoreCase compares code point pairs via Character.toUpperCase/toLowerCase: Unicode-aware but simple — ß ≠ SS, and never locale-specific. ALWAYS pass an explicit Locale to toUpperCase/toLowerCase; the no-arg versions use the JVM default locale and drift by deployment machine.

SwSwift
let a = "Hello", b = "HELLO"

print(a.caseInsensitiveCompare(b) == .orderedSame) // true — Unicode-aware
print(a.lowercased() == b.lowercased())            // true — full case mapping

// Neither is locale-aware; hand the locale in when the text has one:
let tr = Locale(identifier: "tr")
print("I".lowercased(with: tr) == "ı") // true — dotted/dotless i

caseInsensitiveCompare (and compare(options: .caseInsensitive)) use a canonical case-insensitive mapping — right for matching, and grapheme-cluster safe. lowercased()/uppercased() are for display; lowercased(with:) applies a specific locale's rules.

KtKotlin
fun main() {
    println("Hello".equals("HELLO", ignoreCase = true)) // true

    val headers = mapOf("Content-Type" to "application/json")
    println(headers.keys.any { it.equals("content-type", ignoreCase = true) }) // true
}

ignoreCase = true folds with Character.toLowerCase/toUpperCase per code point — Unicode-aware but simple: ß vs SS stays unequal, and there is no locale parameter. For locale rules, use toLowerCase(locale) on BOTH sides first.

RbRuby
puts 'Hello'.casecmp?('HELLO')     # true — Unicode-aware fold
puts 'ä'.casecmp?('Ä')             # true — not just ASCII
puts 'Straße'.casecmp?('STRASSE') # false — no ß→ss expansion

# casecmp (no ?) returns -1/0/1 for sorting instead of a boolean:
puts 'Hello'.casecmp('HELLO').zero? # true

casecmp? (Ruby 2.4+) is the boolean you want; before that, casecmp == 0. It is Unicode simple case folding — accented pairs match, but full-fold expansions (ß=ss) and locale rules (Turkish i) do not; gem 'unicode' or ICU bindings cover those.

ZigZig
const std = @import("std");

/// Case-insensitive equality — ASCII only. UTF-8 bytes above 0x7F
/// pass through untouched, so accented pairs compare unequal.
fn eqlIgnoreCase(a: []const u8, b: []const u8) bool {
    if (a.len != b.len) return false;
    for (a, b) |x, y| {
        if (std.ascii.toLower(x) != std.ascii.toLower(y)) return false;
    }
    return true;
}

pub fn main() void {
    std.debug.print("{}\n", .{eqlIgnoreCase("Hello", "HELLO")}); // true
}

std.ascii.toLower is a table lookup on a single byte — the honest scope of the stdlib here. For Unicode, decode runes (std.unicode.Utf8View) and fold code points with a data table; ICU/utf8proc bindings exist for the full algorithm.