Skip to content

Regex Tokens Explained

Every core regex token: anchors, character classes, quantifiers, groups, and flags. What each one means, with a concrete match. The quick-reference companion to the Regex Tester tool.

Regex tokens compose: stack an anchor, a character class, and a quantifier to build a pattern. ^ pins to a start, \d matches a digit, and {2,4} sets how many times. Combine them as ^\d{2,4}$ to match a whole string of two to four digits. Read a pattern left to right, one token at a time. The companion Regex Tester tool lets you try any of these tokens live against your own text.

Reference table · 55 entriesOpen the regex-tester tool →
55 of 55 rows
Anchors
Matches the position at the start of the string, or the start of a line with the m flag.^He → He
Matches the position at the end of the string, or the end of a line with the m flag.lo$ → lo
Matches a word boundary, the position between a word character and a non-word character.\bcat\b → cat
Matches a non-word boundary, the position between two word characters or two non-word characters.\Bar → ar
Character classes
Matches any single character except a line break, unless the s flag is set.a.c → abc
Matches any digit, 0 through 9.\d → 7
Matches any character that is not a digit.\D → a
Matches any word character: a letter, a digit, or an underscore.\w → _
Matches any character that is not a word character.\W → !
Matches any whitespace character: a space, a tab, or a line break.a\sb → a b
Matches any character that is not whitespace.\S → a
Matches any one of the characters listed inside the brackets.[aeiou] → e
Matches any character that is not one of those listed inside the brackets.[^aeiou] → s
Matches any single character in the given range, from a to z.[a-z] → m
Matches any single character including a line break, the common idiom for a dot that spans lines.[\s\S] → \n
Matches any unicode letter from any script, when the u flag is set.\p{L} → é
Matches any unicode number character from any script, when the u flag is set.\p{N} → ٣
Escape: matches a single line feed character.\n → line feed
Escape: matches a single tab character.\t → tab
Escape: matches a single carriage return character.\r → carriage return
Matches the single character with the given two-digit hexadecimal code.\x41 → A
Matches the single character with the given hexadecimal code point, when the u flag is set.\u{1F600} → 😀
Class intersection, with the v flag: matches characters that are in both sets at once.[[a-z]&&[^aeiou]] → t
Quantifiers
Matches the previous token zero or more times.ab* → abbb
Matches the previous token one or more times.a+ → aaa
Matches the previous token zero or one time, making it optional.colou?r → color
Matches the previous token exactly n times.a{3} → aaa
Matches the previous token n or more times.a{2,} → aaaa
Matches the previous token between n and m times, inclusive.a{2,4} → aaa
Lazy quantifier: matches the previous token zero or more times, taking as few as possible.<.*?> on <a><b> → <a>
Lazy quantifier: matches the previous token one or more times, taking as few as possible.\d+? on 123 → 1
Lazy quantifier: matches the previous token zero or one time, preferring zero.ab?? on ab → a
Lazy quantifier: matches the previous token between n and m times, taking as few as possible.a{1,3}? on aaa → a
Possessive quantifier (PCRE): matches the previous token zero or more times and keeps it, never backtracking to try other splits..*+a on aaa → no match
Groups & alternation
A capturing group: matches the sequence and remembers it for backreferences.(ab)+ → abab
A non-capturing group: matches the sequence without remembering it.(?:ab)+ → abab
Positive lookahead: matches if the next characters are abc, without consuming them.a(?=b) → a
Negative lookahead: matches if the next characters are not abc, without consuming them.a(?!b) → a
Backreference: matches the exact text that capturing group 1 captured.(a)\1 → aa
Alternation: matches either the expression before or the one after the bar.cat|dog → dog
A named capturing group: matches the sequence and remembers it under the given name.(?<year>\d{4}) → year = 2026
Positive lookbehind: matches if the preceding characters are abc, without consuming them.(?<=a)b → b
Negative lookbehind: matches if the preceding characters are not abc, without consuming them.(?<!a)b → b
Named backreference: matches the exact text that the named group captured.(?<w>\w+) \k<w> → the the
Atomic group (PCRE): matches the sequence once and never backtracks into it.(?>a+)ab on aab → no match
Branch reset (PCRE): alternation where every branch reuses the same group numbers.(?|(cat)|(dog)) → \1 = dog
Flags
Global flag: find every match in the string, not just the first./a/g on aaa → a, a, a
Case-insensitive flag: ignore letter case when matching./Hi/i → hi
Multiline flag: treat ^ and $ as the start and end of each line./^a/m matches at the start of each line
Dotall flag: let the dot match a line break too./a.b/s → a\nb
Unicode flag: treat a pattern and its input as unicode code points./\u{1F600}/u → 😀
Sticky flag: match only at the exact position set by lastIndex./ab/y matches only at lastIndex
Indices flag: adds the start and end position of every match to the result./b/d on abc → indices [1, 2]
Unicode sets flag: a stricter u that also allows set operations inside a class./[a-z&&[^aeiou]]/v → t
Extended flag (PCRE): ignores unescaped whitespace and # comments inside the pattern./a b/x → ab