Skip to main content
Back to Blog

Regex Cheat Sheet: Patterns That Hold Up in Production

Copy-ready regex patterns, what each one deliberately does not match, and how to avoid the backtracking trap.

Kashyap Thakar
9 min
RegexValidationDevelopment
Regex Cheat Sheet: Patterns That Hold Up in Production

Most regex references list syntax and stop there. The bugs that reach production come from somewhere else: a pattern that matches too much, a missing anchor, or a nested quantifier that freezes a server on one unlucky input. This sheet pairs each pattern with its limits, and you can paste any of them into the Regex Tester to try it against your own data.

The Syntax You Actually Use

.      any character except newline
\d \w \s  digit, word character, whitespace   (\D \W \S are the inverses)
\b     word boundary (a position, not a character)
^ $    start / end of input (of each line with the m flag)
*  +  ?    0 or more, 1 or more, 0 or 1
{n,m}  between n and m repetitions
*? +?  lazy versions: match as little as possible
(...)  capture group      (?:...)  non-capturing group
(?<name>...)  named group
(?=...) (?!...)  lookahead / negative lookahead
(?<=...) (?<!...)  lookbehind / negative lookbehind

Flags change behaviour more than people expect: g finds every match instead of the first, i ignores case, m makes anchors work per line, s lets . match newlines, and u enables full Unicode handling. Forgetting m is the most common reason a per-line pattern appears to match only the first line of a log.

Patterns With Their Limits

Hex color

^#(?:[0-9a-fA-F]{3}|[0-9a-fA-F]{6})$

Matches #fff and #1a2b3c. It rejects the 4 and 8 digit forms that include alpha; add {4} and {8} alternatives if you need them.

ISO date (shape only)

^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$

This accepts 2026-02-31. A regex can check shape, but it cannot know February has 28 days. Use the pattern to reject obvious junk, then parse with a real date library and treat a failed parse as invalid.

UUID

^[0-9a-f]{8}-[0-9a-f]{4}-[1-8][0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$

Pair it with the i flag. The [1-8] constrains the version nibble and [89ab] the variant, so a random 32-hex string with dashes in the right places will not pass. See UUID versions explained for what those digits mean.

Semantic version

^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-[0-9A-Za-z.-]+)?(?:\+[0-9A-Za-z.-]+)?$

Follows the semver rule that numeric parts have no leading zeros, so 1.02.3 is rejected while 1.2.3-beta.1+build.5 passes.

Extract query parameters

[?&]([^=&#]+)=([^&#]*)

Fine for a quick extraction from a string. For anything that matters, use new URL(...).searchParams, which handles encoding, repeated keys and edge cases that this pattern ignores. The URL Parser shows the same breakdown interactively.

Log line timestamp and level

^(?<ts>\d{4}-\d{2}-\d{2}T[\d:.]+Z)\s+(?<level>DEBUG|INFO|WARN|ERROR)\s+(?<msg>.*)$

Named groups make the result readable: match.groups.level beats counting parentheses. Use the m flag when running it over a multi-line string.

Strip repeated whitespace

str.replace(/\s+/g, " ").trim()

Do Not Validate Emails With a Regex

The full email grammar is far larger than any pattern you will maintain, and strict patterns reject real addresses such as ones with a + tag or a new top-level domain. A sensible check is deliberately loose:

^[^\s@]+@[^\s@]+\.[^\s@]+$

It only catches typos like a missing @. The real validation is sending a confirmation email.

Greedy vs Lazy

Quantifiers are greedy by default: they take as much as they can, then give back. Given <b>one</b> and <b>two</b>, the pattern <b>.*</b> matches the whole string, while <b>.*?</b> matches each tag pair separately. A negated class such as <b>[^<]*</b> is usually both clearer and faster, because the engine never has to backtrack. Parsing real HTML with regex is still a mistake; use a parser.

Catastrophic Backtracking

When a quantified group contains something that can match the same text in more than one way, a failed match can make the engine try an exponential number of splits. The classic shape is a nested quantifier:

^(a+)+$        // against "aaaaaaaaaaaaaaaaaaaaaaaa!" can hang
^(\w+\s?)*$    // same problem on long text ending in a bad character

The input that triggers it is a long run that almost matches and then fails at the end, which is exactly what an attacker sends. Mitigations, in order of preference:

  • Remove the ambiguity: ^a+$ says the same thing without the nesting.
  • Cap input length before matching anything user-supplied.
  • Use an engine with linear-time guarantees (RE2 and Rust's regex crate) for untrusted input.
  • Test with a long near-miss string, not only with inputs that match.

Escaping Pitfalls

In a JavaScript string literal, backslashes need doubling: new RegExp("\\d+") is the same as the literal /\d+/. When building a pattern from user input, escape it first, otherwise a search for a.b also matches axb:

const escape = (s) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
new RegExp(escape(userInput), "i");

A Short Checklist

  • Anchor validation patterns with ^ and $; otherwise they match anywhere in the input.
  • Prefer [^x]* over .*? when you know the terminator.
  • Use non-capturing groups when you do not need the capture.
  • Test the empty string, a very long string, and a near-miss.
  • If you cannot explain a pattern in a sentence, add a comment or split it up.

For a deeper walkthrough of how the engine reads a pattern, see Mastering Regular Expressions.

Written by

Kashyap Thakar

Kashyap is the founder of 11Vertex, a product engineering studio that builds infrastructure, identity, and data systems for startups and growing teams. He writes these guides from problems that came up in client work, and builds the tools on ThenCatch to go with them.

Spotted an error in this article? Corrections are welcome and we update posts rather than quietly pulling them.

Part of the ThenCatch blog. Learn more about us or browse more guides.