1---2name: regex3description: Write correct, efficient regular expressions across different engines.4---56## Greedy vs Lazy78- `.*` is greedy—matches as much as possible; `.*?` is lazy—matches minimum9- Greedy often overshoots: `<.*>` on `<a>b</a>` matches entire string, not `<a>`10- Default quantifiers `+ * {n,}` are greedy—add `?` for lazy: `+?` `*?` `{n,}?`1112## Escaping1314- Metacharacters need escape: `\. \* \+ \? \[ \] \( \) \{ \} \| \\ \^ \$`15- Inside character class `[]`: only `]`, `\`, `^`, `-` need escape (and `^` only at start, `-` only mid)16- Literal backslash: `\\` in regex, but in strings often need `\\\\` (double escape)1718## Anchors1920- `^` start, `$` end—but behavior changes with multiline flag21- Multiline mode: `^` `$` match line starts/ends; without, only string start/end22- `\A` always string start, `\Z` always string end (not all engines)23- Word boundary `\b` matches position, not character—`\bword\b` for whole words2425## Character Classes2627- `[abc]` matches one of a, b, c; `[^abc]` matches anything except a, b, c28- Ranges: `[a-z]` `[0-9]`—but `[a-Z]` is invalid (ASCII order matters)29- Shorthand: `\d` digit, `\w` word char, `\s` whitespace; uppercase negates: `\D` `\W` `\S`30- `.` matches any char except newline—use `[\s\S]` for truly any, or `s` flag if available3132## Groups3334- Capturing `()` vs non-capturing `(?:)`—use `(?:)` when you don't need backreference35- Named groups: `(?<name>...)` or `(?P<name>...)` depending on engine36- Backreferences: `\1` `\2` refer to captured groups in same pattern37- Groups also establish scope for alternation: `cat|dog` vs `ca(t|d)og`3839## Lookahead & Lookbehind4041- Positive lookahead `(?=...)`: assert what follows, don't consume42- Negative lookahead `(?!...)`: assert what doesn't follow43- Positive lookbehind `(?<=...)`: assert what precedes44- Negative lookbehind `(?<!...)`: assert what doesn't precede45- Lookbehinds must be fixed-width in most engines—no `*` or `+` inside4647## Flags4849- `i` case-insensitive, `m` multiline (^$ match lines), `g` global (find all)50- `s` (dotall): `.` matches newline—not supported everywhere51- `u` unicode: enables `\p{}` properties, proper surrogate handling52- Flags syntax varies: `/pattern/flags` (JS), `(?flags)` inline, or function arg (Python `re.I`)5354## Engine Differences5556- JavaScript: no lookbehind until ES2018; no `\A` `\Z`; no possessive quantifiers57- Python `re`: uses `(?P<name>)` for named groups; no `\p{}` without `regex` module58- PCRE (PHP, grep -P): full features; possessive `++` `*+`; recursive patterns59- Go: RE2 engine, no backreferences, no lookahead—guaranteed linear time6061## Performance6263- Catastrophic backtracking: `(a+)+` against `aaaaaaaaaab` is exponential—avoid nested quantifiers64- Possessive quantifiers `++` `*+` prevent backtracking—use when backtracking pointless65- Atomic groups `(?>...)` don't give back chars—similar to possessive66- Anchor patterns when possible—`^prefix` is O(1), unanchored `prefix` is O(n)6768## Common Mistakes6970- Email validation: RFC-compliant regex is 6000+ chars—use simple check or library71- URL matching: edge cases are endless—use URL parser, regex for quick extraction only72- Don't use regex for HTML/XML—use a parser; regex can't handle nesting73- Forgetting to escape user input—regex injection is real; use literal escaping functions7475## Testing7677- Test edge cases: empty string, special chars, unicode, very long input78- Visualize with tools: regex101.com shows matches and explains79- Check which engine documentation you're reading—features vary significantly