1---2name: regex3description: Write correct, efficient regular expressions across different engines.4---5
6## Greedy vs Lazy
7
8- `.*` is greedy—matches as much as possible; `.*?` is lazy—matches minimum
9- Greedy often overshoots: `<.*>` on `<a>b</a>` matches entire string, not `<a>`
10- Default quantifiers `+ * {n,}` are greedy—add `?` for lazy: `+?` `*?` `{n,}?`
11
12## Escaping
13
14- Metacharacters need escape: `\. \* \+ \? \[ \] \( \) \{ \} \| \\ \^ \$`
15- Inside character class `[]`: only `]`, `\`, `^`, `-` need escape (and `^` only at start, `-` only mid)
16- Literal backslash: `\\` in regex, but in strings often need `\\\\` (double escape)
17
18## Anchors
19
20- `^` start, `$` end—but behavior changes with multiline flag
21- Multiline mode: `^` `$` match line starts/ends; without, only string start/end
22- `\A` always string start, `\Z` always string end (not all engines)
23- Word boundary `\b` matches position, not character—`\bword\b` for whole words
24
25## Character Classes
26
27- `[abc]` matches one of a, b, c; `[^abc]` matches anything except a, b, c
28- Ranges: `[a-z]` `[0-9]`—but `[a-Z]` is invalid (ASCII order matters)
29- Shorthand: `\d` digit, `\w` word char, `\s` whitespace; uppercase negates: `\D` `\W` `\S`
30- `.` matches any char except newline—use `[\s\S]` for truly any, or `s` flag if available
31
32## Groups
33
34- Capturing `()` vs non-capturing `(?:)`—use `(?:)` when you don't need backreference
35- Named groups: `(?<name>...)` or `(?P<name>...)` depending on engine
36- Backreferences: `\1` `\2` refer to captured groups in same pattern
37- Groups also establish scope for alternation: `cat|dog` vs `ca(t|d)og`
38
39## Lookahead & Lookbehind
40
41- Positive lookahead `(?=...)`: assert what follows, don't consume
42- Negative lookahead `(?!...)`: assert what doesn't follow
43- Positive lookbehind `(?<=...)`: assert what precedes
44- Negative lookbehind `(?<!...)`: assert what doesn't precede
45- Lookbehinds must be fixed-width in most engines—no `*` or `+` inside
46
47## Flags
48
49- `i` case-insensitive, `m` multiline (^$ match lines), `g` global (find all)
50- `s` (dotall): `.` matches newline—not supported everywhere
51- `u` unicode: enables `\p{}` properties, proper surrogate handling
52- Flags syntax varies: `/pattern/flags` (JS), `(?flags)` inline, or function arg (Python `re.I`)
53
54## Engine Differences
55
56- JavaScript: no lookbehind until ES2018; no `\A` `\Z`; no possessive quantifiers
57- Python `re`: uses `(?P<name>)` for named groups; no `\p{}` without `regex` module
58- PCRE (PHP, grep -P): full features; possessive `++` `*+`; recursive patterns
59- Go: RE2 engine, no backreferences, no lookahead—guaranteed linear time
60
61## Performance
62
63- Catastrophic backtracking: `(a+)+` against `aaaaaaaaaab` is exponential—avoid nested quantifiers
64- Possessive quantifiers `++` `*+` prevent backtracking—use when backtracking pointless
65- Atomic groups `(?>...)` don't give back chars—similar to possessive
66- Anchor patterns when possible—`^prefix` is O(1), unanchored `prefix` is O(n)
67
68## Common Mistakes
69
70- Email validation: RFC-compliant regex is 6000+ chars—use simple check or library
71- URL matching: edge cases are endless—use URL parser, regex for quick extraction only
72- Don't use regex for HTML/XML—use a parser; regex can't handle nesting
73- Forgetting to escape user input—regex injection is real; use literal escaping functions
74
75## Testing
76
77- Test edge cases: empty string, special chars, unicode, very long input
78- Visualize with tools: regex101.com shows matches and explains
79- Check which engine documentation you're reading—features vary significantly