Semgrep Rule Variant Creator
Port existing Semgrep rules to new target languages with mandatory applicability analysis, test-first validation, and an independent four-phase cycle per language.
When to Use
Use this skill when:
- Porting an existing Semgrep rule to one or more target languages (e.g., "port
sql-injection to Go and Java")
- Creating language-specific variants of a universal vulnerability pattern
- Expanding rule coverage across a polyglot codebase
- Translating rules between languages with equivalent constructs
Do NOT use this skill for:
- Creating a new Semgrep rule from scratch — use
semgrep-rule-creator instead
- Running existing rules against code
- Languages where the vulnerability pattern fundamentally doesn't apply
- Minor syntax variations within the same language
Prerequisites
- Semgrep CLI installed and on PATH. Verify:
semgrep --version
- An existing Semgrep rule — either a YAML file path or inline YAML rule content.
- One or more target languages specified by name (e.g.,
go, java, python, javascript, ruby, c, cpp).
- Read the
semgrep-rule-creator skill first — it is the authoritative reference for rule creation fundamentals (taint mode vs pattern matching, test-first methodology, anti-patterns, iteration, optimization). This skill applies those same principles in a new language context.
Official docs — when to load each
| Doc |
When to load |
| Pattern examples |
Before Phase 3 — per-language constructs; do not assume 1:1 syntax. |
| Rule syntax |
YAML operators, metavariables, taint vs pattern mode. |
| Testing rules |
Phase 2–4 annotations (ruleid: / ok:) and --test behavior. |
Input Specification
This skill requires:
- Existing Semgrep rule — YAML file path or YAML rule content.
- Target languages — one or more languages to port to (e.g., "Golang and Java").
Output Specification
For each applicable target language, produce an independent directory:
<original-rule-id>-<language>/
├── <original-rule-id>-<language>.yaml # Ported Semgrep rule
└── <original-rule-id>-<language>.<ext> # Test file with annotations
Example output for porting sql-injection to Go and Java:
sql-injection-golang/
├── sql-injection-golang.yaml
└── sql-injection-golang.go
sql-injection-java/
├── sql-injection-java.yaml
└── sql-injection-java.java
Overview
Each target language goes through an independent four-phase cycle. Complete the full cycle for one language before starting the next — errors compound and become hard to debug if you create all variants first and test later.
FOR EACH target language:
Phase 1: Applicability Analysis → Verdict
Phase 2: Test Creation (Test-First)
Phase 3: Rule Creation
Phase 4: Validation
(Complete full cycle before moving to next language)
Strictness level
This workflow is strict — do not skip steps:
- Applicability analysis is mandatory. Do not assume patterns translate.
- Each language is independent. Complete the full cycle before moving to the next.
- Test-first for each variant. Never write a rule without test cases.
- 100% test pass required. "Most tests pass" is not acceptable.
Procedure
Phase 1: Applicability Analysis
Before porting, determine if the pattern applies to the target language.
- Apply the applicability questions below; do not assume the pattern translates.
- Analyze the original rule's vulnerability class and pattern against the target language:
- Does the vulnerability class exist in the target language?
- Does an equivalent construct exist (function, pattern, library)?
- Are the semantics similar enough for meaningful detection?
- Record a verdict:
APPLICABLE → Proceed with variant creation.
APPLICABLE_WITH_ADAPTATION → Proceed but note significant changes needed.
NOT_APPLICABLE → Skip this language and document why in the output.
- If
NOT_APPLICABLE, stop here for this language and move to the next target language.
Phase 2: Test Creation (Test-First)
Always write tests before the rule. No exceptions.
- Create the output directory for this language:
New-Item -ItemType Directory -Force -Path "<original-rule-id>-<language>"
- Create the test file at
<original-rule-id>-<language>/<original-rule-id>-<language>.<ext> using target-language idioms:
- Minimum 2 vulnerable cases annotated with
// ruleid: <rule-id>
- Minimum 2 safe cases annotated with
// ok: <rule-id>
- Include language-specific edge cases that differ from the original language
- Example test file (Go):
// ruleid: sql-injection-golang
db.Query("SELECT * FROM users WHERE id = " + userInput)
// ok: sql-injection-golang
db.Query("SELECT * FROM users WHERE id = ?", userInput)
Phase 3: Rule Creation
- Open Semgrep pattern examples and pattern syntax for the target language.
- Dump the AST of the test file to understand target-language node structure:
semgrep --dump-ast -l <lang> <original-rule-id>-<language>\<original-rule-id>-<language>.<ext>
- Translate patterns from the original rule to target-language syntax based on the AST dump. Do not translate syntax 1:1 — research target-language idioms.
- Update metadata in the new rule YAML:
rules[].id → <original-rule-id>-<language>
rules[].languages → [<lang>]
rules[].message → adapt if language-specific
- Adapt for idioms — handle language-specific constructs, library equivalents, and data-flow differences. Verify API semantics match; surface similarity hides differences.
- Write the rule to
<original-rule-id>-<language>/<original-rule-id>-<language>.yaml.
Phase 4: Validation
- Follow the validation commands below and Semgrep testing rules. For taint misses, use
--dataflow-traces as in the Quick Reference.
- Validate the YAML:
semgrep --validate --config "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml"
- Run the tests:
semgrep --test --config "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml" "<original-rule-id>-<language>\<original-rule-id>-<language>.<ext>"
- Checkpoint: Output MUST show
All tests passed. If not, iterate on the rule (not the tests) until it passes. "Most tests pass" is not acceptable.
- For taint-mode rules, debug data flow if tests fail:
semgrep --dataflow-traces -f "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml" "<original-rule-id>-<language>\<original-rule-id>-<language>.<ext>"
- After tests pass, optimize the rule — remove redundant patterns, tighten metavariable constraints, and ensure no false positives in the safe cases.
- Move to the next target language and repeat from Phase 1.
Pitfalls
| Rationalization |
Why It Fails |
Correct Approach |
| "Pattern structure is identical" |
Different ASTs across languages |
Always dump AST for target language |
| "Same vulnerability, same detection" |
Data flow differs between languages |
Analyze target language idioms |
| "Rule doesn't need tests since original worked" |
Language edge cases differ |
Write NEW test cases for target |
| "Skip applicability — it obviously applies" |
Some patterns are language-specific |
Complete applicability analysis first |
| "I'll create all variants then test" |
Errors compound, hard to debug |
Complete full cycle per language |
| "Library equivalent is close enough" |
Surface similarity hides differences |
Verify API semantics match |
| "Just translate the syntax 1:1" |
Languages have different idioms |
Research target language patterns |
Additional hard rules:
- Never skip Phase 1 (applicability analysis). A pattern that makes sense in Python may not exist in Go.
- Never write the rule before the test file. Test-first is mandatory.
- Never lower the pass threshold. 100% of test cases must pass.
- Never reuse the original rule's test file for a ported variant — language edge cases differ.
- If required inputs, permissions, safety boundaries, or success criteria are missing, stop and ask for clarification.
Verification
Run these commands to confirm each variant is correct:
# 1. Validate YAML syntax
semgrep --validate --config "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml"
# Expected: Configuration is valid
# 2. Run tests — must show all passed
semgrep --test --config "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml" "<original-rule-id>-<language>\<original-rule-id>-<language>.<ext>"
# Expected: All tests passed
# 3. (Taint rules only) Verify data flow traces
semgrep --dataflow-traces -f "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml" "<original-rule-id>-<language>\<original-rule-id>-<language>.<ext>"
Pass criteria:
semgrep --validate reports the configuration is valid.
semgrep --test reports All tests passed with every ruleid: case flagged and every ok: case clean.
- The output directory structure matches the Output Specification for every applicable language.
- For any
NOT_APPLICABLE language, a documented reason exists in the output.
Quick Reference
| Task |
Command |
| Run tests |
semgrep --test --config rule.yaml test-file |
| Validate YAML |
semgrep --validate --config rule.yaml |
| Dump AST |
semgrep --dump-ast -l <lang> <file> |
| Debug taint flow |
semgrep --dataflow-traces -f rule.yaml file |
Related Skills
semgrep-rule-creator — Authoritative reference for rule creation fundamentals. Consult first when uncertain about rule structure, taint mode vs pattern matching, or anti-patterns.
Documentation
Before porting rules, read relevant Semgrep documentation:
Limitations
- Use this skill only when the task clearly matches the scope described above.
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
1---2name: semgrep-rule-variant-creator3description: Ports an existing Semgrep YAML rule into per-language variants with applicability verdicts, AST dumps, and test-first ruleid/ok files. Trigger on translating a detection to Go, Java, Python, JavaScript, or another target in a polyglot tree. Not for authoring a new rule from scratch (semgrep-rule-creator) and not a scan runner for already published rules.4---5
6# Semgrep Rule Variant Creator
7
8Port existing Semgrep rules to new target languages with mandatory applicability analysis, test-first validation, and an independent four-phase cycle per language.
9
10## When to Use
11
12**Use this skill when:**
13- Porting an existing Semgrep rule to one or more target languages (e.g., "port `sql-injection` to Go and Java")
14- Creating language-specific variants of a universal vulnerability pattern
15- Expanding rule coverage across a polyglot codebase
16- Translating rules between languages with equivalent constructs
17
18**Do NOT use this skill for:**
19- Creating a new Semgrep rule from scratch — use `semgrep-rule-creator` instead
20- Running existing rules against code
21- Languages where the vulnerability pattern fundamentally doesn't apply
22- Minor syntax variations within the same language
23
24## Prerequisites
25
261. **Semgrep CLI installed and on PATH.** Verify:
27 ```powershell
28 semgrep --version
29 ```
302. **An existing Semgrep rule** — either a YAML file path or inline YAML rule content.
313. **One or more target languages** specified by name (e.g., `go`, `java`, `python`, `javascript`, `ruby`, `c`, `cpp`).
324. **Read the `semgrep-rule-creator` skill first** — it is the authoritative reference for rule creation fundamentals (taint mode vs pattern matching, test-first methodology, anti-patterns, iteration, optimization). This skill applies those same principles in a new language context.
33
34### Official docs — when to load each
35
36| Doc | When to load |
37|-----|--------------|
38| [Pattern examples](https://semgrep.dev/docs/writing-rules/pattern-examples) | Before Phase 3 — per-language constructs; do not assume 1:1 syntax. |
39| [Rule syntax](https://semgrep.dev/docs/writing-rules/rule-syntax) | YAML operators, metavariables, taint vs pattern mode. |
40| [Testing rules](https://semgrep.dev/docs/writing-rules/testing-rules) | Phase 2–4 annotations (`ruleid:` / `ok:`) and `--test` behavior. |
41
42## Input Specification
43
44This skill requires:
451. **Existing Semgrep rule** — YAML file path or YAML rule content.
462. **Target languages** — one or more languages to port to (e.g., "Golang and Java").
47
48## Output Specification
49
50For each applicable target language, produce an independent directory:
51
52```
53<original-rule-id>-<language>/
54├── <original-rule-id>-<language>.yaml # Ported Semgrep rule
55└── <original-rule-id>-<language>.<ext> # Test file with annotations
56```
57
58Example output for porting `sql-injection` to Go and Java:
59
60```
61sql-injection-golang/
62├── sql-injection-golang.yaml
63└── sql-injection-golang.go
64
65sql-injection-java/
66├── sql-injection-java.yaml
67└── sql-injection-java.java
68```
69
70## Overview
71
72Each target language goes through an **independent four-phase cycle**. Complete the full cycle for one language before starting the next — errors compound and become hard to debug if you create all variants first and test later.
73
74```
75FOR EACH target language:
76 Phase 1: Applicability Analysis → Verdict
77 Phase 2: Test Creation (Test-First)
78 Phase 3: Rule Creation
79 Phase 4: Validation
80 (Complete full cycle before moving to next language)
81```
82
83### Strictness level
84
85This workflow is **strict** — do not skip steps:
86- **Applicability analysis is mandatory.** Do not assume patterns translate.
87- **Each language is independent.** Complete the full cycle before moving to the next.
88- **Test-first for each variant.** Never write a rule without test cases.
89- **100% test pass required.** "Most tests pass" is not acceptable.
90
91## Procedure
92
93### Phase 1: Applicability Analysis
94
95Before porting, determine if the pattern applies to the target language.
96
971. Apply the applicability questions below; do not assume the pattern translates.
982. Analyze the original rule's vulnerability class and pattern against the target language:
99 - Does the vulnerability class exist in the target language?
100 - Does an equivalent construct exist (function, pattern, library)?
101 - Are the semantics similar enough for meaningful detection?
1023. Record a verdict:
103 - `APPLICABLE` → Proceed with variant creation.
104 - `APPLICABLE_WITH_ADAPTATION` → Proceed but note significant changes needed.
105 - `NOT_APPLICABLE` → Skip this language and document why in the output.
1064. If `NOT_APPLICABLE`, stop here for this language and move to the next target language.
107
108### Phase 2: Test Creation (Test-First)
109
110**Always write tests before the rule.** No exceptions.
111
1121. Create the output directory for this language:
113 ```powershell
114 New-Item -ItemType Directory -Force -Path "<original-rule-id>-<language>"
115 ```
1162. Create the test file at `<original-rule-id>-<language>/<original-rule-id>-<language>.<ext>` using target-language idioms:
117 - Minimum **2 vulnerable cases** annotated with `// ruleid: <rule-id>`
118 - Minimum **2 safe cases** annotated with `// ok: <rule-id>`
119 - Include language-specific edge cases that differ from the original language
1203. Example test file (Go):
121 ```go
122 // ruleid: sql-injection-golang
123 db.Query("SELECT * FROM users WHERE id = " + userInput)
124
125 // ok: sql-injection-golang
126 db.Query("SELECT * FROM users WHERE id = ?", userInput)
127 ```
128
129### Phase 3: Rule Creation
130
1311. Open Semgrep [pattern examples](https://semgrep.dev/docs/writing-rules/pattern-examples) and [pattern syntax](https://semgrep.dev/docs/writing-rules/pattern-syntax) for the target language.
1322. **Dump the AST** of the test file to understand target-language node structure:
133 ```powershell
134 semgrep --dump-ast -l <lang> <original-rule-id>-<language>\<original-rule-id>-<language>.<ext>
135 ```
1363. **Translate patterns** from the original rule to target-language syntax based on the AST dump. Do not translate syntax 1:1 — research target-language idioms.
1374. **Update metadata** in the new rule YAML:
138 - `rules[].id` → `<original-rule-id>-<language>`
139 - `rules[].languages` → `[<lang>]`
140 - `rules[].message` → adapt if language-specific
1415. **Adapt for idioms** — handle language-specific constructs, library equivalents, and data-flow differences. Verify API semantics match; surface similarity hides differences.
1426. Write the rule to `<original-rule-id>-<language>/<original-rule-id>-<language>.yaml`.
143
144### Phase 4: Validation
145
1461. Follow the validation commands below and Semgrep [testing rules](https://semgrep.dev/docs/writing-rules/testing-rules). For taint misses, use `--dataflow-traces` as in the Quick Reference.
1472. **Validate the YAML:**
148 ```powershell
149 semgrep --validate --config "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml"
150 ```
1513. **Run the tests:**
152 ```powershell
153 semgrep --test --config "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml" "<original-rule-id>-<language>\<original-rule-id>-<language>.<ext>"
154 ```
1554. **Checkpoint:** Output MUST show `All tests passed`. If not, iterate on the rule (not the tests) until it passes. "Most tests pass" is not acceptable.
1565. **For taint-mode rules**, debug data flow if tests fail:
157 ```powershell
158 semgrep --dataflow-traces -f "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml" "<original-rule-id>-<language>\<original-rule-id>-<language>.<ext>"
159 ```
1606. **After tests pass**, optimize the rule — remove redundant patterns, tighten metavariable constraints, and ensure no false positives in the safe cases.
1617. **Move to the next target language** and repeat from Phase 1.
162
163## Pitfalls
164
165| Rationalization | Why It Fails | Correct Approach |
166|-----------------|--------------|------------------|
167| "Pattern structure is identical" | Different ASTs across languages | Always dump AST for target language |
168| "Same vulnerability, same detection" | Data flow differs between languages | Analyze target language idioms |
169| "Rule doesn't need tests since original worked" | Language edge cases differ | Write NEW test cases for target |
170| "Skip applicability — it obviously applies" | Some patterns are language-specific | Complete applicability analysis first |
171| "I'll create all variants then test" | Errors compound, hard to debug | Complete full cycle per language |
172| "Library equivalent is close enough" | Surface similarity hides differences | Verify API semantics match |
173| "Just translate the syntax 1:1" | Languages have different idioms | Research target language patterns |
174
175**Additional hard rules:**
176- Never skip Phase 1 (applicability analysis). A pattern that makes sense in Python may not exist in Go.
177- Never write the rule before the test file. Test-first is mandatory.
178- Never lower the pass threshold. 100% of test cases must pass.
179- Never reuse the original rule's test file for a ported variant — language edge cases differ.
180- If required inputs, permissions, safety boundaries, or success criteria are missing, stop and ask for clarification.
181
182## Verification
183
184Run these commands to confirm each variant is correct:
185
186```powershell
187# 1. Validate YAML syntax
188semgrep --validate --config "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml"
189# Expected: Configuration is valid
190
191# 2. Run tests — must show all passed
192semgrep --test --config "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml" "<original-rule-id>-<language>\<original-rule-id>-<language>.<ext>"
193# Expected: All tests passed
194
195# 3. (Taint rules only) Verify data flow traces
196semgrep --dataflow-traces -f "<original-rule-id>-<language>\<original-rule-id>-<language>.yaml" "<original-rule-id>-<language>\<original-rule-id>-<language>.<ext>"
197```
198
199**Pass criteria:**
200- `semgrep --validate` reports the configuration is valid.
201- `semgrep --test` reports `All tests passed` with every `ruleid:` case flagged and every `ok:` case clean.
202- The output directory structure matches the Output Specification for every applicable language.
203- For any `NOT_APPLICABLE` language, a documented reason exists in the output.
204
205## Quick Reference
206
207| Task | Command |
208|------|---------|
209| Run tests | `semgrep --test --config rule.yaml test-file` |
210| Validate YAML | `semgrep --validate --config rule.yaml` |
211| Dump AST | `semgrep --dump-ast -l <lang> <file>` |
212| Debug taint flow | `semgrep --dataflow-traces -f rule.yaml file` |
213
214## Related Skills
215
216- **`semgrep-rule-creator`** — Authoritative reference for rule creation fundamentals. Consult first when uncertain about rule structure, taint mode vs pattern matching, or anti-patterns.
217
218## Documentation
219
220Before porting rules, read relevant Semgrep documentation:
221
222- [Rule Syntax](https://semgrep.dev/docs/writing-rules/rule-syntax) — YAML structure and operators
223- [Pattern Syntax](https://semgrep.dev/docs/writing-rules/pattern-syntax) — Pattern matching and metavariables
224- [Pattern Examples](https://semgrep.dev/docs/writing-rules/pattern-examples) — Per-language pattern references
225- [Testing Rules](https://semgrep.dev/docs/writing-rules/testing-rules) — Testing annotations
226- [Trail of Bits Testing Handbook](https://appsec.guide/docs/static-analysis/semgrep/advanced/) — Advanced patterns
227
228## Limitations
229
230- Use this skill only when the task clearly matches the scope described above.
231- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
232- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.