← back to red-teaming-llms-with-garak

SkillSpector · red-teaming-llms-with-garak

independent scanner by NVIDIA · skill by mukul975 · how it works ↗

WARNINGmax severity: HIGHrisk score: 65

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direc; This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.; YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

scanned 2026-07-07

Findings (3)

HIGHAnti-Refusalconfidence: 0.9

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direc

SKILL.md

HIGHPrompt Injectionconfidence: 0.9

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

SKILL.md

HIGHYARA Matchconfidence: 0.8

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

SKILL.md

What the verdicts mean

SkillSpector reports on SkillMD's shared five-tier scale. See how SkillSpector works ↗.

PASS

Overall severity LOW (risk score in the safe range)

CAUTION

Overall severity MEDIUM

WARNINGthis skill

Overall severity HIGH

FAIL

Overall severity CRITICAL

INCONCLUSIVE

Scan could not complete