← back to karpathy-llm-wiki

SkillSpector · karpathy-llm-wiki

independent scanner by NVIDIA · skill by azusagasaku · how it works ↗

CAUTIONmax severity: MEDIUMrisk score: 30

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak p…; Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or…

scanned 2026-08-23

Findings (3)

HIGHAnti-Refusalconfidence: 0.24

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

examples/2026-03-19-claude-code-statusline-landscape.md

HIGHAnti-Refusalconfidence: 0.24

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

examples/claude-code-statusline-landscape.md

HIGHRogue Agentconfidence: 0.85

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

SKILL.md

What the verdicts mean

SkillSpector reports on SkillMD's shared five-tier scale. See how SkillSpector works ↗.

PASS

Overall severity LOW (risk score in the safe range)

CAUTIONthis skill

Overall severity MEDIUM

WARNING

Overall severity HIGH

FAIL

Overall severity CRITICAL

INCONCLUSIVE

Scan could not complete