← back to research-paper-writing

SkillSpector · research-paper-writing

independent scanner by NVIDIA · skill by FlyFireF · how it works ↗

WARNINGmax severity: HIGHrisk score: 52

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.; Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a di…; Model output is used without validation or sanitization. Unvalidated output injected into downstream contexts (SQL, shell, HTML) enables injection attacks an…; +3 more

scanned 2026-08-22

Findings (6)

MEDIUMMCP Rug Pullconfidence: 0.7

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

SKILL.md

HIGHAnti-Refusalconfidence: 0.24

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

references/autoreason-methodology.md

HIGHOutput Handlingconfidence: 0.255

Model output is used without validation or sanitization. Unvalidated output injected into downstream contexts (SQL, shell, HTML) enables injection attacks and arbitrary code execution.

references/experiment-patterns.md

MEDIUMPrivilege Escalationconfidence: 0.7

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

templates/README.md

MEDIUMRogue Agentconfidence: 0.65

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

SKILL.md

HIGHYARA Matchconfidence: 0.8

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

SKILL.md

What the verdicts mean

SkillSpector reports on SkillMD's shared five-tier scale. See how SkillSpector works ↗.

PASS

Overall severity LOW (risk score in the safe range)

CAUTION

Overall severity MEDIUM

WARNINGthis skill

Overall severity HIGH

FAIL

Overall severity CRITICAL

INCONCLUSIVE

Scan could not complete