SkillSpector · deep-research
independent scanner by NVIDIA · skill by richardnguyen0715 · how it works ↗
Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a di…; Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.; Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, dat…; +2 more
scanned 2026-08-23
Findings (8)
Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.
agents/ethics_review_agent.md
Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.
agents/ethics_review_agent.md
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
agents/timeline_extraction_agent.md
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
agents/ethics_review_agent.md
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
agents/socratic_mentor_agent.md
Skill attempts to fill the context window with filler content, displacing legitimate instructions and safety constraints. This can degrade agent performance or bypass safety boundaries.
agents/bibliography_agent.md
Skill attempts to fill the context window with filler content, displacing legitimate instructions and safety constraints. This can degrade agent performance or bypass safety boundaries.
agents/bibliography_agent.md
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
agents/report_compiler_agent.md
What the verdicts mean
SkillSpector reports on SkillMD's shared five-tier scale. See how SkillSpector works ↗.
Overall severity LOW (risk score in the safe range)
Overall severity MEDIUM
Overall severity HIGH
Overall severity CRITICAL
Scan could not complete