← back to guardian-angel

SkillSpector · guardian-angel

independent scanner by NVIDIA · skill by aAAaqwq · how it works ↗

FAILmax severity: CRITICALrisk score: 100

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, a…; Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a di…; Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.; +10 more

scanned 2026-08-22

Findings (20)

MEDIUMAgent Snoopingconfidence: 0.8

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

plugin/index.ts

HIGHAnti-Refusalconfidence: 0.9

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

SKILL.md

HIGHAnti-Refusalconfidence: 0.22499999999999998

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

references/prompt-injection-defense.md

HIGHAnti-Refusalconfidence: 0.22499999999999998

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

references/prompt-injection-defense.md

MEDIUMData Exfiltrationconfidence: 0.7

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

SKILL.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

PLUGIN-SPEC.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

PLUGIN-SPEC.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

SKILL-v1-backup.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

SKILL-v2.1-backup.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

SKILL.md

MEDIUMExcessive Agencyconfidence: 0.85

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

SKILL.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

drafts/SKILL-v2-draft.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

drafts/SKILL-v2-draft.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

drafts/SKILL-v2-draft.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

drafts/edge-cases-consolidated.md

MEDIUMExcessive Agencyconfidence: 0.75

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

plugin/src/hook.ts

MEDIUMMemory Poisoningconfidence: 0.8

Skill attempts to fill the context window with filler content, displacing legitimate instructions and safety constraints. This can degrade agent performance or bypass safety boundaries.

PLUGIN-SPEC.md

MEDIUMMemory Poisoningconfidence: 0.8

Skill attempts to fill the context window with filler content, displacing legitimate instructions and safety constraints. This can degrade agent performance or bypass safety boundaries.

PLUGIN-SPEC.md

MEDIUMMemory Poisoningconfidence: 0.8

Skill attempts to fill the context window with filler content, displacing legitimate instructions and safety constraints. This can degrade agent performance or bypass safety boundaries.

PLUGIN-SPEC.md

MEDIUMMemory Poisoningconfidence: 0.8

Skill attempts to fill the context window with filler content, displacing legitimate instructions and safety constraints. This can degrade agent performance or bypass safety boundaries.

PLUGIN-SPEC.md

What the verdicts mean

SkillSpector reports on SkillMD's shared five-tier scale. See how SkillSpector works ↗.

PASS

Overall severity LOW (risk score in the safe range)

CAUTION

Overall severity MEDIUM

WARNING

Overall severity HIGH

FAILthis skill

Overall severity CRITICAL

INCONCLUSIVE

Scan could not complete