Jailbreak Classifier
Scores an inbound prompt and returns a risk band so the calling agent can decide whether to escalate to a human.
Usage
python scripts/score_prompt.py < prompt.txt
Scores an inbound prompt for jailbreak/persona-override risk before the agent acts on it.
npx skillmds@latest add gl0di/jailbreak-classifier Scores an inbound prompt and returns a risk band so the calling agent can decide whether to escalate to a human.
python scripts/score_prompt.py < prompt.txt
gl0di/clawseccheck/tree/main/fixtures/clean_b66_script_docstring_classifier/skills/jailbreak-classifier commit da1b16b50b