Machiavelli Safeguard Eval

Evaluates an LLM agent's susceptibility to unethical steering prompts in a text-based adventure game environment. It probes the capability of anomaly detection systems to classify agent trajectories as ethical or unethical based on their interaction traces. Use when the user wants to benchmark on MACHIAVELLI, or asks about evaluating this task. Reports AUPRC.

qhjqhj00 d36531a 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/machiavelli-safeguard-eval commit d36531ad03

Frequently asked questions

npx skillmds add qhjqhj00/machiavelli-safeguard-eval