Calm Audit Eval

Evaluates the ability of a curiosity-driven reinforcement learning auditor to autonomously generate prompts that elicit harmful, toxic, or target-specific outputs from black-box LLMs without parameter access. It measures how efficiently the auditor explores the prompt space to uncover rare or sensitive model behaviors. Use when the user wants to benchmark on Inverse Suffix Generation Task, Toxic Completion Task, or asks about evaluating this task. Reports Auditing Objective.

qhjqhj00 3d3042f 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/calm-audit-eval commit 3d3042f116

Frequently asked questions

npx skillmds add qhjqhj00/calm-audit-eval