Curate Eval

Evaluates conversational AI assistants' ability to maintain user-specific awareness and correctly prioritize safety-critical constraints over conflicting preferences in multi-turn interactions. It probes whether models can distinguish hard safety limits from softer user desires and avoid generic or evasive responses. Use when the user wants to benchmark on CURATe, or asks about evaluating this task. Reports pass rates.

qhjqhj00 fb34cff 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/curate-eval commit fb34cffe3a

Frequently asked questions

npx skillmds add qhjqhj00/curate-eval