Answer Leakage Robustness Eval

This benchmark evaluates the robustness of LLM-based tutoring models against adversarial student agents designed to elicit final answers. It measures how easily tutors disclose solutions under various attack strategies and tracks the dialogue length required for answer leakage across math, multiple-choice, and coding domains. Use when the user wants to benchmark on GSM8K, MMLU, HumanEval, or asks about evaluating this task. Reports answer leakage rate.

qhjqhj00 c0c09e8 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/answer-leakage-robustness-eval commit c0c09e808d

Frequently asked questions

npx skillmds add qhjqhj00/answer-leakage-robustness-eval