Medical Safety Eval

Probes whether black-box behavioral distillation preserves safety alignment in medical LLMs. It measures functional fidelity on benign medical prompts and quantifies safety violations and refusal failures on adversarial inputs using an automated moderation classifier. Use when the user wants to benchmark on Medical QA Datasets (MedQA, PubMedQA, MedMCQA, EMRQA), Handcrafted Red-Teaming Suite, GQ-Generated Harmful Prompts, or asks about evaluating this task. Reports Violation Rate.

qhjqhj00 705f3ac 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medical-safety-eval commit 705f3ac74e

Frequently asked questions

npx skillmds add qhjqhj00/medical-safety-eval