Das Medical Red Teaming Eval

This evaluation probes the robustness, privacy compliance, bias/fairness, and hallucination resistance of medical large language models under dynamic, adversarial stress. It measures how well models maintain safety and accuracy when prompts are iteratively mutated by autonomous agents to exploit vulnerabilities, mimicking real-world clinical interactions rather than static benchmark conditions. Use when the user wants to benchmark on MedQA, Privacy-trap scenarios, Medical bias dataset, or asks about evaluating this task. Reports jailbreak rate.

qhjqhj00 99b8132 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/das-medical-red-teaming-eval commit 99b813237a

Frequently asked questions

npx skillmds add qhjqhj00/das-medical-red-teaming-eval