Medebench Eval

Evaluates the reliability and clinical appropriateness of text-guided medical image editing models. It probes anatomical localization precision, preservation of surrounding clinical context, and overall visual realism across diverse medical imaging modalities and anatomical regions. Use when the user wants to benchmark on MedEBench, or asks about evaluating this task. Reports GPT-4o Editing Accuracy, Masked SSIM, GPT-4o Visual Quality.

qhjqhj00 17519a5 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medebench-eval commit 17519a53e7

Frequently asked questions

npx skillmds add qhjqhj00/medebench-eval