Mtbbench Eval

Evaluates AI agents' ability to perform longitudinal, multimodal clinical decision-making in oncology. Agents must integrate evolving patient data across pathology, genomics, hematology, and imaging over multiple turns to answer diagnostic and prognostic questions, simulating molecular tumor board workflows. Use when the user wants to benchmark on MTBBench-Multimodal (HANCOCK subset), MTBBench-Longitudinal (MSK-CHORD subset), or asks about evaluating this task. Reports accuracy.

qhjqhj00 e4b08a4 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mtbbench-eval commit e4b08a4868

Frequently asked questions

npx skillmds add qhjqhj00/mtbbench-eval