Timer Eval

This benchmark evaluates a model's ability to perform temporal reasoning and extract accurate information from longitudinal electronic health records (EHRs). It probes whether models can correctly synthesize evidence across multiple time-stamped clinical visits, adhere to specified temporal boundaries, and maintain accuracy over long patient timelines. Use when the user wants to benchmark on TIMER-Bench, MedAlign, or asks about evaluating this task. Reports Correct.

qhjqhj00 cd79546 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/timer-eval commit cd795462bf

Frequently asked questions

npx skillmds add qhjqhj00/timer-eval