Timeseries Exam Eval

This benchmark evaluates vision-language models' ability to reason over time series data across medical, financial, and meteorological domains. It probes pattern recognition, anomaly detection, and causal reasoning using synthetically generated multiple-choice questions derived from real-world datasets. Use when the user wants to benchmark on PTB-XL, MIT-BIH, MIMIC-IV Waveform, Yahoo Finance, WeatherBench 2, or asks about evaluating this task. Reports accuracy.

qhjqhj00 176ba57 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/timeseries-exam-eval commit 176ba57274

Frequently asked questions

npx skillmds add qhjqhj00/timeseries-exam-eval