Mctaco Eval

Evaluates a model's ability to reason about temporal commonsense, including event duration, ordering, typical time, frequency, and stationarity. It tests whether systems can correctly classify candidate answers as 'likely' or 'unlikely' given a context sentence and a question. Use when the user wants to benchmark on MCTACO, or asks about evaluating this task. Reports F1.

qhjqhj00 40799fb 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mctaco-eval commit 40799fb1b3

Frequently asked questions

npx skillmds add qhjqhj00/mctaco-eval