Blackswan Eval

Evaluates vision-language models on abductive and defeasible reasoning in videos depicting unpredictable events. It probes the ability to infer hidden causes from limited visual cues and revise hypotheses when new evidence emerges, testing reasoning beyond simple statistical recall. Use when the user wants to benchmark on BlackSwanSuite, or asks about evaluating this task. Reports accuracy.

qhjqhj00 4d3a31e 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/blackswan-eval commit 4d3a31eb7e

Frequently asked questions

npx skillmds add qhjqhj00/blackswan-eval