Vhelm Eval

Holistic evaluation of vision-language models across multiple dimensions including visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. Use when the user wants to benchmark on VHELM Scenarios, or asks about evaluating this task. Reports scenario_score.

qhjqhj00 7c9635f 2.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vhelm-eval commit 7c9635f824

Frequently asked questions

npx skillmds add qhjqhj00/vhelm-eval