Rlaif V Trustworthiness Eval

Evaluates the trustworthiness (hallucination reduction) and helpfulness of multimodal large language models across generative, discriminative, and free-format tasks. Use when the user wants to benchmark on Object HalBench, MMHal-Bench, MHumanEval, AMBER, RefoMB, MMStar, or asks about evaluating this task. Reports response-level hallucination rate, trustworthiness win rate.

qhjqhj00 60ce5cb 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/rlaif-v-trustworthiness-eval commit 60ce5cb44b

Frequently asked questions

npx skillmds add qhjqhj00/rlaif-v-trustworthiness-eval