Crosscheckgpt Eval

This evaluation probes the ability of multimodal foundation models to generate factual content without hallucination across text, image, and audio-visual modalities. It measures how well reference-free ranking methods correlate with human judgments or gold-standard references to rank model outputs by hallucination severity. Use when the user wants to benchmark on WikiBio, MHaluBench, AVHalluBench, or asks about evaluating this task. Reports System($ ho$).

qhjqhj00 95dea68 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/crosscheckgpt-eval commit 95dea68cb8

Frequently asked questions

npx skillmds add qhjqhj00/crosscheckgpt-eval