Ccfqa Eval

This benchmark evaluates the factual accuracy and consistency of multimodal large language models (MLLMs) when answering questions in text or speech modalities across eight languages. It specifically probes cross-lingual transfer capabilities and cross-modal alignment by measuring how well models maintain factual correctness when switching between languages or between text and audio inputs. Use when the user wants to benchmark on CCFQA, or asks about evaluating this task. Reports F1 score.

qhjqhj00 5a1b049 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ccfqa-eval commit 5a1b049b6f

Frequently asked questions

npx skillmds add qhjqhj00/ccfqa-eval