Mrceval Eval

This benchmark evaluates large language models' ability to comprehend passages and answer questions across multiple dimensions, including context understanding, external knowledge integration, and complex reasoning. It probes factual fidelity, counterfactual handling, commonsense, world knowledge, and multi-hop reasoning capabilities. Use when the user wants to benchmark on MRCEval, or asks about evaluating this task. Reports accuracy.

qhjqhj00 93799f6 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mrceval-eval commit 93799f6823

Frequently asked questions

npx skillmds add qhjqhj00/mrceval-eval