Eureka Bench Eval

Granular, capability-level analysis of large foundation models across multimodal reasoning, language understanding, safety, and stability. It dissects performance across fine-grained subcategories (e.g., geometric depth vs. height, single vs. multi-object detection) to reveal persistent failures and complementary strengths across models. Use when the user wants to benchmark on EUREKA-BENCH, or asks about evaluating this task. Reports accuracy.

qhjqhj00 66e6875 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/eureka-bench-eval commit 66e6875a7a

Frequently asked questions

npx skillmds add qhjqhj00/eureka-bench-eval