Carpe Eval

Evaluates the visual classification and vision-language understanding capabilities of large vision-language models (LVLMs) under context-aware ensemble prompting. It probes fine-grained visual recognition, scientific question answering, text-rich VQA, hallucination detection, and multimodal reasoning across diverse benchmarks. Use when the user wants to benchmark on ImageNet, Caltech101, Flower102, Food101, ScienceQA (image subset), TextVQA, POPE, MME, MMBench, CV-Bench, MMVP, or asks about evaluating this task. Reports accuracy / F1 score / scaled MME score.

qhjqhj00 7bc9e6b 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/carpe-eval commit 7bc9e6bc9e

Frequently asked questions

npx skillmds add qhjqhj00/carpe-eval