P Bench Eval

Evaluates multimodal large language models' ability to recognize and respond to specific individuals in images using in-context learning. It probes robustness to complex scenes (multiple people, augmentations) and the capability to correctly reject unanswerable queries. Use when the user wants to benchmark on P-Bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 adadc34 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/p-bench-eval commit adadc34941

Frequently asked questions

npx skillmds add qhjqhj00/p-bench-eval