Multiverse Eval

Evaluates the multi-turn conversational reasoning and sustained dialogue capabilities of Vision-Language Models (VLMs) across diverse domains like mathematics, coding, and creative tasks. It probes how well models leverage dialogue history (in-context learning) and maintain consistency over extended interactions. Use when the user wants to benchmark on MultiVerse, or asks about evaluating this task. Reports checklist-based evaluation.

qhjqhj00 9f9c2de 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multiverse-eval commit 9f9c2de3ae

Frequently asked questions

npx skillmds add qhjqhj00/multiverse-eval