Multi Prompt LLM Eval

This evaluation protocol probes the robustness of large language models to instruction phrasing by measuring performance across multiple semantically equivalent prompts. It assesses whether model rankings and absolute scores remain stable when the same task is presented with different instruction templates. Use when the user wants to benchmark on LMentry, BIG-bench Lite, BIG-bench Hard, or asks about evaluating this task. Reports exact match evaluation.

qhjqhj00 1bcd228 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multi-prompt-llm-eval commit 1bcd228f7e

Frequently asked questions

npx skillmds add qhjqhj00/multi-prompt-llm-eval