Gpo Prompt Optimization Eval

Evaluates the effectiveness of LLM-based prompt optimizers across complex reasoning, knowledge-intensive, and common NLP tasks. It measures how much optimized prompts improve model performance compared to baseline prompts and other optimization methods. Use when the user wants to benchmark on Big-Bench Hard (BBH), GSM8K, MMLU, WSC, WebNLG, or asks about evaluating this task. Reports average accuracy.

qhjqhj00 f169ebb 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/gpo-prompt-optimization-eval commit f169ebb2e6

Frequently asked questions

npx skillmds add qhjqhj00/gpo-prompt-optimization-eval