Ape Prompt Eval

Evaluates the effectiveness of automatically generated prompts (instructions) from the APE framework compared to human-designed or baseline prompts across various natural language processing tasks. Use when the user wants to benchmark on Instruction Induction, BIG-Bench Instruction Induction (BBII), MultiArith, GSM8K, or asks about evaluating this task. Reports zero-shot execution accuracy.

qhjqhj00 6ee5a05 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ape-prompt-eval commit 6ee5a05f03

Frequently asked questions

npx skillmds add qhjqhj00/ape-prompt-eval