Prompt Benchmark

Systematic prompt evaluation framework with MATH, GSM8K, and Game of 24 benchmarks. Use when evaluating prompt effectiveness on standard benchmarks, comparing meta-prompting strategies quantitatively, measuring prompt quality improvements, or validating categorical prompt optimizations against ground truth datasets.

manutej 1c2b074 2 files · 21.5 KB Updated

File contents

manutej/categorical-meta-prompting/tree/main/.claude/skills/prompt-benchmark commit 1c2b074212

Frequently asked questions

npx skillmds@latest add manutej/prompt-benchmark