LLM Benchmark Eval

Evaluates a language model's general capabilities, including in-context learning, instruction following, mathematical reasoning, code generation, and bidirectional reversal reasoning across a suite of standard and custom benchmarks. Use when the user wants to benchmark on MMLU, GSM8K, HumanEval, Chinese Poem Sentence Pairs, or asks about evaluating this task. Reports accuracy.

qhjqhj00 1c22d72 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/llm-benchmark-eval commit 1c22d72a3d

Frequently asked questions

npx skillmds add qhjqhj00/llm-benchmark-eval