Kimi K1.5 Benchmark Eval

Evaluates multimodal reasoning, coding, and instruction-following capabilities across text, code, and vision tasks using a standardized suite of academic benchmarks. Use when the user wants to benchmark on MMLU, IF-Eval, CLUEWSC, C-EVAL, HumanEval-Mul, LiveCodeBench, Codeforces, AIME 2024, MATH-500, MMMU, MATH-Vision, MathVista, or asks about evaluating this task. Reports exact-match accuracy (EM).

qhjqhj00 c4daf9d 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/kimi-k1.5-benchmark-eval commit c4daf9da5e

Frequently asked questions

npx skillmds add qhjqhj00/kimi-k1-5-benchmark-eval