Dianjin R1 Eval

Evaluates large language models' financial reasoning capabilities and general problem-solving skills across multiple benchmarks. It measures how well models can answer domain-specific financial questions and general math/science reasoning tasks, while also assessing compliance rule adherence in Chinese financial contexts. Use when the user wants to benchmark on CFLUE, FinQA, CCC, MATH-500, GPQA-Diamond, or asks about evaluating this task. Reports accuracy.

qhjqhj00 a5bedaa 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dianjin-r1-eval commit a5bedaa17e

Frequently asked questions

npx skillmds add qhjqhj00/dianjin-r1-eval