Bizfinbench Eval

Evaluates LLMs on real-world financial reasoning tasks, including numerical calculation, temporal reasoning, information extraction, prediction recognition, and knowledge-based QA in Chinese. It probes the models' ability to handle noisy, context-dependent financial data and produce structured, reasoned outputs. Use when the user wants to benchmark on BizFinBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 00aa6b5 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bizfinbench-eval commit 00aa6b555e

Frequently asked questions

npx skillmds add qhjqhj00/bizfinbench-eval