Open Finllm Leaderboard Eval

Evaluates the capability of financial LLMs and agents across seven core financial task categories, including information extraction, sentiment analysis, question answering, text generation, risk management, forecasting, and decision-making. The benchmark aggregates 42 existing financial datasets to provide a standardized comparison of model performance and compliance readiness. Use when the user wants to benchmark on Open FinLLM Leaderboard (42 financial datasets), or asks about evaluating this task. Reports average score across all tasks.

qhjqhj00 f45a4e3 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/open-finllm-leaderboard-eval commit f45a4e3ef0

Frequently asked questions

npx skillmds add qhjqhj00/open-finllm-leaderboard-eval