Golden Touchstone Eval

This benchmark evaluates the capability of large language models to perform a wide range of financial natural language processing tasks in both English and Chinese. It probes domain-specific understanding, information extraction, reasoning, and generation across sentiment analysis, classification, entity/relation extraction, summarization, question answering, and stock movement prediction. Use when the user wants to benchmark on FPB, Fiqa-SA, Headlines, FOMC, lendingclub, NER, FinRE, CFA, EDTSUM, Finqa, Convfinqa, DJIA, FinFe-CN, FinNL-CN, FinESE-CN, FinRE-CN, FinQa-CN, or asks about evaluating this task. Reports Weighted-F1.

qhjqhj00 dd71720 6.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/golden-touchstone-eval commit dd71720ca9

Frequently asked questions

npx skillmds add qhjqhj00/golden-touchstone-eval