fingpt-financial-eval
Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications — Djagba et al. (2025) (arXiv:2507.08015, 2025)
What this evaluates
Evaluates a financial domain-specific LLM (FinGPT) across six core NLP tasks: sentiment analysis, text classification, named entity recognition, financial question answering, stock movement prediction, and text summarization. It probes the model's ability to handle domain-specific terminology, numerical reasoning, and structured output generation under instruction-tuning.
Datasets
- FLARE-FPB — total 970; splits: test (970)
- FLARE-FIQASA — total 235; splits: test (235)
- FinGPT Headline Classification — total ?; splits: test (-1)
- FinGPT/fingpt-ner — total 98; splits: test (98); HF
FinGPT/fingpt-ner - ConvFinQA — total 200; splits: test (200); HF
FinGPT/fingpt-convfinqa - FLARE-FinQA — total 50; splits: test (50); HF
ChanceFocus/flare-finqa - CIKM18 (flare-ECTSum) — total ?; splits: test (-1); HF
ChanceFocus/flare-ectsum - StockNet (flare-SM-ACL) — total ?; splits: test (-1); HF
ChanceFocus/flare-sm-acl - BigData22 (flare-SM-BigData) — total ?; splits: test (-1); HF
TheFinAI/flare-sm-bigdata
Metrics
F1-score(primary) — range: [0, 1]- Harmonic mean of precision and recall, calculated over the positive class (yes/no for classification, or per-entity type for NER).
accuracy— range: [0, 1]- Proportion of correctly predicted labels or numerically exact answers out of the total evaluated instances.
macro F1— range: [0, 1]- Unweighted mean of F1-scores computed independently for each class or entity type.
Input / output format
Input: Instruction-style prompts formatted with explicit templates (e.g., [INST]Classify the sentiment of the following financial headline:<HEADLINE>[/INST], Instruction: Please extract entities... Input: <sentence> Answer:). Inputs are lowercased, normalized, and tokenized with padding/truncation to a task-specific max length (32–1012 tokens).
Output: Text generation constrained to specific categories or values: sentiment labels ('yes', 'no', 'unknown'), entity type strings, numerical values extracted via regex, or stock movement directions ('up', 'down'). Outputs are post-processed and mapped to standardized labels for evaluation.
Scoring recipe
def score_classification(pred_text, gold_label):
pred = normalize_output(pred_text) # map to yes/no/unknown
if pred == 'unknown': return None # excluded per protocol
return 1 if pred == gold_label else 0
def score_qa(pred_text, gold_num):
nums = extract_numbers(pred_text) # regex parse
if not nums: return None
best_pred = min(nums, key=lambda x: abs(x - gold_num))
return 1 if abs(best_pred - gold_num) < 1e-3 else 0
# Aggregate precision, recall, F1 over non-None results
Common pitfalls
- Outputs categorized as 'unknown' or ambiguous are explicitly excluded from metric aggregation, which can artificially inflate scores if the exclusion rate is high.
- Numerical QA evaluation requires strict regex parsing and filtering of invalid predictions; including malformed outputs skews accuracy.
- Generation parameters (max tokens, decoding strategy) heavily impact structured tasks like NER; greedy decoding with insufficient token limits causes severe hallucination and drops macro F1.
Evidence (verbatim from paper)
Model performance was assessed using standard classification metrics, including precision, recall, and F1-score, calculated over the yes and no classes. Any outputs categorized as unknown—due to lack of recognizable sentiment indicators –were excluded from the score aggregation to maintain the reliability of the evaluation.
Citation
@misc{djagba2025assessing,
title={Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications},
author={Djagba et al. (2025)},
year={2025},
note={arXiv:2507.08015}
}
- arXiv: 2507.08015