# Fingpt Financial Eval

> fingpt-financial-eval

- Skill: `qhjqhj00/fingpt-financial-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/fingpt-financial-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/fingpt-financial-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/fingpt-financial-eval

---


# fingpt-financial-eval

> Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications — Djagba et al. (2025) (arXiv:2507.08015, 2025)

## What this evaluates

Evaluates a financial domain-specific LLM (FinGPT) across six core NLP tasks: sentiment analysis, text classification, named entity recognition, financial question answering, stock movement prediction, and text summarization. It probes the model's ability to handle domain-specific terminology, numerical reasoning, and structured output generation under instruction-tuning.

## Datasets

- **FLARE-FPB** — total 970; splits: test (970)
- **FLARE-FIQASA** — total 235; splits: test (235)
- **FinGPT Headline Classification** — total ?; splits: test (-1)
- **FinGPT/fingpt-ner** — total 98; splits: test (98); HF `FinGPT/fingpt-ner`
- **ConvFinQA** — total 200; splits: test (200); HF `FinGPT/fingpt-convfinqa`
- **FLARE-FinQA** — total 50; splits: test (50); HF `ChanceFocus/flare-finqa`
- **CIKM18 (flare-ECTSum)** — total ?; splits: test (-1); HF `ChanceFocus/flare-ectsum`
- **StockNet (flare-SM-ACL)** — total ?; splits: test (-1); HF `ChanceFocus/flare-sm-acl`
- **BigData22 (flare-SM-BigData)** — total ?; splits: test (-1); HF `TheFinAI/flare-sm-bigdata`

## Metrics

- `F1-score` **(primary)** — range: [0, 1]
  - Harmonic mean of precision and recall, calculated over the positive class (yes/no for classification, or per-entity type for NER).
- `accuracy` — range: [0, 1]
  - Proportion of correctly predicted labels or numerically exact answers out of the total evaluated instances.
- `macro F1` — range: [0, 1]
  - Unweighted mean of F1-scores computed independently for each class or entity type.

## Input / output format

**Input**: Instruction-style prompts formatted with explicit templates (e.g., `[INST]Classify the sentiment of the following financial headline:<HEADLINE>[/INST]`, `Instruction: Please extract entities... Input: <sentence> Answer:`). Inputs are lowercased, normalized, and tokenized with padding/truncation to a task-specific max length (32–1012 tokens).

**Output**: Text generation constrained to specific categories or values: sentiment labels ('yes', 'no', 'unknown'), entity type strings, numerical values extracted via regex, or stock movement directions ('up', 'down'). Outputs are post-processed and mapped to standardized labels for evaluation.

## Scoring recipe

```python
def score_classification(pred_text, gold_label):
    pred = normalize_output(pred_text)  # map to yes/no/unknown
    if pred == 'unknown': return None  # excluded per protocol
    return 1 if pred == gold_label else 0

def score_qa(pred_text, gold_num):
    nums = extract_numbers(pred_text)  # regex parse
    if not nums: return None
    best_pred = min(nums, key=lambda x: abs(x - gold_num))
    return 1 if abs(best_pred - gold_num) < 1e-3 else 0

# Aggregate precision, recall, F1 over non-None results
```

## Common pitfalls

- Outputs categorized as 'unknown' or ambiguous are explicitly excluded from metric aggregation, which can artificially inflate scores if the exclusion rate is high.
- Numerical QA evaluation requires strict regex parsing and filtering of invalid predictions; including malformed outputs skews accuracy.
- Generation parameters (max tokens, decoding strategy) heavily impact structured tasks like NER; greedy decoding with insufficient token limits causes severe hallucination and drops macro F1.

## Evidence (verbatim from paper)

> Model performance was assessed using standard classification metrics, including precision, recall, and F1-score, calculated over the yes and no classes. Any outputs categorized as unknown—due to lack of recognizable sentiment indicators –were excluded from the score aggregation to maintain the reliability of the evaluation.

## Citation

```bibtex
@misc{djagba2025assessing,
  title={Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications},
  author={Djagba et al. (2025)},
  year={2025},
  note={arXiv:2507.08015}
}
```

- arXiv: 2507.08015

