# Fingtruthqa Eval

> fingtruthqa-eval

- Skill: `qhjqhj00/fingtruthqa-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/fingtruthqa-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/fingtruthqa-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/fingtruthqa-eval

---


# fingtruthqa-eval

> FinTruthQA: A Benchmark Dataset for Evaluating the Quality of Financial Information Disclosure — Xu et al. (2024) (arXiv:2406.12009, 2024)

## What this evaluates

Evaluates AI models' ability to assess the quality of financial information disclosure in Chinese. It probes four capabilities: identifying relevant questions, determining question relevance, evaluating answer readability, and measuring answer relevance to the question.

## Datasets

- **FinTruthQA** — total 6000; splits: test (-1)

## Metrics

- `Accuracy, Precision, Recall, F1-score, Micro/Macro/Weighted F1, Quadratic Weighted Kappa (QWK)` **(primary)** — range: percent
  - Standard classification metrics. Accuracy, Precision, Recall, and F1-score for binary tasks (question identification and relevance). Micro, Macro, and Weighted F1-scores plus Quadratic Weighted Kappa (QWK) for ordinal/multi-class tasks (answer readability and relevance). QWK penalizes misclassifications based on their distance from the true label.

## Input / output format

**Input**: For question identification and relevance: raw Chinese question text. For answer readability and relevance: concatenated question and answer text, separated by [SEP] token, prefixed with [CLS] token.

**Output**: Classification label: binary (relevant/irrelevant or identified/not) for tasks 1-2; ordinal/multi-class label (e.g., 1-3 for readability/relevance levels) for tasks 3-4.

## Scoring recipe

```python
def compute_metrics(preds, golds, task_type):
    if task_type == 'binary':
        acc = accuracy_score(golds, preds)
        prec, rec, f1, _ = precision_recall_fscore_support(golds, preds, average='binary')
        return acc, prec, rec, f1
    else:
        f1_micro = f1_score(golds, preds, average='micro')
        f1_macro = f1_score(golds, preds, average='macro')
        f1_weighted = f1_score(golds, preds, average='weighted')
        qwk = quadratic_weighted_kappa(golds, preds)
        return f1_micro, f1_macro, f1_weighted, qwk
```

## Common pitfalls

- Highly imbalanced label distributions (e.g., 910 positive vs. 7 negative samples for question relevance) can cause models to trivially predict all positives to achieve high recall.
- LLMs trained primarily on English corpora struggle with Chinese financial jargon, causing Chain-of-Thought prompting to degrade performance rather than improve it.
- Continued pre-training on domain-specific data only benefits models already aligned with the financial domain (e.g., FinBERT), while others may suffer from representation mismatch.

## Evidence (verbatim from paper)

> For question identification and question relevance, which are binary classification tasks, we used accuracy, precision, recall, and F1-score as evaluation metrics. For answer readability and answer relevance, which are multi-class classification tasks, we calculated both micro and macro F1-scores, along with the Quadratic Weighted Kappa (QWK), which is well-suited for ordinal classification by penalizing misclassifications based on their distance from the true label.

## Citation

```bibtex
@misc{xu2024fingtruthqa,
  title={FinTruthQA: A Benchmark Dataset for Evaluating the Quality of Financial Information Disclosure},
  author={Xu et al. (2024)},
  year={2024},
  note={arXiv:2406.12009}
}
```

- arXiv: 2406.12009

