# Plutus Ben Eval

> plutus-ben-eval

- Skill: `qhjqhj00/plutus-ben-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/plutus-ben-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/plutus-ben-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/plutus-ben-eval

---


# plutus-ben-eval

> Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance — Xueqing Peng et al. (2025) (arXiv:2502.18772, 2025)

## What this evaluates

This benchmark evaluates large language models on five core Greek financial NLP tasks: numeric and textual named entity recognition, multiple-choice question answering, abstractive summarization, and financial topic classification. It probes models' ability to handle low-resource language morphology, domain-specific financial terminology, and reasoning within Greek financial contexts.

## Datasets

- **GRFinNUM** — total 500; splits: train (320), valid (80), test (100)
- **GRFinNER** — total 500; splits: train (320), valid (80), test (100)
- **GRFinQA** — total 540; splits: train (267), valid (48), test (225)
- **GRFNS-2023** — total 262; splits: train (169), valid (43), test (50)
- **GRMultiFin** — total 268; splits: train (171), valid (43), test (54)

## Metrics

- `Entity F1` — range: [0, 1]
  - Harmonic mean of precision and recall calculated over extracted entity spans. It balances the model's ability to correctly identify numeric or textual financial entities without over- or under-predicting.
- `Accuracy (Acc)` **(primary)** — range: [0, 1]
  - Proportion of correctly predicted answers or topic labels out of the total number of instances. Calculated as exact match for multiple-choice QA and correct class assignment for topic classification.
- `Rouge-1` — range: [0, 1]
  - Unigram overlap score measuring the recall of gold-standard words in the generated summary. Computed as 2 * (precision * recall) / (precision + recall) for unigrams.

## Input / output format

**Input**: Task-specific prompts following a standardized template: a domain expert-crafted instruction followed by the financial input data (e.g., a sentence for NER, a question with multiple choices for QA, a financial document for summarization, or a headline for topic classification).

**Output**: Model-specific outputs: token-level sequence labels (e.g., MONETARY, PERSON, O) for NER; the exact text of the correct answer choice for QA; a generated natural language summary for summarization; or a single predicted topic label from the predefined set for classification.

## Scoring recipe

```python
def compute_metrics(predictions, golds, task):
    if task == 'NER':
        pred_spans = extract_spans(predictions)
        gold_spans = extract_spans(golds)
        tp = len(pred_spans & gold_spans)
        prec = tp / len(pred_spans) if pred_spans else 0
        rec = tp / len(gold_spans) if gold_spans else 0
        return 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0
    elif task in ['QA', 'Topic']:
        correct = sum(1 for p, g in zip(predictions, golds) if p == g)
        return correct / len(golds)
    elif task == 'Summarization':
        return rouge1_score(predictions, golds)
    return 0
```

## Common pitfalls

- Assuming cross-lingual transfer from English models will work seamlessly; Greek's morphological complexity and lack of financial domain adaptation cause significant performance drops.
- Incorrectly splitting long-form financial names with attribution (e.g., 'George Demetriou of Konstantinos') into multiple entities during NER evaluation.
- Using English-centric tokenizers or evaluation thresholds that misalign with Greek character distributions and financial terminology.

## Evidence (verbatim from paper)

> Various evaluation metrics are employed in these benchmarks, including Entity F1, Accuracy (Acc), and Rouge-1 score*(Lin, [2004])**, to assess LLMs’ performance across multiple dimensions: topical content categorization, long-form financial document comprehension, language understanding and reasoning, and both textual and numerical information extraction.

## Citation

```bibtex
@misc{peng2025plutus,
  title={Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance},
  author={Xueqing Peng et al. (2025)},
  year={2025},
  note={arXiv:2502.18772}
}
```

- arXiv: 2502.18772

