# Japanese Financial Bench Eval

> japanese-financial-bench-eval

- Skill: `qhjqhj00/japanese-financial-bench-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/japanese-financial-bench-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/japanese-financial-bench-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/japanese-financial-bench-eval

---


# japanese-financial-bench-eval

> Construction of a Japanese Financial Benchmark for Large Language Models — Hirano et al. (2024) (arXiv:2403.15062, 2024)

## What this evaluates

Evaluates large language models on Japanese financial domain knowledge across five distinct tasks: sentiment analysis, fundamental financial knowledge, CPA auditing, and two levels of financial planner exam questions. It probes the models' ability to understand and reason over domain-specific multiple-choice questions in Japanese.

## Datasets

- **chabsa** — total ?; splits: test (-1); repo https://github.com/pfnet-research/japanese-lm-fin-harness
- **cma Basics** — total ?; splits: test (-1); repo https://github.com/pfnet-research/japanese-lm-fin-harness
- **cpa Audit** — total ?; splits: test (-1); repo https://github.com/pfnet-research/japanese-lm-fin-harness
- **fp2** — total ?; splits: test (-1); repo https://github.com/pfnet-research/japanese-lm-fin-harness
- **security_sales_1** — total ?; splits: test (-1); repo https://github.com/pfnet-research/japanese-lm-fin-harness

## Metrics

- `accuracy` **(primary)** — range: percent
  - Percentage of correctly predicted choices out of total instances. Calculated as (correct predictions / total instances) * 100.

## Input / output format

**Input**: Multiple-choice questions in Japanese, formatted with task-specific prompts and 0-4 shot examples.

**Output**: The model outputs the selected choice (determined by highest likelihood or earliest appearance in generation).

## Scoring recipe

```python
correct = 0
for pred, gold in zip(predictions, gold_labels):
    if pred == gold:
        correct += 1
accuracy = (correct / len(gold_labels)) * 100
```

## Common pitfalls

- Prompt tuning (0-4 shots) was performed per task, risking in-sample leakage.
- OpenAI API models were restricted to 0-shot due to cost, creating an unfair comparison.
- Content filters on OpenAI API blocked some responses, which were counted as incorrect.

## Evidence (verbatim from paper)

> To answer the multiple-choice questions, the likelihoods of the choices in the context were calculated and the choice with the highest likelihood was employed as the output. For GPT3.5 and GPT-4 series, the outputs with the temperature parameter set to 0 were obtained via API, and the choice that appeared earliest in the outputs was used as the output.

## Citation

```bibtex
@misc{hirano2024construction,
  title={Construction of a Japanese Financial Benchmark for Large Language Models},
  author={Hirano et al. (2024)},
  year={2024},
  note={arXiv:2403.15062}
}
```

- arXiv: 2403.15062

