# Jmedbench Eval

> jmedbench-eval

- Skill: `qhjqhj00/jmedbench-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/jmedbench-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/jmedbench-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/jmedbench-eval

---


# jmedbench-eval

> JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models — Jiang et al. (2024) (arXiv:2409.13317, 2024)

## What this evaluates

Evaluates Japanese biomedical large language models across five tasks: multiple-choice question answering, named entity recognition, machine translation, document classification, and semantic text similarity. It probes domain-specific knowledge, multilingual comprehension, and in-context learning capabilities.

## Datasets

- **JMedBench** — total 38130; splits: test (-1)

## Metrics

- `F1-entity` **(primary)** — range: percent
  - Micro-averaged F1 score for entity recognition, computed over predicted and gold entity spans.
- `Accuracy` **(primary)** — range: percent
  - Percentage of correctly predicted labels or multiple-choice answers.
- `BLEU` — range: percent
  - Standard BLEU score for machine translation evaluation.
- `Pearson correlation` — range: [-1, 1]
  - Pearson correlation coefficient between predicted and gold similarity scores.

## Input / output format

**Input**: Text prompt formatted with one of four templates (Standard, English-centric, Instructed, Minimal), optionally containing zero-shot or few-shot demonstrations.

**Output**: Model-generated text response (e.g., selected option, entity tags, translated sentence, class label, or numerical similarity score).

## Scoring recipe

```python
def compute_metrics(predictions, golds, task):
    if task == 'NER':
        return micro_f1(predictions, golds)
    elif task in ['MCQA', 'DC']:
        return accuracy(predictions, golds)
    elif task == 'MT':
        return bleu_score(predictions, golds)
    elif task == 'STS':
        return pearson_corr(predictions, golds)
```

## Common pitfalls

- Few-shot NER evaluation is restricted to 1 shot for long-text datasets (MRNER, NRNER) to avoid exceeding model context windows.
- Continual pre-training on biomedical or Japanese text can cause catastrophic forgetting, degrading multilingual performance (e.g., Meditron).
- Machine translation performance correlates more with language generation ability than understanding, especially for cross-lingual directions.

## Evidence (verbatim from paper)

> We report the highest F1-entity score among four prompt templates as discussed in Section [3.3]. The best and second-best performances are highlighted in bold and underlined, respectively.

## Citation

```bibtex
@misc{jiang2024jmedbench,
  title={JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models},
  author={Jiang et al. (2024)},
  year={2024},
  note={arXiv:2409.13317}
}
```

- arXiv: 2409.13317

