# Biomedical Nlp Eval

> biomedical-nlp-eval

- Skill: `qhjqhj00/biomedical-nlp-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/biomedical-nlp-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/biomedical-nlp-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/biomedical-nlp-eval

---


# biomedical-nlp-eval

> An evaluation of DeepSeek Models in Biomedical Natural Language Processing — Zaifu Zhan et al. (2025) (arXiv:2503.00624, 2025)

## What this evaluates

This benchmark evaluates large language models on four core biomedical natural language processing tasks: event extraction, relation extraction, named entity recognition, and text classification. It probes the models' ability to identify complex biomedical entities, relationships, and events, as well as classify medical texts, highlighting precision-recall trade-offs in domain-specific applications.

## Datasets

- **PHEE** — total ?; splits: test (-1)
- **Genia2013** — total ?; splits: test (-1)
- **Genia2011** — total ?; splits: test (-1)
- **DDI** — total ?; splits: test (-1)
- **GIT** — total ?; splits: test (-1)
- **BioRED** — total ?; splits: test (-1)
- **BC5CDR** — total ?; splits: test (-1)
- **BC2GM** — total ?; splits: test (-1)
- **BC4Chemd** — total ?; splits: test (-1)
- **ADE** — total ?; splits: test (-1)
- **PubMed 20k RCT** — total ?; splits: test (-1)
- **HealthAdvice** — total ?; splits: test (-1)

## Metrics

- `F1` **(primary)** — range: [0, 1]
  - Harmonic mean of precision and recall: F1 = 2 * (precision * recall) / (precision + recall). Used as the headline metric across all four tasks.
- `Precision` — range: [0, 1]
  - Ratio of correctly predicted positive instances to the total number of predicted positive instances.
- `Recall` — range: [0, 1]
  - Ratio of correctly predicted positive instances to the total number of actual positive instances.

## Input / output format

**Input**: Biomedical text snippets or documents. Exact prompt format is not specified in the provided text.

**Output**: Structured predictions corresponding to the task: entity/event/relation spans for extraction tasks, or class labels for text classification. Exact format is not specified in the provided text.

## Scoring recipe

```python
def compute_f1(precision, recall):
    if precision + recall == 0:
        return 0.0
    return 2 * (precision * recall) / (precision + recall)

# For each dataset/task:
# 1. Extract model predictions and gold labels.
# 2. Compute precision and recall (token/span-level for extraction, token-level for classification).
# 3. Apply compute_f1 to get the F1 score.
# 4. Report F1 as the primary metric, with precision and recall for trade-off analysis.
```

## Common pitfalls

- High recall in extraction tasks often masks low precision, leading to many false positives; relying solely on recall overestimates model utility.
- Performance varies drastically across datasets within the same task (e.g., Genia2013 vs PHEE), so reporting only an average F1 can obscure task-specific weaknesses.
- The paper notes that base models (e.g., Deepseek-LLM-7B-base) perform significantly worse than instruction-tuned or distilled variants, so model type must be explicitly tracked.

## Evidence (verbatim from paper)

> Figure 1. Evaluation results for 12 LLMs on 4 tasks across 12 datasets using precision, recall, F1 as the metrics

## Citation

```bibtex
@misc{zhan2025biomedllmeval,
  title={An evaluation of DeepSeek Models in Biomedical Natural Language Processing},
  author={Zaifu Zhan et al. (2025)},
  year={2025},
  note={arXiv:2503.00624}
}
```

- arXiv: 2503.00624

