# Clinical Assertion Detection Eval

> clinical-assertion-detection-eval

- Skill: `qhjqhj00/clinical-assertion-detection-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/clinical-assertion-detection-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/clinical-assertion-detection-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/clinical-assertion-detection-eval

---


# clinical-assertion-detection-eval

> Beyond Negation Detection: Comprehensive Assertion Detection Models for Clinical NLP — Kocaman et al. (2025) (arXiv:2503.17425, 2025)

## What this evaluates

Evaluates a model's ability to classify the assertion status of medical entities in clinical text across six categories: present, absent, possible, hypothetical, conditional, and associated with someone else. It benchmarks fine-tuned LLMs, transformer classifiers, rule-based systems, and commercial APIs to measure domain-specific clinical NLP performance.

## Datasets

- **i2b2 2010** — total ?; splits: test (-1)

## Metrics

- `weighted avg performance` **(primary)** — range: [0, 1]
  - Weighted average of per-category accuracy/F1 scores across the six assertion labels. Calculated by averaging the performance metric for each category, weighted by the number of instances per category.

## Input / output format

**Input**: Clinical text sentences or phrases containing specified medical entities. For cloud API evaluations, text is obfuscated for PHI and medical terms using Healthcare NLP tools.

**Output**: A single classification label from the set: Present, Absent, Possible, Hypothetical, Conditional, Associated with someone else.

## Scoring recipe

```python
def compute_weighted_avg_accuracy(predictions, gold_labels, categories):
    total_weighted_score = 0.0
    total_instances = 0
    for cat in categories:
        cat_preds = [p for p, g in zip(predictions, gold_labels) if g == cat]
        cat_gold = [g for p, g in zip(predictions, gold_labels) if g == cat]
        if not cat_gold:
            continue
        correct = sum(1 for p, g in zip(cat_preds, cat_gold) if p == g)
        cat_acc = correct / len(cat_gold)
        total_weighted_score += cat_acc * len(cat_gold)
        total_instances += len(cat_gold)
    return total_weighted_score / total_instances if total_instances > 0 else 0.0
```

## Common pitfalls

- Cloud API evaluations (AWS/Azure) are restricted to partially or fully overlapped entities with the i2b2 dataset, not the full test set, leading to potential selection bias.
- Conditional and Hypothetical labels are merged/treated as a single label for LLMs and fine-tuned models due to ambiguity, making direct comparison with full 6-class baselines invalid.
- NegEx is a negation-only rule-based system and does not predict the full assertion taxonomy, so its scores only reflect 'Absent' detection capability.

## Evidence (verbatim from paper)

> The evaluation and benchmarking in this study are conducted exclusively on the official 2010 i2b2 dataset (test split), which represents a comprehensive resource for assessing assertion detection frameworks in real-world clinical scenarios. Table 2 presents the experimental results, highlighting the performance of each model across relevant categories. The models in the first section of this table are developed by JSL. In LLM and GPT-4o experiments, hypothetical and conditional labels are merged/treated as a single label.

## Citation

```bibtex
@misc{kocaman2025beyondnegation,
  title={Beyond Negation Detection: Comprehensive Assertion Detection Models for Clinical NLP},
  author={Kocaman et al. (2025)},
  year={2025},
  note={arXiv:2503.17425}
}
```

- arXiv: 2503.17425

