# Covidx Classification Eval

> covidx-classification-eval

- Skill: `qhjqhj00/covidx-classification-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/covidx-classification-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/covidx-classification-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/covidx-classification-eval

---


# covidx-classification-eval

> Identification of images of COVID-19 from Chest X-rays using Deep Learning: Comparing COGNEX VisionPro Deep Learning 1.0 Software with Open Source Convolutional Neural Networks — Arjun Sarkar et al. (2020) (arXiv:2008.00597, 2020)

## What this evaluates

Evaluates deep learning models' ability to classify chest X-ray images into three diagnostic categories: normal, non-COVID-19 pneumonia, and COVID-19. It probes medical image classification performance under realistic class imbalance and tests whether models rely on clinically relevant lung regions or artifacts.

## Datasets

- **COVIDx** — total 13975; splits: train (13675), test (300); repo https://github.com/lindawangg/COVID-Net

## Metrics

- `F-score` **(primary)** — range: percent
  - Harmonic mean of precision and recall: F1 = 2 * (precision * recall) / (precision + recall). Reported as a percentage, typically macro-averaged across the three classes.

## Input / output format

**Input**: Chest X-ray images (either full images or lung-segmented regions)

**Output**: Classification label from three classes: Normal, Non-COVID-19/Pneumonia, or COVID-19

## Scoring recipe

```python
def compute_fscore(predictions, gold):
    classes = [0, 1, 2]
    f1_scores = []
    for c in classes:
        tp = sum(1 for p, g in zip(predictions, gold) if p == c and g == c)
        fp = sum(1 for p, g in zip(predictions, gold) if p == c and g != c)
        fn = sum(1 for p, g in zip(predictions, gold) if p != c and g == c)
        prec = tp / (tp + fp) if (tp + fp) > 0 else 0
        rec = tp / (tp + fn) if (tp + fn) > 0 else 0
        f1 = 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0
        f1_scores.append(f1)
    return sum(f1_scores) / len(f1_scores) * 100
```

## Common pitfalls

- The training set is highly imbalanced (258 COVID-19 vs 7966 Normal), so models may overfit to majority classes without class weighting or augmentation.
- The test set is artificially balanced (100 images per class), which inflates F-score compared to real-world clinical prevalence and may not reflect true diagnostic utility.

## Evidence (verbatim from paper)

> The test set was a balanced set, with each of the three classes having 100 images each [18]. The model achieved an F-score of 94.0% on full images and 95.3% on lung-segmented regions.

## Citation

```bibtex
@misc{sarkar2020cognex,
  title={Identification of images of COVID-19 from Chest X-rays using Deep Learning: Comparing COGNEX VisionPro Deep Learning 1.0 Software with Open Source Convolutional Neural Networks},
  author={Arjun Sarkar et al. (2020)},
  year={2020},
  note={arXiv:2008.00597}
}
```

- arXiv: 2008.00597

