# Loghub Eval

> loghub-eval

- Skill: `qhjqhj00/loghub-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/loghub-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/loghub-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/loghub-eval

---


# loghub-eval

> Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics — Zhu et al. (2020) (arXiv:2008.06448, 2020)

## What this evaluates

Evaluates AI-driven log analytics systems on three core tasks: parsing unstructured log messages into event templates, compressing log data efficiently, and detecting system anomalies using supervised or unsupervised models. It probes how well algorithms generalize across diverse, real-world system logs ranging from distributed systems to mobile apps.

## Datasets

- **Loghub** — total ?; splits: test (-1); repo https://github.com/logpai/loghub

## Metrics

- `Parsing Accuracy (PA)` **(primary)** — range: [0, 1]
  - PA = (# of corrected parsed logs) / (# of total logs). A log message is correctly parsed if its extracted event template matches the same ground truth cluster as the original log.
- `Compression Ratio (CR)` — range: other
  - CR = Original File Size / Compressed File Size. Higher values indicate more effective compression.
- `F-measure` — range: [0, 1]
  - Harmonic mean of precision and recall used to evaluate anomaly detection performance.

## Input / output format

**Input**: Raw log messages or log files; for anomaly detection, a block-ID-by-event count matrix representing system operations per block.

**Output**: Event templates or clusters mapping log messages; compressed binary files; binary anomaly labels (normal/anomalous).

## Scoring recipe

```python
def calc_parsing_accuracy(predictions, gold):
    correct = 0
    for pred_template, gold_template in zip(predictions, gold):
        if pred_template == gold_template:
            correct += 1
    return correct / len(gold)
```

## Common pitfalls

- Parsers struggle with complex logs containing many templates (e.g., Mac, Linux) due to intricate structures.
- Accuracy degrades significantly for rare log messages that violate frequent pattern assumptions.
- Most tools only separate templates from parameters, lacking fine-grained parameter type classification needed for root cause analysis.

## Evidence (verbatim from paper)

> To evaluate the accuracy of different log parsing algorithms. We define the metric of parsing accuracy (PA) as follows: PA = (# of corrected parsed logs) / (# of total logs). After parsing, every log message transforms into an event template, and each template relates to a cluster of log messages sharing the same template. A log message is considered correctly parsed if and only if its event template matches the same cluster of log messages as the groudtruth does.

## Citation

```bibtex
@misc{zhu2020loghub,
  title={Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics},
  author={Zhu et al. (2020)},
  year={2020},
  note={arXiv:2008.06448}
}
```

- arXiv: 2008.06448

