# Harood Ood Eval

> harood-ood-eval

- Skill: `qhjqhj00/harood-ood-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/harood-ood-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/harood-ood-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/harood-ood-eval

---


# harood-ood-eval

> HAROOD: A Benchmark for Out-of-distribution Generalization in Sensor-based Human Activity Recognition — Wang Lu et al. (2025) (arXiv:2512.10807, 2025)

## What this evaluates

Evaluates out-of-distribution (OOD) generalization in sensor-based human activity recognition (HAR) across four realistic domain-shift scenarios: cross-person, cross-position, cross-device, and cross-time. It benchmarks how well 16 OOD methods with CNN or Transformer backbones maintain performance when trained on one domain and tested on unseen domains.

## Datasets

- **DSADS** — total ?; splits: train (-1), val (-1), test (-1)

## Metrics

- `accuracy` **(primary)** — range: [0, 1]
  - Standard classification accuracy: the proportion of correctly predicted class labels out of the total number of instances.

## Input / output format

**Input**: Multi-channel time-series sensor data (raw sequences).

**Output**: Discrete class labels corresponding to human activities.

## Scoring recipe

```python
def compute_accuracy(predictions, gold_labels):
    correct = sum(1 for p, g in zip(predictions, gold_labels) if p == g)
    return correct / len(gold_labels)

# Protocol: Run each hyperparameter config 3 times.
# Select the config with the highest mean validation accuracy.
# Report the mean of the best validation scores across the 3 trials.
```

## Common pitfalls

- Training-domain validation assumes train/test distributions are similar, which often fails under OOD shifts and causes overfitting to the validation set.
- Oracle selection directly uses test-domain performance for model selection, causing information leakage and yielding optimistically biased results.
- Hyperparameter tuning uses an 80/20 train/val split per task, which may not adequately represent the true OOD generalization capability.

## Evidence (verbatim from paper)

> The model is trained using the training subsets and evaluated on the aggregated validation set. The hyperparameters yielding the highest validation accuracy are selected. Final reported performance metrics are the mean of the best validation scores obtained over three trials.

## Citation

```bibtex
@misc{wanglu2025harood,
  title={HAROOD: A Benchmark for Out-of-distribution Generalization in Sensor-based Human Activity Recognition},
  author={Wang Lu et al. (2025)},
  year={2025},
  note={arXiv:2512.10807}
}
```

- arXiv: 2512.10807

