# Kilian Group Arxiv Score

> Compute kilian-group/arxiv_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of kilian-group/arxiv_score.

- Skill: `qhjqhj00/kilian-group-arxiv-score` (Agent Skill)
- Install (CLI): `npx skillmds add qhjqhj00/kilian-group-arxiv-score`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/kilian-group-arxiv-score/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/qhjqhj00/kilian-group-arxiv-score

---


# kilian-group-arxiv-score

> Metric `kilian-group/arxiv_score` from the HuggingFace `evaluate` library.

## When to invoke

User asks to compute `kilian-group/arxiv_score` or wants HF evaluate's canonical version.

## Recipe

```python
import evaluate
metric = evaluate.load("kilian-group/arxiv_score")
result = metric.compute(predictions=preds, references=refs)
print(result)
```

## Don'ts

- Don't assume your in-house `kilian-group/arxiv_score` matches HF — version conventions vary.
- Many evaluate metrics have task-specific arguments (`average=`, `lang=`, `model_type=`); read the metric card before reporting numbers.

