# Btzsc Eval

> btzsc-eval

- Skill: `qhjqhj00/btzsc-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/btzsc-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/btzsc-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/btzsc-eval

---


# btzsc-eval

> BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, Rerankers and LLMs — Aarab (2026) (arXiv:2603.11991, 2026)

## What this evaluates

Evaluates zero-shot text classification capabilities across diverse datasets using four model families: NLI cross-encoders, embedding models, rerankers, and instruction-tuned LLMs. It probes how well models can assign text to predefined categories without task-specific fine-tuning by verbalizing labels as descriptions.

## Datasets

- **BTZSC Benchmark** — total ?; splits: test (-1); repo https://github.com/IliasAarab/btzsc

## Metrics

- `macro F1` **(primary)** — range: [0, 1]
  - Macro-averaged F1 score computed across all classes and datasets, averaging per-class F1 scores equally regardless of class frequency.

## Input / output format

**Input**: Input text paired with a set of verbalized label descriptions (one per class).

**Output**: Predicted class label (for cross-encoders, embeddings, rerankers) or selected multiple-choice option (for LLMs).

## Scoring recipe

```python
def compute_macro_f1(predictions, gold_labels):
    f1_scores = []
    for label in set(gold_labels):
        tp = sum(1 for p, g in zip(predictions, gold_labels) if p == label and g == label)
        fp = sum(1 for p, g in zip(predictions, gold_labels) if p == label and g != label)
        fn = sum(1 for p, g in zip(predictions, gold_labels) if p != label and g == label)
        prec = tp / (tp + fp) if (tp + fp) > 0 else 0.0
        rec = tp / (tp + fn) if (tp + fn) > 0 else 0.0
        f1 = 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0.0
        f1_scores.append(f1)
    return sum(f1_scores) / len(f1_scores)
```

## Common pitfalls

- Label verbalization is mandatory; using raw class names instead of context-rich descriptions will break zero-shot prompting.
- Different model families require distinct inference pipelines (logits for cross-encoders, cosine similarity for embeddings, relevance scores for rerankers, next-token probabilities for LLMs).
- Performance is aggregated across 22 diverse datasets, so reporting a single aggregate number without dataset-level breakdowns obscures domain-specific variations.

## Evidence (verbatim from paper)

> To facilitate zero-shot classification, each class label is verbalized as a short, semantically clear, and context-rich description. ... Results show rerankers (e.g., Qwen3-Reranker-8B) achieve state-of-the-art macro F1 of 0.72, embedding models (e.g., GTE-large-en-v1.5) offer optimal accuracy-latency trade-offs, and LLMs (4–12B params) perform competitively on topic classification but lag behind rerankers.

## Citation

```bibtex
@misc{aarab2026btzsc,
  title={BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, Rerankers and LLMs},
  author={Aarab (2026)},
  year={2026},
  note={arXiv:2603.11991}
}
```

- arXiv: 2603.11991

