Leaderboard Zero Shot Rte Eval

Evaluates whether pre-trained Recognizing Textual Entailment (RTE) models can generalize to unseen task-dataset-metric (TDM) extraction pairs in a zero-shot setting. It probes whether models learn genuine semantic entailment or merely memorize training distribution patterns. Use when the user wants to benchmark on LEADERBOARDS, or asks about evaluating this task. Reports macro F1.

qhjqhj00 89e5427 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/leaderboard-zero-shot-rte-eval commit 89e542711e

Frequently asked questions

npx skillmds add qhjqhj00/leaderboard-zero-shot-rte-eval