Leaderboard Triple Extraction Eval

Evaluates a model's ability to verify whether a candidate (Task, Dataset, Metric) triple is actually used or mentioned in a specific AI research paper. The task is framed as a natural language inference problem where the model must distinguish between valid triples and randomly sampled invalid ones. Use when the user wants to benchmark on AI Research Paper Collection, or asks about evaluating this task. Reports micro-F1.

qhjqhj00 f780638 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/leaderboard-triple-extraction-eval commit f780638832

Frequently asked questions

npx skillmds add qhjqhj00/leaderboard-triple-extraction-eval