Multilingual Tot Sim Eval

Evaluates the fidelity of synthetic Tip-of-the-Tongue (ToT) queries by measuring how well they reproduce the relative ranking of retrieval systems compared to real human-authored ToT queries across four languages. Use when the user wants to benchmark on Multilingual ToT Test Collection, or asks about evaluating this task. Reports Kendall's tau & Pearson's r.

qhjqhj00 70c6cc4 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multilingual-tot-sim-eval commit 70c6cc444c

Frequently asked questions

npx skillmds add qhjqhj00/multilingual-tot-sim-eval