Psiloqa Eval

Evaluates the ability of models to detect span-level hallucinations in multilingual question-answering contexts. It probes cross-lingual generalization and token-level inconsistency detection between generated answers and ground truth. Use when the user wants to benchmark on PsiloQA, or asks about evaluating this task. Reports IoU.

qhjqhj00 6ee1997 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/psiloqa-eval commit 6ee199718c

Frequently asked questions

npx skillmds add qhjqhj00/psiloqa-eval