Trilemma Of Truth Eval

Evaluates large language models' ability to distinguish factually true statements from factually false and unverifiable ('neither') statements. It probes both prompt-based output probabilities and internal hidden activations to measure veracity classification accuracy and uncertainty quantification. Use when the user wants to benchmark on Trilemma of Truth Datasets, or asks about evaluating this task. Reports MCC.

qhjqhj00 a8d6d49 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/trilemma-of-truth-eval commit a8d6d49a0a

Frequently asked questions

npx skillmds add qhjqhj00/trilemma-of-truth-eval