Saplma Truthfulness Eval

Evaluates whether an LLM's internal hidden layer activations can predict the veracity of a given statement. It probes the model's implicit knowledge of truthfulness by training a classifier on neural activations rather than relying on explicit prompting or output probabilities. Use when the user wants to benchmark on True-False Dataset, LLM-Generated Statements, or asks about evaluating this task. Reports accuracy.

qhjqhj00 b39abdd 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/saplma-truthfulness-eval commit b39abdd643

Frequently asked questions

npx skillmds add qhjqhj00/saplma-truthfulness-eval