Gistbench Eval

Evaluates LLMs' ability to extract and verify user interests from interaction histories, focusing on factual grounding, specificity, and strict instruction following across heterogeneous engagement types. Use when the user wants to benchmark on Unspecified real-world engagement datasets, or asks about evaluating this task. Reports IG.

qhjqhj00 fb7d777 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/gistbench-eval commit fb7d777702

Frequently asked questions

npx skillmds add qhjqhj00/gistbench-eval