Horizonbench Eval

Probes long-horizon personalization and belief-update capability. It tests whether models can track evolving user preferences across ~6 months of conversation history and correctly select responses aligned with updated preferences, rather than anchoring on outdated values. Use when the user wants to benchmark on HorizonBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 4fa09a5 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/horizonbench-eval commit 4fa09a5ed7

Frequently asked questions

npx skillmds add qhjqhj00/horizonbench-eval