Datahub Evals

Use this skill to run DataHub's saved evals and report answers for judging. Triggers on: "run our evals", "run the eval suite", "run eval urn:li:eval:...", "how are our evals doing", "check for eval regressions", "upload this answer as an eval result", "score this answer with the DataHub judge", "compare two agents on the same eval". Answers each eval in a fresh agent with the DataHub tools attached, reports the answer through the DataHub Cloud CLI, and reads back the verdict DataHub's own judge produced.

datahub-project Updated

File contents

datahub-project/datahub-skills/tree/main/skills/datahub-evals commit 0a2db4b449

Frequently asked questions

npx skillmds@latest add datahub-project-datahub-skills/datahub-evals