Build Verdict Dataset

Guides coding agents through building the verdict_quality_dataset from W&B Weave research_annotation evidence — querying annotated research_run calls, mapping human QualitySelector verdicts to gold labels, refining row inputs per the rubric, and publishing a new versioned Weave Dataset. Use when the user asks to (re)generate the verdict dataset from annotations, seed a new eval dataset version, or rebuild verdict_quality_dataset.

wandb Updated

File contents

wandb/discovery-forge/tree/main/skills/build-verdict-dataset commit 7337bad925

Frequently asked questions

npx skillmds@latest add wandb/build-verdict-dataset