Offline Eval Improvement

Guides coding agents through Discovery Forge prompt improvement from a specified Weave offline evaluation baseline. Use `wandb-primary` to fetch evaluation results, failed eval rows, dataset refs, prompt refs, and scorer evidence, then use this workflow when improving researcher.md from failed evaluation rows, comparing eval runs, or iterating on a fixed dataset.

wandb 3be00eb 15.9 KB Updated

File contents

wandb/discovery-forge/tree/main/skills/offline-eval-improvement commit 3be00ebe5f

Frequently asked questions

npx skillmds@latest add wandb/offline-eval-improvement