Nasde Benchmark Runner

Run coding agent benchmarks and verify results with nasde. Use this skill when the user wants to: - Run a benchmark (all tasks, single task, specific variant) - Re-run assessment evaluation on existing trial results - Check or verify results in Opik (traces, feedback scores, experiments) - Troubleshoot a failed benchmark run - View or compare trial results Even if the user doesn't say "benchmark" — if they're talking about running evaluations, checking scores, or analyzing agent performance, this skill applies. After every run that uses --with-opik, ALWAYS verify results via Opik REST API — don't wait for the user to ask.

NoesisVision d19b12b 8 files · 62.5 KB Updated

File contents

NoesisVision/nasde-toolkit/tree/main/.claude/skills/nasde-benchmark-runner commit d19b12b9ab

Frequently asked questions

npx skillmds@latest add noesisvision/nasde-benchmark-runner