Improve Primitive From Benchmark

Fully autonomous improvement loop for flowai primitives — run the flowai SWE-bench arm, root-cause failures, generalize fixes beyond code, implement via Acceptance-Test TDD; gated by a critic subagent instead of the user. Use when asked to improve primitives from benchmark results or diagnose benchmark failures.

korchasa Updated

File contents

korchasa/flowai/tree/main/.agents/skills/improve-primitive-from-benchmark commit 2771b0782c

Frequently asked questions

npx skillmds@latest add korchasa/improve-primitive-from-benchmark