Empirical Systems Evaluation

Benchmark multi-agent coordination systems with experiment design, power analysis, human rating protocols, bootstrap confidence intervals, and reproducible reporting. Use for salvage latency, recovery fidelity, coordination overhead, and protocol comparison studies. NOT for ML model benchmarking, web-product A/B testing, survey design, or general-purpose data science.

curiositech Updated 10 repo stars

File contents

curiositech/windags-skills/tree/main/skills/empirical-systems-evaluation commit ac04e4c52d

Frequently asked questions

npx skillmds@latest add curiositech/empirical-systems-evaluation