Experimentation

Designing, running, and interpreting controlled experiments to decide whether a change is genuinely better — online A/B tests, randomized controlled trials, and offline hypothesis tests. Covers hypothesis framing, randomization and unit-of-assignment, sample-size and statistical-power calculation, metric design (primary, guardrail, invariant, and the overall evaluation criterion), variance reduction (CUPED/CUPAC) for sensitivity, significance testing and confidence intervals, sequential and anytime-valid testing for the peeking problem, multiple-comparison correction, the trust checklist led by sample-ratio-mismatch, interference-aware designs (cluster and switchback), long-term holdouts and reverse experiments, bandit/adaptive-allocation boundaries, experiment ethics and blast-radius limits, quasi-experimental fallbacks when randomization is infeasible, and the validity threats that quietly invalidate results (novelty and primacy effects, interference/SUTVA violations, Simpson's paradox, and survivorship).

jacob-balslev 3325dc3 2 files · 48.5 KB Updated

File contents

jacob-balslev/skills/tree/main/skills/software-engineering-method/experimentation commit 3325dc35d4

Frequently asked questions

npx skillmds@latest add jacob-balslev/experimentation