A/B Test Setup
Overview
Use this skill to guide the full experiment lifecycle: hypothesis, design, sample size, implementation, analysis, and playbook documentation. Keep tests focused, measurable, and resistant to common errors like peeking early or testing too many changes at once.
Workflow
- Define the business goal, user segment, current baseline, target metric, and guardrail metrics.
- Write a specific hypothesis:
If we change X for audience Y, metric Z will improve because...
- Design the test:
- Control and variant.
- Primary metric.
- Secondary and guardrail metrics.
- Traffic split, eligibility, exclusions, and duration.
- Estimate sample size or minimum detectable effect when baseline traffic and conversion rates are available.
- Create an implementation checklist:
- Tracking, randomization, QA, exposure logging, analytics events, and rollback.
- Define decision rules before launch:
- Ship, revert, iterate, or continue testing.
- Analyze results after the test reaches the agreed sample size.
- Document what changed, what was learned, and follow-up experiments.
Test Backlog Pattern
When building a backlog, score each idea with ICE:
- Impact: expected business or user benefit.
- Confidence: evidence quality.
- Effort: complexity and implementation cost.
Prioritize tests that combine high impact, credible evidence, and low operational risk.
Example Prompts
I want to A/B test our signup CTA button. Current conversion rate is 3.2%, 8,000 visitors/month. Help me design the test, calculate the required sample size, and define what success looks like.
Our A/B test just hit sample size. Here are the results [paste metrics]. Is this statistically significant? Should we ship the variant, revert, or keep testing?
Build a prioritized A/B test backlog for our onboarding flow. Use ICE scoring. Sources to mine: our drop-off analytics, last month's support tickets, and these 3 heatmap observations.
1---2name: ab-test-setup3description: Design, analyze, and document A/B tests for conversion, onboarding, pricing, lifecycle, and product experiments. Use when the user asks for `/ab-test-setup`, experiment design, sample size, statistical significance, A/B test analysis, ICE-scored test backlogs, or avoiding common testing mistakes.4---56# A/B Test Setup78## Overview910Use this skill to guide the full experiment lifecycle: hypothesis, design, sample size, implementation, analysis, and playbook documentation. Keep tests focused, measurable, and resistant to common errors like peeking early or testing too many changes at once.1112## Workflow13141. Define the business goal, user segment, current baseline, target metric, and guardrail metrics.151. Write a specific hypothesis:16 - `If we change X for audience Y, metric Z will improve because...`171. Design the test:18 - Control and variant.19 - Primary metric.20 - Secondary and guardrail metrics.21 - Traffic split, eligibility, exclusions, and duration.221. Estimate sample size or minimum detectable effect when baseline traffic and conversion rates are available.231. Create an implementation checklist:24 - Tracking, randomization, QA, exposure logging, analytics events, and rollback.251. Define decision rules before launch:26 - Ship, revert, iterate, or continue testing.271. Analyze results after the test reaches the agreed sample size.281. Document what changed, what was learned, and follow-up experiments.2930## Test Backlog Pattern3132When building a backlog, score each idea with ICE:3334- Impact: expected business or user benefit.35- Confidence: evidence quality.36- Effort: complexity and implementation cost.3738Prioritize tests that combine high impact, credible evidence, and low operational risk.3940## Example Prompts4142- `I want to A/B test our signup CTA button. Current conversion rate is 3.2%, 8,000 visitors/month. Help me design the test, calculate the required sample size, and define what success looks like.`43- `Our A/B test just hit sample size. Here are the results [paste metrics]. Is this statistically significant? Should we ship the variant, revert, or keep testing?`44- `Build a prioritized A/B test backlog for our onboarding flow. Use ICE scoring. Sources to mine: our drop-off analytics, last month's support tickets, and these 3 heatmap observations.`