Measurement & Experimentation Ops
The other skills act on measured differences; this decides whether a difference
is real or noise before they do.
Pick the testing mode by the decision at stake
Three modes, different evidence bars — match to the cost of being wrong:
- Causal: estimates incrementality (no design proves causality without
assumptions). Two sub-modes — a randomized experiment (one treatment vs a
non-overlapping control/holdout, pre-sized) is the strongest; quasi-experimental
causal estimation (GeoLift synthetic control, pre/post) is the fallback when you
can't randomize. Use for expensive, hard-to-reverse bets: offer, funnel, landing
page, "does this channel even lift sales." Cost: volume + discipline + often a
Meta rep.
- Screening (directional): many concepts in one ad set / parallel ABO cells;
delivery is UNEQUAL by design, so a "winner" is a hypothesis, not a proof.
Use for high-throughput creative hunting where being fast beats being certain.
Never present a screen result as validated.
- Infrastructure (isolate infra variance): hold the CREATIVE fixed, vary one
infra axis (domain / proxy cluster / account batch) across a balanced set to
attribute delivery/ban/CPM differences to infra, not creative. The grey
inversion of a normal test — see meta-grey-ops/06.
Feasibility gate (grey reality — check BEFORE promising a clean test)
Causal measurement often isn't available on grey/small-account buys: too little
volume to power a holdout, accounts die mid-test, no clean pixel signal, no rep
for a sandboxed Conversion Lift. When you can't run causal, SAY SO and drop to
the best affordable proxy (geo holdout, pre/post with tracker truth, screening)
— label it directional, don't dress a screen up as a lift study. Choosing the
honest weaker method beats a "causal" test that's silently contaminated.
Validity traps (each one silently flips a conclusion)
- SRM: check the RANDOMIZED-UNIT split (a 50/50 arriving 55/45 = broken
randomization/logging → invalid) — on assignment counts, NOT on
spend/impressions/conversions (those diverging is a delivery effect, not SRM).
- Peeking: Meta's A/B "end test early if a winner is found" — leave off and run
the pre-set window unless Meta's sequential rule is verified (unpublished) (02).
- Contamination: overlapping audiences between cells — Advantage+ broad
bleeding into manual cells; duplicated winners cannibalizing in the auction →
not clean groups. Use the A/B tool's non-overlapping split, or geo separation.
- Conversion lag: judging before the payout event matures counts spend against
unripe conversions → every fresh cohort looks like a loser. Window ≥ lag; nowcast
if you must decide early (tracker-ops/03).
- Multiple testing: screening tolerates chance winners (you re-test anyway); a
causal decision needs the bar corrected for the number of comparisons.
- Underpowered: "no significant difference" ≠ "no effect" — size first (01).
Route references
| Need |
Reference |
| Sizing (MDE/power as decision rules), SRM, peeking, contamination, lag, inconclusive handling |
references/01-experiment-design.md |
Meta tools: A/B Test, ad_study API, Conversion Lift, Brand Lift, GeoLift, Robyn/MMM, Andromeda implication |
references/02-meta-measurement-tools.md |
| Google tools: Experiments/drafts, PMax experiments, Conversion Lift, Meridian MMM, brand-search incrementality, plus two 2026 confounders — read before attributing any Google result to your own change |
references/03-google-measurement-tools.md |
Buy mechanics → meta-ads (its /09 owns single-account diagnosis & test-design
intake) or google-ads (its /08 owns the diagnostic tree and unit economics); this
skill owns the validity/incrementality layer above both. Counting &
cohort truth → tracker-ops. Portfolio decisions on the result → senior-buyer-ops.
1---2name: measurement-experimentation-ops3description: Decide whether a media-buying result is real before scaling it: testing-mode choice (causal / screening / infrastructure), validity traps (SRM, peeking, contamination, lag, multiple testing), and the platforms' measurement tools — Meta (A/B Test, ad_study API, Conversion Lift, GeoLift, Robyn) and Google (Experiments, Conversion Lift, Meridian MMM, brand-search incrementality). Pairs with the media-buying set.4---56# Measurement & Experimentation Ops78The other skills act on measured differences; this decides whether a difference9is real or noise before they do.1011## Pick the testing mode by the decision at stake1213Three modes, different evidence bars — match to the cost of being wrong:14151. **Causal**: estimates incrementality (no design *proves* causality without16 assumptions). Two sub-modes — a randomized experiment (one treatment vs a17 non-overlapping control/holdout, pre-sized) is the strongest; quasi-experimental18 causal estimation (GeoLift synthetic control, pre/post) is the fallback when you19 can't randomize. Use for expensive, hard-to-reverse bets: offer, funnel, landing20 page, "does this channel even lift sales." Cost: volume + discipline + often a21 Meta rep.222. **Screening** (directional): many concepts in one ad set / parallel ABO cells;23 delivery is UNEQUAL by design, so a "winner" is a hypothesis, not a proof.24 Use for high-throughput creative hunting where being fast beats being certain.25 Never present a screen result as validated.263. **Infrastructure** (isolate infra variance): hold the CREATIVE fixed, vary one27 infra axis (domain / proxy cluster / account batch) across a balanced set to28 attribute delivery/ban/CPM differences to infra, not creative. The grey29 inversion of a normal test — see meta-grey-ops/06.3031## Feasibility gate (grey reality — check BEFORE promising a clean test)3233Causal measurement often isn't available on grey/small-account buys: too little34volume to power a holdout, accounts die mid-test, no clean pixel signal, no rep35for a sandboxed Conversion Lift. When you can't run causal, SAY SO and drop to36the best affordable proxy (geo holdout, pre/post with tracker truth, screening)37— label it directional, don't dress a screen up as a lift study. Choosing the38honest weaker method beats a "causal" test that's silently contaminated.3940## Validity traps (each one silently flips a conclusion)4142- **SRM:** check the RANDOMIZED-UNIT split (a 50/50 arriving 55/45 = broken43 randomization/logging → invalid) — on assignment counts, NOT on44 spend/impressions/conversions (those diverging is a delivery effect, not SRM).45- **Peeking:** Meta's A/B "end test early if a winner is found" — leave off and run46 the pre-set window unless Meta's sequential rule is verified (unpublished) (02).47- **Contamination:** overlapping audiences between cells — Advantage+ broad48 bleeding into manual cells; duplicated winners cannibalizing in the auction →49 not clean groups. Use the A/B tool's non-overlapping split, or geo separation.50- **Conversion lag:** judging before the payout event matures counts spend against51 unripe conversions → every fresh cohort looks like a loser. Window ≥ lag; nowcast52 if you must decide early (tracker-ops/03).53- **Multiple testing:** screening tolerates chance winners (you re-test anyway); a54 causal decision needs the bar corrected for the number of comparisons.55- **Underpowered:** "no significant difference" ≠ "no effect" — size first (01).5657## Route references5859| Need | Reference |60|---|---|61| Sizing (MDE/power as decision rules), SRM, peeking, contamination, lag, inconclusive handling | `references/01-experiment-design.md` |62| Meta tools: A/B Test, `ad_study` API, Conversion Lift, Brand Lift, GeoLift, Robyn/MMM, Andromeda implication | `references/02-meta-measurement-tools.md` |63| Google tools: Experiments/drafts, PMax experiments, Conversion Lift, Meridian MMM, brand-search incrementality, plus two 2026 confounders — read before attributing any Google result to your own change | `references/03-google-measurement-tools.md` |6465Buy mechanics → meta-ads (its /09 owns single-account diagnosis & test-design66intake) or google-ads (its /08 owns the diagnostic tree and unit economics); this67skill owns the validity/incrementality layer above both. Counting &68cohort truth → tracker-ops. Portfolio decisions on the result → senior-buyer-ops.