Marketing Experimentation
Overview
Use this skill when you need to validate marketing concepts or business ideas through rigorous experimental cycles. This skill orchestrates the complete Build-Measure-Learn cycle from concept to data-driven signal.
When to use this skill:
- You have a marketing concept or business idea that needs validation
- You want to test multiple related hypotheses systematically
- You need to integrate results across multiple experiments
- You're designing the next iteration based on experimental evidence
- You're conducting qualitative market research combined with quantitative testing
What this skill does:
- Validates concepts through market research before experimentation
- Generates multiple testable hypotheses from marketing ideas
- Coordinates multiple experiments: quantitative (hypothesis-testing) and qualitative (qualitative-research)
- Synthesizes results across experiments using interpreting-results and creating-visualizations
- Produces clear signals (positive/negative/null/mixed) for each campaign
- Generates actionable next-iteration ideas based on experimental evidence
What this skill does NOT do:
- Design individual experiments (delegates to hypothesis-testing or qualitative-research)
- Execute statistical analysis directly (uses hypothesis-testing for quantitative rigor)
- Conduct interviews/surveys/observations directly (uses qualitative-research for qualitative rigor)
- Operationalize successful ideas (focuses on validation, not scaling)
- Platform-specific implementation (tool-agnostic techniques only)
Integration with existing skills:
- Delegates to
hypothesis-testing for quantitative experiment design and execution (metrics, A/B tests, statistical analysis)
- Delegates to
qualitative-research for qualitative experiment design and execution (interviews, surveys, focus groups, observations)
- Uses
interpreting-results to synthesize findings across multiple experiments
- Uses
creating-visualizations to communicate aggregate results
- Invokes
market-researcher agent for concept validation via internet research
Multi-conversation persistence:
This skill is designed for campaigns spanning days or weeks. Each phase documents completely enough that new conversations can resume after extended breaks. The experiment tracker (04-experiment-tracker.md) serves as the living coordination hub.
Prerequisites
Required skills:
hypothesis-testing - Quantitative experiment design and execution (invoked for metric-based experiments)
qualitative-research - Qualitative experiment design and execution (invoked for interviews, surveys, focus groups, observations)
interpreting-results - Result synthesis and pattern identification (invoked in Phase 5)
creating-visualizations - Aggregate result visualization (invoked in Phase 5)
Required agents:
market-researcher - Concept validation via internet research (invoked in Phase 1)
Required knowledge:
- Understanding of Lean Startup Build-Measure-Learn cycle
- Familiarity with marketing tactics (landing pages, ads, email, content)
- Basic experimental design principles (control/treatment, signals, metrics)
- Understanding of qualitative vs quantitative research methods
Data requirements:
- None initially (market research is qualitative)
- Data requirements emerge from experiment design in Phase 4
- Quantitative experiments (hypothesis-testing): SQL databases, analytics data, A/B test results
- Qualitative experiments (qualitative-research): Interview transcripts, survey responses, observation notes
Mandatory Process Structure
CRITICAL: This is a 6-phase process skill. You MUST complete all phases in order. Use TodoWrite to track progress through each phase.
TodoWrite template:
When starting a marketing-experimentation session, create these todos:
- [ ] Phase 1: Discovery & Asset Inventory
- [ ] Phase 2: Hypothesis Generation
- [ ] Phase 3: Prioritization
- [ ] Phase 4: Experiment Coordination
- [ ] Phase 5: Cross-Experiment Synthesis
- [ ] Phase 6: Iteration Planning
Workspace structure:
All work for a marketing-experimentation session is saved to:
analysis/marketing-experimentation/[campaign-name]/
├── 01-discovery.md
├── 02-hypothesis-generation.md
├── 03-prioritization.md
├── 04-experiment-tracker.md
├── 05-synthesis.md
├── 06-iteration-plan.md
└── experiments/
├── [experiment-1]/ # hypothesis-testing session
├── [experiment-2]/ # hypothesis-testing session
└── [experiment-3]/ # hypothesis-testing session
Phase progression rules:
- Each phase has a CHECKPOINT with verification requirements
- You MUST satisfy all checkpoint requirements before proceeding to the next phase
- Document every decision with rationale in the numbered markdown files
- Commit markdown files after each phase completes
- The experiment tracker (04-experiment-tracker.md) is a LIVING DOCUMENT - update it throughout Phase 4
Multi-conversation resumption:
- At the start of any conversation, check if an experiment tracker exists for this campaign
- If it exists, read it first to understand current experiment status
- Update the tracker as experiments progress
- All phases should be complete enough to resume after days or weeks
Phase 1: Discovery & Asset Inventory
CHECKPOINT: Before proceeding, you MUST have:
Instructions
Gather the business concept
- Ask user to describe the marketing concept or business idea to validate
- What problem does it solve? Who is the target audience?
- What's the desired outcome? (awareness, leads, conversions, etc.)
- What stage is this idea at? (new concept, existing campaign, iteration)
Invoke market-researcher agent for concept validation
Dispatch the market-researcher agent with the concept description:
- Agent will research market demand signals
- Agent will identify similar solutions and competitors
- Agent will analyze audience needs and pain points
- Agent will find validation evidence (case studies, reviews, testimonials)
Document agent findings in 01-discovery.md under "Market Research Findings"
Conduct asset inventory
Work with user to inventory existing assets that could be leveraged:
Content Assets:
- Blog posts, case studies, whitepapers
- Video content, webinars, tutorials
- Social media presence and following
- Email lists and subscriber segments
Campaign Assets:
- Existing ad campaigns and performance data
- Landing pages and conversion rates
- Email campaigns and open/click rates
- SEO performance and keyword rankings
Audience Assets:
- Customer segments and personas
- Audience data (demographics, behaviors, preferences)
- Customer feedback and reviews
- Support tickets and common questions
Data Assets:
- Analytics platforms (Google Analytics, Mixpanel, etc.)
- CRM data (Salesforce, HubSpot, etc.)
- Ad platform data (Google Ads, Facebook Ads, etc.)
- Email platform data (Mailchimp, SendGrid, etc.)
Define success criteria and validation signals
Work with user to define:
- What metrics indicate success? (CTR, conversion rate, CAC, LTV, etc.)
- What magnitude of change is meaningful? (practical significance thresholds)
- What signal types are acceptable?
- Positive: Validates concept, proceed to scale
- Negative: Invalidates concept, pivot or abandon
- Null: Inconclusive, needs refinement or more data
- Mixed: Some aspects work, some don't, iterate strategically
Document known constraints
Capture any constraints that will affect experimentation:
- Budget constraints (ad spend limits, tool costs)
- Time constraints (launch deadlines, seasonal factors)
- Resource constraints (team capacity, content production)
- Technical constraints (platform limitations, integration issues)
- Regulatory constraints (GDPR, CCPA, industry regulations)
Create 01-discovery.md with: ./templates/01-discovery.md
STOP and get user confirmation
- Review discovery findings with user
- Confirm asset inventory is complete
- Confirm success criteria are appropriate
- Do NOT proceed to Phase 2 until confirmed
Common Rationalization: "I'll skip discovery and go straight to testing - the concept is obvious"
Reality: Discovery surfaces assumptions, constraints, and existing assets that dramatically affect experiment design. Always start with discovery.
Common Rationalization: "I don't need market research - I already know this market"
Reality: The market-researcher agent provides current, data-driven validation signals that prevent building experiments around false assumptions. Always validate.
Common Rationalization: "Asset inventory is busywork - I'll figure out what's available as I go"
Reality: Existing assets can dramatically reduce experiment cost and time. Inventorying first prevents reinventing wheels and enables building on proven foundations.
Phase 2: Hypothesis Generation
CHECKPOINT: Before proceeding, you MUST have:
Instructions
Generate 5-10 testable hypotheses
For each hypothesis, use this format:
Hypothesis [N]: [Brief statement]
- Tactic/Channel: [landing page | ad campaign | email sequence | content marketing | social media | SEO | etc.]
- Expected Outcome: [Specific, measurable result]
- Rationale: [Why we believe this will work based on discovery findings]
- Variables to Test: [What will we manipulate/measure]
Example hypothesis:
Hypothesis 1: Value proposition clarity drives conversion
- Tactic/Channel: Landing page A/B test
- Expected Outcome: 15%+ increase in conversion rate from landing page variant with simplified value proposition
- Rationale: Market research showed audience confusion about product benefits. Discovery found existing landing page has 8 different value propositions competing for attention.
- Variables to Test: Headline clarity, benefit hierarchy, CTA prominence
Ensure tactic coverage
Verify hypotheses cover multiple marketing tactics:
Acquisition Tactics:
- Landing pages (conversion optimization, value prop testing, layout)
- Ad campaigns (targeting, creative, messaging, platforms)
- Content marketing (blog posts, videos, webinars, lead magnets)
- SEO (keyword targeting, content optimization, technical SEO)
Activation Tactics:
- Email sequences (onboarding, nurture, activation)
- Product tours (in-app guidance, feature discovery)
- Social proof (testimonials, case studies, reviews)
Retention Tactics:
- Email campaigns (engagement, re-activation, upsell)
- Content (newsletters, educational content, community)
Don't generate 10 ad hypotheses. Aim for diversity across tactics.
Reference experimentation frameworks
Lean Startup Build-Measure-Learn:
- Build: What's the minimum viable test? (landing page, ad, email, etc.)
- Measure: What metrics indicate success/failure?
- Learn: What will we learn regardless of outcome?
AARRR Pirate Metrics:
- Acquisition: How do users find us?
- Activation: Do they have a great first experience?
- Retention: Do they come back?
- Referral: Do they tell others?
- Revenue: Do they pay?
Map each hypothesis to one or more AARRR stages.
ICE/RICE Prioritization (used in Phase 3):
- Impact: How much will this move the metric?
- Confidence: How sure are we this will work?
- Ease: How easy is this to implement?
- Reach: How many users will this affect? (RICE only)
Create 02-hypothesis-generation.md with: ./templates/02-hypothesis-generation.md
STOP and get user confirmation
- Review all hypotheses with user
- Confirm hypotheses are testable and meaningful
- Confirm tactic coverage is appropriate
- Do NOT proceed to Phase 3 until confirmed
Common Rationalization: "I'll generate hypotheses as I build experiments - more efficient"
Reality: Generating hypotheses before prioritization enables strategic selection of highest-impact tests. Generating ad-hoc leads to testing whatever's easiest, not what matters most.
Common Rationalization: "I'll focus all hypotheses on one tactic (ads) since that's what we know"
Reality: Tactic diversity reveals which channels work for this concept. Single-tactic testing creates blind spots and missed opportunities.
Common Rationalization: "I'll write vague hypotheses and refine them during experiment design"
Reality: Vague hypotheses lead to vague experiments that produce vague results. Specific hypotheses with expected outcomes enable clear signal detection.
Common Rationalization: "More hypotheses = better coverage, I'll generate 20+"
Reality: Too many hypotheses dilute focus and create analysis paralysis in prioritization. 5-10 high-quality hypotheses enable strategic selection of 2-4 tests.
Phase 3: Prioritization
CHECKPOINT: Before proceeding, you MUST have:
Instructions
CRITICAL: You MUST use computational methods (Python scripts) to calculate scores. Do NOT estimate or manually calculate scores.
Choose prioritization framework
ICE Framework (simpler, faster):
- Impact: How much will this move the success metric? (1-10 scale)
- Confidence: How confident are we this will work? (1-10 scale)
- Ease: How easy is this to implement? (1-10 scale)
- Score: (Impact × Confidence) / Ease
RICE Framework (more comprehensive):
- Reach: How many users will this affect? (absolute number or percentage)
- Impact: How much will this move the metric per user? (1-10 scale: 0.25=minimal, 3=massive)
- Confidence: How confident are we in our estimates? (percentage: 50%, 80%, 100%)
- Effort: Person-weeks to implement (absolute number)
- Score: (Reach × Impact × Confidence) / Effort
Choose ICE for speed, RICE for precision when reach varies significantly.
Score each hypothesis using Python script
For ICE Framework:
Create a Python script to compute and sort ICE scores:
#!/usr/bin/env python3
"""
ICE Score Calculator for Marketing Experimentation
Computes ICE scores: (Impact × Confidence) / Ease
Sorts hypotheses by score (highest to lowest)
"""
hypotheses = [
{
"id": "H1",
"name": "Value proposition clarity drives conversion",
"impact": 8,
"confidence": 7,
"ease": 9
},
{
"id": "H2",
"name": "Ad targeting refinement",
"impact": 7,
"confidence": 6,
"ease": 5
},
{
"id": "H3",
"name": "Email sequence optimization",
"impact": 6,
"confidence": 8,
"ease": 8
},
{
"id": "H4",
"name": "Content marketing expansion",
"impact": 5,
"confidence": 4,
"ease": 3
},
]
# Calculate ICE scores
for h in hypotheses:
h['ice_score'] = (h['impact'] * h['confidence']) / h['ease']
# Sort by ICE score (descending)
sorted_hypotheses = sorted(hypotheses, key=lambda x: x['ice_score'], reverse=True)
# Print results table
print("| Hypothesis | Impact | Confidence | Ease | ICE Score | Rank |")
print("|------------|--------|------------|------|-----------|------|")
for rank, h in enumerate(sorted_hypotheses, 1):
print(f"| {h['id']}: {h['name'][:30]} | {h['impact']} | {h['confidence']} | {h['ease']} | {h['ice_score']:.2f} | {rank} |")
Usage:
python3 ice_calculator.py
For RICE Framework:
Create a Python script to compute and sort RICE scores:
#!/usr/bin/env python3
"""
RICE Score Calculator for Marketing Experimentation
Computes RICE scores: (Reach × Impact × Confidence) / Effort
Sorts hypotheses by score (highest to lowest)
"""
hypotheses = [
{
"id": "H1",
"name": "Value proposition clarity drives conversion",
"reach": 10000, # users affected
"impact": 3, # 0.25=minimal, 1=low, 2=medium, 3=high, 5=massive
"confidence": 80, # percentage (50, 80, 100)
"effort": 2 # person-weeks
},
{
"id": "H2",
"name": "Ad targeting refinement",
"reach": 50000,
"impact": 1,
"confidence": 50,
"effort": 4
},
{
"id": "H3",
"name": "Email sequence optimization",
"reach": 5000,
"impact": 2,
"confidence": 80,
"effort": 3
},
{
"id": "H4",
"name": "Content marketing expansion",
"reach": 20000,
"impact": 1,
"confidence": 50,
"effort": 8
},
]
# Calculate RICE scores
for h in hypotheses:
# Convert confidence percentage to decimal
confidence_decimal = h['confidence'] / 100
h['rice_score'] = (h['reach'] * h['impact'] * confidence_decimal) / h['effort']
# Sort by RICE score (descending)
sorted_hypotheses = sorted(hypotheses, key=lambda x: x['rice_score'], reverse=True)
# Print results table
print("| Hypothesis | Reach | Impact | Confidence | Effort | RICE Score | Rank |")
print("|------------|-------|--------|------------|--------|------------|------|")
for rank, h in enumerate(sorted_hypotheses, 1):
print(f"| {h['id']}: {h['name'][:30]} | {h['reach']} | {h['impact']} | {h['confidence']}% | {h['effort']}w | {h['rice_score']:.2f} | {rank} |")
Usage:
python3 rice_calculator.py
Scoring Guidance:
Impact (1-10 for ICE, 0.25-5 for RICE):
- ICE: 1-3 minimal, 4-6 moderate, 7-8 significant, 9-10 transformative
- RICE: 0.25 minimal, 1 low, 2 medium, 3 high, 5 massive
Confidence (1-10 for ICE, 50-100% for RICE):
- ICE: 1-3 speculative, 4-6 uncertain, 7-8 likely, 9-10 validated
- RICE: 50% low confidence, 80% high confidence, 100% certainty
Ease (1-10 for ICE):
- 1-3: Complex, significant resources
- 4-6: Moderate effort, some obstacles
- 7-8: Straightforward, few dependencies
- 9-10: Trivial, immediate execution
Effort (person-weeks for RICE):
- Estimate total person-weeks required
- Include design, implementation, monitoring time
- Examples: 1w (simple landing page), 4w (complex ad campaign), 8w (content series)
Run scoring script and document results
- Create the Python script (ice_calculator.py or rice_calculator.py)
- Update the
hypotheses list with actual hypothesis data from Phase 2
- Run the script:
python3 [ice|rice]_calculator.py
- Copy the output table into
03-prioritization.md
- Include the script in the markdown file for reproducibility:
## Prioritization Calculation
**Method:** ICE Framework
**Calculation Script:**
```python
[paste full script here]
Results:
[paste output table here]
Select 2-4 highest-priority hypotheses
Considerations for selection:
- Don't test everything: Focus on highest-scoring 2-4 hypotheses
- Resource constraints: Match selection to available time/budget/capacity
- Learning value: Sometimes lower-scoring hypothesis with high uncertainty is worth testing
- Dependencies: Test prerequisites before dependent hypotheses
- Sequencing: Consider whether experiments need to run sequentially or can run in parallel
Selection criteria:
- Primary: Top 2-4 by ICE/RICE score from computational results
- Secondary: Balance quick wins vs. high-impact long-term bets
- Tertiary: Ensure tactic diversity (don't test 3 ad variants if other tactics untested)
Document experiment sequence
Determine execution strategy:
- Parallel: Multiple experiments running simultaneously (faster results, higher resource needs)
- Sequential: One experiment at a time (slower, easier to manage)
- Hybrid: Run independent experiments in parallel, sequence dependent ones
Example sequence plan:
Week 1-2: Launch H1 (landing page) and H3 (email) in parallel
Week 3-4: Analyze H1 and H3 results
Week 5-6: Launch H2 (ads) based on H1 learnings
Week 7-8: Analyze H2 results
Create 03-prioritization.md with: ./templates/03-prioritization.md
STOP and get user confirmation
- Review computed scores and prioritization with user
- Confirm selected hypotheses are appropriate
- Confirm experiment sequence is feasible
- Do NOT proceed to Phase 4 until confirmed
Common Rationalization: "I'll test all hypotheses - don't want to miss opportunities"
Reality: Resource constraints make testing everything impossible. Prioritization ensures highest-value experiments get resources. Unfocused testing produces weak signals across too many fronts.
Common Rationalization: "Scoring is subjective and arbitrary - I'll just pick what feels right"
Reality: Scoring frameworks force explicit reasoning about trade-offs. "Feels right" selections optimize for recency bias and personal preference, not business value. Computational methods ensure consistency.
Common Rationalization: "I'll skip prioritization and go straight to easiest test"
Reality: Easiest test rarely equals highest value. Prioritization prevents optimizing for ease at the expense of impact.
Common Rationalization: "I'll estimate scores mentally instead of running the script"
Reality: Manual estimation introduces calculation errors and inconsistency. Python scripts ensure exact, reproducible results that can be audited and verified.
Phase 4: Experiment Coordination
CHECKPOINT: Before proceeding, you MUST have:
Instructions
CRITICAL: This phase is designed for multi-conversation workflows. The experiment tracker is a LIVING DOCUMENT that you will update throughout experimentation. New conversations should ALWAYS read this file first.
Determine experiment type for each hypothesis
CRITICAL: Before creating the tracker, classify each hypothesis as Quantitative or Qualitative.
Quantitative experiments use hypothesis-testing skill:
- Measure numeric metrics: CTR, conversion rate, bounce rate, time on page, revenue, CAC, LTV, etc.
- Rely on existing data sources: Google Analytics, ad platforms, CRM, email platforms, database queries
- Test using A/B tests, multivariate tests, or time-series analysis
- Require statistical significance testing
- Examples:
- Landing page A/B test measuring conversion rate
- Ad campaign comparing CTR across different creatives
- Email sequence measuring open rate and click-through rate
- SEO experiment tracking organic traffic changes
Qualitative experiments use qualitative-research skill:
- Gather non-numeric insights: opinions, experiences, needs, pain points, motivations
- Collect through interviews, surveys, focus groups, or observations
- Analyze using thematic analysis rather than statistical tests
- Focus on understanding why behaviors occur
- Examples:
- Customer discovery interviews to understand pain points
- Open-ended survey asking about product needs
- Focus group discussing ad creative perceptions
- Observational study of how users interact with product
Decision criteria:
Ask: "What do we need to learn?"
- If answer is "Does X increase metric Y by Z%?" → Quantitative (hypothesis-testing)
- If answer is "Why do users do X?" or "What do users think about Y?" → Qualitative (qualitative-research)
Ask: "What data will we collect?"
- If answer is "Metrics from analytics" → Quantitative (hypothesis-testing)
- If answer is "Interview transcripts, survey responses, or observation notes" → Qualitative (qualitative-research)
Mixed methods:
- Some hypotheses may require BOTH quantitative and qualitative experiments
- Example: "Value prop clarity drives conversion" could test:
- Quantitatively: A/B test landing page, measure conversion rate (hypothesis-testing)
- Qualitatively: Interview users about which value prop resonates (qualitative-research)
- Track these as separate experiments with linked hypotheses in the tracker
Create experiment tracker
The tracker is your coordination hub for managing multiple experiments over time.
Create 04-experiment-tracker.md with: ./templates/04-experiment-tracker.md
Tracker format:
For each selected hypothesis, create an entry:
### Experiment 1: [Hypothesis Brief Name]
**Status:** [Planned | In Progress | Complete]
**Hypothesis:** [Full hypothesis statement from Phase 2]
**Tactic/Channel:** [landing page | ads | email | etc.]
**Priority Score:** [ICE/RICE score from Phase 3]
**Start Date:** [YYYY-MM-DD or "Not started"]
**Completion Date:** [YYYY-MM-DD or "In progress"]
**Location:** `analysis/marketing-experimentation/[campaign-name]/experiments/[experiment-name]/`
**Signal:** [Positive | Negative | Null | Mixed | "Not analyzed"]
**Key Findings:** [Brief summary when complete, "TBD" otherwise]
Invoke appropriate skill for each experiment
For each hypothesis marked "Planned" or "In Progress":
Step 1: Read the hypothesis details from 02-hypothesis-generation.md
Step 2a: If Quantitative experiment, invoke hypothesis-testing skill:
Use hypothesis-testing skill to test: [Hypothesis statement]
Context for hypothesis-testing:
- Session name: [descriptive-name-for-experiment]
- Save location: analysis/marketing-experimentation/[campaign-name]/experiments/[experiment-name]/
- Success criteria: [From Phase 1 discovery]
- Expected outcome: [From Phase 2 hypothesis]
- Metric to measure: [CTR, conversion rate, etc.]
- Data source: [Google Analytics, database, etc.]
Step 2b: If Qualitative experiment, invoke qualitative-research skill:
Use qualitative-research skill to conduct: [Hypothesis statement]
Context for qualitative-research:
- Session name: [descriptive-name-for-experiment]
- Save location: analysis/marketing-experimentation/[campaign-name]/experiments/[experiment-name]/
- Research question: [From Phase 2 hypothesis]
- Collection method: [Interviews | Surveys | Focus Groups | Observations]
- Success criteria: [What insights validate/invalidate hypothesis?]
Step 3: Update experiment tracker:
- Change status from "Planned" to "In Progress"
- Add start date
- Update location with actual path
- Document experiment type (Quantitative or Qualitative)
Step 4a: Let hypothesis-testing skill complete its 5-phase workflow:
- Phase 1: Hypothesis Formulation
- Phase 2: Test Design
- Phase 3: Data Analysis
- Phase 4: Statistical Interpretation
- Phase 5: Conclusion
Step 4b: Let qualitative-research skill complete its 6-phase workflow:
- Phase 1: Research Design
- Phase 2: Data Collection
- Phase 3: Data Familiarization
- Phase 4: Systematic Coding
- Phase 5: Theme Development
- Phase 6: Synthesis & Reporting
Step 5: When skill completes, update tracker:
- Change status to "Complete"
- Add completion date
- Document signal (Positive/Negative/Null/Mixed)
- Summarize key findings
Step 6: Commit tracker updates after each status change
Handle multi-conversation resumption
At the start of EVERY conversation during Phase 4:
- Check if
04-experiment-tracker.md exists
- If it exists, READ IT FIRST before doing anything else
- Review experiment status:
- Planned: Ready to launch
- In Progress: Check hypothesis-testing session for current phase
- Complete: Ready for synthesis (Phase 5)
- Ask user which experiment to continue or which new experiment to launch
- Update tracker with new status/dates/findings
- Commit tracker updates
Example resumption:
I've read the experiment tracker. Current status:
- Experiment 1 (H1: Value prop): Complete, Positive signal
- Experiment 2 (H3: Email sequence): In Progress, currently in hypothesis-testing Phase 3
- Experiment 3 (H2: Ad targeting): Planned, not yet started
What would you like to do?
a) Continue Experiment 2 (in hypothesis-testing Phase 3)
b) Start Experiment 3
c) Move to synthesis (Phase 5) since Experiment 1 is complete
Coordinate parallel vs. sequential experiments
Parallel execution (multiple experiments simultaneously):
- Launch multiple hypothesis-testing sessions
- Track each separately in experiment tracker
- Update tracker as each progresses independently
- Requires managing multiple analysis directories
Sequential execution (one at a time):
- Complete one experiment fully before starting next
- Simpler tracking, easier to manage
- Can incorporate learnings between experiments
Hybrid execution:
- Run independent experiments in parallel
- Sequence dependent experiments (e.g., H2 depends on H1 insights)
Progress through all experiments
Continue invoking hypothesis-testing and updating the tracker until:
- All selected experiments have status "Complete"
- All experiments have documented signals
- All findings are summarized in tracker
Only when ALL experiments are complete should you proceed to Phase 5.
Common Rationalization: "I'll keep experiment details in my head - the tracker is just busywork"
Reality: Multi-day campaigns lose context between conversations. The tracker is the ONLY source of truth that persists across sessions. Without it, you'll re-ask questions and lose progress.
Common Rationalization: "I'll wait until all experiments finish before updating the tracker"
Reality: Batch updates create opportunity for lost data. Update the tracker IMMEDIATELY after status changes. Real-time tracking prevents confusion and missed experiments.
Common Rationalization: "I'll design the experiment myself instead of using hypothesis-testing"
Reality: hypothesis-testing skill provides rigorous experimental design, statistical analysis, and signal detection. Skipping it produces weak experiments with ambiguous results.
Common Rationalization: "All experiments are done, I don't need to update the tracker before synthesis"
Reality: The tracker is your input to Phase 5. Incomplete tracker means incomplete synthesis. Update ALL fields (status, dates, signals, findings) before proceeding.
Phase 5: Cross-Experiment Synthesis
CHECKPOINT: Before proceeding, you MUST have:
Instructions
CRITICAL: This phase synthesizes results ACROSS multiple experiments. Do NOT proceed until ALL Phase 4 experiments are complete with documented signals.
Verify experiment completion
Read 04-experiment-tracker.md and verify:
- All experiments have Status = "Complete"
- All experiments have Signal documented (Positive/Negative/Null/Mixed)
- All experiments have Key Findings summarized
If any experiments are incomplete, return to Phase 4 to finish them.
Create aggregate results table
Compile findings from all experiments into a summary table:
Example Aggregate Table (Mixed Quantitative & Qualitative):
| Experiment |
Type |
Hypothesis |
Tactic |
Signal |
Key Finding |
Confidence |
| E1 |
Quant |
Value prop clarity |
Landing page A/B test |
Positive |
Conversion rate +18% (p<0.05) |
High |
| E2 |
Qual |
Customer pain points |
Discovery interviews |
Positive |
8 of 10 cited onboarding complexity |
High |
| E3 |
Quant |
Ad targeting |
Ads |
Null |
CTR +2% (not sig., p=0.12) |
Medium |
| E4 |
Qual |
Ad message resonance |
Focus groups |
Negative |
6 of 8 found messaging confusing |
High |
For Quantitative experiments (hypothesis-testing):
- Signal classification (Positive/Negative/Null/Mixed)
- Key metric measured (CTR, conversion rate, etc.)
- Magnitude of effect with statistical significance (e.g., "+18%, p<0.05")
- Confidence level from statistical analysis
For Qualitative experiments (qualitative-research):
- Signal classification (Positive/Negative/Null/Mixed)
- Key themes identified with prevalence (e.g., "8 of 10 participants mentioned X")
- Representative quotes or patterns
- Confidence assessment (credibility, dependability, transferability)
Invoke presenting-data skill for comprehensive synthesis
Use the presenting-data skill to create complete synthesis with visualizations and presentation materials:
Use presenting-data skill to synthesize marketing experimentation results:
Context:
- Campaign: [campaign name]
- Experiments completed: [count]
- Results table: [paste aggregate table]
- Audience: [stakeholders/decision-makers]
- Format: [markdown report | slides | whitepaper]
- Focus: Pattern identification across experiments (what works, what doesn't, what's unclear)
presenting-data skill will handle:
- Pattern identification (using interpreting-results internally)
- Visualization creation (using creating-visualizations internally)
- Synthesis documentation (markdown, slides, or whitepaper format)
- Citation of sources (individual hypothesis-testing sessions)
- Reproducibility (references to experiment locations)
Focus areas for synthesis:
- What worked: Experiments with Positive signals (both quantitative and qualitative)
- What didn't work: Experiments with Negative signals (both quantitative and qualitative)
- What's unclear: Experiments with Null or Mixed signals
- Cross-experiment patterns: Do results cluster by tactic? By audience? By timing? Do quantitative and qualitative findings align or conflict?
- Triangulation: Do qualitative findings explain quantitative results? (e.g., interviews reveal WHY conversion rate increased)
- Confounding factors: Are there external factors affecting multiple experiments?
- Confidence assessment: Which findings are robust? Which are uncertain? How do qualitative and quantitative confidence levels compare?
Document patterns and insights
The presenting-data skill will create 05-synthesis.md (or slides/whitepaper) with:
What Worked (Positive Signals):
- List experiments with positive results
- Explain WHY these worked (based on analysis)
- Identify commonalities across successful experiments
What Didn't Work (Negative Signals):
- List experiments with negative results
- Explain WHY these failed (based on analysis)
- Identify lessons learned
What's Unclear (Null/Mixed Signals):
- List experiments with inconclusive results
- Explain potential reasons (insufficient power, confounding factors, etc.)
- Identify what additional investigation is needed
Cross-Experiment Patterns:
- Do results cluster by tactic, audience, timing, or other factors?
- Are there confounding variables affecting multiple experiments?
- What overarching insights emerge?
Visualizations (created by presenting-data):
- Signal distribution (bar chart: Positive/Negative/Null/Mixed counts)
- Effect sizes (bar chart: metric changes by experiment)
- Confidence levels (scatter plot: effect size vs. confidence)
- Tactic performance (grouped by channel/tactic)
Classify overall campaign signal
Based on aggregate analysis (from presenting-data output), classify the campaign:
Positive: Campaign validates concept, proceed to scaling
- Multiple experiments show positive signals
- Successful tactics identified for scale-up
- Clear path to ROI improvement
Negative: Campaign invalidates concept, pivot or abandon
- Multiple experiments show negative signals
- No successful tactics identified
- Concept doesn't resonate with audience
Null: Campaign results inconclusive, needs refinement
- Most experiments show null signals
- Insufficient power or confounding factors
- Needs redesigned experiments or longer observation
Mixed: Some aspects work, some don't, iterate strategically
- Mix of positive and negative signals across experiments
- Some tactics work, others don't
- Selective scaling + pivots needed
Review presenting-data output and finalize synthesis
After presenting-data skill completes:
- Review generated synthesis document
- Verify all experiments are covered
- Confirm visualizations are appropriate
- Ensure signal classification is documented
- Make any necessary edits for clarity
STOP and get user confirmation
- Review synthesis findings with user
- Confirm pattern interpretations are accurate
- Confirm overall signal classification is appropriate
- Do NOT proceed to Phase 6 until confirmed
Common Rationalization: "I'll synthesize results mentally - no need to document patterns"
Reality: Mental synthesis loses details and creates false confidence. Documented synthesis with presenting-data skill ensures intellectual honesty and identifies confounding factors you'd otherwise miss.
Common Rationalization: "I'll
…(truncated)
1---2name: marketing-experimentation3description: Systematic marketing experimentation process - discover concepts, generate hypotheses, coordinate multiple experiments, synthesize results, generate next-iteration ideas through rigorous validation cycles4---5
6# Marketing Experimentation
7
8## Overview
9
10Use this skill when you need to validate marketing concepts or business ideas through rigorous experimental cycles. This skill orchestrates the complete Build-Measure-Learn cycle from concept to data-driven signal.
11
12**When to use this skill:**
13- You have a marketing concept or business idea that needs validation
14- You want to test multiple related hypotheses systematically
15- You need to integrate results across multiple experiments
16- You're designing the next iteration based on experimental evidence
17- You're conducting qualitative market research combined with quantitative testing
18
19**What this skill does:**
20- Validates concepts through market research before experimentation
21- Generates multiple testable hypotheses from marketing ideas
22- Coordinates multiple experiments: quantitative (hypothesis-testing) and qualitative (qualitative-research)
23- Synthesizes results across experiments using interpreting-results and creating-visualizations
24- Produces clear signals (positive/negative/null/mixed) for each campaign
25- Generates actionable next-iteration ideas based on experimental evidence
26
27**What this skill does NOT do:**
28- Design individual experiments (delegates to hypothesis-testing or qualitative-research)
29- Execute statistical analysis directly (uses hypothesis-testing for quantitative rigor)
30- Conduct interviews/surveys/observations directly (uses qualitative-research for qualitative rigor)
31- Operationalize successful ideas (focuses on validation, not scaling)
32- Platform-specific implementation (tool-agnostic techniques only)
33
34**Integration with existing skills:**
35- **Delegates to `hypothesis-testing`** for quantitative experiment design and execution (metrics, A/B tests, statistical analysis)
36- **Delegates to `qualitative-research`** for qualitative experiment design and execution (interviews, surveys, focus groups, observations)
37- **Uses `interpreting-results`** to synthesize findings across multiple experiments
38- **Uses `creating-visualizations`** to communicate aggregate results
39- **Invokes `market-researcher` agent** for concept validation via internet research
40
41**Multi-conversation persistence:**
42This skill is designed for campaigns spanning days or weeks. Each phase documents completely enough that new conversations can resume after extended breaks. The experiment tracker (04-experiment-tracker.md) serves as the living coordination hub.
43
44## Prerequisites
45
46**Required skills:**
47- `hypothesis-testing` - Quantitative experiment design and execution (invoked for metric-based experiments)
48- `qualitative-research` - Qualitative experiment design and execution (invoked for interviews, surveys, focus groups, observations)
49- `interpreting-results` - Result synthesis and pattern identification (invoked in Phase 5)
50- `creating-visualizations` - Aggregate result visualization (invoked in Phase 5)
51
52**Required agents:**
53- `market-researcher` - Concept validation via internet research (invoked in Phase 1)
54
55**Required knowledge:**
56- Understanding of Lean Startup Build-Measure-Learn cycle
57- Familiarity with marketing tactics (landing pages, ads, email, content)
58- Basic experimental design principles (control/treatment, signals, metrics)
59- Understanding of qualitative vs quantitative research methods
60
61**Data requirements:**
62- None initially (market research is qualitative)
63- Data requirements emerge from experiment design in Phase 4
64- Quantitative experiments (hypothesis-testing): SQL databases, analytics data, A/B test results
65- Qualitative experiments (qualitative-research): Interview transcripts, survey responses, observation notes
66
67## Mandatory Process Structure
68
69**CRITICAL:** This is a 6-phase process skill. You MUST complete all phases in order. Use TodoWrite to track progress through each phase.
70
71**TodoWrite template:**
72
73When starting a marketing-experimentation session, create these todos:
74
75```markdown
76- [ ] Phase 1: Discovery & Asset Inventory
77- [ ] Phase 2: Hypothesis Generation
78- [ ] Phase 3: Prioritization
79- [ ] Phase 4: Experiment Coordination
80- [ ] Phase 5: Cross-Experiment Synthesis
81- [ ] Phase 6: Iteration Planning
82```
83
84**Workspace structure:**
85
86All work for a marketing-experimentation session is saved to:
87```
88analysis/marketing-experimentation/[campaign-name]/
89├── 01-discovery.md
90├── 02-hypothesis-generation.md
91├── 03-prioritization.md
92├── 04-experiment-tracker.md
93├── 05-synthesis.md
94├── 06-iteration-plan.md
95└── experiments/
96 ├── [experiment-1]/ # hypothesis-testing session
97 ├── [experiment-2]/ # hypothesis-testing session
98 └── [experiment-3]/ # hypothesis-testing session
99```
100
101**Phase progression rules:**
1021. Each phase has a CHECKPOINT with verification requirements
1032. You MUST satisfy all checkpoint requirements before proceeding to the next phase
1043. Document every decision with rationale in the numbered markdown files
1054. Commit markdown files after each phase completes
1065. The experiment tracker (04-experiment-tracker.md) is a LIVING DOCUMENT - update it throughout Phase 4
107
108**Multi-conversation resumption:**
109- At the start of any conversation, check if an experiment tracker exists for this campaign
110- If it exists, read it first to understand current experiment status
111- Update the tracker as experiments progress
112- All phases should be complete enough to resume after days or weeks
113
114---
115
116## Phase 1: Discovery & Asset Inventory
117
118**CHECKPOINT:** Before proceeding, you MUST have:
119- [ ] Gathered business concept description from user
120- [ ] Invoked market-researcher agent and documented findings
121- [ ] Completed asset inventory (content, campaigns, audiences, data)
122- [ ] Defined success criteria and validation signals
123- [ ] Documented known constraints
124- [ ] Saved to `01-discovery.md`
125
126### Instructions
127
1281. **Gather the business concept**
129 - Ask user to describe the marketing concept or business idea to validate
130 - What problem does it solve? Who is the target audience?
131 - What's the desired outcome? (awareness, leads, conversions, etc.)
132 - What stage is this idea at? (new concept, existing campaign, iteration)
133
1342. **Invoke market-researcher agent for concept validation**
135
136 Dispatch the `market-researcher` agent with the concept description:
137 - Agent will research market demand signals
138 - Agent will identify similar solutions and competitors
139 - Agent will analyze audience needs and pain points
140 - Agent will find validation evidence (case studies, reviews, testimonials)
141
142 Document agent findings in `01-discovery.md` under "Market Research Findings"
143
1443. **Conduct asset inventory**
145
146 Work with user to inventory existing assets that could be leveraged:
147
148 **Content Assets:**
149 - Blog posts, case studies, whitepapers
150 - Video content, webinars, tutorials
151 - Social media presence and following
152 - Email lists and subscriber segments
153
154 **Campaign Assets:**
155 - Existing ad campaigns and performance data
156 - Landing pages and conversion rates
157 - Email campaigns and open/click rates
158 - SEO performance and keyword rankings
159
160 **Audience Assets:**
161 - Customer segments and personas
162 - Audience data (demographics, behaviors, preferences)
163 - Customer feedback and reviews
164 - Support tickets and common questions
165
166 **Data Assets:**
167 - Analytics platforms (Google Analytics, Mixpanel, etc.)
168 - CRM data (Salesforce, HubSpot, etc.)
169 - Ad platform data (Google Ads, Facebook Ads, etc.)
170 - Email platform data (Mailchimp, SendGrid, etc.)
171
1724. **Define success criteria and validation signals**
173
174 Work with user to define:
175 - What metrics indicate success? (CTR, conversion rate, CAC, LTV, etc.)
176 - What magnitude of change is meaningful? (practical significance thresholds)
177 - What signal types are acceptable?
178 - Positive: Validates concept, proceed to scale
179 - Negative: Invalidates concept, pivot or abandon
180 - Null: Inconclusive, needs refinement or more data
181 - Mixed: Some aspects work, some don't, iterate strategically
182
1835. **Document known constraints**
184
185 Capture any constraints that will affect experimentation:
186 - Budget constraints (ad spend limits, tool costs)
187 - Time constraints (launch deadlines, seasonal factors)
188 - Resource constraints (team capacity, content production)
189 - Technical constraints (platform limitations, integration issues)
190 - Regulatory constraints (GDPR, CCPA, industry regulations)
191
1926. **Create `01-discovery.md`** with: `./templates/01-discovery.md`
193
1947. **STOP and get user confirmation**
195 - Review discovery findings with user
196 - Confirm asset inventory is complete
197 - Confirm success criteria are appropriate
198 - Do NOT proceed to Phase 2 until confirmed
199
200**Common Rationalization:** "I'll skip discovery and go straight to testing - the concept is obvious"
201**Reality:** Discovery surfaces assumptions, constraints, and existing assets that dramatically affect experiment design. Always start with discovery.
202
203**Common Rationalization:** "I don't need market research - I already know this market"
204**Reality:** The market-researcher agent provides current, data-driven validation signals that prevent building experiments around false assumptions. Always validate.
205
206**Common Rationalization:** "Asset inventory is busywork - I'll figure out what's available as I go"
207**Reality:** Existing assets can dramatically reduce experiment cost and time. Inventorying first prevents reinventing wheels and enables building on proven foundations.
208
209---
210
211## Phase 2: Hypothesis Generation
212
213**CHECKPOINT:** Before proceeding, you MUST have:
214- [ ] Generated 5-10 testable hypotheses from the concept
215- [ ] Each hypothesis maps to a specific tactic/channel
216- [ ] Each hypothesis has expected outcome and rationale
217- [ ] Hypotheses cover multiple tactics (not all ads or all email)
218- [ ] Referenced relevant frameworks (Lean Startup, AARRR, ICE/RICE)
219- [ ] Saved to `02-hypothesis-generation.md`
220
221### Instructions
222
2231. **Generate 5-10 testable hypotheses**
224
225 For each hypothesis, use this format:
226
227 **Hypothesis [N]: [Brief statement]**
228 - **Tactic/Channel:** [landing page | ad campaign | email sequence | content marketing | social media | SEO | etc.]
229 - **Expected Outcome:** [Specific, measurable result]
230 - **Rationale:** [Why we believe this will work based on discovery findings]
231 - **Variables to Test:** [What will we manipulate/measure]
232
233 **Example hypothesis:**
234
235 **Hypothesis 1: Value proposition clarity drives conversion**
236 - **Tactic/Channel:** Landing page A/B test
237 - **Expected Outcome:** 15%+ increase in conversion rate from landing page variant with simplified value proposition
238 - **Rationale:** Market research showed audience confusion about product benefits. Discovery found existing landing page has 8 different value propositions competing for attention.
239 - **Variables to Test:** Headline clarity, benefit hierarchy, CTA prominence
240
2412. **Ensure tactic coverage**
242
243 Verify hypotheses cover multiple marketing tactics:
244
245 **Acquisition Tactics:**
246 - Landing pages (conversion optimization, value prop testing, layout)
247 - Ad campaigns (targeting, creative, messaging, platforms)
248 - Content marketing (blog posts, videos, webinars, lead magnets)
249 - SEO (keyword targeting, content optimization, technical SEO)
250
251 **Activation Tactics:**
252 - Email sequences (onboarding, nurture, activation)
253 - Product tours (in-app guidance, feature discovery)
254 - Social proof (testimonials, case studies, reviews)
255
256 **Retention Tactics:**
257 - Email campaigns (engagement, re-activation, upsell)
258 - Content (newsletters, educational content, community)
259
260 Don't generate 10 ad hypotheses. Aim for diversity across tactics.
261
2623. **Reference experimentation frameworks**
263
264 **Lean Startup Build-Measure-Learn:**
265 - Build: What's the minimum viable test? (landing page, ad, email, etc.)
266 - Measure: What metrics indicate success/failure?
267 - Learn: What will we learn regardless of outcome?
268
269 **AARRR Pirate Metrics:**
270 - Acquisition: How do users find us?
271 - Activation: Do they have a great first experience?
272 - Retention: Do they come back?
273 - Referral: Do they tell others?
274 - Revenue: Do they pay?
275
276 Map each hypothesis to one or more AARRR stages.
277
278 **ICE/RICE Prioritization (used in Phase 3):**
279 - Impact: How much will this move the metric?
280 - Confidence: How sure are we this will work?
281 - Ease: How easy is this to implement?
282 - Reach: How many users will this affect? (RICE only)
283
2844. **Create `02-hypothesis-generation.md`** with: `./templates/02-hypothesis-generation.md`
285
2865. **STOP and get user confirmation**
287 - Review all hypotheses with user
288 - Confirm hypotheses are testable and meaningful
289 - Confirm tactic coverage is appropriate
290 - Do NOT proceed to Phase 3 until confirmed
291
292**Common Rationalization:** "I'll generate hypotheses as I build experiments - more efficient"
293**Reality:** Generating hypotheses before prioritization enables strategic selection of highest-impact tests. Generating ad-hoc leads to testing whatever's easiest, not what matters most.
294
295**Common Rationalization:** "I'll focus all hypotheses on one tactic (ads) since that's what we know"
296**Reality:** Tactic diversity reveals which channels work for this concept. Single-tactic testing creates blind spots and missed opportunities.
297
298**Common Rationalization:** "I'll write vague hypotheses and refine them during experiment design"
299**Reality:** Vague hypotheses lead to vague experiments that produce vague results. Specific hypotheses with expected outcomes enable clear signal detection.
300
301**Common Rationalization:** "More hypotheses = better coverage, I'll generate 20+"
302**Reality:** Too many hypotheses dilute focus and create analysis paralysis in prioritization. 5-10 high-quality hypotheses enable strategic selection of 2-4 tests.
303
304---
305
306## Phase 3: Prioritization
307
308**CHECKPOINT:** Before proceeding, you MUST have:
309- [ ] Scored all hypotheses using ICE or RICE framework with computational method
310- [ ] Created prioritized backlog (highest to lowest score) using Python script
311- [ ] Selected 2-4 highest-priority hypotheses for testing
312- [ ] Documented prioritization rationale
313- [ ] Identified any dependencies or sequencing requirements
314- [ ] Saved to `03-prioritization.md`
315
316### Instructions
317
318**CRITICAL:** You MUST use computational methods (Python scripts) to calculate scores. Do NOT estimate or manually calculate scores.
319
3201. **Choose prioritization framework**
321
322 **ICE Framework** (simpler, faster):
323 - **Impact:** How much will this move the success metric? (1-10 scale)
324 - **Confidence:** How confident are we this will work? (1-10 scale)
325 - **Ease:** How easy is this to implement? (1-10 scale)
326 - **Score:** (Impact × Confidence) / Ease
327
328 **RICE Framework** (more comprehensive):
329 - **Reach:** How many users will this affect? (absolute number or percentage)
330 - **Impact:** How much will this move the metric per user? (1-10 scale: 0.25=minimal, 3=massive)
331 - **Confidence:** How confident are we in our estimates? (percentage: 50%, 80%, 100%)
332 - **Effort:** Person-weeks to implement (absolute number)
333 - **Score:** (Reach × Impact × Confidence) / Effort
334
335 Choose ICE for speed, RICE for precision when reach varies significantly.
336
3372. **Score each hypothesis using Python script**
338
339 **For ICE Framework:**
340
341 Create a Python script to compute and sort ICE scores:
342
343 ```python
344 #!/usr/bin/env python3
345 """
346 ICE Score Calculator for Marketing Experimentation
347
348 Computes ICE scores: (Impact × Confidence) / Ease
349 Sorts hypotheses by score (highest to lowest)
350 """
351
352 hypotheses = [
353 {
354 "id": "H1",
355 "name": "Value proposition clarity drives conversion",
356 "impact": 8,
357 "confidence": 7,
358 "ease": 9
359 },
360 {
361 "id": "H2",
362 "name": "Ad targeting refinement",
363 "impact": 7,
364 "confidence": 6,
365 "ease": 5
366 },
367 {
368 "id": "H3",
369 "name": "Email sequence optimization",
370 "impact": 6,
371 "confidence": 8,
372 "ease": 8
373 },
374 {
375 "id": "H4",
376 "name": "Content marketing expansion",
377 "impact": 5,
378 "confidence": 4,
379 "ease": 3
380 },
381 ]
382
383 # Calculate ICE scores
384 for h in hypotheses:
385 h['ice_score'] = (h['impact'] * h['confidence']) / h['ease']
386
387 # Sort by ICE score (descending)
388 sorted_hypotheses = sorted(hypotheses, key=lambda x: x['ice_score'], reverse=True)
389
390 # Print results table
391 print("| Hypothesis | Impact | Confidence | Ease | ICE Score | Rank |")
392 print("|------------|--------|------------|------|-----------|------|")
393 for rank, h in enumerate(sorted_hypotheses, 1):
394 print(f"| {h['id']}: {h['name'][:30]} | {h['impact']} | {h['confidence']} | {h['ease']} | {h['ice_score']:.2f} | {rank} |")
395 ```
396
397 **Usage:**
398 ```bash
399 python3 ice_calculator.py
400 ```
401
402 **For RICE Framework:**
403
404 Create a Python script to compute and sort RICE scores:
405
406 ```python
407 #!/usr/bin/env python3
408 """
409 RICE Score Calculator for Marketing Experimentation
410
411 Computes RICE scores: (Reach × Impact × Confidence) / Effort
412 Sorts hypotheses by score (highest to lowest)
413 """
414
415 hypotheses = [
416 {
417 "id": "H1",
418 "name": "Value proposition clarity drives conversion",
419 "reach": 10000, # users affected
420 "impact": 3, # 0.25=minimal, 1=low, 2=medium, 3=high, 5=massive
421 "confidence": 80, # percentage (50, 80, 100)
422 "effort": 2 # person-weeks
423 },
424 {
425 "id": "H2",
426 "name": "Ad targeting refinement",
427 "reach": 50000,
428 "impact": 1,
429 "confidence": 50,
430 "effort": 4
431 },
432 {
433 "id": "H3",
434 "name": "Email sequence optimization",
435 "reach": 5000,
436 "impact": 2,
437 "confidence": 80,
438 "effort": 3
439 },
440 {
441 "id": "H4",
442 "name": "Content marketing expansion",
443 "reach": 20000,
444 "impact": 1,
445 "confidence": 50,
446 "effort": 8
447 },
448 ]
449
450 # Calculate RICE scores
451 for h in hypotheses:
452 # Convert confidence percentage to decimal
453 confidence_decimal = h['confidence'] / 100
454 h['rice_score'] = (h['reach'] * h['impact'] * confidence_decimal) / h['effort']
455
456 # Sort by RICE score (descending)
457 sorted_hypotheses = sorted(hypotheses, key=lambda x: x['rice_score'], reverse=True)
458
459 # Print results table
460 print("| Hypothesis | Reach | Impact | Confidence | Effort | RICE Score | Rank |")
461 print("|------------|-------|--------|------------|--------|------------|------|")
462 for rank, h in enumerate(sorted_hypotheses, 1):
463 print(f"| {h['id']}: {h['name'][:30]} | {h['reach']} | {h['impact']} | {h['confidence']}% | {h['effort']}w | {h['rice_score']:.2f} | {rank} |")
464 ```
465
466 **Usage:**
467 ```bash
468 python3 rice_calculator.py
469 ```
470
471 **Scoring Guidance:**
472
473 **Impact (1-10 for ICE, 0.25-5 for RICE):**
474 - ICE: 1-3 minimal, 4-6 moderate, 7-8 significant, 9-10 transformative
475 - RICE: 0.25 minimal, 1 low, 2 medium, 3 high, 5 massive
476
477 **Confidence (1-10 for ICE, 50-100% for RICE):**
478 - ICE: 1-3 speculative, 4-6 uncertain, 7-8 likely, 9-10 validated
479 - RICE: 50% low confidence, 80% high confidence, 100% certainty
480
481 **Ease (1-10 for ICE):**
482 - 1-3: Complex, significant resources
483 - 4-6: Moderate effort, some obstacles
484 - 7-8: Straightforward, few dependencies
485 - 9-10: Trivial, immediate execution
486
487 **Effort (person-weeks for RICE):**
488 - Estimate total person-weeks required
489 - Include design, implementation, monitoring time
490 - Examples: 1w (simple landing page), 4w (complex ad campaign), 8w (content series)
491
4923. **Run scoring script and document results**
493
494 1. Create the Python script (ice_calculator.py or rice_calculator.py)
495 2. Update the `hypotheses` list with actual hypothesis data from Phase 2
496 3. Run the script: `python3 [ice|rice]_calculator.py`
497 4. Copy the output table into `03-prioritization.md`
498 5. Include the script in the markdown file for reproducibility:
499
500 ```markdown
501 ## Prioritization Calculation
502
503 **Method:** ICE Framework
504
505 **Calculation Script:**
506 ```python
507 [paste full script here]
508 ```
509
510 **Results:**
511
512 [paste output table here]
513 ```
514
5154. **Select 2-4 highest-priority hypotheses**
516
517 Considerations for selection:
518 - **Don't test everything:** Focus on highest-scoring 2-4 hypotheses
519 - **Resource constraints:** Match selection to available time/budget/capacity
520 - **Learning value:** Sometimes lower-scoring hypothesis with high uncertainty is worth testing
521 - **Dependencies:** Test prerequisites before dependent hypotheses
522 - **Sequencing:** Consider whether experiments need to run sequentially or can run in parallel
523
524 **Selection criteria:**
525 - Primary: Top 2-4 by ICE/RICE score from computational results
526 - Secondary: Balance quick wins vs. high-impact long-term bets
527 - Tertiary: Ensure tactic diversity (don't test 3 ad variants if other tactics untested)
528
5295. **Document experiment sequence**
530
531 Determine execution strategy:
532 - **Parallel:** Multiple experiments running simultaneously (faster results, higher resource needs)
533 - **Sequential:** One experiment at a time (slower, easier to manage)
534 - **Hybrid:** Run independent experiments in parallel, sequence dependent ones
535
536 **Example sequence plan:**
537 ```
538 Week 1-2: Launch H1 (landing page) and H3 (email) in parallel
539 Week 3-4: Analyze H1 and H3 results
540 Week 5-6: Launch H2 (ads) based on H1 learnings
541 Week 7-8: Analyze H2 results
542 ```
543
5446. **Create `03-prioritization.md`** with: `./templates/03-prioritization.md`
545
5467. **STOP and get user confirmation**
547 - Review computed scores and prioritization with user
548 - Confirm selected hypotheses are appropriate
549 - Confirm experiment sequence is feasible
550 - Do NOT proceed to Phase 4 until confirmed
551
552**Common Rationalization:** "I'll test all hypotheses - don't want to miss opportunities"
553**Reality:** Resource constraints make testing everything impossible. Prioritization ensures highest-value experiments get resources. Unfocused testing produces weak signals across too many fronts.
554
555**Common Rationalization:** "Scoring is subjective and arbitrary - I'll just pick what feels right"
556**Reality:** Scoring frameworks force explicit reasoning about trade-offs. "Feels right" selections optimize for recency bias and personal preference, not business value. Computational methods ensure consistency.
557
558**Common Rationalization:** "I'll skip prioritization and go straight to easiest test"
559**Reality:** Easiest test rarely equals highest value. Prioritization prevents optimizing for ease at the expense of impact.
560
561**Common Rationalization:** "I'll estimate scores mentally instead of running the script"
562**Reality:** Manual estimation introduces calculation errors and inconsistency. Python scripts ensure exact, reproducible results that can be audited and verified.
563
564---
565
566## Phase 4: Experiment Coordination
567
568**CHECKPOINT:** Before proceeding, you MUST have:
569- [ ] Created experiment tracker with all selected hypotheses
570- [ ] Determined experiment type for each hypothesis (Quantitative or Qualitative)
571- [ ] Invoked appropriate skill for each hypothesis (hypothesis-testing OR qualitative-research)
572- [ ] Updated tracker with experiment status (Planned, In Progress, Complete)
573- [ ] Documented location of each experiment session
574- [ ] Saved experiment tracker to `04-experiment-tracker.md`
575- [ ] Note: This phase may span multiple days/weeks and conversations
576
577### Instructions
578
579**CRITICAL:** This phase is designed for multi-conversation workflows. The experiment tracker is a LIVING DOCUMENT that you will update throughout experimentation. New conversations should ALWAYS read this file first.
580
5811. **Determine experiment type for each hypothesis**
582
583 **CRITICAL:** Before creating the tracker, classify each hypothesis as Quantitative or Qualitative.
584
585 **Quantitative experiments** use hypothesis-testing skill:
586 - Measure **numeric metrics**: CTR, conversion rate, bounce rate, time on page, revenue, CAC, LTV, etc.
587 - Rely on **existing data sources**: Google Analytics, ad platforms, CRM, email platforms, database queries
588 - Test using **A/B tests, multivariate tests, or time-series analysis**
589 - Require **statistical significance testing**
590 - **Examples:**
591 - Landing page A/B test measuring conversion rate
592 - Ad campaign comparing CTR across different creatives
593 - Email sequence measuring open rate and click-through rate
594 - SEO experiment tracking organic traffic changes
595
596 **Qualitative experiments** use qualitative-research skill:
597 - Gather **non-numeric insights**: opinions, experiences, needs, pain points, motivations
598 - Collect through **interviews, surveys, focus groups, or observations**
599 - Analyze using **thematic analysis** rather than statistical tests
600 - Focus on **understanding why** behaviors occur
601 - **Examples:**
602 - Customer discovery interviews to understand pain points
603 - Open-ended survey asking about product needs
604 - Focus group discussing ad creative perceptions
605 - Observational study of how users interact with product
606
607 **Decision criteria:**
608
609 Ask: "What do we need to learn?"
610 - If answer is "Does X increase metric Y by Z%?" → **Quantitative** (hypothesis-testing)
611 - If answer is "Why do users do X?" or "What do users think about Y?" → **Qualitative** (qualitative-research)
612
613 Ask: "What data will we collect?"
614 - If answer is "Metrics from analytics" → **Quantitative** (hypothesis-testing)
615 - If answer is "Interview transcripts, survey responses, or observation notes" → **Qualitative** (qualitative-research)
616
617 **Mixed methods:**
618 - Some hypotheses may require BOTH quantitative and qualitative experiments
619 - Example: "Value prop clarity drives conversion" could test:
620 - Quantitatively: A/B test landing page, measure conversion rate (hypothesis-testing)
621 - Qualitatively: Interview users about which value prop resonates (qualitative-research)
622 - Track these as separate experiments with linked hypotheses in the tracker
623
6242. **Create experiment tracker**
625
626 The tracker is your coordination hub for managing multiple experiments over time.
627
628 Create `04-experiment-tracker.md` with: `./templates/04-experiment-tracker.md`
629
630 **Tracker format:**
631
632 For each selected hypothesis, create an entry:
633
634 ```markdown
635 ### Experiment 1: [Hypothesis Brief Name]
636
637 **Status:** [Planned | In Progress | Complete]
638 **Hypothesis:** [Full hypothesis statement from Phase 2]
639 **Tactic/Channel:** [landing page | ads | email | etc.]
640 **Priority Score:** [ICE/RICE score from Phase 3]
641 **Start Date:** [YYYY-MM-DD or "Not started"]
642 **Completion Date:** [YYYY-MM-DD or "In progress"]
643 **Location:** `analysis/marketing-experimentation/[campaign-name]/experiments/[experiment-name]/`
644 **Signal:** [Positive | Negative | Null | Mixed | "Not analyzed"]
645 **Key Findings:** [Brief summary when complete, "TBD" otherwise]
646 ```
647
6482. **Invoke appropriate skill for each experiment**
649
650 For each hypothesis marked "Planned" or "In Progress":
651
652 **Step 1:** Read the hypothesis details from `02-hypothesis-generation.md`
653
654 **Step 2a:** If **Quantitative experiment**, invoke `hypothesis-testing` skill:
655 ```markdown
656 Use hypothesis-testing skill to test: [Hypothesis statement]
657
658 Context for hypothesis-testing:
659 - Session name: [descriptive-name-for-experiment]
660 - Save location: analysis/marketing-experimentation/[campaign-name]/experiments/[experiment-name]/
661 - Success criteria: [From Phase 1 discovery]
662 - Expected outcome: [From Phase 2 hypothesis]
663 - Metric to measure: [CTR, conversion rate, etc.]
664 - Data source: [Google Analytics, database, etc.]
665 ```
666
667 **Step 2b:** If **Qualitative experiment**, invoke `qualitative-research` skill:
668 ```markdown
669 Use qualitative-research skill to conduct: [Hypothesis statement]
670
671 Context for qualitative-research:
672 - Session name: [descriptive-name-for-experiment]
673 - Save location: analysis/marketing-experimentation/[campaign-name]/experiments/[experiment-name]/
674 - Research question: [From Phase 2 hypothesis]
675 - Collection method: [Interviews | Surveys | Focus Groups | Observations]
676 - Success criteria: [What insights validate/invalidate hypothesis?]
677 ```
678
679 **Step 3:** Update experiment tracker:
680 - Change status from "Planned" to "In Progress"
681 - Add start date
682 - Update location with actual path
683 - Document experiment type (Quantitative or Qualitative)
684
685 **Step 4a:** Let **hypothesis-testing** skill complete its 5-phase workflow:
686 - Phase 1: Hypothesis Formulation
687 - Phase 2: Test Design
688 - Phase 3: Data Analysis
689 - Phase 4: Statistical Interpretation
690 - Phase 5: Conclusion
691
692 **Step 4b:** Let **qualitative-research** skill complete its 6-phase workflow:
693 - Phase 1: Research Design
694 - Phase 2: Data Collection
695 - Phase 3: Data Familiarization
696 - Phase 4: Systematic Coding
697 - Phase 5: Theme Development
698 - Phase 6: Synthesis & Reporting
699
700 **Step 5:** When skill completes, update tracker:
701 - Change status to "Complete"
702 - Add completion date
703 - Document signal (Positive/Negative/Null/Mixed)
704 - Summarize key findings
705
706 **Step 6:** Commit tracker updates after each status change
707
7083. **Handle multi-conversation resumption**
709
710 **At the start of EVERY conversation during Phase 4:**
711
712 1. Check if `04-experiment-tracker.md` exists
713 2. If it exists, READ IT FIRST before doing anything else
714 3. Review experiment status:
715 - Planned: Ready to launch
716 - In Progress: Check hypothesis-testing session for current phase
717 - Complete: Ready for synthesis (Phase 5)
718 4. Ask user which experiment to continue or which new experiment to launch
719 5. Update tracker with new status/dates/findings
720 6. Commit tracker updates
721
722 **Example resumption:**
723 ```markdown
724 I've read the experiment tracker. Current status:
725 - Experiment 1 (H1: Value prop): Complete, Positive signal
726 - Experiment 2 (H3: Email sequence): In Progress, currently in hypothesis-testing Phase 3
727 - Experiment 3 (H2: Ad targeting): Planned, not yet started
728
729 What would you like to do?
730 a) Continue Experiment 2 (in hypothesis-testing Phase 3)
731 b) Start Experiment 3
732 c) Move to synthesis (Phase 5) since Experiment 1 is complete
733 ```
734
7354. **Coordinate parallel vs. sequential experiments**
736
737 **Parallel execution (multiple experiments simultaneously):**
738 - Launch multiple hypothesis-testing sessions
739 - Track each separately in experiment tracker
740 - Update tracker as each progresses independently
741 - Requires managing multiple analysis directories
742
743 **Sequential execution (one at a time):**
744 - Complete one experiment fully before starting next
745 - Simpler tracking, easier to manage
746 - Can incorporate learnings between experiments
747
748 **Hybrid execution:**
749 - Run independent experiments in parallel
750 - Sequence dependent experiments (e.g., H2 depends on H1 insights)
751
7525. **Progress through all experiments**
753
754 Continue invoking hypothesis-testing and updating the tracker until:
755 - All selected experiments have status "Complete"
756 - All experiments have documented signals
757 - All findings are summarized in tracker
758
759 Only when ALL experiments are complete should you proceed to Phase 5.
760
761**Common Rationalization:** "I'll keep experiment details in my head - the tracker is just busywork"
762**Reality:** Multi-day campaigns lose context between conversations. The tracker is the ONLY source of truth that persists across sessions. Without it, you'll re-ask questions and lose progress.
763
764**Common Rationalization:** "I'll wait until all experiments finish before updating the tracker"
765**Reality:** Batch updates create opportunity for lost data. Update the tracker IMMEDIATELY after status changes. Real-time tracking prevents confusion and missed experiments.
766
767**Common Rationalization:** "I'll design the experiment myself instead of using hypothesis-testing"
768**Reality:** hypothesis-testing skill provides rigorous experimental design, statistical analysis, and signal detection. Skipping it produces weak experiments with ambiguous results.
769
770**Common Rationalization:** "All experiments are done, I don't need to update the tracker before synthesis"
771**Reality:** The tracker is your input to Phase 5. Incomplete tracker means incomplete synthesis. Update ALL fields (status, dates, signals, findings) before proceeding.
772
773---
774
775## Phase 5: Cross-Experiment Synthesis
776
777**CHECKPOINT:** Before proceeding, you MUST have:
778- [ ] ALL experiments from Phase 4 marked "Complete" with signals documented
779- [ ] Created aggregate results table across all experiments
780- [ ] Invoked presenting-data skill to synthesize findings with visualizations
781- [ ] Documented what worked, what didn't, and what's unclear
782- [ ] Classified overall campaign signal (Positive/Negative/Null/Mixed)
783- [ ] Saved to `05-synthesis.md`
784
785### Instructions
786
787**CRITICAL:** This phase synthesizes results ACROSS multiple experiments. Do NOT proceed until ALL Phase 4 experiments are complete with documented signals.
788
7891. **Verify experiment completion**
790
791 Read `04-experiment-tracker.md` and verify:
792 - All experiments have Status = "Complete"
793 - All experiments have Signal documented (Positive/Negative/Null/Mixed)
794 - All experiments have Key Findings summarized
795
796 If any experiments are incomplete, return to Phase 4 to finish them.
797
7982. **Create aggregate results table**
799
800 Compile findings from all experiments into a summary table:
801
802 **Example Aggregate Table (Mixed Quantitative & Qualitative):**
803
804 | Experiment | Type | Hypothesis | Tactic | Signal | Key Finding | Confidence |
805 |------------|------|------------|--------|--------|-------------|------------|
806 | E1 | Quant | Value prop clarity | Landing page A/B test | Positive | Conversion rate +18% (p<0.05) | High |
807 | E2 | Qual | Customer pain points | Discovery interviews | Positive | 8 of 10 cited onboarding complexity | High |
808 | E3 | Quant | Ad targeting | Ads | Null | CTR +2% (not sig., p=0.12) | Medium |
809 | E4 | Qual | Ad message resonance | Focus groups | Negative | 6 of 8 found messaging confusing | High |
810
811 **For Quantitative experiments (hypothesis-testing):**
812 - Signal classification (Positive/Negative/Null/Mixed)
813 - Key metric measured (CTR, conversion rate, etc.)
814 - Magnitude of effect with statistical significance (e.g., "+18%, p<0.05")
815 - Confidence level from statistical analysis
816
817 **For Qualitative experiments (qualitative-research):**
818 - Signal classification (Positive/Negative/Null/Mixed)
819 - Key themes identified with prevalence (e.g., "8 of 10 participants mentioned X")
820 - Representative quotes or patterns
821 - Confidence assessment (credibility, dependability, transferability)
822
8233. **Invoke presenting-data skill for comprehensive synthesis**
824
825 Use the `presenting-data` skill to create complete synthesis with visualizations and presentation materials:
826
827 ```markdown
828 Use presenting-data skill to synthesize marketing experimentation results:
829
830 Context:
831 - Campaign: [campaign name]
832 - Experiments completed: [count]
833 - Results table: [paste aggregate table]
834 - Audience: [stakeholders/decision-makers]
835 - Format: [markdown report | slides | whitepaper]
836 - Focus: Pattern identification across experiments (what works, what doesn't, what's unclear)
837 ```
838
839 **presenting-data skill will handle:**
840 - Pattern identification (using interpreting-results internally)
841 - Visualization creation (using creating-visualizations internally)
842 - Synthesis documentation (markdown, slides, or whitepaper format)
843 - Citation of sources (individual hypothesis-testing sessions)
844 - Reproducibility (references to experiment locations)
845
846 **Focus areas for synthesis:**
847 - **What worked:** Experiments with Positive signals (both quantitative and qualitative)
848 - **What didn't work:** Experiments with Negative signals (both quantitative and qualitative)
849 - **What's unclear:** Experiments with Null or Mixed signals
850 - **Cross-experiment patterns:** Do results cluster by tactic? By audience? By timing? Do quantitative and qualitative findings align or conflict?
851 - **Triangulation:** Do qualitative findings explain quantitative results? (e.g., interviews reveal WHY conversion rate increased)
852 - **Confounding factors:** Are there external factors affecting multiple experiments?
853 - **Confidence assessment:** Which findings are robust? Which are uncertain? How do qualitative and quantitative confidence levels compare?
854
8554. **Document patterns and insights**
856
857 The presenting-data skill will create `05-synthesis.md` (or slides/whitepaper) with:
858
859 **What Worked (Positive Signals):**
860 - List experiments with positive results
861 - Explain WHY these worked (based on analysis)
862 - Identify commonalities across successful experiments
863
864 **What Didn't Work (Negative Signals):**
865 - List experiments with negative results
866 - Explain WHY these failed (based on analysis)
867 - Identify lessons learned
868
869 **What's Unclear (Null/Mixed Signals):**
870 - List experiments with inconclusive results
871 - Explain potential reasons (insufficient power, confounding factors, etc.)
872 - Identify what additional investigation is needed
873
874 **Cross-Experiment Patterns:**
875 - Do results cluster by tactic, audience, timing, or other factors?
876 - Are there confounding variables affecting multiple experiments?
877 - What overarching insights emerge?
878
879 **Visualizations (created by presenting-data):**
880 - Signal distribution (bar chart: Positive/Negative/Null/Mixed counts)
881 - Effect sizes (bar chart: metric changes by experiment)
882 - Confidence levels (scatter plot: effect size vs. confidence)
883 - Tactic performance (grouped by channel/tactic)
884
8855. **Classify overall campaign signal**
886
887 Based on aggregate analysis (from presenting-data output), classify the campaign:
888
889 **Positive:** Campaign validates concept, proceed to scaling
890 - Multiple experiments show positive signals
891 - Successful tactics identified for scale-up
892 - Clear path to ROI improvement
893
894 **Negative:** Campaign invalidates concept, pivot or abandon
895 - Multiple experiments show negative signals
896 - No successful tactics identified
897 - Concept doesn't resonate with audience
898
899 **Null:** Campaign results inconclusive, needs refinement
900 - Most experiments show null signals
901 - Insufficient power or confounding factors
902 - Needs redesigned experiments or longer observation
903
904 **Mixed:** Some aspects work, some don't, iterate strategically
905 - Mix of positive and negative signals across experiments
906 - Some tactics work, others don't
907 - Selective scaling + pivots needed
908
9096. **Review presenting-data output and finalize synthesis**
910
911 After presenting-data skill completes:
912 - Review generated synthesis document
913 - Verify all experiments are covered
914 - Confirm visualizations are appropriate
915 - Ensure signal classification is documented
916 - Make any necessary edits for clarity
917
9187. **STOP and get user confirmation**
919 - Review synthesis findings with user
920 - Confirm pattern interpretations are accurate
921 - Confirm overall signal classification is appropriate
922 - Do NOT proceed to Phase 6 until confirmed
923
924**Common Rationalization:** "I'll synthesize results mentally - no need to document patterns"
925**Reality:** Mental synthesis loses details and creates false confidence. Documented synthesis with presenting-data skill ensures intellectual honesty and identifies confounding factors you'd otherwise miss.
926
927**Common Rationalization:** "I'll
928
929…(truncated)